Blog

Notes from building it

What the app is for, what it honestly does and does not do, and the hard parts of building it — clocks, phones that disagree, and what it takes to make a room sound like one speaker. Tap any of them to read the whole thing.

13 September 2026

Six phones, no speaker

Everybody has a phone in their pocket. That is already a sound system, if they can be persuaded to agree about the time.

Read the post

Here is the situation this app was built for, and it happens constantly.

Somebody’s roof. Eight people. A cake, or a birthday, or nothing in particular. Somebody says put some music on, and the only thing anybody has is a phone. One phone, played through its own speaker, in the open air, with eight people talking over it. You know exactly how that sounds. Thin at first, then annoying, then somebody gives up and turns it off.

The strange part is that the room is not short of speakers. There are eight of them. They are just all playing nothing.

Why you cannot simply press play together

The obvious thing — everyone opens the same song and hits play on three — fails within seconds, and it fails for a reason worth understanding, because it is the same reason this is harder than it looks.

Two phones playing the same track a tenth of a second apart do not sound like one loud phone. They sound like a slap-back echo, and your ear notices it instantly because your ear is extraordinarily good at exactly this. Below about 30 milliseconds of separation you stop hearing an echo and start hearing a hollow, phasey quality — comb filtering, where some frequencies reinforce and others cancel. It sounds like the music is coming through a pipe.

Thumbs cannot hit play within 30 milliseconds of each other. Not close. Human reaction time to a spoken cue is somewhere around 200 milliseconds and varies by more than the whole budget.

What actually has to happen

So nothing in WatchBuddies ever says “play”. Each phone is told when the song started, in one shared frame of reference, and works out for itself where the playhead ought to be at this instant. Then it keeps checking, and nudges its own playback rate by fractions of a percent to stay there — the same trick NTP uses to keep a computer’s clock honest, applied to a song instead of a clock.

The upshot is that a phone joining four minutes late does not start at the beginning. It starts four minutes in, in step, and you do not hear it arrive.

What six phones actually sound like

Not six times louder. Sound does not work that way: doubling the number of equally loud sources adds about 3 decibels, so six phones are roughly 8 dB up on one — noticeably louder, not transformatively so.

The transformation is somewhere else, and it surprised us. Six phones scattered around a roof do not sound like a louder phone. They sound like the music is in the space rather than coming out of an object. Everybody is near a source, so nobody is straining. Nobody is crowded around a table with a phone on it. The thin, directional, one-small-speaker quality disappears, because it was never about volume — it was about the sound arriving from one point.

And the bass is still bad. Phone speakers cannot move enough air to make low frequencies, and six of them cannot either. Anybody who tells you otherwise is selling something. What you get is a room that sounds filled instead of a table that sounds tinny, which for eight people on a roof is the entire problem solved.

The cases it is genuinely good at

Rooftops and terraces, where there is no power socket and no speaker. Hostel and dorm rooms, where somebody always has music and nobody has a sound system. Long train journeys and road trips, where a group wants the same thing in their ears. Study rooms and small offices. Picnics. Any gathering that formed in the last ten minutes and will be over in two hours.

What it is not good at: replacing a real speaker when there is one. If somebody owns a decent Bluetooth speaker, use it. This exists for the far more common evening when nobody does.

11 September 2026

Why this exists

A rooftop in West Bengal, a dead speaker, and five people holding the answer.

Read the post

The idea arrived the way most of them do, which is to say it arrived as an irritation.

A rooftop, an evening that had been planned for a week, and a Bluetooth speaker that turned out to be flat. Not low. Flat, and the charger was an hour away in somebody else’s flat. So we did what everybody does: put a phone on the parapet wall, turned it all the way up, and stood near it.

It was fine for about four minutes. Then people drifted to the edges to talk, and the music became a noise happening at one end of a roof, and eventually somebody put it on their own phone too so they could hear it where they were standing.

Which is when it got annoying in a much more interesting way. Two phones, same song, about a second apart. It sounded terrible — genuinely worse than one phone — and somebody said the obvious thing: why can’t they just be in sync.

Because it is a clock problem, not a music problem

I went and found out. The answer is that it is not difficult in the way it first appears — the file is not the problem, the network is not really the problem — it is that no two phones agree what time it is, and they need to agree to within a few thousandths of a second before your ear will accept the result.

That is a solved problem in the abstract. NTP has kept the world’s computers in step for forty years by measuring round trips and estimating offset. Nobody had pointed it at a song on a group of phones in one room and made it simple enough that eight people could join by typing six characters.

So that is what WatchBuddies is. One person starts a party and reads out a code. Everybody types it. The phones sort out the time between themselves, and the roof has music on it.

What was actually hard

Not the syncing, in the end. The syncing is mathematics and the mathematics is known.

What was hard was everything around it. Phones whose clocks drift at different rates depending on temperature. Android devices that report an audio position that lags reality by an amount they will not tell you. A person who joins fourteen minutes into a song and must not hear a single moment of catching up. Somebody walking out of Wi-Fi range and back in. A phone that locks its screen and quietly deprioritises your audio thread.

And the least technical problem of all, which took the longest: making the joining bit so short that somebody who has had two drinks, in a dark room, with a phone at 4% battery, can do it without being talked through it. Six characters. No account needed to join. Ten minutes free before anybody is asked for anything.

That is the whole product. The clock is just what makes it possible.

9 September 2026

What ₹11 buys, and why it is ₹11

Only the person starting the party pays. Here is the arithmetic behind the number.

Read the post

A week of WatchBuddies costs ₹11. A month is ₹39, a year is ₹449, and if you would rather never think about it again there is a single payment of ₹2999 that does not expire.

People ask how it can be that cheap, usually with a suspicion that something is being harvested instead. It is worth answering properly, because the answer is just arithmetic.

Only one person pays

This is the part that changes everything. A party has one host and up to a roomful of guests, and the guests never pay and never even make an account. They type six characters and they are in. Ten minutes free, and a free sign-up carries them on from there.

So the price is not ₹11 per person at the party. It is ₹11 for the party. Eight people on a roof for a Saturday evening, and the bill is eleven rupees, paid by whoever happened to open the app first.

What it actually costs us to run

Very little, and deliberately so. A party’s traffic is a few kilobytes a second of clock probes and state — not audio. Uploaded tracks are fetched once by the people in one room and deleted within hours. There is no recommendation engine, no analytics pipeline, no ad stack, nothing running that is not answering somebody in a room right now. The app has zero runtime dependencies, which sounds like an aesthetic choice and is mostly a cost one.

A server that would be bored hosting a small blog can hold a great many parties at once. That is why the number can be eleven.

On the alternatives

The idea is not ours. Apps that sync music across phones have existed for years, and the best-known of them deserves credit for making anyone believe it was possible at all. What they generally are not is cheap — the category is priced in the hundreds of rupees a month, and it is worth checking today’s figure for yourself rather than taking ours. Against that, ₹39 a month is a different order of thing entirely.

We are also not pretending to be the same product. Those apps have years of features we do not have. What we have is the part that matters on a roof, priced like something you would not think twice about.

Things we will not do to keep it there

No advertising. Nothing that follows you between visits. No selling what you listen to, which would be the obvious way to make a free tier pay and is the reason there is not one.

What there is instead: three days of everything, free, with no card asked for. If the app does not survive one real evening with real friends, it has not earned eleven rupees and you should not pay it.

And cancelling

Two taps on your account page. It stops at the end of the period you have paid for, you keep what you bought, and we email you three days before every renewal with the amount and the date, so a payment is never a surprise. A subscription you cannot leave easily is not a subscription, it is a trap, and we would rather you came back than felt caught.

7 September 2026

Does it actually sound good? An honest answer

Yes, and not for the reason people expect. Also: here is what it will never fix.

Read the post

The honest version, including the parts that are not flattering.

It is not six times louder

Sound intensity does not add the way people assume. Two equally loud sources are about 3 decibels up on one, four are about 6, eight are about 9. Six phones land somewhere around 8 dB over a single phone, which is clearly louder but nothing like six times.

If you were hoping to fill a garden the way a proper speaker does, six phones will not do it, and anybody claiming otherwise is hoping you will not check.

What it does instead is better than louder

The real change is spatial. One phone is a point source: the music comes out of an object, it gets quieter fast as you move away, and everyone ends up orbiting a table. Six phones spread around a space produce something closer to a diffuse field. Everybody is near a source. Nobody strains, nobody has to stand in the right place, and the music stops being a thing in the corner and becomes the sound of the room.

That is the effect people react to, and it is not the one they were expecting.

The bass will still be bad

A phone speaker is a few millimetres of cone. It cannot move enough air to produce low frequencies, and six of them cannot either — you get six copies of the same missing bottom end. If the track lives on its low end, it will sound thin. That is physics, and no amount of software fixes it.

What we can control, and do

Timing. The whole engineering effort goes into keeping every phone within a few milliseconds of the others, because that is the difference between “the room is playing music” and “something is wrong with the music”. Past roughly 30 milliseconds of spread you hear comb filtering — a hollow, phasey, through-a-pipe quality — and past 50 or so it becomes an audible echo. Both are far more unpleasant than a single quiet phone, which is why so many attempts at this idea are abandoned after one try.

The app shows you the real figure while a party is running, measured on the device rather than assumed, so you are not taking our word for it.

How to get the best out of it

Spread the phones out rather than piling them on one table — the spatial effect is the point, and it needs distance. Put them on hard surfaces: a wooden table or a steel railing adds a surprising amount of body, a cushion swallows it. Turn every phone to full and leave them there. And if one phone is much louder than the rest, move it further away rather than turning it down, because distance costs you less than a quiet source does.

One phone left face-down on a soft chair is doing nothing for anybody. Face them up and out.

3 September 2026

A party belongs near the people in it

The distance problem in this app was never bandwidth. It is the clock.

Read the post

Every phone in a party works out where the music should be by measuring its own offset from one server, the way NTP does. Send it the instant the song started and it can calculate, for itself, exactly where the playhead ought to be right now. Nothing ever has to say “play”.

That measurement is only ever as good as the round trip it rides on. A phone in London talking to a server in Mumbai is 200 milliseconds away, and more to the point the variation in those 200 milliseconds is large. Jitter is the one thing a sync algorithm cannot see through: it can average out noise given time, but noise is exactly what it is trying to measure against.

The obvious answer is a CDN, and it is the wrong one. A CDN caches content, and content was never the problem — an uploaded track is fetched by six people in one room, once, and deleted within hours. There is nothing to cache. What needs to be close is not the file. It is the clock.

Why this turned out to be easy

A party is a room. Everybody in one is in one place. That single fact removes almost all of the difficulty, because it means regions never have to talk to each other. There is no shared database, nothing to replicate, no consensus to reach, no clock to reconcile across an ocean. Each region is a complete, independent copy of the whole app.

Two pieces make it work.

Where a new party goes. The app asks every region for the time and keeps whichever answers soonest. Not a guess from an IP address — a measurement. Geo-IP gets a VPN wrong, gets an oddly routed mobile network wrong, and gets any country whose servers sit outside its borders wrong. Asking “which of you is nearest” gets all three right, and it costs one extra request because it is the same probe the clock runs anyway.

How a join finds it. The first character of the party code says which region holds it. 2K7QMX is India, 3ABCDE is the United States. Any phone, anywhere, reads the code and talks to the right server directly. No lookup, no directory service, no round trip to find out where to go.

The whole routing layer is about ninety lines. With one region configured it behaves exactly as a single server always did, which is the property that mattered most: nothing about it is speculative infrastructure waiting for a second deployment to justify itself.

3 September 2026

A countdown that a restart cannot reset

Guests get ten minutes without an account. Keeping that clock on the phone would mean it did not work at all.

Read the post

The feature is simple to describe. Somebody handed a link or a QR code can listen for ten minutes without signing up for anything. When the ten minutes are up, a free account carries them on.

The difficulty is entirely in making it survive being closed.

A countdown held in the phone resets the moment the app is killed, the tab is closed, the page is reloaded, or the browser's storage is cleared. Which is to say: it does not work. Anybody who wanted more time would find the way to get it within about four seconds, and would not even need to be trying.

Two keys, and the earlier one wins

So the clock lives on the server, and the phone is only ever told what is left. That much is obvious. The interesting part is what the clock is keyed on, because no single key is good enough.

A ticket in a cookie. Precise, survives closing the app, and is the normal case. Also trivially cleared.

A one-way scramble of the network address and the browser string. Coarse, but it is what catches the cleared cookie and the private window, because neither of those changes the router.

Neither works alone. A cookie is cleared; an address is shared by a whole household. Taking the earlier first-seen of the two means the honest case is measured precisely by the cookie, and the evasive case still runs into the fingerprint.

A household behind one router shares the ten minutes. That is a real cost and it is the right way round: it errs toward asking somebody to sign up, never toward handing out unlimited free time. Getting that direction backwards is how a limit becomes decorative.

Wall clock, not accumulated minutes

The countdown runs from first sight rather than adding up active time. It is simpler, it cannot be gamed by closing the app between songs, and it is what a countdown means to the person watching one. The number on screen is re-anchored to the server on every heartbeat and only ticks locally in between, so a phone whose own clock is wrong, or which was closed for an hour, cannot end up believing it has more time than it does.

2 September 2026

Why a song is downloaded once, not twenty times

A four minute song is 85 MB once decoded. An ordinary phone holds about two. The prefetch window wanted four.

Read the post

This one cost people real money, and the arithmetic is the whole story.

decodeAudioData turns a compressed file into raw 32-bit samples. Stereo at 44.1 kHz is roughly 21 MB of memory for every minute of music, whatever the file weighed on the way in. So a four minute song is about 85 MB of memory. The prefetch window wanted four songs ready at once, which is 339 MB. The memory budget on an ordinary phone is 160 MB.

It never fitted. It could not ever have fitted.

What happened next: the app downloaded a track, decoded it, and the memory manager immediately threw it away to get back under budget. A moment later the next state update arrived, noticed the track was missing, and downloaded it again. Round and round, every few seconds, for as long as the party lasted. Measured on a real browser with real files: 81 requests for 4 songs.

Keep the file, not just the audio

The fix is small and it is not the one people reach for first. The player used to hold only the decoded buffer, so when memory pressure took it back, the only way to get it again was the network.

Now the compressed file it was decoded from stays behind. A four minute song is about 5 MB as a file and 85 MB as audio, so holding the file costs a sixtieth of holding the audio — and coming back costs a re-decode instead of a download.

One detail worth knowing if you ever do this: decodeAudioData takes the buffer away from you. It is detached and unreadable afterwards. So it gets a copy and we keep ours, which is the only reason a second decode is possible without a second download.

And ask for what fits

The second half: the number of tracks decoded at once is now worked out from the memory budget and the length of the songs actually in the queue, rather than fixed at four. On an ordinary phone that comes to two — the one playing and the one after it. The rest of the window is still fetched, because the file is small, so when the window slides the next song is already there.

81 requests became 4. Three tests now fail loudly if anyone ever makes it 5.

30 August 2026

Nothing here ever says “play now”

Messages arrive late, and differently late on every phone. So we never send one.

Read the post

The naive way to sync music across phones is to send every device a message that says “start now”. It does not work, and it cannot be made to work, because that message does not arrive at the same time on every phone. It arrives 40 milliseconds later on one, 180 on another, and 12 on the one sitting next to the router. You have not synchronised anything. You have distributed your network's jitter directly into the audio.

So WatchBuddies never sends that message. There is no “play” command anywhere in it.

An instant, not an instruction

What the server sends is an anchor: this track was at this position, at this instant, on the shared clock. Every phone then works out for itself where the music should be right now:

position = anchorPosition + (serverNow − anchorAt) / 1000

That is the entire idea. A phone that receives the anchor 200 milliseconds late calculates a position 200 milliseconds further along, and starts there. It is not behind. It was never going to be behind, because nothing about the calculation depends on when the message arrived.

A slow phone is not a late phone. That sentence is the whole design.

The shared clock

All of which rests on every device agreeing what “now” is. Each one measures its own offset from the server the way NTP does, with a four-timestamp exchange, and keeps measuring so the estimate improves and follows drift.

It does not matter whether the server's clock is right. It only has to be the same for everyone. If the server were ten minutes wrong, every party would still be in perfect sync, because every device measures against the same wrong clock and the error cancels out exactly.

28 August 2026

Two clocks, and the drift between them

No two sound cards run at exactly the same speed. Over four minutes that difference becomes audible.

Read the post

Getting every phone to start together is the part people expect to be hard. It is not the hard part.

The hard part is that they do not stay together. Two phones told to play the same four minute song, started at precisely the same instant, will finish at measurably different times, because the crystal oscillator clocking each sound card is not running at exactly 44,100 Hz. It is running at 44,100 give or take a few parts per million, and those few parts per million are different in every device.

A drift of 30 parts per million is 7 milliseconds over four minutes. That is audible in a quiet room as a smearing of transients. 100 ppm is 24 milliseconds, which sounds like a slap echo.

Chasing is not enough

The obvious correction is to measure how far off you are and nudge the playback rate to catch up. It half works. The problem is that a pure chase is always reacting to an error that has already happened, and as soon as it catches up the drift starts opening the gap again. The correction saturates and never settles.

What is needed is for the phone to learn its own crystal, not just react to the symptom.

The integral term

So the controller keeps a running term — rateBias — that accumulates the persistent part of the error. Once it settles, the audio is running at the rate that cancels the crystal difference, and the proportional part of the correction has almost nothing left to do. The phone has effectively worked out that its sound card runs 34 parts per million fast, and is permanently compensating.

Two details that matter more than they look. Corrections are only applied every 250 milliseconds, because correcting continuously makes every device chase its own measurement noise and they drift apart in a new and more interesting way. And errors below a threshold are ignored entirely: a millisecond is inaudible, and correcting it only adds movement.

The number is on screen in Sync & settings, under “Sound card”. Watching it converge on a real phone was, personally, the most satisfying moment of building this.

24 August 2026

Writing a QR encoder by hand, and the bug that hid in it

The format information bits were being written in reverse. The code looked perfect and no scanner on earth would read it.

Read the post

The join code is shown as a QR, and the encoder that draws it is written from scratch rather than pulled from a library. There is one reason for that: the app has to work with no connection. The interface is inside the APK and behind a service worker precisely so a party in a basement still opens, and fetching a script off a CDN to draw the join code would undo that on the one screen that most needs it.

It does byte mode, error correction level M, versions 1 to 10. Level M rather than L because this gets scanned across a room, on a screen with fingerprints on it, held by somebody who has had a drink.

The bug

A finished QR carries fifteen bits of format information — the error correction level and which mask pattern was applied — written twice, in two different places, in two different orientations. I had the placement loop right and the bit order backwards.

The result is a code that looks completely correct. All the finder patterns are there, the timing patterns are right, the data modules are genuinely correct data. Hold it up to any scanner in the world and nothing happens, because the scanner reads the format bits first, gets nonsense, and stops.

You cannot see this by looking. That is the whole lesson.

Three ways of checking

What found it was refusing to trust the thing at all. Every codeword was compared against Python's qrcode library. Then every module of the finished matrix was compared against the same. Then each one was read back by OpenCV's scanner.

The first two checks are what caught it, because the third only tells you that something is wrong, not where. Eleven cases now pass all three, and three golden matrices are in the test suite so it cannot come back. The browser test goes further and screenshots the QR a real browser rendered, then decodes the screenshot.

If you are ever tempted to write one of these: the specification is readable and the maths is enjoyable, but budget most of your time for verification rather than implementation. The implementation is a weekend. Knowing it is right is the work.