Table of Contents
Why adding more electrodes will not fix them
I watched Cyberpunk: Edgerunners last month, and over the past week I listened to the eight hour Neuralink episode of the Lex Fridman podcast across a couple of evenings.
In the show there is a kid called David who buys a piece of hardware called a Sandevistan off a dead soldier and slots it into a port at the base of his skull. It lets him slow down time. He crosses a room before anybody in it has finished turning their head.
The series spends most of its running time on the hardware. How much chrome a body can hold before it rejects it, what the good implants cost, who you buy them from, what happens to people who install too much.
It never shows anybody teaching their implant what they meant. David thinks about moving and the chrome moves. Nobody in Night City sits through a setup session so the hardware can learn to read them.
It is fiction, so I cannot complain that it skipped a step. But that step is the one I could not stop thinking about.
A brain implant is a strip of metal and polymer sitting in tissue. It picks up voltages. Voltages are not instructions, and nothing about a rising voltage on electrode 340 says move left rather than clench your jaw. Somebody has to write the software that turns one into the other. To write it they need examples of what David’s brain looked like when he meant to move left.
In the back half of the recording, Bliss Chapman, who runs Neuralink’s decoding software, mentions almost in passing that this is not a problem you solve by adding channels. He says it again later in slightly different words. His job is the one that would get easiest if more electrodes were the answer, so it is a strange thing for him to be arguing.
To see why he is right, start with the person the argument is about.
Noland Arbaugh was twenty two, working at a summer camp, when he dove into a lake in 2016 and hit something under the surface. The impact broke his neck at the fourth and fifth vertebrae. He has been paralyzed from the shoulders down ever since, with no movement in either arm or leg. He can speak and turn his head a little. He also gets muscle spasms he cannot control, which rules out more assistive devices than you would expect.
For the eight years after that, the way he used a computer was a stick held in his teeth, tapping at a tablet propped in front of him. It works. It is also exhausting, it gives him pressure sores, and it stops him talking while he does it. Somebody else has to put it in his mouth and take it out again. If his mother has gone to bed, he is finished using computers for the night.
In January 2024 he became the first person to receive Neuralink’s N1 implant. The device is about the size of a coin and sits in a hole cut in his skull, with sixty four flexible threads pushed a few millimeters into the part of his brain that handles hand movement. He now moves a cursor by thinking about it, at roughly 8.5 bits per second.
Bits per second here has nothing to do with internet speeds. The measure scores how fast somebody can hit targets on a screen. It comes from an old rule in interface research about how hitting a small target far away is harder than hitting a big one close by. Your score is how much of that difficulty you get through per second.
The old record for a human with a brain implant was 4.6. An average person using a mouse scores about 10. One of the Neuralink engineers, who plays the same benchmark on his apartment floor at two in the morning after eating a lot of peanut butter, scores 17.
Noland got to 8.5 within a few months of surgery. That is most of the way to an average person with a mouse, and half of what somebody gets who has practiced the same benchmark for years.
A neuron sits with a voltage across its outer membrane, around 70 millivolts, negative on the inside. Signals arrive from other cells and nudge that voltage up or down. If enough arrive close enough together, the voltage crosses a threshold and the cell does something sudden that always looks the same. Channels open, sodium floods in, the voltage swings positive, overshoots, and drops back.
The whole event takes about a millisecond. It is called an action potential, or a spike. Two researchers worked out the mechanism in a squid’s giant nerve fibre in 1952, and their model is still the one everybody uses.
A spike is all or nothing. It does not come in different sizes. A neuron cannot fire a bigger spike to mean something more strongly, the way you might say a word louder. The only two things it can change are how often it fires and exactly when.
So when a paper says a cell encodes something, it usually means the rate. Work through the 1980s showed that cells in the motor cortex each have a direction they prefer. A cell fires fastest when the arm reaches one particular way, and slows as the reach turns away.
The same firing rate could mean several different reaches, so one cell on its own tells you very little. Take a few hundred cells, weight each one’s preferred direction by how hard it is firing, add the arrows together, and the sum points roughly where the arm is going. That sum is called a population vector.
An older experiment gets at the shape of the problem more directly. In 1969 somebody put an electrode into a monkey’s motor cortex, picked a single cell out of the recording, and wired that cell’s firing rate to two things. One was a meter the animal could watch. The other was a dispenser that handed out food pellets when the rate went up.
The monkeys learned to do it, which is not surprising on its own. What is surprising is that after a few sessions they could take a cell isolated for the first time that morning, a cell nobody had ever recorded from, and drive its rate up by 50 to 500 percent within a single session. They were learning a general skill and pointing it at whatever cell the electrode happened to sit beside that day.
The meter is the part I keep coming back to. The monkey could not feel that neuron. It had no idea which cell it was or what the cell normally did. It got better because it could watch a needle move.
A spike lasts about a millisecond, so to catch its shape you sample much faster than that. The sampling theorem says you need samples arriving more than twice as fast as the fastest thing in the signal. Neuralink’s implant reads all 1,024 electrodes 20,000 times a second, which puts about twenty samples inside each spike. That part is an engineering cost you pay in power and heat.
Distance is the actual difficulty.
The voltage a spike produces outside the cell drops off steeply as you move away, and the neighboring cells do not stop firing while you listen. Past roughly 100 microns, a cell stops being separable from the crowd. Not because it has gone quiet. Because everything else at that distance has blurred into a floor of noise you cannot pull it out of.
A micron is a thousandth of a millimetre, and a human hair is 80 to 100 of them across. So the useful listening range around an electrode is about one hair’s width. Inside a cube 100 microns on a side there are around forty neurons, and an electrode in the middle of that cube can usually pick out two or three.
A stadium is a fair picture of it. From outside the building you hear the crowd, so you know roughly whether the game is going well. You cannot hear the score, and you cannot hear what any one person is saying.
An EEG cap, reading voltages through the scalp, is out on the street. An ECoG grid, laid on the surface of the brain under the skull, is up in the stands. An electrode pushed into the tissue is a microphone in the huddle.
Once you are going inside, the question is what you leave in there. For thirty years the answer was the Utah array. A bed of rigid silicon spikes about four millimeters square, usually ninety six of them, pushed into the cortex with a small pneumatic inserter, recording only at the tips, with wires running out through a plug in the skin.
Nearly all the published human work, including the BrainGate trials, came off those arrays. They also fail in a particular way.
The brain is not still. It pulses with your heartbeat, it shifts when you breathe, and it moves when you turn your head. A rigid silicon spike does not move with it. So the tissue walls the spike off with scar, the scar pushes the neurons back from the tip, and the recording dies. One group at Caltech went as far as testing whether a chemotherapy drug could suppress the scarring, which tells you how stuck the problem was.
Neuralink’s answer is to make the thing floppy. Each thread is a strip of polymer with metal traces running down it, 16 microns wide at the tip, widening to 84, and less than 5 microns thick. Thinner than the hair I used as a measuring stick two paragraphs ago, and far too limp for a surgeon to place by hand.
A robot does it, catching a loop at the end of each thread on a needle and pushing them in one at a time, steering around the blood vessels it can see through a microscope. Sixty four threads, sixteen electrodes on each, 1,024 in total.
Figure 1. The threads, at the scale they are made.
Seven months after implantation in an animal, stained slices show neurons pressed against the threads with almost no scar collagen near them.
I would not lean on that too hard. A histology slide in a company blog post is not a peer reviewed study of how the device holds up over years, and seven months in a pig is not ten years in a person.
Getting the electrodes into the right place creates a second problem, because now something has to be done with everything they hear. A thousand channels, 20,000 times a second, at 10 bits a sample, comes to about 200 megabits per second. None of that is allowed to leave the head.
The reason is heat. The implant sits in a body running at 37 degrees and is not allowed to warm the tissue around it by more than about two degrees, because past that you start cooking the thing you came to listen to.
Two degrees sets a power budget. That budget decides how much computation the chip can do, how strong its radio can be, and how much shielding goes around the charging coil so the battery does not double as a heater.
So the chip throws away almost all of its own data before transmitting. Each channel runs a spike detector, which works something like a matched filter. You know roughly what a spike looks like, a dip of a certain depth with a certain recovery over a certain duration, so you slide that shape along the signal and emit a bit whenever it matches.
Two hundred megabits become about one. It happens in under a microsecond per channel, because the delay budget has no room for anything slower.
Figure 2. What the chip records against what it sends.
What gets thrown away is not nothing. You lose spike sorting, which is the practice of comparing waveform shapes to work out which spikes came from which neuron. You lose the waveform itself, which is how anyone checks afterwards whether a detected spike was real or an electrical artifact. What you keep is a count, per channel, per window of time.
Throwing that away is a bet on rate coding. Bats appear to use spike timing for navigation, and there are parts of hearing where microseconds carry the message. Motor cortex has behaved so far.
The same measurement has a gentler version, and it rescues Noland’s implant when the threads start pulling out. Instead of detecting individual spikes, you measure how much energy sits in the frequency band where spikes live, roughly 300 to 1,000 hertz, and record that one number per channel. By 2020 it had been shown that this spiking band power tracks movement about as well as counting spikes does, and costs far less power to compute.
Compressing that aggressively might sound like it would make the system sluggish. The opposite is true.
From a spike firing in Noland’s cortex to the cursor moving takes about 22 milliseconds. A good gaming mouse is around 5. The path from your own motor cortex to your own hand moving is about 75.
Figure 3. Intention to cursor, against intention to hand.
The implant beats the arm because it taps the line further up. The cortex represents a movement before the body performs it. Between the cortical command and the hand moving sit the spinal cord, the peripheral nerve, the junction between nerve and muscle, and the muscle itself. Every one adds delay.
Noland says the cursor sometimes seems to move before he means it to. I think he is watching his own intention arrive without the lag his body used to add.
Twenty two is not a physical limit either. Bluetooth Low Energy sets the current floor at 7.5 milliseconds between updates, with packets batched around 15. Under that the screen becomes the problem, since a 120 hertz monitor paints a new frame every 8.3 milliseconds.
The stated goal at Neuralink is not a usable mouse but the best mouse there is. Attached to it is a prediction that within ten years competitive gaming will be dominated by people with paralysis.
I wrote that off as a good line for a podcast. Then I tried to find the step that was wrong and could not. They get the intention tens of milliseconds before the muscle would have moved, which is time nobody with hands can recover. They have more free hours to practice than anyone holding down a job. And nobody knows where the ceiling is. I still think it is a stretch, and I cannot tell you which part of it is false.
Everything to here is instrumentation, and instrumentation improves when a company spends money on it. Reliably, on a schedule, in a way you can put on a slide for investors.
The software half does not improve like that.
A decoder is the software that takes spike counts off the implant and puts out a cursor velocity. Stripped of the neuroscience vocabulary it is supervised learning, the same basic shape as a spam filter.
Supervised learning needs examples, and every example has two halves. The input is a short window of spike counts, maybe twenty milliseconds wide, across a thousand channels. The second half is the label, the correct answer for that window. You fit a function mapping the first to the second, then check whether it works on examples it has never seen.
The population vector is a weighted sum, which is linear regression under another name. Then the Kalman filter, which adds the assumption that a cursor’s position is smooth over time. Then neural networks, eventually. In 2022 a shallow feedforward network beat a refit Kalman filter on finger movement by 36 percent throughput.
I had not expected the date to be that late. Several earlier attempts with convolutional networks and expanded Kalman variants failed to beat plain linear decoding at all. Linear methods held the state of the art in motor decoding for most of a decade after they had stopped holding it anywhere else in machine learning.
Vision and language went the other way, with deeper models pulling ahead early and never giving the lead back. The difficulty here is sitting somewhere other than the model.
Put an implant in a monkey with a working arm and the label half is free. The monkey moves a joystick, you record the joystick position, and you know exactly what the animal was doing during every window of neural data.
Put an implant in a person who is paralyzed and there is nothing to record. The intention happens, the spinal cord fails to pass it on, and it leaves no trace outside the skull. You have the input half of every training example and none of the output half.
So the labels get manufactured. You run a calibration task. A cursor slides across the screen, you ask the person to follow it with their intention, and then you write down that they intended to move right at the displayed speed for every millisecond the cursor was moving right.
Nobody measured that label. You assumed it, because you asked the person to do something and they said they would.
Different tasks assert different things. Following a moving cursor assumes their intended speed tracks the displayed speed moment by moment, which cannot be true in detail. Moving an invisible cursor two hundred pixels right assumes only that they can hold a steady intention for a couple of seconds, which is a smaller and therefore safer claim.
Whichever task you choose, you build its assumption into the decoder, and nothing downstream can tell that an assumption is what it was. Sixteen thousand electrodes will measure the wrong thing with beautiful precision.
A second problem sits underneath.
Offline validation cannot catch it. The standard defense against fooling yourself is to hold out data and test on it, but the held out data was labeled by the same task carrying the same assumption, so it is wrong in the same direction as the training data. Cross validation catches a bad fit. It cannot catch a bad premise. The one test that can is putting the model in front of the person and watching them try to use it, because their success or failure owes nothing to your assumption. That test is slow, needs the participant in the room, and cannot be run overnight on a cluster, which is why people keep skipping it.
If manufactured labels sound like an unfortunate compromise, a result from 2012 makes it stranger.
In monkeys you do have ground truth, so the obvious thing is to train the decoder to predict the hand movement you measured. One experiment did something else. It took cursor velocities recorded during an earlier session and changed them twice before using them as labels. Every velocity arrow was rotated to point straight at the target. Whenever the cursor was already on the target, the velocity was replaced with zero.
Both changes are false as descriptions of what happened. The cursor was wandering and correcting the whole way, and the intended velocity during a hold was not exactly zero.
Yet the decoder trained on that fiction outperformed the one trained on the real recording. By enough that the method, called ReFIT, became what everyone compared themselves against for years.
The measured path contains tremor, overshoot, correction, and arm inertia that the cortex never commanded. You are trying to predict the command. What you recorded is the consequence of the command after a body has interpreted it.
So the more faithfully you record what the body did, the further you drift from what the person meant. In the one setting where ground truth is available, it is actively misleading.
Noland found the human version himself. For his first few weeks he drove the cursor by attempted movement, physically trying to move a hand that no longer receives the command, while the decoder read the intention on its way out.
Then one day, in the middle of the target game, he stopped trying to move his hand and simply thought about the cursor going somewhere. It went. His reaction is on tape and it is not printable.
Nobody told him to try that. And there is a wrinkle: calibration done with attempted movement produces models that support the direct kind of control, while calibration with imagined movement does not work nearly as well. The signal you train on is not the signal that ends up driving the thing. I do not know why, and as far as I can tell neither does anyone else.
Now two adaptive systems point at each other. Software being fitted to a brain, and a brain learning to drive the software. This was set out properly in 2014. Write it as a joint optimization, brain as encoder and software as decoder with both halves moving, and an awkward consequence falls out. The pair can settle into a bad equilibrium and sit there. Neither half is broken.
Figure 4. The loop, drawn as a control system.
Noland’s version of stuck looks like this. He finds an angle to hold an imagined motion at that works around some quirk in that day’s model, and posts good numbers with it. The next morning he cannot find it again. Nothing anyone wrote down says which half moved.
The loop also breaks the metric. The team once swapped a fully connected decoder for a small convolution over time on each channel. Swapping in a convolution is standard for time series data. Fewer parameters, better fit, better generalisation on held out data. You would wave it through in code review.
Every offline number improved. In his hands it got worse.
This is the same trap that shows up everywhere else in applied machine learning, where the metric you can compute cheaply and the outcome you actually want come apart. Offline you measure average error. What a person feels is where the error sits in time. A decoder wrong smoothly can be steered around, the way you steer a car that pulls slightly left. A decoder with identical average error delivered in sudden jumps is unusable.
Figure 5. Predicted against observed performance.
None of this is a Neuralink discovery. A closed loop human simulator was built in 2011 for exactly this reason, because offline analysis kept mispredicting online performance. They put human subjects in the loop driving previously recorded monkey data, so decoders could be screened before anyone spent animal time on them.
By 2019 the loop itself was being modelled as a feedback control system, so decoder parameters could be chosen by simulating closed loop behaviour instead of grid searching over offline error.
The field has known that offline metrics lie for more than a decade. It keeps relearning it, because offline metrics are the cheap ones.
The same trouble shows up in a second form. A real user does not experience all mistakes as equal.
Velocity errors are cheap. Position is the running total of the output, the person corrects continuously, and a decoder right on average still gets to the target. Clicks are expensive. A click resolves instantly, usually cannot be taken back, and a false one closes the tab with the work in it.
The expensive errors arrive when they cost most. People slow down and hold still just before they click, so in the training data holding still and intending to click look almost identical. Early on the two ate each other in both directions. Noland’s click collapsed whenever he moved the cursor, and his movement collapsed whenever he tried to click and drag.
There are two currencies for paying this down. Time is the dwell cursor, where holding still over a target for three tenths of a second counts as a click. Dwell is reliable and taxes every selection by that third of a second. It also means he keeps the cursor drifting constantly so it does not click things he never asked for. He caught himself making the motion in bed one night with the implant switched off.
Interface is the better value. The application looks at what is on screen, works out that a close button is a small target, and widens its catch radius when the cursor behaves as though it is heading there. Scroll bars get turned into something the cursor snaps onto, with momentum, because a scroll driven by a noisy signal without those two properties makes the page shake.
Your phone has done a version of this for fifteen years, silently resizing the hit areas of letter keys based on what word a language model thinks you are typing. Noisy input, predictable intent, and nobody calls it cheating.
What that changes is the goal. You do not need a decoder that is right so much as a system where being wrong is cheap.
The human brain is about ten times the size of the animal brains Neuralink had worked in, and it moves a lot more. About four weeks after surgery, the threads started pulling out of Noland’s cortex.
Three signals at once. His performance was falling, the impedance measured at the electrodes changed, and because each thread records at sixteen depths the team could watch the signal march up the thread as it withdrew. A single depth array does not give you that third signal.
The obvious response is to open him up and put the threads back. Instead they changed what the device measures. Rather than detecting individual spikes, the implant switched to reading spiking band power, which keeps working at distances where you can no longer resolve any single cell.
It shipped as a firmware update, over the air, into a device inside a man’s head. They retrained the decoder, control came back, and then it went past where it had been.
Figure 6. The four published bits per second numbers.
Before the retraction his best was 7.5 bits per second, with two separate click types available. After the fix, with only the dwell click and its third of a second tax, he set 8, then 8.5.
Nothing was repaired. The information was still coming out of electrodes that had physically moved to the wrong place. The extra channels did not make the reading more accurate. They gave the engineers somewhere else to look on the day the first method failed.
The retraction was a loud version of something quiet that happens all the time. A neuron’s baseline firing rate today is not what it was yesterday. Electrodes settle by a few microns, tissue responds, the recording drifts. What the user experiences is the cursor sliding steadily toward one edge of the screen.
Noland fixes that by hand. He flicks the cursor to the side of the screen, a panel opens, and he tunes the bias out along with gain, smoothing and friction. He has become good at it. What that describes is a medical device requiring its user to be a part time control engineer.
The standard alternative is worse. Stop using the computer, run another calibration block, generate fresh manufactured labels, retrain. Every recalibration is time the person is not using the computer they got an implant to use.
Figure 7. A year of drift, relabeled and not.
A group then did something about it that lands directly on the labeling problem, and it is my favourite idea in the field.
Their system decodes attempted handwriting into text, then runs a language model over the output to fix errors. So far that is ordinary autocorrect. The clever move is what happens next. The corrected text goes back in as the label. The model’s best guess at what the person meant to write becomes the training target for the next sentence.
They call these pseudo labels, and the decoder updates continuously while the person is simply using the device. They ran it with one participant for 403 days with no supervised calibration at all, holding accuracy at 93.84 percent.
The labels still do not exist and are still being manufactured. What changed is where the assumption comes from. A calibration task assumes the person did what the screen told them. This assumes the person was writing English.
English is far more constrained than a cursor trajectory, and the assumption is available every second the device is in use rather than only during a setup block.
That result is the strongest argument against everything I have written here, and I want to put it at full strength. If a decoder can hold 93.84 percent for 403 days with nobody running a calibration session, then the labeling problem is not a wall. It is something you route around with a good enough prior, and the binding constraint goes back to being bandwidth, which is exactly what more electrodes buy.
My answer is that the prior did the work, and the prior was English. A language model knows which letter sequences are possible because millions of people wrote them down, and none of that was collected for this purpose. Handwriting and speech both sit on top of that free structure. General motor control does not. There is no corpus of what a hand meant to do next, no dictionary of valid reaches, nothing that says this trajectory is plausible and that one is not. Where a strong prior exists the labeling problem softens, and where it does not exist it stays exactly as hard.
Which leaves the question of what the roadmap of ever more electrodes is for. A thousand channels now, three to six thousand next, sixteen thousand after. Channel count against control quality is reportedly a log curve, so every doubling buys less than the last.
Accuracy is the least of it. Reliability is better: if drift is roughly independent across channels, more channels average it away, and the system gets steadier without anybody touching the decoder.
Actions are the largest. Electrodes sit beside cells representing different intended movements, and a thousand of them cover the hand region densely while covering almost nothing else. Sixteen thousand would reach further across the body map, and every imagined motion you can resolve reliably becomes a button. A mouse beats an eye tracker because it has buttons, not because it points more precisely.
Every one of those buttons arrives with the same problem the cursor had. An array that can resolve an imagined toe curl still has to be told what a toe curl means, and there is no more ground truth for a toe curl than there was for cursor velocity. You would run another calibration task, assert another set of labels, and inherit another assumption. More electrodes multiply the number of things you have to guess about.
Everything so far has been about Neuralink because that is what the recording was about. The clearest way to show the argument is not about Neuralink is a company doing almost the opposite.
Synchron’s device never goes through the skull. Their Stentrode is a stent, the mesh tube used to prop open blocked arteries, threaded up through the jugular vein on a catheter and parked in a blood vessel running over the motor cortex. It reads through the vessel wall. It carries sixteen electrodes against the N1’s 1,024.
It works because they spend the difference on interface. In May 2025 Apple shipped a protocol called BCI HID that makes a brain signal a native input type on the iPhone, alongside touch and voice, and Synchron integrated with it first.
The device also sends live screen context back, so the decoder knows what is on the screen before it decides what the person meant.
That is the consequence of sixteen channels in a blood vessel. You cannot pull a smooth two dimensional velocity out of that, so the interaction becomes selection rather than free movement, built on the Switch Control accessibility feature that has been in the operating system for years, walking a highlight through a list of things the person might have meant.
Paradromics pushes into the cortex with high channel counts and aims at speech. Precision Neuroscience lays a thin film on the surface without penetrating it. Three very different bets about how much of a skull you have to open, and all three land on the same question the moment the hardware works, which is what the person was trying to do.
In 2021 attempted handwriting got decoded. The participant imagined writing letters one at a time, and the system reached 90 characters per minute at about 94 percent raw accuracy, or 99 percent with autocorrect running on top. That is close to smartphone typing speed for someone his age, produced by a man who had not moved a hand in years.
The question I found more interesting is why handwriting rather than typing. Imagined letter shapes are easier to tell apart in the neural data than imagined key presses, because the movements differ more. They chose the task whose labels separate more cleanly and called it an interface decision. That choice moved performance further than any change of model architecture in the paper.
Then speech, in 2023, out of the same lab. A participant with ALS who can no longer speak intelligibly attempted to speak, and the system reached 62 words per minute, with a 9.1 percent word error rate on a fifty word vocabulary and 23.8 percent on 125,000 words. Ordinary conversation runs around 160.
The labels there are manufactured the same way. You show the person a sentence, ask them to attempt it, and assume that is what they attempted at roughly the pace on screen. Then a language model cleans up the output.
Put the three side by side and the same arrangement shows up. Handwriting has an autocorrect. Speech has a language model. The cursor has a close button that quietly gets bigger.
If labels are the bottleneck, the obvious machine learning answer is to need fewer of them.
The obstacle is that every recording session is its own small world. Different electrodes, different neurons, a different animal, a different number of channels in a different order. A model fitted to Monday’s array means nothing on Tuesday’s, so each session got its own model trained on its own labels, which caps the dataset at whatever one person can produce before they get tired.
POYO gets around that by treating individual spikes as tokens rather than binning them into rates, then running cross attention onto a fixed set of latent tokens, so recordings with different channel counts can all feed one model. NDT2 does a related thing with a transformer pretrained across sessions, subjects and tasks.
Figure 8. Pretraining against single-session training.
What that buys, if it works, is a decoder arriving already knowing roughly what motor cortex looks like, needing only a little person specific data. Calibration would go from forty minutes to a few.
Figure 9. A pretrained decoder on a new day.
I am more cautious for one reason. The strong results are largely on monkey datasets, which is exactly the setting where real labels exist and the problem does not bite. Whether a model pretrained on animals with working arms transfers to a person whose labels were invented is a different question, and not the one those papers set out to answer.
About 300,000 people in the United States are living with a spinal cord injury, and about 18,000 new ones happen every year. Tetraplegia accounts for around 60 percent, so call it 180,000 people who cannot use their hands. Widen it to include stroke, multiple sclerosis and ALS, and a large household survey put the number living with some form of paralysis at about 5.4 million.
Eye trackers exist and work. Mouth sticks exist and work. They also require the user held in position in front of a camera, or an object placed between their teeth by somebody else. Noland has spasms, so anything depending on him staying in frame fails, and anything rigid in his mouth is a hazard.
Every one of those has to be set up for him by another person, which puts his access to a computer on somebody else’s schedule.
The implant is inside him. He can use it at two in the morning with the house asleep. He can text a friend without his mother in the loop. He plays a target selection game until four or five because he wants his own number to go up, and what stops him is the battery percentage in the corner of the screen.
Almost all of this was measured in one person, and an unusual one. Noland is articulate, competitive, endlessly available, and he spent years before the surgery trying to move in ways that produced nothing visible, which is close to ideal preparation for driving this device. Most of the design decisions were tuned against him.
That is less true than a year ago. Neuralink said in January that 21 people had received the implant, across the United States, Canada, the United Kingdom and the UAE. Several are reported in the same 8 to 10 range. The company also says the thread anchoring problem has not recurred, though I could not find that in anything peer reviewed.
Four numbers will tell you whether this is working, and they matter in this order.
Whether multiple click types come back. Dwell puts a hard floor of three tenths of a second under every selection, and the current record was set while paying it. A decoded left and right click that does not contaminate the movement signal is the change most likely to push past 10.
Calibration time. Around forty minutes for a good model against a stated target of under seven. Forty minutes of setup makes it a laboratory instrument.
Whether the route to direct control transfers. Noland took weeks and found it himself, by accident, in the middle of a game. If new participants get there in days because somebody mapped the route, the knowledge compounds. The growing participant count means this has an answer now, and I have not seen anyone publish it.
Hours of independent use with nobody from the company in the loop. The other three are proxies for this one.
So here is where all of that leaves things.
The hardware side is close to finished. The electrodes work, the surgery is routine enough to do in a few hours, the radio fits inside a two degree heat budget, and the whole path from a spike to a cursor takes 22 milliseconds. A device has now sat in a man’s head for two years and got better over that time rather than worse.
The part that is not finished is the step in the middle. Somebody has to tell the software what the person meant, and there is no way to ask them. The person cannot move, so nothing they intend leaves a mark anyone can record. Every training example is missing its answer.
So every result in this essay comes from finding a better way to guess that answer. Rotate the recorded velocities so they point at the target. Let an autocorrect rewrite the handwriting and train on the corrected version. Send the decoder a copy of what is on the screen so it has fewer options to choose between. Not one of those is a better electrode. They are all better guesses.
That is why the channel count is not the thing to watch. Going from a thousand electrodes to sixteen thousand gives you a more precise measurement of a signal whose meaning you are still assuming. It buys reliability, and it buys more buttons, and both are worth having. It does not buy you the answer key.
David never had this problem. He slots the Sandevistan into the port at the base of his skull and it reads him correctly from the first second, because the show does not have to explain where that came from. In the real version, working out what somebody meant is most of the job.
reading list
Everything above is my own argument. Everything below is where it came from.
Lex Fridman Podcast #438, with Elon Musk, DJ Seo, Matthew MacDougall, Bliss Chapman and Noland Arbaugh. Source of the bits per second figures, the surgical account, the thread retraction, the convolution result, and the remark about channel count this essay is built on.
An Integrated Brain-Machine Interface Platform With Thousands of Channels. Elon Musk and Neuralink, Journal of Medical Internet Research, 2019. Thread dimensions, channel counts, the robot. Figure 1 in this piece is its Figure 1, reproduced under CC BY-ND 4.0.
Neuralink PRIME Study updates. Company material, not peer reviewed. The histology images and the participant counts come from here and should be read accordingly.
Operant Conditioning of Cortical Unit Activity. Eberhard Fetz, Science 163, 1969. The monkey driving a single newly isolated neuron against a meter. Four pages.
Population coding and directional tuning in motor cortex. Apostolos Georgopoulos and colleagues, through the 1980s. The ancestor of every decoder described here.
The Hodgkin-Huxley model. Alan Hodgkin and Andrew Huxley on the ionic basis of the action potential, 1952.
The origin of extracellular fields and currents. Gyorgy Buzsaki, Costas Anastassiou and Christof Koch, Nature Reviews Neuroscience 13, 2012. Why the recording radius is roughly a hundred microns.
A low-power band of neuronal spiking activity dominated by local single units. Samuel Nason, Alex Vaskov, Matthew Willsey and colleagues, Nature Biomedical Engineering, 2020. Spiking band power, three years before it rescued the first human implant.
A high-performance neural prosthesis enabled by control algorithm design. Vikash Gilja and colleagues, Nature Neuroscience 15, 2012. ReFIT, and the rotated velocity vectors that beat the true recording.
A closed-loop human simulator for investigating the role of feedback control. John Cunningham, Paul Nuyujukian and colleagues, Journal of Neurophysiology 105, 2011.
Combining Decoder Design and Neural Adaptation in Brain-Machine Interfaces. Krishna Shenoy and Jose Carmena, Neuron, 2014. The two-learners framing.
Encoder-Decoder Optimization for Brain-Computer Interfaces. Josh Merel, Donald Pianto, John Cunningham and Liam Paninski, PLoS Computational Biology, 2015.
Principled BCI Decoder Design and Parameter Selection Using a Feedback Control Model. Francis Willett and colleagues, Scientific Reports, 2019. Figures 4 and 5 in this piece are its Figures 2 and 4, reproduced under CC BY 4.0.
High-performance brain-to-text communication via handwriting. Francis Willett, Donald Avansino, Leigh Hochberg, Jaimie Henderson and Krishna Shenoy, Nature 593, 2021. Ninety characters a minute.
A high-performance speech neuroprosthesis. Francis Willett, Erin Kunz, Chaofei Fan and colleagues, Nature 620, 2023. Sixty two words a minute, with a language model doing part of the work.
Real-time brain-machine interface in non-human primates using a shallow feedforward decoder. Matthew Willsey and colleagues, Nature Communications, 2022.
Plug-and-Play Stability for Intracortical Brain-Computer Interfaces. Chaofei Fan and colleagues, arXiv:2311 .03611, 2023. The 403 days without a calibration session, and the pseudo labels.
A Unified, Scalable Framework for Neural Population Decoding. Mehdi Azabou and colleagues, arXiv:2310 .16046, 2023. POYO.
Neural Data Transformer 2: Multi-context Pretraining for Neural Spiking Activity. Joel Ye, Jennifer Collinger, Leila Wehbe and Robert Gaunt, bioRxiv 2023.09.18.558113. Figures 8 and 9 in this piece are its Figures 3 and 5, reproduced under CC BY-NC 4.0.
Diffusion-Based Generation of Neural Activity from Disentangled Latent Codes. Jonathan McCart, Andrew Sedler, Christopher Versteeg and colleagues, arXiv:2407 .21195, 2024.
Spinal cord injury facts and figures. National Spinal Cord Injury Statistical Center, plus the Reeve Foundation paralysis prevalence survey for the wider number.
Synchron, Paradromics and Precision Neuroscience, including the Apple BCI HID integration. Company announcements and press material. None of it peer reviewed.