What did the mouse see? Your first neural decoder

In lesson 2 you asked whether a neuron responds to the stimulus. This lesson turns the question around: given only the neurons, can you tell what the stimulus was? You will turn a recording into a table, guess the stimulus from one neuron, discover why that guess cannot be trusted, then train a classifier on 66 neurons at once and find that it gets the answer right on 339 trials out of 340. Along the way you will build the three habits that separate machine learning from wishful thinking: hold out test data, cross-validate, and check what luck alone would score. At the end you will point the decoder at eight brain areas and at every moment of the trial, and watch the information appear.

This is the first step toward a brain-computer interface. A BCI is a loop (here is the map), and the stage in the middle of it, where neural activity is turned into a guess about what the user wants, is exactly what you are about to build.

flowchart LR

A[Spike counts: neurons x trials x bins] --> B[Table: one row per trial, one column per neuron]

B --> C[One neuron + a threshold]

C --> D[Train / test split]

D --> E[Population decoder: 66 neurons]

E --> F[Cross-validate]

F --> G[Shuffle the labels: what does luck score?]

G --> H[Every area, every moment]
The lesson in one picture. Each box is a step below.

The data

Same file as lesson 2: steinmetz_session11.zip, one session from Steinmetz and colleagues (2019), with 698 neurons, 340 trials and spike counts in 10-millisecond bins. On each trial a striped pattern appeared on the right screen at one of four contrasts, or the screen stayed blank. The question for the decoder is the simplest one possible: was there anything on the right screen, or not?

Step 1: turn the recording into a table

Every classifier ever built wants the same thing: a table with one row per example and one column per measurement, plus a list of the right answers. In machine learning the table is called X and the answers are called y. Here, an example is a trial, a measurement is one neuron’s spike count in the 50 to 250 ms window from lesson 2, and the answer is 1 if the right screen showed a stimulus and 0 if it was blank. Two lines of NumPy build the whole table.

import numpy as np
import matplotlib.pyplot as plt

dat = np.load("steinmetz_session11.zip", allow_pickle=True)
spks = dat["spks"]                        # neurons x trials x time bins: spike counts in 10 ms bins
area = dat["brain_area"]                  # one label per neuron
contrast = dat["contrast_right"]          # stimulus contrast on the right screen, one value per trial
dt = float(dat["bin_size"])               # 0.01 seconds
t = np.arange(spks.shape[2]) * dt - 0.5   # time of each bin relative to stimulus onset

post = (t >= 0.05) & (t < 0.25)           # the response window from lesson 2
visp = np.where(area == "VISp")[0]        # primary visual cortex neurons

X = spks[visp][:, :, post].sum(axis=2).T  # trials x neurons: each neuron's spike count in the window
y = (contrast > 0).astype(int)            # 1 if there was a stimulus on the right, 0 if the screen was blank

print(X.shape, "trials x neurons")
print("stimulus on", y.sum(), "of", len(y), "trials")
print("trial 0:", X[0, :12], "... label", y[0])
print("trial 3:", X[3, :12], "... label", y[3])
# (340, 66) trials x neurons
# stimulus on 173 of 340 trials
# trial 0: [0 0 0 0 0 0 0 0 0 0 0 0] ... label 0
# trial 3: [2 2 4 2 0 0 2 0 0 0 2 0] ... label 1

340 rows, 66 columns. Trial 0 was a blank screen and the first dozen visual cortex neurons were silent. Trial 3 had a stimulus and they were not. The .T at the end of the X line transposes the array so that trials are rows, which is the convention every machine learning tool expects. Everything from here on works on X and y; the raw recording is not needed again until step 7.

Step 2: one neuron, one threshold

Start with the neuron we know best. Neuron 141 fired 66 spikes per second to a full-contrast stimulus and 8 to a blank screen. So here is a decoder: count its spikes in the window, and if there are more than some number, say the stimulus was there. Before choosing the number, look at the two piles of trials.

one = spks[141][:, post].sum(axis=1)      # neuron 141's spike count in the window, on every trial

fig, ax = plt.subplots(figsize=(10, 3.8))
bins = np.arange(one.max() + 2) - 0.5     # one bar per whole number of spikes
ax.hist(one[y == 0], bins=bins, alpha=0.6, label="blank screen")
ax.hist(one[y == 1], bins=bins, alpha=0.6, label="stimulus")
ax.set(xlabel="spikes from neuron 141, 50 to 250 ms after onset", ylabel="number of trials")
ax.legend()
plt.show()
Neuron 141's spike count on every trial, split by whether there was a stimulus. The piles are clearly different and clearly overlap: a low-contrast stimulus often produces only a few spikes, and a blank screen occasionally produces several.
Neuron 141’s spike count on every trial, split by whether there was a stimulus. The piles are clearly different and clearly overlap: a low-contrast stimulus often produces only a few spikes, and a blank screen occasionally produces several.

A decoder is a rule that turns a row of the table into a guess, and its accuracy is the fraction of trials it gets right. Try every threshold from 1 to 9.

def accuracy(guess, truth):
    """Fraction of trials where the guess matches the truth."""
    return np.mean(guess == truth)

for k in range(1, 10):
    guess = one > k                       # True where the neuron fired more than k spikes
    print(f"more than {k} spikes means stimulus: {accuracy(guess, y):.1%} correct")
# more than 1 spikes means stimulus: 81.8% correct
# more than 2 spikes means stimulus: 85.0% correct
# more than 3 spikes means stimulus: 87.9% correct
# more than 4 spikes means stimulus: 87.9% correct
# more than 5 spikes means stimulus: 88.2% correct
# more than 6 spikes means stimulus: 85.3% correct
# more than 7 spikes means stimulus: 82.6% correct
# more than 8 spikes means stimulus: 79.4% correct
# more than 9 spikes means stimulus: 76.5% correct

88 percent from a single neuron and a single number. That is genuinely good. It is also, as it stands, slightly dishonest, and the next step is about why.

Step 3: never grade yourself on the questions you studied

We picked the threshold by trying all of them on the 340 trials and keeping the best. Then we reported the accuracy on the same 340 trials. That is grading yourself on the exam you used to revise. With one number to choose it barely matters, but with 66 weights to choose, as in the next step, a decoder can memorise the quirks of the trials it was trained on and score brilliantly on them while knowing nothing that transfers to a new trial. The fix is a rule so important that it is the one thing to take away from this lesson if you take away nothing else: choose the rule on some trials, and measure it on different ones.

rng.permutation shuffles the trial numbers, and slicing gives us 240 training trials and 100 test trials that the decoder never sees until it is judged.

rng = np.random.default_rng(0)
shuffled = rng.permutation(len(y))                  # the trial numbers 0..339 in random order
train, test = shuffled[:240], shuffled[240:]        # 240 trials to learn from, 100 to be examined on

scores = [accuracy(one[train] > k, y[train]) for k in range(15)]
best_k = int(np.argmax(scores))                     # the threshold that did best on the training trials

print(f"rule learned from training trials: more than {best_k} spikes means stimulus")
print(f"training trials: {scores[best_k]:.1%} correct")
print(f"test trials:     {accuracy(one[test] > best_k, y[test]):.1%} correct")
# rule learned from training trials: more than 3 spikes means stimulus
# training trials: 88.8% correct
# test trials:     86.0% correct

The rule chosen on the training trials scores 88.8 percent on them and 86 percent on the held-out test trials. The drop is small here because the rule is simple. Watch for it in everything you do from now on: the gap between training and test accuracy is the size of the lie you would have told yourself.

Step 4: 66 neurons, and a model neuron to read them

Neuron 141 is one of 66 in primary visual cortex, and every one of them saw the stimulus. Combining them needs a rule with more than one number: give each neuron a weight, multiply its spike count by that weight, add everything up, and say “stimulus” if the sum is above zero. Neurons that fire more for the stimulus should get positive weights, neurons that fire less should get negative ones, and neurons that do not care should get weights near zero. The only question is how to find the 66 weights, and the answer is to learn them from the training trials, one small correction at a time.

Look at the shape of that rule before reading the code. Inputs arrive, each is multiplied by a weight, they are summed, and the sum is pushed through a threshold. That is a neuron. Not the leaky integrate-and-fire neuron of lesson 1 but the other kind, the one Frank Rosenblatt built out of motors and potentiometers in 1958 and called a perceptron, and which is still the basic unit of every deep network. We are going to decode 66 real neurons with one artificial one.

def train_decoder(X, y, lr=0.01, steps=1000):
    """Logistic regression, trained by gradient descent.
    X: trials x neurons spike counts. y: 0 or 1 per trial.
    Returns one weight per neuron and a bias."""
    w = np.zeros(X.shape[1])                        # start with every weight at zero
    b = 0.0
    for _ in range(steps):
        p = 1 / (1 + np.exp(-(X @ w + b)))          # weighted sum of the counts, squashed to a probability
        error = p - y                               # how wrong it was on each trial, and in which direction
        w -= lr * (X.T @ error) / len(y)            # nudge each weight against its share of the error
        b -= lr * error.mean()
    return w, b

def predict(X, w, b):
    """1 where the weighted sum says stimulus, 0 where it says blank."""
    return (X @ w + b > 0).astype(int)

w, b = train_decoder(X[train], y[train])
print(f"training trials: {accuracy(predict(X[train], w, b), y[train]):.1%} correct")
print(f"test trials:     {accuracy(predict(X[test], w, b), y[test]):.1%} correct")
# training trials: 99.6% correct
# test trials:     100.0% correct

Read train_decoder as a loop of three moves, repeated a thousand times. X @ w + b computes the weighted sum for every trial at once; @ is matrix multiplication, and it does in one symbol what would otherwise be a loop over trials inside a loop over neurons. The 1 / (1 + np.exp(-...)) wrapper squashes each sum to a number between 0 and 1, the decoder’s probability that the stimulus was there. error is the gap between that probability and the truth. And the update line moves each neuron’s weight in the direction that would have shrunk the error, by an amount proportional to how much that neuron fired. The learning rate lr keeps the steps small.

That update rule is worth a second look, because it is local: the change to a neuron’s weight depends only on that neuron’s own activity and the error. A synapse could implement it. This method has a name, logistic regression, and a longer history than the name suggests. In a library like scikit-learn it is one line. We wrote it out so that there is nothing hidden.

99.6 percent on the training trials and 100 percent on the 100 test trials. The population knows something no single neuron knows. To see how confident it is, look at the weighted sum itself, on the test trials only.

evidence = X[test] @ w + b                          # the decoder's weighted sum on each test trial

fig, ax = plt.subplots(figsize=(10, 3.8))
bins = np.linspace(evidence.min(), evidence.max(), 40)
ax.hist(evidence[y[test] == 0], bins=bins, alpha=0.6, label="blank screen")
ax.hist(evidence[y[test] == 1], bins=bins, alpha=0.6, label="stimulus")
ax.axvline(0, color="orange")                       # the decision boundary
ax.set(xlabel="decoder's weighted sum", ylabel="number of test trials")
ax.legend()
plt.show()
The decoder's weighted sum on the 100 held-out trials. Compare with the first figure: the two piles no longer touch. A trial's distance from the boundary is how sure the decoder is.
The decoder’s weighted sum on the 100 held-out trials. Compare with the first figure: the two piles no longer touch. A trial’s distance from the boundary is how sure the decoder is.

Step 5: cross-validation

One split of 240 and 100 gives one number, and 100 test trials is a small exam: one lucky trial moves the score by a full percentage point. Cross-validation fixes this by dividing the trials into five groups and letting each group be the test set in turn, training on the other four. Every trial gets tested exactly once, by a decoder that never saw it, and the five scores are averaged. This is the standard way to report a decoder’s accuracy, and the function below is the one we will use for the rest of the lesson.

def cross_validate(X, y, folds=5, seed=0):
    """Average test accuracy when every trial takes one turn in the test set."""
    rng = np.random.default_rng(seed)
    parts = np.array_split(rng.permutation(len(y)), folds)   # five random, equal groups of trials
    scores = []
    for part in parts:
        is_test = np.zeros(len(y), dtype=bool)
        is_test[part] = True                                 # this group is the test set this time
        w, b = train_decoder(X[~is_test], y[~is_test])       # train on everything else
        scores.append(accuracy(predict(X[is_test], w, b), y[is_test]))
    return np.mean(scores)

print(f"VISp population, cross-validated: {cross_validate(X, y):.1%} correct")
print(f"neuron 141 alone, cross-validated: {cross_validate(X[:, visp == 141], y):.1%} correct")
# VISp population, cross-validated: 99.7% correct
# neuron 141 alone, cross-validated: 87.9% correct

99.7 percent: 339 of 340 trials. The single-neuron rule from step 2, run through the same machinery, gets 87.9. The next two steps make sure that 99.7 means what we think it means.

Step 6: what would luck score?

Fifty percent is not the only kind of chance. A decoder can find structure in things that have nothing to do with the stimulus: a slow drift in firing rates over the session, say, or a neuron that fires more in the second half of the experiment when more of the stimulus trials happened to be. The permutation test from lesson 2 handles this here too. Shuffle the labels so that each trial keeps its spike counts but is assigned a random answer, run the whole cross-validation, and see what accuracy comes out. Do it 200 times. This is the slow part of the lesson; it takes a minute or two.

rng = np.random.default_rng(1)
chance = np.array([cross_validate(X, rng.permutation(y)) for _ in range(200)])

print(f"shuffled labels: mean {chance.mean():.1%}, best of 200 shuffles {chance.max():.1%}")
print(f"95% of shuffles score below {np.percentile(chance, 95):.1%}")
# shuffled labels: mean 49.6%, best of 200 shuffles 58.8%
# 95% of shuffles score below 54.7%

With the labels scrambled the decoder averages 49.6 percent, and its best result in 200 attempts is 58.8. Our real score is 99.7. Anything under about 55 percent could be luck; that line is going to matter in the next step, where the numbers are not so clear-cut.

Step 7: which areas know?

Wrap the whole pipeline in a function that takes a list of neurons, and run it on each brain area. The seven areas from lesson 2 are here, plus MD, the mediodorsal thalamus, which has the most neurons of any area in this recording.

def decode(neurons, window=post):
    """Cross-validated accuracy of a decoder that reads these neurons in this time window."""
    counts = spks[neurons][:, :, window].sum(axis=2).T
    return cross_validate(counts, y)

names = ["VISp", "VISam", "LGd", "MD", "CA1", "DG", "MOs", "ACA"]
by_area = []
for a in names:
    idx = np.where(area == a)[0]
    by_area.append(decode(idx))
    print(f"{a:5s} {len(idx):3d} neurons, {by_area[-1]:.0%} correct")

fig, ax = plt.subplots(figsize=(10, 3.8))
ax.bar(names, by_area)
ax.axhline(np.percentile(chance, 95), color="orange")       # anything below this line could be luck
ax.set(ylabel="decoding accuracy", ylim=(0, 1))
plt.show()
# VISp   66 neurons, 100% correct
# VISam  79 neurons, 71% correct
# LGd    11 neurons, 57% correct
# MD    126 neurons, 77% correct
# CA1    50 neurons, 52% correct
# DG     65 neurons, 63% correct
# MOs     6 neurons, 57% correct
# ACA    16 neurons, 61% correct
Decoding accuracy by area. Primary visual cortex is near perfect. CA1, in grey, is the only area that scores below what shuffled labels achieve. Small samples again: LGd has 11 neurons and MOs has 6.
Decoding accuracy by area. Primary visual cortex is near perfect. CA1, in grey, is the only area that scores below what shuffled labels achieve. Small samples again: LGd has 11 neurons and MOs has 6.

Primary visual cortex is essentially perfect. VISam, a higher visual area, gets 71 percent, and CA1, which in lesson 2 had not one neuron responding to the stimulus, is at chance. So far the map from lesson 2 is confirmed. Then there is MD at 77 percent, the second best area in the recording, and DG at 63, even though lesson 2 found that only two percent of its neurons respond to the stimulus. A decoder pools weak signals that a single-neuron test misses, and part of the explanation is that. But look at what else these trials differ in. On stimulus trials the mouse almost always turns the wheel; on blank trials it often does nothing. An area that carries no visual information at all, but knows what the mouse is about to do, will decode the stimulus above chance, because in this task the two are correlated. Restrict the analysis to trials where the mouse made the same response, and MD’s advantage over simply guessing the commoner label shrinks to a few points.

This is the most important caveat in decoding, and it applies to every result in the field, including the ones in brain-computer interfaces. A decoder tells you that the information is present in an area. It does not tell you why it is there, or what the area is doing with it, or whether it would still be there if the animal’s behaviour were different. Steinmetz and colleagues built their whole paper around this problem, and found that signals related to movement are present nearly everywhere in the mouse brain.

Step 8: when does the brain know?

So far the decoder has read one fixed window. Slide a 100 ms window across the trial instead, training a fresh decoder at each position, and you get the accuracy as a function of time. This is the neural equivalent of watching the information arrive.

starts = np.arange(10, 140, 5)                      # window start, in bins: every 50 ms
acc_time = []
for s in starts:
    window = np.zeros(len(t), dtype=bool)
    window[s:s + 10] = True                         # a 100 ms window starting at bin s
    acc_time.append(decode(visp, window))

fig, ax = plt.subplots(figsize=(10, 3.8))
ax.plot(t[starts] + 0.05, acc_time, marker="o")     # plot each window at its centre
ax.axvline(0, color="orange")                       # stimulus onset
ax.axhline(0.5, color="gray", ls=":")               # chance
ax.set(xlabel="centre of 100 ms window, time from stimulus onset (s)", ylabel="decoding accuracy", ylim=(0.4, 1))
plt.show()
VISp decoder accuracy from a 100 ms window at each position. Chance before the stimulus, 97 percent in the first 100 ms after it, and above 90 percent for most of the next half second. The dip around 200 ms is the trough between the visual transient and the second hump you saw in neuron 141's PSTH.
VISp decoder accuracy from a 100 ms window at each position. Chance before the stimulus, 97 percent in the first 100 ms after it, and above 90 percent for most of the next half second. The dip around 200 ms is the trough between the visual transient and the second hump you saw in neuron 141’s PSTH.

Before the stimulus, chance. That is a sanity check as much as a result: a decoder that scores well before the stimulus appears has found a leak, and you should go looking for it. In the first 100 ms after onset the accuracy jumps to 97 percent, which is the transient response from lesson 2 doing its work. It peaks at 99 in the 50 to 150 ms window, dips around 200 ms, and then stays above 90 percent for most of the next half second, partly because the stimulus is still on the screen and partly, as step 7 warned, because the mouse is now moving.

What you just did

  • Turned a neural recording into the trials-by-features table that every machine learning method starts from.
  • Built a one-neuron decoder and learned why accuracy must be measured on held-out trials.
  • Wrote logistic regression from scratch, trained it by gradient descent, and understood it as a model neuron with 66 learnable synapses.
  • Cross-validated it, and established what luck alone would score by shuffling the labels.
  • Decoded the stimulus from eight brain areas and from every moment of the trial, and met the central caveat of the field: information present is not the same as information used.

Every brain-computer interface, from the cursor-control systems in clinical trials to the speech decoders in the news, is this pipeline with more neurons, a fancier classifier, and a loop that feeds the guess back to the user. You have now built the part in the middle.

Exercises

  1. Decode the left stimulus instead: y = (dat["contrast_left"] > 0).astype(int). The probes are in the left hemisphere. What does VISp score now, and what does that tell you?
  2. Train the decoder on all 340 VISp trials and sort the neurons by the size of their weight. Is neuron 141 at the top? The neuron with the largest weight has a negative sign. Plot its PSTH for blank and stimulus trials, as in lesson 2, and work out what the decoder is using it for.
  3. How many neurons do you need? Pick random subsets of 1, 2, 5, 10, 20 and 40 VISp neurons, cross-validate each, and plot accuracy against the number of neurons. Where does it saturate?
  4. Harder, and the real first step of a BCI: decode what the mouse is about to do. Keep only the trials where it turned the wheel (dat["response"] != 0), label them by direction, and use a window from 250 to 750 ms. Try VISp, VISam and MD. Which area knows which way the mouse will turn, and how early can you decode it?

Data: Steinmetz, Zatka-Haas, Carandini & Harris, “Distributed coding of choice, action and engagement across the mouse brain”, Nature 2019, CC-BY 4.0, via Neuromatch Academy. The code in this lesson runs top to bottom as a single script; the shuffle test in step 6 is the only slow part.

Machine learning inside the microscope: Nikon’s NIS.ai and the open tools behind it

This started life in 2020 as a one-line note pointing at a page on Nikon’s website. The page is still worth pointing at, because it is a clear example of something that has quietly happened across microscopy: the machine learning is now inside the microscope software, sold as a set of modules, and a great many neuroscience images are passing through convolutional networks before anyone looks at them. This is a short guide to what those modules do, what they are under the hood, which free tools do the same job, and the one thing you must check before trusting any of them. Updated October 2026.

What Nikon ships

NIS.ai is a family of add-ons for NIS-Elements, the acquisition and analysis software that drives Nikon microscopes. Each module is a neural network that takes an image in and gives an image, or a set of labelled regions, back. Some come pre-trained; the more interesting ones learn from your own data. Nikon’s page is the authority on current pricing and bundling, which I will not repeat here because it changes; at the time of writing Denoise.ai ships with the AR package and the rest are licensed separately, and all of them want a GPU.

NIS.ai moduleWhat it doesTrained onThe open-source equivalent
Denoise.aiRemoves shot noise from confocal imagesPre-trained by NikonNoise2Void, CARE
Clarify.aiRemoves out-of-focus blur from widefield fluorescencePre-trained by NikonDeconvolution; CARE
Enhance.aiRestores detail in under-exposed imagesYour own paired imagesCARE
Convert.aiPredicts one channel from another, e.g. a nuclear stain from brightfieldYour own paired imagesIn-silico labelling; CARE’s image-to-image mode
Segment.aiFinds structures that thresholding cannot, from hand-traced examplesYour own traced imagesCellpose, StarDist, ilastik

Two of these are the kind of thing you would expect a camera company to ship. Denoising and deblurring have a long history in microscopy, deconvolution in particular, and a network trained on a lot of Nikon images is a sensible modern replacement. The other three are more surprising, and more important.

What they are, underneath

Every NIS.ai module is a descendant of a handful of academic papers from 2015 to 2021, all of them open, most of them with code. Knowing the lineage tells you what the tool can and cannot do.

  • U-Net (Ronneberger, Fischer and Brox, 2015). A convolutional network shaped like a U: it shrinks the image to understand it, then grows it back to the original size to label every pixel. Written for biomedical segmentation, and still the backbone of nearly every image-to-image tool below, including, almost certainly, Nikon’s.
  • CARE, content-aware image restoration (Weigert and colleagues, 2018). Train a U-Net on pairs of bad and good images of the same sample (low light and high light, low resolution and high resolution) and it learns to turn bad into good. This is Enhance.ai’s recipe, and the CSBDeep package is the open version.
  • Noise2Void (Krull, Buchholz and Jug, 2019). Denoising without clean examples: the network learns to predict each pixel from its neighbours, which it can only do for the signal and not for the noise. Pre-trained denoisers like Denoise.ai are a cousin of this idea.
  • In-silico labelling (Christiansen and colleagues, 2018; Ounkomol and colleagues, 2018). Predict a fluorescent stain from a transmitted-light image, because the structure the stain reveals is faintly present in the unstained picture. This is Convert.ai, and it is the module most likely to make a biologist gasp and a statistician wince, for the same reason.
  • Cellpose (Stringer, Wang, Michaelos and Pachitariu, 2021) and StarDist. General-purpose cell and nucleus segmentation that works out of the box on most images and can be fine-tuned on yours in a GUI. Segment.ai does the same job inside NIS-Elements. Cellpose comes from the Janelia group whose data appears in the data page, and it is the tool I would reach for first.

The thing to check

An image-to-image network produces a plausible image. That is its job, and it is also the problem. A restored or converted image can contain structure that was never in the sample: a denoiser that has seen a lot of mitochondria will draw mitochondria into noise, and a channel-prediction network will confidently paint a nucleus where the training set says one usually is. Nothing in the output flags this. The people who wrote CARE say so plainly in the paper, and the rule they give is the right one:

  • Keep the raw images, always, and do your quantification on them or on something you can trace back to them.
  • Validate on data where you know the answer: hold out real pairs, as in lesson 3, and look at the worst cases, not the average.
  • Treat restored images as a way to see, and segmentation masks as a hypothesis to check, not as measurements.
  • When the method goes in a paper, say which network, which version, and what it was trained on. A reviewer cannot evaluate “AI denoising”.

If you want to try this without a Nikon licence

Everything in the table has a free equivalent that runs on a laptop, or in a browser, and that you can read the source of.

  • ZeroCostDL4Mic (von Chamier and colleagues, 2021): Google Colab notebooks that train and run CARE, Noise2Void, StarDist, Cellpose, U-Net and more on your own images, with no installation and a free GPU. The best first step.
  • Cellpose: pip install cellpose, open the GUI, drop in an image. Fine-tune on a few of your own corrected masks when the default model is not good enough.
  • napari: the Python image viewer that most of these tools plug into, and the place to look at your results slice by slice.
  • BioImage Model Zoo: pre-trained models for microscopy, in a format that runs in Fiji, ilastik, napari and DeepImageJ.
  • For calcium imaging specifically, Suite2p and CaImAn do registration, cell detection and spike inference end to end, and are what the labs producing the datasets on this site actually use.

The vendor tools earn their price by being there at the moment of acquisition, with support, inside the software the microscope already runs. The open tools earn their place by being inspectable, citable and free, which is what a result needs to be when it leaves the lab. Most imaging groups I know end up using both, and the ones who understand the lineage above use both well.