When FFT Beats Averagine for Peptide Isotopic Patterns in MS for Labs

A peptide isotopic pattern is a monoisotopic peak plus an envelope of heavier isotopologues, with spacing between peaks equal to the neutron mass divided by charge, a concept explained in detail in peptide sequence characterization methods for researchers. To identify or quantify a peptide from this pattern, you simulate a theoretical envelope, detect the matching cluster in your spectrum, fit it against the data using a method like NNLS or RMSE, and validate the fit against known standards before trusting the result.
TL;DR:
- Accurate charge state determination from peak spacing is essential before converting m/z to neutral mass, especially for complex or overlapping isotope envelopes.
- FFT and BRAIN methods provide exact isotope distributions from elemental composition, while Averagine offers rapid but approximate estimates suitable for mass-only screening.
- Linear combination fitting, such as RMSE minimization, is necessary for reliable monoisotopic mass extraction in noisy spectra or when dealing with labeled and complex samples.
- Validating fitting results against known standards or synthetic mixtures, and reporting fit metrics like RMSE, ensures reproducibility and trustworthiness of quantitative isotope pattern analysis.
Table of Contents
- The physics and chemistry behind peptide isotope envelopes
- Exact and heuristic methods for simulating isotope distributions
- How deconvolution algorithms recover monoisotopic masses
- Fitting labeled peptides and detecting metabolic scrambling
- Building an end-to-end workflow with the right tools
- Validating your isotope-fitting results before you publish
- Where lab-verified reference data strengthens isotope-pattern work
- When to trust the pipeline and when to look at the spectrum yourself
- Validate your isotope-fitting work with lab-verified peptide data
- FAQ
- Sources
The physics and chemistry behind peptide isotope envelopes
Every peptide ion produces a cluster of peaks instead of a single line because carbon, nitrogen, hydrogen, oxygen, and sulfur each occur naturally as a mix of isotopes. The lightest combination, built entirely from the most abundant isotopes (12C, 1H, 14N, 16O), forms the monoisotopic peak. Heavier peaks in the envelope come from molecules that happen to contain one or more heavy isotopes, chiefly 13C, which occurs at roughly 1.1% natural abundance per carbon atom according to pyopenms documentation on charge and isotope deconvolution.
For a small peptide with few carbons, the monoisotopic peak dominates the envelope. As peptide mass grows, so does the carbon count, and the odds that at least one carbon in the molecule is 13C rise accordingly. At some point the second or third isotopologue peak becomes taller than the monoisotopic peak itself, a shift that the same pyopenms documentation describes as a direct consequence of atom count scaling with mass. This matters practically: software that blindly picks “the tallest peak” as monoisotopic will misassign larger peptides.

Charge state changes everything about how this envelope looks in your spectrum. Peak spacing in m/z equals the neutron mass divided by the charge, so a singly charged peptide ion shows peaks spaced about 1 Da apart, a doubly charged ion shows spacing near 0.5 Da, and a triply charged ion compresses that further to about 0.335 Da. The pyopenms charge and isotope deconvolution guide puts the precise isotope spacing for singly charged peptide ions at approximately 1.0033 Da, derived from the mass difference between 13C and 12C divided by charge. That predictable spacing is what lets deisotoping algorithms assign charge state correctly and separate overlapping envelopes from different charge states that would otherwise look like noise.
A few practical consequences follow from this theory:
- Larger peptides (roughly above 3,000 to 5,000 Da, depending on elemental composition) often have a monoisotopic peak too small to detect reliably above noise.
- Charge state must be determined from peak spacing before you can convert m/z back to neutral mass with any confidence.
- Sulfur-containing peptides show a distinct 34S contribution that shifts the envelope shape compared to sulfur-free sequences of similar mass.
- Instrument resolution sets a hard floor on how close two isotope peaks (or two charge states) can be before they merge into one apparent peak.
When the monoisotopic peak truly is not detectable, the practical fix is not to guess which peak is “number one.” It is to simulate the full theoretical envelope and fit the entire measured pattern against it, which recovers the monoisotopic mass as a model parameter rather than as an observed peak.
Exact and heuristic methods for simulating isotope distributions
Once you know why the envelope looks the way it does, the next question is how to generate a theoretical version of it to compare against real data. Two philosophies dominate: exact calculation from elemental composition, and fast approximation from average peptide composition.
Exact methods start from the peptide’s actual elemental formula and compute the full convolution of isotope probabilities across every atom. Fast Fourier Transform (FFT) based approaches and BRAIN (a polynomial recurrence method) both fall into this category. The pyisogen project on PyPI documents both implementations inside the IsoGen toolbox and notes that FFT is treated as the default exact calculation, limited mainly by the accuracy of the input formula rather than by any approximation in the math itself. BRAIN achieves the same exactness through a different computational route, using recurrence relations over isotope polynomials instead of transforming into frequency space. Both methods give you the correct theoretical distribution for a known sequence, including unusual elemental compositions, post-translational modifications, or isotope labels, provided you supply the right formula.
Averagine takes a different approach entirely. Instead of using the peptide’s real elemental composition, it substitutes an average amino acid composition scaled to the observed mass, then computes an isotope distribution from that average formula. This is fast and works well at scale because it needs no sequence information, only a measured mass. The tradeoff is accuracy: Averagine assumes a “typical” peptide, so it degrades for sequences that deviate from that average, including isotopically labeled peptides, heavily modified peptides, or sequences with unusual amino acid skew. Documentation for ms_deisotope notes explicitly that when accuracy matters for labeled peptides, glycosylation, or other nonstandard compositions, a sequence-based elemental composition approach or FFT and BRAIN should replace the Averagine shortcut.
A newer category uses trained neural-network models to predict isotope distributions directly from sequence or composition features. The pyisogen package includes pretrained peptide models alongside its FFT and BRAIN implementations, letting a researcher choose between an exact calculation and a learned approximation depending on throughput needs. Retraining or fine-tuning a neural model becomes worthwhile mainly when your peptide population differs systematically from the training distribution, such as a heavily labeled sample set or a non-tryptic digest with unusual composition.
- FFT and BRAIN give exact theoretical distributions from known elemental composition, including labeled or modified peptides.
- Averagine gives fast approximate distributions from mass alone but degrades for labeled, modified, or compositionally unusual peptides.
- Neural-network models like those in IsoGen offer a learned middle ground, useful for high-throughput screening where retraining can adapt to a specific sample population.
- Choose exact methods (FFT/BRAIN) when sequence is known and accuracy matters; choose Averagine or NN models for rapid, large-scale screening where approximate masses suffice.
Pro Tip: Default to FFT or BRAIN whenever you have sequence information, and reserve Averagine for situations where you only have an observed mass and need a fast first-pass estimate.
How deconvolution algorithms recover monoisotopic masses
Simulating a theoretical pattern only gets you half the pipeline. The other half is pattern matching: taking a real spectrum, with its noise, overlapping charge states, and co-eluting species, and extracting the monoisotopic mass and abundance that best explain what you observed.
Several algorithm families handle this problem, each suited to different noise regimes and spectral complexity:
- Non-negative least squares (NNLS) and least absolute deviation (LAD) template matching fit raw spectral intensities against a library of theoretical isotope-pattern templates, solving for which templates (and at what abundance) best reconstruct the observed data. Because the fit enforces non-negative coefficients, it naturally handles overlapping isotope patterns from different charge states or co-eluting peptides without needing heavy regularization, provided a reasonable intensity threshold is applied to exclude pure noise peaks.
- Intensity-ratio approaches compare the relative heights of adjacent isotope peaks against the theoretical ratios predicted for a given composition. These methods are computationally cheap and work well at high signal-to-noise, but they lose robustness as spectra get noisier, since a single mis-measured peak height can throw off the ratio comparison.
- RMSE-based linear-combination fitting builds a model spectrum from a weighted sum of simulated component patterns (for example, labeled and unlabeled versions of the same peptide) and adjusts the weights to minimize root-mean-square error against the measured spectrum. This approach generalizes well to mixtures because it does not assume a single pure species is present.
- Deep-learning feature detectors operate directly on raw LC-MS maps rather than isolated spectra. A CNN and RNN combination, as described in work on DeepIso-style architectures, can detect isotopic features and estimate charge state across retention time and m/z simultaneously, which helps in complex samples where traditional peak-picking struggles to separate overlapping elution profiles.
In practice, NNLS or LAD template matching with sensible thresholding has become a common default because it generalizes across heterogeneous spectra better than approaches relying on heavy regularization, which tend to overfit on some sample types and underfit on others. Intensity-ratio methods remain useful as a fast sanity check, but they should not be the sole basis for quantification in noisy data. RMSE linear-combination fitting becomes necessary, not just preferable, once you move into labeled or mixed-state samples, which is the subject of the next section.
Fitting labeled peptides and detecting metabolic scrambling
Isotope labeling with 13C or 15N introduces a problem that simple monoisotopic peak-picking cannot solve: the “monoisotopic” peak of a labeled peptide is no longer the lightest peak in a simple sense, because the label itself shifts the entire envelope toward heavier mass, and a population of molecules can contain a mixture of labeling states simultaneously.
Monoisotopic peak ratios, which work reasonably well for comparing two unlabeled peptide pools, fail here because a partially labeled sample produces an envelope that is really a blend of several distinct theoretical envelopes (fully unlabeled, partially labeled, and fully labeled) superimposed on one another. Picking a single peak and comparing its height to another single peak throws away the information contained in the rest of the envelope shape.

The working method, demonstrated in a 2014 study on quantifying peptide m/z distributions from 13C-labeled cultures using high-resolution mass spectrometry, is to simulate both the labeled and unlabeled theoretical isotope distributions separately, then fit the measured spectrum as a linear combination of those simulated patterns. Minimizing the root-mean-square error (RMSE) between this combined model and the observed spectrum yields both the percent incorporation of the label and, critically, can reveal more than two labeling states if the data supports it, which is how that study detected metabolic scrambling in HEK293-expressed proteins that simpler monoisotopic-ratio methods missed entirely.
Several practices improve the reliability of this fitting process:
- Simulate every plausible labeling state in advance (unlabeled, each expected enrichment level, and any suspected scrambled state) rather than assuming only two components exist.
- Use ultrahigh-resolution fine structure, where instrument resolution is sufficient, to separate overlapping isotopologues that a lower-resolution instrument would blend into one peak, a capability documented in work on measuring 15N and 13C enrichment levels in selectively labeled proteins.
- Pair MS1 isotope fitting with MS/MS fragment data when available, since fragment-level isotope patterns can disambiguate which part of the sequence carries the label.
- Design labeling experiments with a clear unlabeled control and, where feasible, a known mixed-ratio standard to calibrate the fitting pipeline before applying it to unknowns.
Metabolic scrambling, where a label migrates to unexpected positions or distributes across more carbon atoms than intended, is exactly the kind of signal that linear-combination RMSE fitting is built to catch, because it shows up as a poor fit to the simple two-state model and a much better fit once additional scrambled states are added to the simulation.
Building an end-to-end workflow with the right tools
A reproducible isotope-pattern workflow breaks into four steps: simulate the theoretical pattern, detect candidate clusters in the raw spectrum, fit the model against the data, and validate the result. Different tools fit each step better than others, and most research pipelines end up combining two or three of them rather than relying on a single package.
For theoretical simulation, pyisogen is a direct choice when you want FFT or BRAIN exact calculations, or a pretrained neural-network model, from a known sequence or elemental composition. The package exposes all three methods under one interface, so switching between an exact calculation and a fast approximation is a parameter change rather than a different codebase.
For cluster detection and deisotoping within a full LC-MS processing pipeline, pyopenms provides a Deisotoper class that groups raw peaks into isotope clusters and assigns charge state based on the expected spacing discussed earlier. Parameter choices here, particularly the mass tolerance window and the minimum number of isotope peaks required to call a cluster, directly affect how aggressively the tool merges or splits overlapping envelopes, so these values are worth logging alongside your results rather than leaving at default.
For Averagine-based modeling and general-purpose deisotoping, ms_deisotope offers prebuilt averagine models tuned for peptides as well as glycans, along with functions that return a theoretical isotopic cluster for a given m/z and charge. It is a reasonable default for high-throughput screening where sequence information is not yet available and speed matters more than per-peptide accuracy.
For fragment-level isotope work, particularly MS/MS-based validation, the Predator Protein Fragment Calculator, documented in the Hunt Lab guide to de novo peptide sequence analysis, breaks a sequence into elemental compositions at the fragment level and returns isotopologue m/z and abundance values across charge states and post-translational modifications. This is the tool to reach for when MS1-level fitting is ambiguous and you need fragment ion isotope patterns to confirm an assignment.
A practical workflow sequence looks like this:
- Simulate the theoretical envelope for each candidate peptide using pyisogen’s FFT or BRAIN method when sequence is known, or an Averagine model from ms_deisotope when it is not.
- Detect candidate isotope clusters in the raw spectrum using pyopenms’s Deisotoper, logging the mass tolerance and minimum peak count parameters used.
- Fit the detected cluster against the simulated model using NNLS, LAD, or RMSE linear-combination fitting depending on whether the sample may contain mixed labeling states.
- Validate ambiguous assignments at the fragment level using Predator before finalizing an identification.
Pro Tip: Log every parameter you set, mass tolerance, minimum peak count, charge state range, alongside your results; a fit that looks clean with one parameter set can shift meaningfully with another, and reviewers will ask.
Instrument resolution and mass accuracy set the ceiling on what any of these tools can recover. A low-resolution instrument that cannot separate adjacent isotope peaks at a given charge state will force you toward Averagine-level approximations regardless of which software you use, since the fine structure that exact methods rely on simply is not resolved in the data.
Validating your isotope-fitting results before you publish
A fit that looks visually reasonable is not the same as a validated one. Before relying on an isotope-pattern fit for a quantitative claim, run it through a short checklist that keeps the method testable and reproducible for anyone who reads the paper later.
- Report the RMSE or equivalent fit-quality metric for every assignment, along with a residual plot showing where the model over- or under-predicts intensity across the envelope.
- Build synthetic mixtures with known labeling ratios, or use spike-in standards of known composition, and confirm the fitting pipeline recovers the expected ratios before applying it to unknowns.
- Include at least one representative spectrum overlay, measured data against the fitted model, so readers can see the fit quality directly rather than trusting a summary number alone.
- Make code or a notebook available alongside parameter values, since isotope-fitting results are sensitive to tolerance and threshold choices that are easy to omit from a methods section.
- Watch specifically for overfitting to noise, where adding more component states to a linear-combination fit always reduces RMSE even when the additional states are not chemically real, and for bias introduced by the instrument’s point-spread function, which can distort peak shape in ways that look like a real isotope effect but are not.
A fit validated only on the same data used to develop it tells you little about how it will behave on the next sample. Spike-in controls and synthetic mixtures with a known ground truth are what separate a method you can trust from one that merely looks plausible on a plot.
Where lab-verified reference data strengthens isotope-pattern work
Simulated envelopes and fitting algorithms are only as trustworthy as the reference material used to validate them. At Boren Health, we maintain independent HPLC and mass spectrometry assessments across a large set of verified peptide vendors, giving researchers access to purity and identity data generated from blind sample purchases rather than vendor-supplied figures.
That kind of independently verified reference material has a direct use in isotope-pattern validation: a peptide with known, lab-confirmed purity and identity makes a reliable spike-in standard for testing whether a fitting pipeline recovers the expected isotope distribution and labeling ratio before you trust it on an unknown sample. We publish results including failures, so the reference data reflects the actual range of purity you might encounter in a real procurement, not just a best-case sample.
When to trust the pipeline and when to look at the spectrum yourself
Automated isotope-fitting pipelines earn their keep on routine, high-throughput work: well-characterized peptides, standard charge states, decent signal-to-noise, and no reason to suspect an unusual labeling or modification state. Let the software run and spot-check a sample of outputs rather than every single one.
Manual inspection becomes worthwhile the moment a pattern looks wrong in a way you cannot immediately explain: low signal-to-noise that makes peak ratios unreliable, an envelope with unexpected asymmetry, or any sample where a label, an unanticipated post-translational modification, or a nonstandard composition could be shifting the pattern away from what Averagine or even a sequence-based model expects. In those cases, pull up the raw spectrum, check the fit residuals directly, and consider fragment-level validation before trusting the MS1 fit alone.
Sharing your parameter sets and a representative raw spectrum alongside processed results costs little and saves the next researcher, including a future version of yourself, from re-deriving which tolerance settings actually mattered.
— Ross
Validate your isotope-fitting work with lab-verified peptide data
Fitting a labeled or unlabeled isotope pattern is only as reliable as the reference standard behind it. Boren Health compares peptide vendors using independent HPLC and mass spectrometry data, real-time pricing, and purity rankings across more than 200 verified vendors, so you can source a reference peptide with a documented, lab-confirmed identity rather than relying on a vendor’s own certificate of analysis.

- Browse peptides by research goal, including options researched for weight loss, muscle growth, and recovery, each with independent purity data attached.
- Compare vendors directly on lab-verified purity before choosing a source for a spike-in standard or reference material.
- Check current vendor pricing and plans for access to the full comparison platform.
FAQ
What is the isotope pattern in mass spectrometry?
An isotope pattern, or isotopic envelope, is the cluster of peaks a single molecular species produces in a mass spectrum because its constituent elements occur naturally as mixtures of isotopes. For peptides, the lightest peak in this cluster, called the monoisotopic peak, is built entirely from the most abundant isotopes (12C, 1H, 14N, 16O), and the peaks that follow it in m/z correspond to molecules containing one or more heavier isotopes, mainly 13C.
Which peptides are best for MS?
There is no single “best” peptide for mass spectrometry in general. The choice depends on your analytical question, and tools like the Predator fragment calculator and sequence-based isotope simulators work across standard, modified, and labeled peptides as long as the elemental composition is specified accurately.
What shouldn’t you mix with peptides?
This article focuses on isotope-pattern analysis in mass spectrometry rather than peptide handling or storage guidance, so we don’t have a sourced answer here. For questions about peptide stability, storage, or compatibility, consult a vendor’s own documentation or a lab safety reference.
How is mass spectrometry used for peptide identification?
Peptide identification by mass spectrometry works by comparing a measured isotope envelope, and often fragment ion spectra from MS/MS, against theoretical patterns simulated for candidate sequences. Matching the monoisotopic mass, charge state spacing, and overall envelope shape against a simulated pattern, refined through deconvolution and fitting algorithms, lets researchers assign an identity with a quantifiable confidence level rather than a single peak guess.
How do you handle isotopically labeled peptides in MS analysis?
Labeled peptides require simulating both the labeled and unlabeled theoretical distributions separately, then fitting the measured spectrum as a linear combination of those simulated patterns using RMSE minimization. This approach, demonstrated in a study on 13C-labeled cultures, can recover percent incorporation and detect metabolic scrambling that simple monoisotopic peak-ratio methods miss.