What composition predicts
The proportion of charged residues is the strongest single predictor of aqueous solubility. A sequence more than about a quarter charged residues generally dissolves in water without help; one below about a tenth, with a high proportion of hydrophobic residues, generally does not.
The individual counts matter too. Methionine and tryptophan oxidise. Asparagine and glutamine deamidate, particularly when followed by glycine. Cysteine forms disulfides, wanted or otherwise. Each of these is a specific degradation route that the composition tells you to expect.
Residues that create liabilities
A methionine near the N-terminus is the most common oxidation site in a synthetic peptide, and shows up in mass spectrometry as a plus 16 dalton species and in reverse-phase chromatography as a peak eluting slightly earlier than the parent.
An Asn-Gly pair is the classic deamidation motif, because the small glycine side chain leaves room for the cyclic intermediate to form. It adds one dalton and, again, elutes slightly earlier.
Residues that make measurement possible
Tryptophan and tyrosine are the only two residues that absorb usefully at 280 nanometres. A sequence containing neither cannot be quantified by ultraviolet absorbance at that wavelength, whatever the concentration.
This is worth checking before buying a peptide you intend to quantify spectrophotometrically. Many short research peptides have no aromatic residue at all.
Grouping residues by property
The four-way split used here, nonpolar, polar, positive and negative, is the standard teaching classification and is enough for most handling questions. Aromatic residues are called out separately because they carry the ultraviolet signal.
The boundaries are conventional rather than sharp. Glycine has no side chain to classify, proline constrains the backbone rather than contributing a chemistry, and tyrosine is aromatic, weakly polar and weakly acidic all at once. Treat the groups as a summary, not a taxonomy.
Composition against sequence
Composition discards the order. Two peptides with identical composition can behave completely differently if one clusters its hydrophobic residues into a block and the other distributes them evenly.
The sequence visualizer keeps the order and colours it by property, and the hydrophobicity plotter shows the running average along the chain. Composition is the summary; those two are the detail.
How the composition is calculated
A count, a division, and a lookup. The only decision in it is what counts as a residue, and the answer is the twenty standard codes and nothing else.
count(aa) = occurrences of that residue in the parsed sequence percent(aa) = count(aa) / total valid residues x 100 group total = SUM of counts for the residues in that group MW = SUM( free amino acid mass - 18.02 ) + 18.02
- Parse to the twenty standard residues. Non-standard characters are removed before counting. They used to survive into the denominator, which made the percentages fail to sum to 100 whenever a sequence contained one.
- Count each residue. A single pass over the sequence. The counts are the raw data every other figure on the page derives from.
- Divide by the valid residue count. Percentages are taken against the number of residues actually counted, so they always sum to 100.
- Aggregate by property group. Each residue belongs to exactly one group, so the group totals sum to the sequence length and can be read as a composition profile.
- Sum the mass contributions. Each residue contributes its free amino acid mass less one water, and one water is added for the chain's free ends, which is the same calculation the molecular weight tool performs.
What this method cannot tell you
- •Composition discards sequence order, and order determines structure. Two peptides with the same composition can behave nothing alike.
- •The four property groups are a conventional simplification. Glycine, proline and tyrosine each sit awkwardly in any four-way scheme.
- •It counts the twenty standard residues only. D-amino acids, non-standard residues and modifications are invisible to it.
- •A high charged-residue fraction predicts solubility only in general terms. A specific peptide can defy the trend.
Amino acid composition: frequently asked questions
It predicts bulk behaviour: how likely the peptide is to dissolve in water, what charge it carries, whether it can be measured by ultraviolet absorbance, and which degradation routes are open to it.
The proportion of charged residues is the single most useful figure for the solubility question.
The charged ones: aspartate, glutamate, lysine, arginine and histidine. Polar residues such as serine, threonine, asparagine and glutamine help but contribute less.
A sequence more than roughly 25 percent charged residues usually dissolves in water without assistance.
The hydrophobic ones: isoleucine, leucine, valine, phenylalanine, tryptophan, methionine, alanine and proline.
A sequence more than about half hydrophobic residues, with few charged ones, generally needs a co-solvent. The solubility predictor combines both proportions into a single recommendation.
Because they are the two residues most prone to oxidation, which is the most common degradation route for a synthetic peptide in solution.
Oxidation adds 16 daltons per oxygen and shows up in reverse-phase chromatography as a peak eluting slightly earlier than the parent, because the oxidised form is more polar.
An asparagine or glutamine followed by a small residue, most notoriously Asn-Gly, where the amide side chain converts to a carboxylic acid.
It adds one dalton and introduces a negative charge. The reaction is faster at higher pH and higher temperature, which is one reason peptide solutions are stored cold and slightly acidic.
Because cysteines form disulfide bridges, and whether they form the right ones determines whether the peptide is correctly folded.
Free cysteines can also form unwanted intermolecular bridges, producing dimers and higher aggregates. The disulfide bond calculator enumerates the possible pairings for a given cysteine count.
Only if it contains tryptophan or tyrosine. Those two residues carry essentially all of the absorbance at that wavelength.
If the count for both is zero, ultraviolet quantitation at 280 nm is not available. Measuring at 205 or 214 nm, where the peptide bond itself absorbs, or using a colorimetric assay such as BCA, are the alternatives.
By the dominant character of the side chain:
- •Nonpolar: A, V, L, I, M, P, G
- •Polar: S, T, C, N, Q
- •Aromatic: F, W, Y
- •Positive: K, R, H
- •Negative: D, E
The boundaries are conventional. Glycine has no side chain, proline constrains the backbone rather than contributing a chemistry, and tyrosine is aromatic, weakly polar and weakly acidic at once.
Its side chain is a hydrocarbon ring, so nonpolar is the closest fit. What matters more about proline is structural: the ring locks the backbone and prevents the peptide from forming a regular helix through that position.
A proline-rich sequence behaves differently from its composition alone would suggest, which is a good example of why order matters.
Weakly. Some residues favour helices and others sheets, but structure depends on the order and on the surrounding environment, and composition throws the order away.
A peptide of fewer than about fifteen residues generally has no stable secondary structure in water regardless of composition.
Directly. The counts of the acidic residues and the basic residues set where the net charge crosses zero.
More acidic than basic residues gives a pI below 7; the reverse gives a pI above it. The charge at pH calculator computes the exact value and plots the full titration curve.
Because they are taken against the number of valid residues counted, not against the number of characters typed.
This used to be wrong. Non-standard characters survived into the denominator, so a sequence containing an X reported percentages that summed to less than 100.
No. The composition is fixed by whatever the peptide is meant to do; the point of looking at it is to know what to expect, not to change it.
A hydrophobic peptide is not a badly designed one. It is one that needs a co-solvent and low-binding plasticware.
Considerably. Long runs of the same residue, beta-branched residues such as valine and isoleucine, and stretches with strong aggregation tendency all lower coupling efficiency during solid-phase synthesis.
That difficulty shows up downstream as deletion sequences: chains missing one residue, which appear as impurity peaks in the chromatogram.
About 110 daltons across the twenty, though the range runs from glycine at 57 to tryptophan at 186.
The 110 figure is only useful as a rough estimate. The exact sum from the composition is what this tool reports.
No. Identical composition with different order gives different molecules, and amino acid analysis, which measures composition experimentally, cannot distinguish them.
Confirming identity requires sequencing or fragmentation mass spectrometry, which reads the order rather than the totals.
Related Products
Related Research News
What Is a Peptide? (Part 2): Understanding Amino Acids and the Distinction from Proteins
Explore the molecular definition of peptides, amino acid structures, and their distinction from proteins in research contexts.
Medicare Pays Billions For Obesity’s Consequences. Its GLP-1 Coverage Gap
Medicare spends billions treating obesity-related diseases but does not cover GLP-1 receptor agonists for weight loss alone. This analysis explores the financial and clinical implications, including recent developments in triple peptide agonists like ASC37 from Ascletis.
Why 99% Purity Means Something Different on a Tripeptide and a 44-mer
At 99% per coupling a tripeptide comes off the resin 98% correct and a 44-residue chain comes off at 65%. What had to be removed, why deletion sequences are the hardest impurity to separate, and the trifluoroacetate counterion that is in the vial and not on the paperwork.



