Research use onlyQualified buyers
Restate Health

Reading Peptide Nomenclature: Sequences, Fragments and Modifications

Product names are trade names; sequences are the actual identity. Covers the one and three-letter codes, which direction a sequence is written in, what bracketed fragment ranges mean, and how modifications and D-amino acids are notated.

6 min readUpdated

A ball-and-stick molecular chain model against dark grey

A trade name identifies a product. A sequence identifies a molecule. Where a catalog lists both, the sequence is the one that can be checked against a mass spectrum, and learning to read it takes about ten minutes.

Two codes for the same twenty residues

Amino acids are written either as three-letter abbreviations or single letters. Three-letter codes are readable and verbose: Gly, His, Lys. Single letters are compact and are what sequences of any length use: G, H, K.

Most single letters are the first letter of the name. The collisions are where it gets memorised rather than deduced, since several residues begin with the same letter and only one of each group keeps it. The practical approach is to keep a reference table rather than to reconstruct the logic.

Direction

A sequence is written from the N-terminus to the C-terminus, left to right. This is universal and it is not arbitrary: it matches the direction in which proteins are biologically synthesised.

Note that solid-phase chemical synthesis runs the other way, building from the C-terminus toward the N-terminus. A sequence as written is therefore the reverse of the order in which the chemist actually assembled it, which occasionally confuses a first reading of a manufacturing record.

Fragment notation

A bracketed range after a parent name indicates which residues of that parent the compound consists of.

  • ACTH(4-10) means residues 4 through 10 of adrenocorticotropic hormone, a seven-residue fragment.
  • HGH fragment 176-191 means residues 176 to 191 of growth hormone, a sixteen-residue fragment.
  • GRF(1-29) means the first 29 residues of growth hormone releasing factor.

Numbering always refers to the parent, which is why a fragment's own residue count and its range numbers are different figures. This notation is the most informative thing on many catalog entries, because it states exactly what portion of a known molecule is in the vial.

Modifications

Modifications are usually written as prefixes or suffixes, and each has a mass consequence a certificate should reflect.

  • N-acetyl or Ac- indicates an acetyl group on the N-terminus, adding 42 daltons.
  • Amide or -NH2 indicates the C-terminal carboxyl replaced by an amide, reducing mass by about one dalton.
  • A lowercase d or a D- prefix on a residue indicates the D stereoisomer rather than the usual L form. This changes no mass at all, which is precisely why mass spectrometry cannot detect a stereochemical error.
  • Position numbers in a modification name indicate where it sits, as in a substitution at position 2 or 8.

Why this is worth knowing

Two reasons, both practical. A sequence plus the modification notation gives a theoretical mass, which is the number a mass spectrometry result is checked against; without the notation the check cannot be made. And distinct products frequently share a trade name while differing in modification, as with the amidate and N-acetyl variants of the same parent sequence. The notation is what separates them.

A catalog entry that states a sequence is making a checkable claim. One that states only a trade name is not.

Where notation causes real purchasing errors

Three places. A fragment designation that names residue positions within a parent sequence, where two suppliers may number from different starting points. A modification written as a prefix rather than shown in the sequence, which changes the mass without changing the name. And a salt form appended inconsistently, which changes the powder weight.

Each of those produces a mass difference, which is why reconciling the certificate mass against the sequence catches more naming problems than reading the name carefully does. Molecular weight sets out the arithmetic.

Reading a modification from a formula

End modifications are the commonest cause of a mass that does not match a bare sequence. An acetyl group on the N-terminus and an amide at the C-terminus each shift the figure by a known amount, and both are routine rather than exotic.

A formula that differs from the sum of the residues by a small, explainable composition change is usually a modification rather than an error. One that differs by an amount nothing accounts for is a question for the supplier.

One-letter and three-letter codes, and when each is used

Three-letter codes are easier to read without error and are the convention for short sequences, which is why most research peptide listings use them. One-letter codes are compact and are the convention for longer chains and for database entries, where a thirty-residue sequence in three-letter form becomes unwieldy.

Both are read in the same direction, from the N-terminus on the left to the C-terminus on the right. That direction is a convention rather than a property of the molecule, and a sequence written in reverse is a different peptide with the same composition and the same mass.

Modifications that change the mass without changing the name

Acetylation at the N-terminus, amidation at the C-terminus, and acylation with a fatty chain are the three most common, and none of them necessarily appears in the name a supplier uses. Each shifts the mass by a known amount, which is why reconciling the certificate figure against the bare sequence catches them.

A disulfide bond works in the opposite direction, removing two hydrogens and lowering the mass by about two. A calculated figure that sits two above the published one is usually a bridge rather than an error.

In short

A name is a label and a mass is a measurement. Where the two disagree, the measurement is the one to work from, and reconciling them takes about two minutes.

For where this sits among the other molecules a catalogue carries, not everything in a peptide catalog is a peptide covers how the classes differ.

This guide is general reference for research buyers. Materials supplied by Restate Health are for laboratory research use only and are not for human or veterinary use.

Common questions

Which direction is a peptide sequence written in?

From the N-terminus to the C-terminus, left to right, matching the direction of biological synthesis. Solid-phase chemical synthesis runs the opposite way, building from the C-terminus, so a written sequence is the reverse of the assembly order.

What does a notation like ACTH(4-10) mean?

That the compound consists of residues 4 through 10 of adrenocorticotropic hormone, a seven-residue fragment. The numbers always refer to positions in the parent molecule, so a fragment's own length and its range numbers are different figures.

How are modifications written?

Usually as prefixes or suffixes: N-acetyl or Ac- for an acetylated N-terminus, adding 42 daltons; amide or -NH2 for a C-terminal amide, reducing mass by about one dalton; and a D- prefix for the D stereoisomer of a residue, which changes no mass at all.

Why does the sequence matter more than the product name?

Because a sequence plus modification notation yields a theoretical mass, which is what a mass spectrometry result is checked against. Distinct products also frequently share a trade name while differing in modification, and the notation is what separates them.

All products are supplied strictly for laboratory research and development purposes. They are not for human or veterinary use and are not intended to diagnose, treat, cure, or prevent any disease or medical condition.