One chain of amino acids finds one shape, reliably, in milliseconds. Predicting that shape is folding. Choosing a shape first and asking which chains would settle into it is inverse folding. The two directions are not mirror images — they fail and succeed for different reasons.
Each proposed mutation is kept if it makes the target shape lower in energy and makes competing shapes worse. That second half is the part people forget.
This is a two-dimensional toy — twenty-four beads, two chemistries, Brownian dynamics with a hydrophobic attraction. It is not a protein. But the two things it does get right are the two things that matter: sequence alone determines where the chain ends up, and the design problem is solved by burying the greasy residues and leaving the polar ones facing out.
The same two objects, two different questions. Sequence space is enormous and discrete; structure space is continuous and, in practice, surprisingly small — nature reuses a few thousand folds. That asymmetry is the whole story.
| Folding | Inverse folding | |
|---|---|---|
| Given | A chain of letters | A backbone you want |
| Wanted | The coordinates it settles into | Letters that settle into it |
| Mapping | Essentially one answer per input | Astronomically many valid answers |
| What makes it work | Coevolution across millions of homologous sequences; geometry learned once, reused everywhere | Local environment is nearly enough — burial, neighbour geometry, backbone angles predict the residue |
| What makes it hard | The signal is non-local: two letters far apart in the chain decide each other's fate | Negative design. The chosen sequence must prefer your fold over every other fold it could adopt |
| How you know you're right | Predicted confidence (pLDDT, PAE), then a crystal or a cryo-EM map | Fold the design back and check it returns. Then express it and see if it behaves |
| Rough difficulty | Was open for fifty years | Solved well enough that recovery rates ~50% beat nature's own choices at stability |
Sketch or generate a backbone with the geometry the job needs — a pocket, an interface, a channel.
Ask which sequences would hold that backbone. Get hundreds of candidates in seconds.
Run each candidate through structure prediction as if it were an unknown natural protein.
Keep only those that come back to the shape you asked for, with high confidence and low error.
Order the DNA, express it, measure. A few percent working is a very good day.
The forward model is the cheap referee for the inverse one. Design proposes; prediction disposes — and because prediction never saw the design during training, agreement between them is real evidence rather than a tautology.