Protein Space Illuminated

Reflecting work in the DeGrado Lab

Published here August 4, 2026

The emergence of novel versus known three-dimensional structures from random sequences

Rose Yang, Hyunjun Yang, Anton Davydenko, Zack Mawald, Rian Kormos, Dru Myerscough, Yibing Wu, and William F. DeGrado

PNAS 2026, 123, e2535076123. https://doi.org/10.1073/pnas.2535076123

View Original Publication


The question of how proteins first acquired stable three-dimensional structures sits at the heart of molecular evolution. One leading hypothesis holds that modern proteins descended from short ancestral peptides that duplicated and diversified over time, with tandem repetition offering a shortcut through the vast space of possible sequences. Testing this idea at scale has been difficult, but the arrival of fast, accurate structure-prediction software now makes it possible to screen millions of sequences computationally and ask, with quantitative precision, what fraction of random repeat sequences can fold, and into what shapes.

Researchers in the DeGrado Group at the University of California San Francisco and the Yang Group at Brandeis University, published in PNAS, generated approximately 120-residue proteins built from tandem repeats ranging from 5 to 120 residues in length and submitted roughly one million sequences per repeat length to structure prediction via RaptorX, retaining models with a mean pLDDT above 90. Representative cluster centroids were then re-predicted with AlphaFold3. Two amino acid alphabets were tested: the full 20-residue canonical set and a reduced 10-residue "primitive" set reflecting amino acids thought to be available early in molecular evolution. Sequences passing the foldability threshold were clustered by structural similarity at 1.0 Å RMSD, and promising helical-screw sequences were expressed in Escherichia coli, purified by immobilized metal affinity chromatography and size-exclusion chromatography, and characterized by circular dichroism spectroscopy and X-ray crystallography.

Foldability varied sharply with repeat length. Pentameric repeats folded at 3.6%, 10-residue repeats reached a peak of 17.2%, and foldability fell to 0.01% for 60-residue repeats and to 0.001 parts per million for fully random 120-residue sequences. Across the 5- to 20-residue range, the dominant predicted structures were β-solenoids whose geometry matched known proteins in the Protein Data Bank, including right-handed square solenoids from 5-mers, triangular solenoids from 6-mers, and rectangular solenoids from 8- and 10-mers. Notably, 9-mer and 10-mer clusters contained Asp residues pointing inward near mainchain carbonyl groups, an arrangement that recapitulates the Ca2+-coordinating Repeats in ToXin, RTX, motifs found in bacterial secretion proteins. Helical hairpins and bundles appeared far less frequently in strict tandem repeats, but introducing small insertions or deletions between blocks of four heptad repeats raised the proportion of sequences folding into helical bundles by 300- to 500-fold, reaching 0.93% for a two-residue glycine insertion.

Beyond the catalogue of known folds, the survey revealed a previously rare supersecondary structure the authors call the α-helical screw: a series of interconnected helices wound tightly around a central axis. Helical screws appeared most frequently in 8-mer repeats, with 133 occurrences, and in several other repeat lengths, yet a TM-align search of the Protein Data Bank returned only a single natural example of the 8-mer variant, and a Foldseek search of the AlphaFold Database UniProt set of 590,183 structures found only three additional examples. To confirm that this motif is physically realizable, the team designed six sequence variants, HeliScrew1 through HeliScrew6, around the cluster centroid structure, HeliScrew7, using the generative model Chroma. All seven constructs expressed robustly and eluted as apparent monomers by size-exclusion chromatography. HeliScrew4 yielded crystals diffracting to 2.20 Å resolution; the solved structure agreed with the Chroma design model at a Cα backbone RMSD of 1.05 Å and with the original cluster centroid at 1.38 Å.

These results carry two distinct implications. First, they provide quantitative support for the hypothesis that protein evolution proceeded via the duplication and diversification of short peptide sequences: tandem repetition raises foldability by many orders of magnitude relative to fully random sequences, and the dominant structural outcomes, β-solenoids, closely parallel the repeat architectures seen across enzymes, binding proteins, and structural proteins in modern organisms. Second, the discovery and experimental validation of the α-helical screw demonstrates that structure-prediction networks trained on known proteins can identify designable folds that lie outside the distribution of their training data, opening a route to proteins with architectures not yet sampled by evolution and potentially useful in contexts where natural protein scaffolds offer no precedent.