Engineering De Novo Binders Against Staphylococcal Enterotoxin B2 with AffiniBind
A full AffiniBind campaign against staphylococcal enterotoxin B2: surface fingerprinting, epitope resolution, parallel binder generation, and a twelve-dimension developability assessment that ranks 11 scored candidates across two epitopes.
Staphylococcal enterotoxin B2 (SEB2) belongs to a family of exotoxins that act as superantigens, cross-linking major histocompatibility complex class II molecules with T-cell receptors and driving massive, nonspecific cytokine release. The clinical consequence is a rapid-onset toxic syndrome that can progress to shock, and the toxin's stability under thermal and proteolytic stress makes neutralization by conventional small-molecule approaches difficult. A high-affinity protein binder that occludes the receptor-binding surface of SEB2 would offer a direct route to toxin neutralization, and the computational design of such a binder is precisely the problem AffiniBind was built to address.
This campaign targeted the SEB2 structure deposited under PDB ID STAPHYLOCOCCAL ENTEROTOXIN B2. The goal was not to replace experimentation but to compress the search space: identify the most probable binding surfaces, generate candidate binders against those surfaces in parallel, and rank the candidates by a developability score that anticipates manufacturing and biophysical failure modes before a single construct is ordered.
Surface Fingerprinting: Where the Protein Invites Contact
The campaign began with whole-surface characterization of SEB2. Of the 238 total protein residues, 122 met the surface criterion at a relative solvent-accessible surface area (SASA) cutoff of 0.25, yielding a surface fraction of 0.5126. The mean total SASA across the structure was 50.88 Ų, with a mean relative SASA of 0.3068 and a mean residue depth of 6.3 Å. These values describe a compact globular toxin with roughly half its residues exposed to solvent — a surface large enough to support multiple distinct binding modes, but not so expansive that patch identification becomes degenerate.
Three complementary fingerprinting models — ScanNet, MPBind, and MaSIF — were run over the surface, and their outputs were combined with APBS electrostatic character and a learned scorer (db5_xgboost_v1) to produce a single binding propensity per residue. The top of that ranking is dominated by aromatic and hydrophobic residues clustered in two regions of the structure.
The highest-propensity residue in the entire structure is A:TYR:94, with a binding propensity of 0.9263. Its component scores tell a coherent story: ScanNet at 0.6884, MaSIF at 0.7546, and a relatively modest MPBind contribution of 0.3500. The residue is polar by electrostatic character (score 0.0814), highly exposed (relative SASA 0.6592), and sits at a residue depth of 7.2170 Å. Immediately behind it is A:TYR:91 at 0.9130, with the highest relative SASA in the top set (0.9398) and a MaSIF score of 0.7529.
The next tier is more hydrophobic. A:ILE:36 (0.8978), A:LEU:45 (0.8946), A:VAL:26 (0.8929), and A:PHE:44 (0.8912) all carry neutral electrostatic character and substantial burial depth, suggesting a contiguous hydrophobic patch that could drive high-affinity contacts if matched with complementary binder surface. A:ASP:147 (0.8939) is the notable exception — a negatively charged residue with strong ScanNet (0.5562) and MaSIF (0.7209) support, indicating that electrostatics are not the sole determinant of predicted binding propensity here.
The full top-15 list also includes A:TYR:90 (0.8911), A:TYR:46 (0.8833), A:GLU:22 (0.8723), A:LEU:20 (0.8702), A:ASN:23 (0.8701), A:TYR:167 (0.8617), A:TYR:129 (0.8611), and A:PHE:177 (0.8448). The aromatic density is striking: tyrosines and phenylalanines account for eight of the top fifteen residues, and their MaSIF scores are consistently the strongest component, ranging from 0.5421 to 0.7883.
Residue
ScanNet
MPBind
MaSIF
XGBoost Propensity
Rel. SASA
Depth (Å)
A:TYR:94
0.6884
0.3500
0.7546
0.9263
0.6592
7.2170
A:TYR:91
0.5973
0.2200
0.7529
0.9130
0.9398
7.6070
Patch Identification: From Residues to Epitopes
The pipeline does not design against isolated high-scoring residues. It groups them into contiguous surface patches using a neighbor cutoff of 12.0 Å and a patch radius of 8.0 Å, requiring a minimum patch size of 5 residues. The top-ranked patches by size were patch_0005 (40 residues), patch_0009 (43 residues), and patch_0016 (64 residues). These patches were then resolved into two epitopes for binder generation.
Epitope 01 attracted 9 of the 11 scored candidates and has a mean unified score of 0.7017. Its representative candidate is SEB_01_All_Generators_bindcraft_binder_006. The target residue list begins with A:23, A:42, A:43, A:44, A:45, A:46, A:47, A:48, A:50, A:65, A:67, A:68, A:69, A:70, A:71, A:72, A:78, A:89, A:90, and A:91, with 30 additional residues beyond those listed. This epitope captures the hydrophobic-aromatic cluster identified in the top binding residues — A:LEU:45, A:PHE:44, A:TYR:46, A:TYR:90, and A:TYR:91 all fall within it.
Epitope 02 attracted 2 candidates and has a higher mean unified score of 0.7492. Its representative is SEB_01_All_Generators_pxdesign_binder_001. The target residue list begins with A:17, A:18, A:19, A:20, A:21, A:22, A:23, A:25, A:26, A:27, A:28, A:29, A:30, A:31, A:32, A:43, A:44, A:45, A:46, and A:47, with 40 additional residues beyond those listed. This epitope overlaps Epitope 01 at the A:43–A:47 segment but extends toward the N-terminal region, capturing A:LEU:20, A:GLU:22, A:ASN:23, and A:VAL:26.
Binder Generation: Three Tools, One Target
The campaign generated 29 candidates in parallel using FreeBindCraft, which couples ProteinMPNN sequence design with AlphaFold2 structure validation. The source-tool split was 9 from bindcraft, 10 from boltzgen, and 10 from pxdesign. After prefiltering, 18 candidates were excluded before unified scoring, leaving 11 scored candidates with non-null unified scores.
The prefilter is the first explicit gate. It is not a soft ranking step — candidates that fail it do not receive a unified score and do not appear in the ranked list. The 18 excluded candidates represent 62% of the initial generation, which is a substantial attrition rate and reflects the stringency of the developability and structural-quality filters applied before scoring.
The per-tool medians over scored candidates reveal sharp differences in generator behavior. bindcraft contributed 9 of the 11 scored candidates, with a median unified score of 0.7064, median PRODIGY ΔG of −9.2000 kcal/mol, and median predicted Kd of 1.80e-07 M. Its median interface pLDDT was 88.74 and median scaffold pLDDT 90.51 — the highest scaffold confidence of any tool.
boltzgen produced only 1 scored candidate, but that candidate carries the strongest predicted binding affinity in the entire campaign: PRODIGY ΔG of −13.6000 kcal/mol and a predicted Kd of 9.80e-11 M. Its median solubility score of 0.7944 is also the highest of the three tools. However, its median interface pLDDT is 65.39 and median scaffold pLDDT is 58.00 — dramatically lower than the other generators, suggesting that the boltzgen candidate's predicted affinity may be accompanied by structural uncertainty that the unified score penalizes.
pxdesign also produced only 1 scored candidate, but that candidate ranks #1 overall with a unified score of 0.7987. Its median PRODIGY ΔG of −10.3000 kcal/mol and predicted Kd of 2.70e-08 M are competitive, and its median SAP3D spatial score of −3.5302 is the most favorable of any tool, indicating low aggregation propensity. Its median flexible loop fraction of 0.0784 is the lowest, suggesting a rigid, well-packed scaffold.
Metric (median)
bindcraft (n=9)
boltzgen (n=1)
pxdesign (n=1)
Unified score
0.7064
0.6996
0.7987
PRODIGY ΔG (kcal/mol)
−9.2000
−13.6000
−10.3000
PRODIGY Kd (M)
1.80e-07
9.80e-11
2.70e-08
Solubility score
0.6330
0.7944
0.7526
SAP3D spatial score
Developability Assessment: Twelve Dimensions, One Score
The unified score is not a single affinity prediction. It integrates twelve developability dimensions: solubility, aggregation risk via 3D-SAP, charge state via PROPKA, binding affinity via PRODIGY, packing density, protease susceptibility, hydrophobic scaffold patch size, buried unsatisfied polar count, flexible loop fraction, interface and scaffold pLDDT, and generator confidence metrics (ipTM, pLDDT, RMSD). Each dimension is a data contract — a specific, auditable number attached to each candidate.
The top eight candidates are all from bindcraft except the #1-ranked pxdesign candidate. The ranking is not simply a sort by predicted affinity. The #4 candidate (SEB_01_All_Generators_bindcraft_binder_009) has the strongest predicted binding of any bindcraft candidate — PRODIGY ΔG of −12.4000 kcal/mol and predicted Kd of 8.20e-10 M — but ranks below three candidates with weaker predicted affinity because its SAP3D spatial score of −1.5489 is less favorable and its scaffold carries a hydrophobic patch size of 87.0 Ų, a flag shared by all eight bindcraft candidates in the top list.
The #1 candidate, SEB_01_All_Generators_pxdesign_binder_001, earns its position through balance rather than dominance in any single category. Its PRODIGY ΔG of −10.3000 kcal/mol and predicted Kd of 2.70e-08 M are competitive but not the best in the set. What distinguishes it is the combination of a favorable SAP3D spatial score of −3.5302, a solubility score of 0.7526, a low flexible loop fraction of 0.0784, zero buried unsatisfied polar residues, and a hydrophobic scaffold patch size of 0.0 Ų. Its interface and scaffold pLDDT values are both 83.00, which are lower than the bindcraft candidates but still within a range the pipeline treats as acceptable.
The bindcraft candidates cluster tightly. Their unified scores span only 0.6997 to 0.7218, and their developability profiles are nearly identical: all eight have hydrophobic scaffold patch sizes of 87.0 Ų, all have zero protease-exposed motifs, and seven of eight have zero buried unsatisfied polar residues. The two exceptions are SEB_01_All_Generators_bindcraft_binder_004 and _003, each with one buried unsatisfied polar residue. Their flexible loop fractions range from 0.1600 to 0.2333, and their average packing densities range from 12.3333 to 14.5333.
The interface core counts tell a structural story. The #4 candidate has the largest interface core at 39 residues, with 23 rim residues — a total of 62 interface residues, the largest in the top set. The #1 pxdesign candidate has 27 core and 24 rim residues, for 51 total. The #8 candidate has 33 core and 15 rim residues, for 48 total. Larger interfaces are not automatically better; they can increase the entropic cost of binding and the risk of off-target contacts, but they also provide more opportunities for specific interactions.
Candidate
Unified
ΔG (kcal/mol)
Kd (M)
Solubility
SAP3D
pI
Interface pLDDT
Scaffold pLDDT
Hydrophobic Patch (Ų)
pxdesign_binder_001
0.7987
−10.3000
2.70e-08
0.7526
−3.5302
9.8000
83.0000
83.0000
0.0000
Ranking and Clustering: What the Unified Score Actually Rewards
The unified score is a weighted combination of the twelve developability dimensions, but the weights are not arbitrary. They reflect the pipeline's engineering-first philosophy: a candidate that binds tightly but aggregates during expression is not a useful reagent. A candidate that expresses well but has a flexible, protease-susceptible scaffold will not survive manufacturing.
The ranking reveals the score's priorities. The #1 candidate does not have the best predicted affinity, the best interface pLDDT, or the best scaffold pLDDT. It wins because it has no hydrophobic scaffold patch, the most favorable SAP3D score, a low flexible loop fraction, and zero buried unsatisfied polar residues. These are the dimensions that predict expression yield, storage stability, and manufacturing reproducibility — the failure modes that kill campaigns after the binding assay.
The #4 candidate illustrates the counterfactual. Its PRODIGY ΔG of −12.4000 kcal/mol and predicted Kd of 8.20e-10 M are the best in the bindcraft set and second only to the boltzgen candidate overall. Its scaffold pLDDT of 92.56 is the highest of any candidate in the top eight. Yet it ranks #4 because its SAP3D spatial score of −1.5489 is the least favorable of the bindcraft candidates, and its hydrophobic scaffold patch of 87.0 Ų is shared by all bindcraft candidates but absent in the pxdesign candidate.
The epitope clustering adds another layer. Epitope 01 has 9 candidates with a mean unified score of 0.7017; Epitope 02 has 2 candidates with a mean unified score of 0.7492. The higher mean for Epitope 02 is driven entirely by the #1 pxdesign candidate, since the other Epitope 02 candidate is not in the top eight. The epitope assignment is not a ranking criterion — it is a structural annotation that tells the experimentalist which surface each binder is predicted to engage.
Validation Plan: Ready to Run, Not Yet Run
The campaign's ordering stage produced order files for two planned experimental validation arms. These experiments are available and ready to run — they have not been executed, and no Kd values, titration curves, or expression yields exist from this campaign.
The first arm is a yeast surface display (YSD) screen using IDT oPools oligo pools covering all 29 generated candidates, tiered by confidence, for gap-repair cloning into the pCTCON2 vector. This screen is designed to test binding and expression simultaneously: candidates that display well and bind SEB2 will be enriched, while candidates that misfold or aggregate on the yeast surface will drop out.
The second arm is an IVTT TR-FRET assay using 96-well IDT gBlock constructs for the 11 scored candidates, with the T7 promoter + GG + SD + Kozak + ATG + FLAG construct context using the E. coli MFC codon table. This cell-free assay is designed to measure binding affinity in a format that avoids the confounding effects of cellular expression and secretion.
Both order files and their protocol documents are downloadable from the experiments tab. The tiering in the YSD screen and the selection of 11 candidates for the TR-FRET assay are direct outputs of the unified ranking — the pipeline's computational gates determine which candidates receive experimental resources first.
Honest Assessment
This campaign demonstrates both the strengths and the current limits of the AffiniBind pipeline. The surface fingerprinting stage produced a coherent, interpretable map of SEB2's binding surface, with a clear aromatic-hydrophobic cluster centered on A:TYR:94, A:TYR:91, A:ILE:36, A:LEU:45, and A:PHE:44. The patch identification resolved this surface into two overlapping epitopes, and the binder generation stage produced 11 scored candidates with predicted affinities spanning from 9.80e-11 M to 4.10e-07 M. The unified ranking is transparent and auditable: every candidate carries twelve developability metrics, and the ranking logic is visible in the data.
What worked well is the discrimination between generators. The pxdesign candidate's favorable SAP3D score, zero hydrophobic scaffold patch, and low flexible loop fraction are exactly the properties the pipeline was designed to reward, and it ranks #1 despite having lower interface and scaffold pLDDT than the bindcraft candidates. The bindcraft candidates cluster tightly, which is expected from a single generator but also means the ranking within that cluster is driven by small differences in packing density, buried unsatisfied polar count, and flexible loop fraction — dimensions that may not be experimentally resolvable at this stage.
What the pipeline struggled with is generator diversity among scored candidates. Of the 11 scored candidates, 9 are from bindcraft, 1 from boltzgen, and 1 from pxdesign. The boltzgen candidate has the best predicted affinity in the campaign but the lowest interface and scaffold pLDDT, suggesting that boltzgen's structural confidence metrics are pulling its unified score down. The pxdesign candidate has the best predicted affinity in the campaign but the lowest interface and scaffold pLDDT, suggesting that boltzgen's structural confidence metrics are pulling its unified score down. The pxdesign candidate ranks #1 but is the only pxdesign candidate that survived prefiltering, so its strong performance is a single data point, not a pattern. A follow-up campaign should investigate why 9 of 10 boltzgen candidates and 9 of 10 pxdesign candidates failed the prefilter — whether the failures were driven by structural quality, developability flags, or generator-specific artifacts — and whether the prefilter thresholds are calibrated appropriately for non-bindcraft generators.
The hydrophobic scaffold patch size of 87.0 Ų shared by all eight bindcraft candidates is a notable liability. It is not a disqualifying flag in this campaign, but it is a consistent signature of the bindcraft generator's scaffold design. If the YSD screen shows aggregation or poor display for the bindcraft candidates, this metric will be the first place to look. The pxdesign candidate's zero hydrophobic scaffold patch is a structural advantage that the unified score rewards, but it remains to be seen whether that advantage translates into better experimental behavior.
The epitope overlap between Epitope 01 and Epitope 02 is both an opportunity and a risk. The shared A:43–A:47 segment includes three top-ten binding-propensity residues, and a binder that engages this segment could occlude both epitopes. But the overlap also means that candidates from the two epitopes are not fully orthogonal — if the shared segment is not actually the functional neutralization surface, both epitopes could be wrong in the same way. The planned YSD screen will not resolve this question directly, because it tests binding to the whole toxin, not to specific epitopes. A follow-up campaign should include competition assays or epitope-mapping experiments to determine whether the two epitopes are functionally distinct.
The validation plan is the campaign's most important deliverable. The order files for the YSD screen and the IVTT TR-FRET assay are available and ready to run, and they are the direct output of the computational ranking. The 29-candidate YSD screen is tiered by confidence, which means the top-ranked candidates will be tested first, and the 11-candidate TR-FRET assay covers the full scored set. These experiments will produce the first experimental Kd values and expression yields for this campaign, and they will either validate the unified score's predictions or reveal which developability dimensions are over- or under-weighted.
A follow-up campaign would target several specific improvements. First, the prefilter attrition rate of 62% is high, and understanding why 18 of 29 candidates failed before unified scoring is essential for improving generator yield. Second, the boltzgen candidate's predicted affinity of 9.80e-11 M is the best in the campaign, but its low pLDDT values suggest structural uncertainty that the current scoring may over-penalize — a follow-up should test whether boltzgen candidates with high predicted affinity but low pLDDT can be rescued by structural refinement or by relaxing the pLDDT threshold. Third, the hydrophobic scaffold patch size of 87.0 Ų in all bindcraft candidates is a systematic generator bias that should be addressed either by scaffold redesign or by adding a hydrophobic patch penalty to the prefilter. Fourth, the positively charged pxdesign candidate (net charge +4.0348 at pH 7, pI 9.8000) should be tested for nonspecific binding in serum, because its charge state is an outlier in the top set and could affect its behavior in any assay that includes complex biological media.
The campaign's core strength is its transparency. Every candidate carries twelve developability metrics, every stage produces an auditable data contract, and the ranking logic is visible in the data. The unified score is not a black box — it is a weighted combination of dimensions that the pipeline's engineers can inspect, adjust, and validate against experimental results. When the YSD and TR-FRET data come back, they will close the loop: the computational predictions will be tested against reality, and the weights in the unified score will be updated accordingly. That is the engineering-first promise of AffiniBind, and this campaign against SEB2 is a clean, complete test of that promise.