Humanness Analysis
Evaluates antibody humanness using OASign (Observed Antibody Space Immunogenicity) and PFA (Positional Frequency Analysis) to quantify how "human-like" a sequence is based on natural antibody repertoires.
[!IMPORTANT] Integrated Immunogenicity Assessment: Sapiens, OASign, and RPEMHC must be evaluated in conjunction for a complete developability profile. OASign and Sapiens quantify sequence humanness (which corresponds to T-cell receptor self-tolerance), whereas RPEMHC predicts physical peptide presentation to human MHC-II. A therapeutic candidate is only safe if it scores well on both metrics: a "Good" (Green) RPEMHC score on a sequence with low humanness (such as a mouse sequence) is still highly immunogenic in patients because there is no self-tolerance for the presented peptides.
This tool provides a granular, residue-level view of humanness, helping identify specific non-human regions that may require engineering or humanization.
Accessing the Tool
Select an antibody in the Project View. Go to the Analysis menu and select Humanness. This will open the Humanness Analysis workspace.

Using the Tool
The Humanness Analysis workspace provides an integrated, side-by-side view combining OASign repertoire sequence humanness with RPEMHC Class II MHC presentation risk alongside granular residue-level diagnostics.
1. Unified Aggregate Scoring (Two Rows)
- Row 1 β OASign Repertoire Humanness:
- VL / VH OASign (%): The fraction of 9-mer peptides in the variable region that appear in natural human antibody repertoires.
- Fv OASign (%): The peptide-weighted aggregate humanness of the paired variable domain.
- Germline Gene Boxes: Automatically displays the closest matching human V and J germlines.
- Row 2 β RPEMHC ADA Immunogenicity Risk:
- VL / VH RPEMHC: The 4-decimal probability (\(0.0000 β 1.0000\)) that the chain variable domain triggers helper T-cell activation.
- Fv RPEMHC: The paired Fv anti-drug antibody (ADA) probability, color-coded by clinical risk (\(\le 0.35\) Green, \(\le 0.45\) Orange, \(> 0.45\) Red).
- Allele Context: Displays the top driving HLA Class II allotype for the active panel.
2. High-Risk MHC-II Epitopes (T-Cell Hotspots)
- Hotspot Extraction & 1-Click Noise Filter: Detects high-affinity 15-mer sliding windows (\(\ge 0.638\)) across the Fv domain.
- Human Germline Context: Badges each epitope as π’
100% Germline(tolerated self-peptide), π‘Somatic/Point Mutation, π£CDR3 Junction, or π΄Non-Germline. - 1-Click Noise Reduction: Use the
Non-Germline Onlyfilter to eliminate noise and focus strictly on actionable de-immunization liabilities. - Interactive Multi-Column Sorting: Click any column header (Chain, Region, Positions, Peptide, Germline Status, Top Allele, Peak Binding Score, Population Breadth) to cycle Ascending \(\rightarrow\) Descending \(\rightarrow\) Remove. Supports composite multi-column sorting with precedence badges (
[1],[2], etc.) and default binding score ranking.
3. Dual-Track Sequence Alignment Cards
- Query (OASign) Track: Colored by non-human 9-mer clusters (white for 100% human, yellow \(\rightarrow\) red for non-human stretches). Hovering over any residue displays overlapping 9-mers and mutations required to reach human sequences.
- Query (RPEMHC) Track: Colored by MHC-II groove binding affinity for the active allele panel (\(\ge 0.638\) orange, \(\ge 0.80\) red).
- Matched Germline Tracks: Displays aligned human germline reference sequences with non-germline mutations highlighted in red.
4. Residue-Level Analysis (Dual-Pane)
The diagnostic tables in the VL and VH panes display both metrics simultaneously:
- OASign (%): Percentage of human 9-mers overlapping this residue.
- MHC-II Score: Maximum predicted Class II binding score overlapping this residue, with tooltips showing the top presenting HLA allele.
- AbLang Diff: Difference between actual residue likelihood and highest-scoring alternative.
- PFA [Highest]: Positional frequency analysis against human germlines.
- Engineering Modal: Clicking any residue opens the Engineering modal with predictive de-immunization scans and live score previews.
- Unified Excel Export: One-click export generates a comprehensive workbook with summary cards, dual OASign/RPEMHC columns, 20 allele scores, and an Fv Hotspots sheet.
Strategy: Triage positions using the Heatmap. Use PFA and AbLang scores to find weak spots. Open the Engineering modal on high-risk residues to inspect the predictive scans and select human-favored substitutions that lower predicted MHC-II binding scores without disrupting structural properties.
Scoring Thresholds
AbLead uses standardized clinical thresholds to categorize OASign and RPEMHC results:
| OASign Score | Category | UI Color (colors.py) |
Description |
|---|---|---|---|
| 85% β 100% | High / Good | Green (#88B46C) |
Matches the mean score of natural human germline antibodies. |
| 75% β 84% | Medium / Low Risk | Yellow (#F8D548) |
Typical range for clinically successful humanized antibodies. |
| < 75% | Low / High Risk | Red (#CF4A3C) |
Below standard clinical humanization norms. |
| RPEMHC Probability (Single Chain / Fv) | Category | UI Color (colors.py) |
Description |
|---|---|---|---|
RPEMHC VL / VH \(\le 0.2000\) (Fv \(\le 0.3500\)) |
Good / Low Risk | Green (#88B46C) |
Low population-weighted helper T-cell activation probability. |
RPEMHC VL / VH \(0.2000β0.3000\) (Fv \(0.3500β0.4500\)) |
Warning / Moderate | Yellow (#F8D548) |
Moderate clinical risk range / candidate T-cell neo-epitopes. |
RPEMHC VL / VH \(> 0.3000\) (Fv \(> 0.4500\)) |
Severe / High Risk | Red (#CF4A3C) |
High helper T-cell activation probability / elevated ADA risk (flags Bococizumab at 0.4630 in RED). |
| PFA Score | Category | UI Color | Description |
|---|---|---|---|
| ≥ 90% | Very High | Green | Highly conserved human germline residue. |
| ≥ 80% | High | Orange | Common human germline residue. |
| ≥ 70% | Medium | Yellow | Moderately common human germline residue. |
| < 70% | Unusual | Red | Infrequent or unusual residue at this position. |
| AbLang Diff | Category | UI Color | Description |
|---|---|---|---|
| < 3 | Good | Green | High naturalness; residue matches or is close to optimal. |
| 3 β 5 | Medium | Orange | Moderate divergence from natural antibody profiles. |
| ≥ 5 | High | Red | Significant divergence; residue may be unstable or non-natural. |
Scientific & Empirical Basis for Thresholds
- OASign Thresholds (\(\ge 0.8500\), \(0.7500β0.8400\), \(< 0.7500\)): Derived from the BioPhi / OASis benchmark study (Prihoda et al., mAbs 2022). Natural human antibodies average \(\ge 85\%\) repertoire identity. Clinically approved humanized antibodies (e.g. Trastuzumab, Bevacizumab, Pertuzumab) score between \(75\%\) and \(85\%\), while non-human/murine sequences fall below \(75\%\).
- RPEMHC Population Probability Thresholds (VL/VH \(\le 0.2000\), Fv \(\le 0.3500\)): Calibrated using first-principles biophysical groove binding activation and global HLA Class II human population genetics benchmarks (Paul et al., Immunity 2013). Clinically approved low-ADA therapeutics display clean profiles (\(\le 0.3500\) Fv / \(\le 0.2000\) VL/VH), while high-risk candidates (e.g. Bococizumab at \(0.4630\) Fv) are flagged with severe developability penalties.
- PFA Conserved Residue Cutoffs (\(\ge 90\%\), \(\ge 80\%\), \(\ge 70\%\), \(< 70\%\)): Based on IMGT human germline amino acid distributions across matched V and J gene families. Residues \(<70\%\) represent non-canonical framework variations, murine backmutations, or somatic mutations.
- AbLang Evolutionary Likelihood (\(\Delta < 3.0\), \(3.0β5.0\), \(\ge 5.0\)): Calibrated against transformer language model log-likelihood distributions, where \(\Delta \ge 5.0\) flags structurally unfavorable substitutions.
The 4 Quadrants of Immunogenicity Risk
Clinical immunogenicity (Anti-Drug Antibody / ADA development) is governed by two orthogonal biological components:
- Biophysical Groove Presentation (RPEMHC): Does the peptide anchor into human HLA Class II pockets with sufficient affinity to be displayed by antigen-presenting cells?
- Repertoire Tolerance vs. Foreignness (OASign / PFA): Is the presented peptide recognized as natural human "Self" (tolerized during thymic negative selection) or "Foreign" (stimulating naive CD4+ T-cells)?
| High OASign (\(\ge 0.8500\) Human Repertoire) | Low / Moderate OASign (\(< 0.8000\) Non-Human / Foreign) | |
|---|---|---|
| Low RPEMHC (\(\le 0.3500\)) | Ideal Clinical Profile (e.g., Trastuzumab, Bevacizumab) β’ Fully tolerized by human immune system β’ Low Class II presentation density β’ Lowest statistical risk of clinical ADA |
De-Immunized Non-Human / Chimeric β’ Sequence is foreign/non-human β’ Lacks physical MHC-II groove anchors β’ Minimal presentation to helper T-cells |
| High RPEMHC (\(> 0.4500\)) | Self-Presenter (e.g., Renvistobart) β’ High biophysical groove affinity β’ Composed of natural human repertoire 9-mers β’ Tolerized by central/peripheral immune mechanisms |
Clinical Danger Zone (e.g., Bococizumab) β’ Strong MHC-II peptide presentation (\(0.4630\), Red) β’ Contains non-human / foreign CDR neo-epitopes β’ High risk of helper T-cell priming & neutralizing ADAs |
Clinical Case Study: Bococizumab (The Danger of Global Score Averaging)
Bococizumab (Pfizerβs humanized anti-PCSK9 mAb) provides a classic clinical example of why granular epitope evaluation is essential:
- Clinical Outcome: In late-stage Phase 3 trials (the SPIRE program, encompassing over 27,000 patients), Pfizer was forced to discontinue development after discovering that 48% of patients developed anti-drug antibodies (ADAs), with 29% developing high-titer neutralizing ADAs that substantially eroded LDL-C lowering efficacy over time and caused elevated injection-site reactions.
- First-Principles Immunogenic Profile: Under AbLead's population-weighted probability model, Bococizumab is unequivocally flagged as a high-risk candidate:
OASign Fv: \(0.8134\) (standard humanized range; yellow)RPEMHC VL: \(0.2540\) (warning / moderate; yellow)RPEMHC VH: \(0.2360\) (warning / moderate; yellow)RPEMHC Fv: \(0.4630\) (severe / high risk; red)
- The Immunological Mechanism: The human immune system does not evaluate global sequence averages. Bococizumabβs CDR loops (derived from a murine precursor) contained non-human somatic mutations that formed high-affinity 15-mer peptide anchors (\(\ge 0.800\)) presented across multiple core HLA-DRB1 alleles. Because these foreign CDR peptides had never undergone thymic negative selection in humans, dendritic cells presented them to naive CD4+ T cells, priming T follicular helper (\(T_{\text{FH}}\)) cells and driving robust B-cell class switching to neutralizing anti-bococizumab IgG antibodies.
- Key Engineering Takeaway: Global Fv scores provide high-throughput triage, but lead candidates must always be evaluated at the residue and sliding-window level using the RPEMHC 2D Matrix Heatmap, Linear Sequence Track, and Engineering Modal to detect and eliminate isolated CDR helper T-cell hotspots prior to clinical commitment.
High-Risk MHC-II Epitopes (Hotspots Card)
To prevent localized immunogenic liabilities from being masked by whole-chain averages, the Humanness workspace in RPEMHC Mode includes a dedicated High-Risk MHC-II Epitopes (Hotspots) summary table:
- Detection Criteria: Automatically extracts 15-mer peptide windows exhibiting high groove binding affinity (\(\ge 0.638\)) or broad population promiscuity (\(\ge 4\) presenting HLA Class II alleles).
- Human Germline Context & Self-Tolerance: Distinguishes 100% human germline peptides (which are naturally tolerized by central immune tolerance) from true neo-epitope liabilities containing somatic hypermutations (SHM), point mutations, or CDR3 junctions.
- Germline Status Badges & Tooltips: Displays clear visual origin badges (π’
100% Germline, π‘1 Mut: Pos, π£CDR3 Junction, π΄Non-Germline). Hovering over any badge or peptide provides an interactive tooltip showing the query sequence vs. matched human germline sequence, assigned V/J gene family, and exact amino acid mutations. - 1-Click Noise-Reduction Filter: Allows users to filter the table instantly between All Hotspots, Non-Germline Only (focusing de-immunization engineering purely on actionable liabilities), and Germline Only.
- Clustered Peak Reporting: Groups contiguous overlapping 15-mers into distinct physical epitope clusters, reporting the peak 15-mer peptide sequence, top presenting HLA allotype, peak binding score (color-coded by risk severity), and total population breadth (out of 20 alleles).
- Interactive Triage: Clicking any hotspot row immediately scrolls both the sequence alignment and residue table to that exact position and applies an outline highlight to facilitate immediate de-immunization engineering.
Positional Max Risk Evaluation
When viewing sequence alignments and residue diagnostics under Max Risk β Tier 1 Global Panel or Max Risk β All 20 Alleles:
- Positional Maximum Evaluation: Rather than fixing the entire sequence to the single allele that had the highest overall chain total sum, each individual amino acid position evaluates the true maximum binding score across all panel alleles at that specific position.
- Allele Attribution Tooltips: Hovering over any residue box in the alignment or the MHC-II score cell in the table displays the exact HLA Class II allotype driving the peak risk (e.g.
MHC-II Score (Max Tier 1: DRB1*04:05): 1.0030), immediately revealing localized neo-epitopes.
Methodology
The AbLead Humanness analysis suite combines repertoire database lookups, germline positional frequencies, masked language models, and deep neural network Class II MHC binding predictors to create an integrated developability and immunogenicity assessment.
1. OASign: Repertoire-Based Sequence Humanness
OASign measures sequence humanness by comparing antibody peptide fragments directly against the Observed Antibody Space (OAS), a comprehensive repository of over one billion natural human antibody sequences derived from diverse healthy donors.
- Fv Domain Trimming: Sequences are automatically identified and trimmed to the Variable Domain (\(\text{Fv}\)) using AntPack annotation. Constant domains are excluded from repertoire scoring to prevent false inflation from invariant constant sequences.
- Sliding 9-mer Peptide Decomposition: The variable region sequence of length \(L\) is chopped into \(L - 8\) overlapping 9-mer peptide windows. A 9-mer window size is biologically calibrated to match the typical footprint of linear B-cell epitopes and linear sequence identity segments.
- Hash-Based Repertoire Querying: Each 9-mer is queried against a high-performance hash index of human antibody repertoires. A 9-mer is scored as Human (1.0) if it appears in the curated human OAS database with sufficient donor support, and Non-Human (0.0) otherwise.
- Mutational Distance to Human: For any non-human 9-mer, OASign computes the minimum single amino acid substitution required to convert the peptide into a known human repertoire 9-mer, identifying minimal-mutation humanization pathways.
-
Chain & Fv Humanness Scores: The percentage of human 9-mers in the variable domain:
\[\text{Score}_{\text{Chain}} = \frac{\sum_{i=1}^{N_{\text{peptides}}} I(\text{peptide}_i \in \text{OAS})}{N_{\text{peptides}}} \times 100\%\]\[\text{Score}_{\text{Fv}} = \frac{N_{\text{pep, VL}} \cdot \text{Score}_{\text{VL}} + N_{\text{pep, VH}} \cdot \text{Score}_{\text{VH}}}{N_{\text{pep, VL}} + N_{\text{pep, VH}}}\]
2. Positional Frequency Analysis (PFA)
Positional Frequency Analysis evaluates how standard or unusual an amino acid is at a specific topological position compared to natural human antibody germlines.
- IMGT Alignment: Sequences are aligned against IMGT human germline reference databases for heavy (
IGHV,IGHJ) and light (IGKV,IGLV,IGKJ,IGLJ) chains. - Positional Matrix Lookups: For each aligned IMGT position, the system looks up the frequency distribution of all 20 amino acids across human germlines belonging to the matching V and J gene families.
- Frequency Output: Reports the percentage of human germline antibodies carrying the query residue alongside the dominant human amino acid at that position (e.g.
0.5% [90.7% T]). Positions with low PFA scores (\(<70\%\)) represent somatic hypermutations, non-canonical framework substitutions, or murine backmutations that may destabilize the antibody or act as foreign epitopes.
3. AbLang & AbLang-2 Language Model Likelihood
AbLang and AbLang-2 are transformer-based masked language models trained on hundreds of millions of antibody sequences to learn the structural and evolutionary rules of antibody folding.
- Contextual Residue Likelihood: Given the full variable domain context, the language model predicts the probability distribution of amino acids at each individual position.
-
Difference Metric (\(\Delta \text{Likelihood}\)): Evaluates the drop in log-likelihood from the model-optimal residue:
\[\Delta \text{Likelihood} = \text{Score}_{\text{optimal}} - \text{Score}_{\text{actual}}\]A low difference (\(\Delta < 3.0\)) indicates that the residue is favored by evolutionary antibody grammar, whereas a high difference (\(\Delta \ge 5.0\)) flags positions that diverge significantly from natural antibody profiles and may introduce conformational instability or poor expression.
4. RPEMHC: Relative Position Encoding for MHC-II Binding
RPEMHC predicts the physical chemical presentation of 15-mer antibody peptides to human Major Histocompatibility Complex Class II (MHC-II) molecules.
- 15-mer Sliding Windows: Antibody sequences are scanned using a sliding window of 15 amino acids (\(L - 14\) windows per chain), capturing the full peptide length that extends beyond the open-ended MHC-II binding groove.
- Relative Position Encoding: Unlike standard sequence encoding, relative position encoding explicitly represents the spatial distance between amino acid pairs and the polymorphic MHC-II binding pockets (P1, P4, P6, P7, P9), capturing cooperative anchoring interactions.
- 20-Allele Panel Evaluation: Evaluates binding affinity across both Tier 1 Core DRB1 alleles (>95% global population coverage) and Tier 2 Extended alleles (DRB3/4/5, DQ, DP).
- Thresholding & Cumulative Epitope Load: Peptides scoring \(> 0.50\) are flagged as presenting helper T-cell epitopes. The per-allele and aggregate construct scores (\(\sum \text{Scores} > 0.50\)) quantify total immunogenic burden to guide de-immunization engineering.
Understanding HLA Class II Alleles & Nomenclature
1. WHO / IMGT-HLA Nomenclature (e.g., HLA-DRB1*04:01)
Allele designations follow the standard World Health Organization (WHO) and IMGT-HLA nomenclature:
- Gene Locus: Prefix (e.g.,
HLA-DRB1*,HLA-DQA1*,HLA-DPA1*) designates the Class II \(\alpha\) or \(\beta\) chain gene. - Field 1 (e.g.,
04): The broad serological antigen group / allele family (e.g., DR4). - Field 2 (e.g.,
:01,:05): The specific protein allotype / subtype. Specific amino acid polymorphisms between subtypes alter the chemical topology and charge distribution of peptide-binding pockets (P1, P4, P6, P7, P9), resulting in distinct epitope repertoires.
2. Tier 1 vs. Tier 2 Allele Panels
- Tier 1 (Core Global Reference Panel): 11 high-prevalence
HLA-DRB1alleles (01:01,03:01,04:01,04:05,07:01,08:02,09:01,11:01,12:01,13:02,15:01) representing over 95% cumulative coverage across diverse global human populations. - Tier 2 (Extended Panel): 9 additional high-impact alleles across secondary
DRB3/4/5,DQA1/DQB1, andDPA1/DPB1loci for complete HLA Class II immunogenicity risk profiling.
3. HLA Class II Reference Panel
| Allele (WHO / IMGT) | Panel / Tier | Estimated Population Frequency | Binding Pocket Motif Preferences | Clinical / Therapeutic Risk Associations |
|---|---|---|---|---|
| HLA-DRB1*01:01 | Tier 1 Core | ~10β15% Global (EUR ~15β20%, ASN ~5β10%) | P1: Large hydrophobic (F, Y, W, L, I); P4/P6/P9: Aliphatic | Broad immunogenic peptide presentation; common ADA epitope presenter |
| HLA-DRB1*03:01 | Tier 1 Core | ~10β15% Global (EUR ~12β18%, AFR ~10β15%) | P1: Hydrophobic (L, I, M, F); P4: Acidic (D, E); P6: Basic/polar | Associated with SLE, Type 1 diabetes, autoimmune thyroiditis, biological drug ADA |
| HLA-DRB1*04:01 | Tier 1 Core | ~8β14% Global (EUR ~15β20%, AMR ~10β15%) | P1: Aromatic/aliphatic (F, Y, W, L); P4: Basic/polar; P9: Hydrophobic | Shared epitope for Rheumatoid Arthritis; high clinical ADA risk in TNFα biologics |
| HLA-DRB1*04:05 | Tier 1 Core | ~5β10% Global (East Asian ~15β25%, EUR ~2β4%) | P1: Aromatic/aliphatic; P4: Acidic/polar (pocket 4 subtype difference) | Key Asian-predominant RA & therapeutic immunogenicity risk allele |
| HLA-DRB1*07:01 | Tier 1 Core | ~15β22% Global (EUR ~25%, Middle Eastern ~25%) | P1: Large aromatic (F, Y, W); P4: Large hydrophobic; P9: Aliphatic | High frequency globally; co-expressed with DRB4*01:01; frequent mAb ADA epitope |
| HLA-DRB1*08:02 | Tier 1 Core | ~3β8% Global (Hispanic/Amerindian ~10β20%, ASN ~5β10%) | P1: Aliphatic/aromatic; P4: Acidic/polar; P6: Polar; P9: Hydrophobic | Important representation for Hispanic/Latin American and Indigenous cohorts |
| HLA-DRB1*09:01 | Tier 1 Core | ~5β10% Global (East Asian ~15β20%, EUR ~2β5%) | P1: Large hydrophobic; P4: Polar/aliphatic; P9: Small neutral | Prominent in East Asian populations; associated with autoimmune hepatitis & ADA |
| HLA-DRB1*11:01 | Tier 1 Core | ~10β18% Global (Mediterranean ~20β25%, AFR ~10β15%) | P1: Aliphatic/aromatic; P4: Small/polar; P6: Basic/polar; P9: Aliphatic | High Mediterranean and Hispanic frequency; co-expressed with DRB3*02:02 |
| HLA-DRB1*12:01 | Tier 1 Core | ~3β8% Global (East Asian ~8β15%, AFR ~5β10%) | P1: Hydrophobic; P4: Polar/basic; P9: Aliphatic | Asian and African cohort diversity representation |
| HLA-DRB1*13:02 | Tier 1 Core | ~6β12% Global (AFR ~8β15%, EUR ~6β10%) | P1: Hydrophobic; P4: Acidic/polar; P6: Polar; P9: Aliphatic | Broad peptide binding capacity; protective in certain viral/autoimmune contexts |
| HLA-DRB1*15:01 | Tier 1 Core | ~12β20% Global (EUR ~20%, ASN ~15β20%) | P1: Hydrophobic/aromatic; P4: Aliphatic; P7: Polar; P9: Small neutral | Strong association with Multiple Sclerosis; major driver of therapeutic protein ADA |
| HLA-DRB3*01:01 | Tier 2 (DRB3) | ~15β30% (Haplotype-linked with DR3, DR11, DR13, DR14) | P1: Hydrophobic; P4: Basic; P6/P9: Aliphatic | Co-expressed secondary β-chain; contributes additive Class II presentation |
| HLA-DRB3*02:02 | Tier 2 (DRB3) | ~20β35% (Co-expressed with DR11, DR13, DR14) | P1: Aliphatic/aromatic; P4: Polar; P9: Aliphatic | High expression in individuals carrying DRB1*11:01 / DRB1*13:01 haplotypes |
| HLA-DRB4*01:01 | Tier 2 (DRB4) | ~30β45% (Haplotype-linked with DR4, DR7, DR9) | P1: Hydrophobic; P4: Neutral/aliphatic; P7/P9: Polar | Highly prevalent secondary DRB molecule in DR4/DR7 carriers; prominent ADA mediator |
| HLA-DRB5*01:01 | Tier 2 (DRB5) | ~15β25% (Haplotype-linked with DR15, DR16) | P1: Large aromatic; P4: Polar/aliphatic; P9: Hydrophobic | Co-expressed with DRB1*15:01; implicated in MS and myelin peptide presentation |
| HLA-DQA1*05:01 / DQB1*02:01 | Tier 2 (DQ2.5) | ~10β20% Global (EUR ~20%, Hispanic ~15%) | P1/P9: Hydrophobic; P4/P6/P7: Negative charge (Glu, Asp, deamidated Gln) | Canonical DQ2.5 heterodimer; strong celiac disease & deamidated therapeutic epitope risk |
| HLA-DQA1*03:01 / DQB1*03:02 | Tier 2 (DQ8) | ~10β18% Global (AMR ~20β30%, EUR ~15%) | P1/P9: Hydrophobic/acidic; P4: Negative/polar | Canonical DQ8 heterodimer; associated with autoimmune diabetes and specific ADA |
| HLA-DQA1*01:02 / DQB1*06:02 | Tier 2 (DQ6.2) | ~15β25% Global (EUR ~20%, ASN ~15%) | P1: Aliphatic; P4: Hydrophobic; P9: Aliphatic | Strongest known genetic association with narcolepsy; high population prevalence |
| HLA-DPA1*01:03 / DPB1*04:01 | Tier 2 (DP401) | ~35β50% Global (Highest frequency DP allele globally) | P1: Large aromatic/hydrophobic; P6: Acidic/polar; P9: Hydrophobic | Predominant DP allele globally (>40% of humans); broad peptide presentation capacity |
| HLA-DPA1*01:03 / DPB1*04:02 | Tier 2 (DP402) | ~15β25% Global | P1: Aromatic/hydrophobic; P6: Polar; P9: Hydrophobic | Secondary major DP allele; essential for complete HLA Class II risk coverage |
Constant Domains: Biophysical Binding Affinity vs. Immunological Self-Tolerance
When evaluating full-length constructs containing constant regions (\(\text{CH}_1\), Hinge, \(\text{CH}_2\), \(\text{CH}_3\), \(\text{C}_\kappa\), \(\text{C}_\lambda\)), users may note high predicted MHC-II binding scores at certain constant domain positions:
-
Biophysical Groove Affinity: Neural network MHC-II predictors (including RPEMHC and NetMHCIIpan) evaluate the physical chemical affinity of 15-mer peptide windows into polymorphic HLA Class II binding pockets (P1, P4, P6, P9). Native human constant sequences naturally contain hydrophobic or aromatic core residues (e.g. Leu, Val, Phe, Tyr) that satisfy MHC-II pocket geometry, resulting in high predicted binding affinity.
-
Central & Peripheral Immune Tolerance: In humans, the immune system is centrally tolerized to native germline constant sequences during thymic development. Negative selection in the thymus eliminates autoreactive CD4+ T-cell clones with high affinity for self-peptides. Consequently, physical MHC-II groove binding in native human constant regions does not translate into clinical helper T-cell activation or anti-drug antibody (ADA) responses.
-
Variable Domain (Fv) Focus: Clinical immunogenicity risk and de-immunization engineering are focused on the Variable Domain (\(\text{Fv}\))βwhere somatic hypermutations, non-human CDR grafts, framework germline deviations, and V/J junctional sequences create foreign neo-epitopes that can trigger ADA responses in patients.
References
- The closest germlines are from the IMGT database. See IMGT.
- Humanness is determined using OASign against curated human 9-mer sequences from the Observed Antibody Space (OAS).
- RPEMHC (MHC-II binding prediction) is based on the DeepMHCII / RPEMHC ensemble deep learning model for predicting class II peptide-MHC binding affinity. Described in Wang et al., Bioinformatics (2024).
- OASign is based on the OASis method in BioPhi. See David Prihoda, Jad Maamary, Andrew Waight, Veronica Juan, Laurence Fayadat-Dilman, Daniel Svozil & Danny A. Bitton (2022) BioPhi: A platform for antibody design, humanization, and humanness evaluation based on natural antibody repertoires and deep learning, mAbs, 14:1, DOI: https://doi.org/10.1080/19420862.2021.2020203.