Skip to content

Humanness Analysis

Evaluates antibody humanness using OASign (Observed Antibody Space Immunogenicity) and PFA (Positional Frequency Analysis) to quantify how "human-like" a sequence is based on natural antibody repertoires.

[!IMPORTANT] Integrated Immunogenicity Assessment: Sapiens, OASign, and RPEMHC must be evaluated in conjunction for a complete developability profile. OASign and Sapiens quantify sequence humanness (which corresponds to T-cell receptor self-tolerance), whereas RPEMHC predicts physical peptide presentation to human MHC-II. A therapeutic candidate is only safe if it scores well on both metrics: a "Good" (Green) RPEMHC score on a sequence with low humanness (such as a mouse sequence) is still highly immunogenic in patients because there is no self-tolerance for the presented peptides.

This tool provides a granular, residue-level view of humanness, helping identify specific non-human regions that may require engineering or humanization.


Accessing the Tool

Select an antibody in the Project View. Go to the Analysis menu and select Humanness. This will open the Humanness Analysis workspace.

Humanness Analysis


Using the Tool

The Humanness Analysis workspace provides an integrated, side-by-side view combining OASign repertoire sequence humanness with RPEMHC Class II MHC presentation risk alongside granular residue-level diagnostics.

1. Unified Aggregate Scoring (Two Rows)

  • Row 1 β€” OASign Repertoire Humanness:
    • VL / VH OASign (%): The fraction of 9-mer peptides in the variable region that appear in natural human antibody repertoires.
    • Fv OASign (%): The peptide-weighted aggregate humanness of the paired variable domain.
    • Germline Gene Boxes: Automatically displays the closest matching human V and J germlines.
  • Row 2 β€” RPEMHC ADA Immunogenicity Risk:
    • VL / VH RPEMHC: The 4-decimal probability (\(0.0000 – 1.0000\)) that the chain variable domain triggers helper T-cell activation.
    • Fv RPEMHC: The paired Fv anti-drug antibody (ADA) probability, color-coded by clinical risk (\(\le 0.35\) Green, \(\le 0.45\) Orange, \(> 0.45\) Red).
    • Allele Context: Displays the top driving HLA Class II allotype for the active panel.

2. High-Risk MHC-II Epitopes (T-Cell Hotspots)

  • Hotspot Extraction & 1-Click Noise Filter: Detects high-affinity 15-mer sliding windows (\(\ge 0.638\)) across the Fv domain.
  • Human Germline Context: Badges each epitope as 🟒 100% Germline (tolerated self-peptide), 🟑 Somatic/Point Mutation, 🟣 CDR3 Junction, or πŸ”΄ Non-Germline.
  • 1-Click Noise Reduction: Use the Non-Germline Only filter to eliminate noise and focus strictly on actionable de-immunization liabilities.
  • Interactive Multi-Column Sorting: Click any column header (Chain, Region, Positions, Peptide, Germline Status, Top Allele, Peak Binding Score, Population Breadth) to cycle Ascending \(\rightarrow\) Descending \(\rightarrow\) Remove. Supports composite multi-column sorting with precedence badges ([1], [2], etc.) and default binding score ranking.

3. Dual-Track Sequence Alignment Cards

  • Query (OASign) Track: Colored by non-human 9-mer clusters (white for 100% human, yellow \(\rightarrow\) red for non-human stretches). Hovering over any residue displays overlapping 9-mers and mutations required to reach human sequences.
  • Query (RPEMHC) Track: Colored by MHC-II groove binding affinity for the active allele panel (\(\ge 0.638\) orange, \(\ge 0.80\) red).
  • Matched Germline Tracks: Displays aligned human germline reference sequences with non-germline mutations highlighted in red.

4. Residue-Level Analysis (Dual-Pane)

The diagnostic tables in the VL and VH panes display both metrics simultaneously:

  • OASign (%): Percentage of human 9-mers overlapping this residue.
  • MHC-II Score: Maximum predicted Class II binding score overlapping this residue, with tooltips showing the top presenting HLA allele.
  • AbLang Diff: Difference between actual residue likelihood and highest-scoring alternative.
  • PFA [Highest]: Positional frequency analysis against human germlines.
  • Engineering Modal: Clicking any residue opens the Engineering modal with predictive de-immunization scans and live score previews.
  • Unified Excel Export: One-click export generates a comprehensive workbook with summary cards, dual OASign/RPEMHC columns, 20 allele scores, and an Fv Hotspots sheet.

Strategy: Triage positions using the Heatmap. Use PFA and AbLang scores to find weak spots. Open the Engineering modal on high-risk residues to inspect the predictive scans and select human-favored substitutions that lower predicted MHC-II binding scores without disrupting structural properties.


Scoring Thresholds

AbLead uses standardized clinical thresholds to categorize OASign and RPEMHC results:

OASign Score Category UI Color (colors.py) Description
85% – 100% High / Good Green (#88B46C) Matches the mean score of natural human germline antibodies.
75% – 84% Medium / Low Risk Yellow (#F8D548) Typical range for clinically successful humanized antibodies.
< 75% Low / High Risk Red (#CF4A3C) Below standard clinical humanization norms.
RPEMHC Probability (Single Chain / Fv) Category UI Color (colors.py) Description
RPEMHC VL / VH \(\le 0.2000\) (Fv \(\le 0.3500\)) Good / Low Risk Green (#88B46C) Low population-weighted helper T-cell activation probability.
RPEMHC VL / VH \(0.2000–0.3000\) (Fv \(0.3500–0.4500\)) Warning / Moderate Yellow (#F8D548) Moderate clinical risk range / candidate T-cell neo-epitopes.
RPEMHC VL / VH \(> 0.3000\) (Fv \(> 0.4500\)) Severe / High Risk Red (#CF4A3C) High helper T-cell activation probability / elevated ADA risk (flags Bococizumab at 0.4630 in RED).
PFA Score Category UI Color Description
≥ 90% Very High Green Highly conserved human germline residue.
≥ 80% High Orange Common human germline residue.
≥ 70% Medium Yellow Moderately common human germline residue.
< 70% Unusual Red Infrequent or unusual residue at this position.
AbLang Diff Category UI Color Description
< 3 Good Green High naturalness; residue matches or is close to optimal.
3 – 5 Medium Orange Moderate divergence from natural antibody profiles.
≥ 5 High Red Significant divergence; residue may be unstable or non-natural.

Scientific & Empirical Basis for Thresholds

  • OASign Thresholds (\(\ge 0.8500\), \(0.7500–0.8400\), \(< 0.7500\)): Derived from the BioPhi / OASis benchmark study (Prihoda et al., mAbs 2022). Natural human antibodies average \(\ge 85\%\) repertoire identity. Clinically approved humanized antibodies (e.g. Trastuzumab, Bevacizumab, Pertuzumab) score between \(75\%\) and \(85\%\), while non-human/murine sequences fall below \(75\%\).
  • RPEMHC Population Probability Thresholds (VL/VH \(\le 0.2000\), Fv \(\le 0.3500\)): Calibrated using first-principles biophysical groove binding activation and global HLA Class II human population genetics benchmarks (Paul et al., Immunity 2013). Clinically approved low-ADA therapeutics display clean profiles (\(\le 0.3500\) Fv / \(\le 0.2000\) VL/VH), while high-risk candidates (e.g. Bococizumab at \(0.4630\) Fv) are flagged with severe developability penalties.
  • PFA Conserved Residue Cutoffs (\(\ge 90\%\), \(\ge 80\%\), \(\ge 70\%\), \(< 70\%\)): Based on IMGT human germline amino acid distributions across matched V and J gene families. Residues \(<70\%\) represent non-canonical framework variations, murine backmutations, or somatic mutations.
  • AbLang Evolutionary Likelihood (\(\Delta < 3.0\), \(3.0–5.0\), \(\ge 5.0\)): Calibrated against transformer language model log-likelihood distributions, where \(\Delta \ge 5.0\) flags structurally unfavorable substitutions.

The 4 Quadrants of Immunogenicity Risk

Clinical immunogenicity (Anti-Drug Antibody / ADA development) is governed by two orthogonal biological components:

  1. Biophysical Groove Presentation (RPEMHC): Does the peptide anchor into human HLA Class II pockets with sufficient affinity to be displayed by antigen-presenting cells?
  2. Repertoire Tolerance vs. Foreignness (OASign / PFA): Is the presented peptide recognized as natural human "Self" (tolerized during thymic negative selection) or "Foreign" (stimulating naive CD4+ T-cells)?
High OASign (\(\ge 0.8500\) Human Repertoire) Low / Moderate OASign (\(< 0.8000\) Non-Human / Foreign)
Low RPEMHC (\(\le 0.3500\)) Ideal Clinical Profile (e.g., Trastuzumab, Bevacizumab)
β€’ Fully tolerized by human immune system
β€’ Low Class II presentation density
β€’ Lowest statistical risk of clinical ADA
De-Immunized Non-Human / Chimeric
β€’ Sequence is foreign/non-human
β€’ Lacks physical MHC-II groove anchors
β€’ Minimal presentation to helper T-cells
High RPEMHC (\(> 0.4500\)) Self-Presenter (e.g., Renvistobart)
β€’ High biophysical groove affinity
β€’ Composed of natural human repertoire 9-mers
β€’ Tolerized by central/peripheral immune mechanisms
Clinical Danger Zone (e.g., Bococizumab)
β€’ Strong MHC-II peptide presentation (\(0.4630\), Red)
β€’ Contains non-human / foreign CDR neo-epitopes
β€’ High risk of helper T-cell priming & neutralizing ADAs

Clinical Case Study: Bococizumab (The Danger of Global Score Averaging)

Bococizumab (Pfizer’s humanized anti-PCSK9 mAb) provides a classic clinical example of why granular epitope evaluation is essential:

  • Clinical Outcome: In late-stage Phase 3 trials (the SPIRE program, encompassing over 27,000 patients), Pfizer was forced to discontinue development after discovering that 48% of patients developed anti-drug antibodies (ADAs), with 29% developing high-titer neutralizing ADAs that substantially eroded LDL-C lowering efficacy over time and caused elevated injection-site reactions.
  • First-Principles Immunogenic Profile: Under AbLead's population-weighted probability model, Bococizumab is unequivocally flagged as a high-risk candidate:
    • OASign Fv: \(0.8134\) (standard humanized range; yellow)
    • RPEMHC VL: \(0.2540\) (warning / moderate; yellow)
    • RPEMHC VH: \(0.2360\) (warning / moderate; yellow)
    • RPEMHC Fv: \(0.4630\) (severe / high risk; red)
  • The Immunological Mechanism: The human immune system does not evaluate global sequence averages. Bococizumab’s CDR loops (derived from a murine precursor) contained non-human somatic mutations that formed high-affinity 15-mer peptide anchors (\(\ge 0.800\)) presented across multiple core HLA-DRB1 alleles. Because these foreign CDR peptides had never undergone thymic negative selection in humans, dendritic cells presented them to naive CD4+ T cells, priming T follicular helper (\(T_{\text{FH}}\)) cells and driving robust B-cell class switching to neutralizing anti-bococizumab IgG antibodies.
  • Key Engineering Takeaway: Global Fv scores provide high-throughput triage, but lead candidates must always be evaluated at the residue and sliding-window level using the RPEMHC 2D Matrix Heatmap, Linear Sequence Track, and Engineering Modal to detect and eliminate isolated CDR helper T-cell hotspots prior to clinical commitment.

High-Risk MHC-II Epitopes (Hotspots Card)

To prevent localized immunogenic liabilities from being masked by whole-chain averages, the Humanness workspace in RPEMHC Mode includes a dedicated High-Risk MHC-II Epitopes (Hotspots) summary table:

  • Detection Criteria: Automatically extracts 15-mer peptide windows exhibiting high groove binding affinity (\(\ge 0.638\)) or broad population promiscuity (\(\ge 4\) presenting HLA Class II alleles).
  • Human Germline Context & Self-Tolerance: Distinguishes 100% human germline peptides (which are naturally tolerized by central immune tolerance) from true neo-epitope liabilities containing somatic hypermutations (SHM), point mutations, or CDR3 junctions.
  • Germline Status Badges & Tooltips: Displays clear visual origin badges (🟒 100% Germline, 🟑 1 Mut: Pos, 🟣 CDR3 Junction, πŸ”΄ Non-Germline). Hovering over any badge or peptide provides an interactive tooltip showing the query sequence vs. matched human germline sequence, assigned V/J gene family, and exact amino acid mutations.
  • 1-Click Noise-Reduction Filter: Allows users to filter the table instantly between All Hotspots, Non-Germline Only (focusing de-immunization engineering purely on actionable liabilities), and Germline Only.
  • Clustered Peak Reporting: Groups contiguous overlapping 15-mers into distinct physical epitope clusters, reporting the peak 15-mer peptide sequence, top presenting HLA allotype, peak binding score (color-coded by risk severity), and total population breadth (out of 20 alleles).
  • Interactive Triage: Clicking any hotspot row immediately scrolls both the sequence alignment and residue table to that exact position and applies an outline highlight to facilitate immediate de-immunization engineering.

Positional Max Risk Evaluation

When viewing sequence alignments and residue diagnostics under Max Risk β€” Tier 1 Global Panel or Max Risk β€” All 20 Alleles:

  • Positional Maximum Evaluation: Rather than fixing the entire sequence to the single allele that had the highest overall chain total sum, each individual amino acid position evaluates the true maximum binding score across all panel alleles at that specific position.
  • Allele Attribution Tooltips: Hovering over any residue box in the alignment or the MHC-II score cell in the table displays the exact HLA Class II allotype driving the peak risk (e.g. MHC-II Score (Max Tier 1: DRB1*04:05): 1.0030), immediately revealing localized neo-epitopes.

Methodology

The AbLead Humanness analysis suite combines repertoire database lookups, germline positional frequencies, masked language models, and deep neural network Class II MHC binding predictors to create an integrated developability and immunogenicity assessment.

1. OASign: Repertoire-Based Sequence Humanness

OASign measures sequence humanness by comparing antibody peptide fragments directly against the Observed Antibody Space (OAS), a comprehensive repository of over one billion natural human antibody sequences derived from diverse healthy donors.

  • Fv Domain Trimming: Sequences are automatically identified and trimmed to the Variable Domain (\(\text{Fv}\)) using AntPack annotation. Constant domains are excluded from repertoire scoring to prevent false inflation from invariant constant sequences.
  • Sliding 9-mer Peptide Decomposition: The variable region sequence of length \(L\) is chopped into \(L - 8\) overlapping 9-mer peptide windows. A 9-mer window size is biologically calibrated to match the typical footprint of linear B-cell epitopes and linear sequence identity segments.
  • Hash-Based Repertoire Querying: Each 9-mer is queried against a high-performance hash index of human antibody repertoires. A 9-mer is scored as Human (1.0) if it appears in the curated human OAS database with sufficient donor support, and Non-Human (0.0) otherwise.
  • Mutational Distance to Human: For any non-human 9-mer, OASign computes the minimum single amino acid substitution required to convert the peptide into a known human repertoire 9-mer, identifying minimal-mutation humanization pathways.
  • Chain & Fv Humanness Scores: The percentage of human 9-mers in the variable domain:

    \[\text{Score}_{\text{Chain}} = \frac{\sum_{i=1}^{N_{\text{peptides}}} I(\text{peptide}_i \in \text{OAS})}{N_{\text{peptides}}} \times 100\%\]
    \[\text{Score}_{\text{Fv}} = \frac{N_{\text{pep, VL}} \cdot \text{Score}_{\text{VL}} + N_{\text{pep, VH}} \cdot \text{Score}_{\text{VH}}}{N_{\text{pep, VL}} + N_{\text{pep, VH}}}\]

2. Positional Frequency Analysis (PFA)

Positional Frequency Analysis evaluates how standard or unusual an amino acid is at a specific topological position compared to natural human antibody germlines.

  • IMGT Alignment: Sequences are aligned against IMGT human germline reference databases for heavy (IGHV, IGHJ) and light (IGKV, IGLV, IGKJ, IGLJ) chains.
  • Positional Matrix Lookups: For each aligned IMGT position, the system looks up the frequency distribution of all 20 amino acids across human germlines belonging to the matching V and J gene families.
  • Frequency Output: Reports the percentage of human germline antibodies carrying the query residue alongside the dominant human amino acid at that position (e.g. 0.5% [90.7% T]). Positions with low PFA scores (\(<70\%\)) represent somatic hypermutations, non-canonical framework substitutions, or murine backmutations that may destabilize the antibody or act as foreign epitopes.

3. AbLang & AbLang-2 Language Model Likelihood

AbLang and AbLang-2 are transformer-based masked language models trained on hundreds of millions of antibody sequences to learn the structural and evolutionary rules of antibody folding.

  • Contextual Residue Likelihood: Given the full variable domain context, the language model predicts the probability distribution of amino acids at each individual position.
  • Difference Metric (\(\Delta \text{Likelihood}\)): Evaluates the drop in log-likelihood from the model-optimal residue:

    \[\Delta \text{Likelihood} = \text{Score}_{\text{optimal}} - \text{Score}_{\text{actual}}\]

    A low difference (\(\Delta < 3.0\)) indicates that the residue is favored by evolutionary antibody grammar, whereas a high difference (\(\Delta \ge 5.0\)) flags positions that diverge significantly from natural antibody profiles and may introduce conformational instability or poor expression.

4. RPEMHC: Relative Position Encoding for MHC-II Binding

RPEMHC predicts the physical chemical presentation of 15-mer antibody peptides to human Major Histocompatibility Complex Class II (MHC-II) molecules.

  • 15-mer Sliding Windows: Antibody sequences are scanned using a sliding window of 15 amino acids (\(L - 14\) windows per chain), capturing the full peptide length that extends beyond the open-ended MHC-II binding groove.
  • Relative Position Encoding: Unlike standard sequence encoding, relative position encoding explicitly represents the spatial distance between amino acid pairs and the polymorphic MHC-II binding pockets (P1, P4, P6, P7, P9), capturing cooperative anchoring interactions.
  • 20-Allele Panel Evaluation: Evaluates binding affinity across both Tier 1 Core DRB1 alleles (>95% global population coverage) and Tier 2 Extended alleles (DRB3/4/5, DQ, DP).
  • Thresholding & Cumulative Epitope Load: Peptides scoring \(> 0.50\) are flagged as presenting helper T-cell epitopes. The per-allele and aggregate construct scores (\(\sum \text{Scores} > 0.50\)) quantify total immunogenic burden to guide de-immunization engineering.

Understanding HLA Class II Alleles & Nomenclature

1. WHO / IMGT-HLA Nomenclature (e.g., HLA-DRB1*04:01)

Allele designations follow the standard World Health Organization (WHO) and IMGT-HLA nomenclature:

  • Gene Locus: Prefix (e.g., HLA-DRB1*, HLA-DQA1*, HLA-DPA1*) designates the Class II \(\alpha\) or \(\beta\) chain gene.
  • Field 1 (e.g., 04): The broad serological antigen group / allele family (e.g., DR4).
  • Field 2 (e.g., :01, :05): The specific protein allotype / subtype. Specific amino acid polymorphisms between subtypes alter the chemical topology and charge distribution of peptide-binding pockets (P1, P4, P6, P7, P9), resulting in distinct epitope repertoires.

2. Tier 1 vs. Tier 2 Allele Panels

  • Tier 1 (Core Global Reference Panel): 11 high-prevalence HLA-DRB1 alleles (01:01, 03:01, 04:01, 04:05, 07:01, 08:02, 09:01, 11:01, 12:01, 13:02, 15:01) representing over 95% cumulative coverage across diverse global human populations.
  • Tier 2 (Extended Panel): 9 additional high-impact alleles across secondary DRB3/4/5, DQA1/DQB1, and DPA1/DPB1 loci for complete HLA Class II immunogenicity risk profiling.

3. HLA Class II Reference Panel

Allele (WHO / IMGT) Panel / Tier Estimated Population Frequency Binding Pocket Motif Preferences Clinical / Therapeutic Risk Associations
HLA-DRB1*01:01 Tier 1 Core ~10–15% Global (EUR ~15–20%, ASN ~5–10%) P1: Large hydrophobic (F, Y, W, L, I); P4/P6/P9: Aliphatic Broad immunogenic peptide presentation; common ADA epitope presenter
HLA-DRB1*03:01 Tier 1 Core ~10–15% Global (EUR ~12–18%, AFR ~10–15%) P1: Hydrophobic (L, I, M, F); P4: Acidic (D, E); P6: Basic/polar Associated with SLE, Type 1 diabetes, autoimmune thyroiditis, biological drug ADA
HLA-DRB1*04:01 Tier 1 Core ~8–14% Global (EUR ~15–20%, AMR ~10–15%) P1: Aromatic/aliphatic (F, Y, W, L); P4: Basic/polar; P9: Hydrophobic Shared epitope for Rheumatoid Arthritis; high clinical ADA risk in TNFα biologics
HLA-DRB1*04:05 Tier 1 Core ~5–10% Global (East Asian ~15–25%, EUR ~2–4%) P1: Aromatic/aliphatic; P4: Acidic/polar (pocket 4 subtype difference) Key Asian-predominant RA & therapeutic immunogenicity risk allele
HLA-DRB1*07:01 Tier 1 Core ~15–22% Global (EUR ~25%, Middle Eastern ~25%) P1: Large aromatic (F, Y, W); P4: Large hydrophobic; P9: Aliphatic High frequency globally; co-expressed with DRB4*01:01; frequent mAb ADA epitope
HLA-DRB1*08:02 Tier 1 Core ~3–8% Global (Hispanic/Amerindian ~10–20%, ASN ~5–10%) P1: Aliphatic/aromatic; P4: Acidic/polar; P6: Polar; P9: Hydrophobic Important representation for Hispanic/Latin American and Indigenous cohorts
HLA-DRB1*09:01 Tier 1 Core ~5–10% Global (East Asian ~15–20%, EUR ~2–5%) P1: Large hydrophobic; P4: Polar/aliphatic; P9: Small neutral Prominent in East Asian populations; associated with autoimmune hepatitis & ADA
HLA-DRB1*11:01 Tier 1 Core ~10–18% Global (Mediterranean ~20–25%, AFR ~10–15%) P1: Aliphatic/aromatic; P4: Small/polar; P6: Basic/polar; P9: Aliphatic High Mediterranean and Hispanic frequency; co-expressed with DRB3*02:02
HLA-DRB1*12:01 Tier 1 Core ~3–8% Global (East Asian ~8–15%, AFR ~5–10%) P1: Hydrophobic; P4: Polar/basic; P9: Aliphatic Asian and African cohort diversity representation
HLA-DRB1*13:02 Tier 1 Core ~6–12% Global (AFR ~8–15%, EUR ~6–10%) P1: Hydrophobic; P4: Acidic/polar; P6: Polar; P9: Aliphatic Broad peptide binding capacity; protective in certain viral/autoimmune contexts
HLA-DRB1*15:01 Tier 1 Core ~12–20% Global (EUR ~20%, ASN ~15–20%) P1: Hydrophobic/aromatic; P4: Aliphatic; P7: Polar; P9: Small neutral Strong association with Multiple Sclerosis; major driver of therapeutic protein ADA
HLA-DRB3*01:01 Tier 2 (DRB3) ~15–30% (Haplotype-linked with DR3, DR11, DR13, DR14) P1: Hydrophobic; P4: Basic; P6/P9: Aliphatic Co-expressed secondary β-chain; contributes additive Class II presentation
HLA-DRB3*02:02 Tier 2 (DRB3) ~20–35% (Co-expressed with DR11, DR13, DR14) P1: Aliphatic/aromatic; P4: Polar; P9: Aliphatic High expression in individuals carrying DRB1*11:01 / DRB1*13:01 haplotypes
HLA-DRB4*01:01 Tier 2 (DRB4) ~30–45% (Haplotype-linked with DR4, DR7, DR9) P1: Hydrophobic; P4: Neutral/aliphatic; P7/P9: Polar Highly prevalent secondary DRB molecule in DR4/DR7 carriers; prominent ADA mediator
HLA-DRB5*01:01 Tier 2 (DRB5) ~15–25% (Haplotype-linked with DR15, DR16) P1: Large aromatic; P4: Polar/aliphatic; P9: Hydrophobic Co-expressed with DRB1*15:01; implicated in MS and myelin peptide presentation
HLA-DQA1*05:01 / DQB1*02:01 Tier 2 (DQ2.5) ~10–20% Global (EUR ~20%, Hispanic ~15%) P1/P9: Hydrophobic; P4/P6/P7: Negative charge (Glu, Asp, deamidated Gln) Canonical DQ2.5 heterodimer; strong celiac disease & deamidated therapeutic epitope risk
HLA-DQA1*03:01 / DQB1*03:02 Tier 2 (DQ8) ~10–18% Global (AMR ~20–30%, EUR ~15%) P1/P9: Hydrophobic/acidic; P4: Negative/polar Canonical DQ8 heterodimer; associated with autoimmune diabetes and specific ADA
HLA-DQA1*01:02 / DQB1*06:02 Tier 2 (DQ6.2) ~15–25% Global (EUR ~20%, ASN ~15%) P1: Aliphatic; P4: Hydrophobic; P9: Aliphatic Strongest known genetic association with narcolepsy; high population prevalence
HLA-DPA1*01:03 / DPB1*04:01 Tier 2 (DP401) ~35–50% Global (Highest frequency DP allele globally) P1: Large aromatic/hydrophobic; P6: Acidic/polar; P9: Hydrophobic Predominant DP allele globally (>40% of humans); broad peptide presentation capacity
HLA-DPA1*01:03 / DPB1*04:02 Tier 2 (DP402) ~15–25% Global P1: Aromatic/hydrophobic; P6: Polar; P9: Hydrophobic Secondary major DP allele; essential for complete HLA Class II risk coverage

Constant Domains: Biophysical Binding Affinity vs. Immunological Self-Tolerance

When evaluating full-length constructs containing constant regions (\(\text{CH}_1\), Hinge, \(\text{CH}_2\), \(\text{CH}_3\), \(\text{C}_\kappa\), \(\text{C}_\lambda\)), users may note high predicted MHC-II binding scores at certain constant domain positions:

  • Biophysical Groove Affinity: Neural network MHC-II predictors (including RPEMHC and NetMHCIIpan) evaluate the physical chemical affinity of 15-mer peptide windows into polymorphic HLA Class II binding pockets (P1, P4, P6, P9). Native human constant sequences naturally contain hydrophobic or aromatic core residues (e.g. Leu, Val, Phe, Tyr) that satisfy MHC-II pocket geometry, resulting in high predicted binding affinity.

  • Central & Peripheral Immune Tolerance: In humans, the immune system is centrally tolerized to native germline constant sequences during thymic development. Negative selection in the thymus eliminates autoreactive CD4+ T-cell clones with high affinity for self-peptides. Consequently, physical MHC-II groove binding in native human constant regions does not translate into clinical helper T-cell activation or anti-drug antibody (ADA) responses.

  • Variable Domain (Fv) Focus: Clinical immunogenicity risk and de-immunization engineering are focused on the Variable Domain (\(\text{Fv}\))β€”where somatic hypermutations, non-human CDR grafts, framework germline deviations, and V/J junctional sequences create foreign neo-epitopes that can trigger ADA responses in patients.


References

  • The closest germlines are from the IMGT database. See IMGT.
  • Humanness is determined using OASign against curated human 9-mer sequences from the Observed Antibody Space (OAS).
  • RPEMHC (MHC-II binding prediction) is based on the DeepMHCII / RPEMHC ensemble deep learning model for predicting class II peptide-MHC binding affinity. Described in Wang et al., Bioinformatics (2024).
  • OASign is based on the OASis method in BioPhi. See David Prihoda, Jad Maamary, Andrew Waight, Veronica Juan, Laurence Fayadat-Dilman, Daniel Svozil & Danny A. Bitton (2022) BioPhi: A platform for antibody design, humanization, and humanness evaluation based on natural antibody repertoires and deep learning, mAbs, 14:1, DOI: https://doi.org/10.1080/19420862.2021.2020203.