RPEMHC Immunogenicity Analysis Tool
The RPEMHC Immunogenicity Tool provides specialized MHC-II binding allele risk analysis for antibody and protein constructs, leveraging the RPEMHC deep learning model for 15-mer peptide sliding window evaluation.
Accessing the Tool
Select an antibody in the Project View. Go to the Analysis menu and select RPEMHC. This will open the RPEMHC Analysis workspace.

Key Features
1. Multi-Chain & Multi-Record Support
- Construct Variety: Supports single-chain (VHH, scFv), paired Fv (Light / Heavy), and multi-chain multispecific constructs (such as 4-chain CrossMab antibodies:
CrossMab Light Chain 1,CrossMab Heavy Chain 1,Standard Light Chain 2,Standard Heavy Chain 2). - FASTA Parsing: Automatically parses multi-record FASTA headers (
>Header) stored in heavy or light chain fields.
2. Signal Peptide Trimming
- Leader Sequence Trimming: Includes a Trim Signal Peptides option (checked by default) using AntPack alignment to automatically identify and strip N-terminal leader signal sequences prior to running 15-mer RPEMHC window scans.
3. Expanded 20-Allele Global Panel Screening
- Alleles Covered: Evaluates binding scores across an expanded panel of 20 human HLA Class II alleles, spanning both Tier 1 (11 core HLA-DRB1 alleles) and Tier 2 (9 HLA-DRB3/4/5, HLA-DQA1/DQB1, and HLA-DPA1/DPB1 alleles).
- Dynamic Views: Dropdown controls allow filtering between:
- Max Risk — Tier 1 Global Panel (11 Core DRB1): Default global reference panel representing >95% cumulative global human population coverage.
- Max Risk — All 20 Alleles (Tiers 1 & 2): Full evaluation across all 20 HLA-DRB1, DRB3/4/5, DQ, and DP alleles.
- Specific Allotypes: Isolates epitope predictions to a single HLA Class II allele (e.g.
HLA-DRB1*04:01,HLA-DQA1*05:01-DQB1*02:01).
4. Construct Summary Cards
Top-level cards displaying total construct score (sum of chain scores where binding prediction \(> 0.5\)) and individual chain RPEMHC scores colored by standardized AbLead severity:
- Good / Low Risk (\(\le 28.0\)): Green (
#88B46C) - Warning / Moderate Risk (\(28.0 - 35.0\)): Orange (
#E29848) - Severe / High Risk (\(> 35.0\)): Red (
#CF4A3C)
5. 2D Sequence \(\times\) Alleles Matrix Heatmap & Publication Figures
- 2D Heatmap Matrix: Displays 15-mer sliding peptide windows along the vertical axis and all 20 HLA Class II alleles across columns, grouped by Tier 1 and Tier 2 with a dedicated Total Binders column.
- Sticky Hierarchy: Multi-tier sticky headers pin super-tier headers (
top: 0) and rotated allele column headers (top: 26px) while keeping sequence window labels pinned (left: 0) during horizontal and vertical scrolling. - Interactive Tooltips: Hover over matrix cells to inspect position ranges, IMGT coordinates, 15-mer peptide sequences, exact predicted binding scores, and the predicted 9-mer binding core.
- Publication Figure Export: Export high-resolution 300-DPI PNG and vector SVG figures for individual chains directly from the chain section title bars (
Save Chain (PNG)andSVG).
6. Linear Sequence Immunogenicity Track & Diagnostic Table
- Multi-View Switcher: Toggle between
All Views,2D Matrix Heatmap,Sequence Track, andDiagnostic Tableusing the toolbar selector. - Linear Sequence Strip Plot: Monospace sequence track with rotated linear and IMGT coordinate labels above severity-filled amino acid boxes.
- Linear Track Image Export: Export publication-grade wrapped strip plots (50 residues per block) with mature linear coordinates, IMGT annotations, severity styling, and 4-tier assessment legends via
Save Track (PNG)andSVG. - Residue Diagnostic Table: Granular tabular view displaying linear coordinates, IMGT positions, regions, amino acids with ABHAND coloring, active scores, max alleles, individual scores across all 20 alleles, and highest-risk 15-mer window descriptions. Includes a one-click
Export Table (Excel)button on each chain. - Interactive Navigation: Clicking any sequence residue box in the linear track smoothly scrolls to and highlights the corresponding row in the residue diagnostic table.
7. Multi-Chain Heatmapped Excel Export
- Excel Workbooks: Generates formatted, heatmapped Excel workbooks containing a top-level RPEMHC Summary sheet alongside individual residue diagnostic sheets and 15-mer sliding window matrix sheets for every chain in the construct, with openpyxl cell severity fills, ABHAND amino acid colors, and strict Light-before-Heavy chain ordering.
Scoring Metrics & Color Scales
The RPEMHC workspace uses three distinct scoring layers to evaluate immunogenicity at the residue, window, and construct levels.
1. Residue & Allele Binding Affinity (Biophysical Risk Score)
Evaluates the predicted physical chemical binding affinity of 15-mer peptide windows into HLA Class II binding pockets. Used for individual cells in the 2D Matrix Heatmap, residue boxes in the Linear Sequence Track, and rows in the Residue Diagnostic Tables:
| Binding Affinity Score | Assessment Category | UI / Export Color | Interpretation |
|---|---|---|---|
| \(< 0.426\) | Good / Baseline | White (#ffffff) |
Weak or no predicted binding; standard germline baseline. |
| \(0.426 – 0.638\) | Low Risk | Yellow (#F8D548) |
Moderate binding potential; low probability of clinical presentation. |
| \(0.638 – 0.800\) | Medium Risk | Orange (#E29848) |
Elevated binding affinity; candidate helper T-cell epitope. |
| \(\ge 0.800\) | High Risk | Red (#CF4A3C) |
Strong physical groove binder; high-risk T-cell activation trigger. |
2. 2D Matrix Total Binders Column (Population Promiscuity)
Counts how many of the 20 tested HLA Class II alleles bind a given 15-mer window with affinity score \(\ge 0.50\). This quantifies epitope promiscuity (how broadly the peptide is presented across the global human population):
| Total Presenting Alleles (\(\ge 0.50\)) | Promiscuity Category | UI / Export Color | Population Risk Context |
|---|---|---|---|
| \(0\) Alleles | Non-Binder | White (#ffffff) |
No significant presentation across global reference panel. |
| \(1 – 3\) Alleles | Low Promiscuity | Yellow (#F8D548) |
Restricted presentation; affects small subsets of patients. |
| \(4 – 7\) Alleles | Moderate Promiscuity | Orange (#E29848) |
Broad presentation across multiple HLA Class II allotypes. |
| \(\ge 8\) Alleles | Epitope Hotspot | Red (#CF4A3C) |
High promiscuity hotspot; broadly immunogenic across >80% of humans. |
3. Population-Weighted ADA Probability Scores (Summary Cards on Top)
Calculates the first-principles biophysical probability (\(0.0000 – 1.0000\)) that the antibody construct triggers helper T-cell activation and downstream anti-drug antibodies (ADA) in the human population, weighted by global HLA Class II allele population frequencies (\(f_a\)). Formatted consistently as a 4-decimal float across all Humanness metrics:
| RPEMHC Metric | Risk Level | Card Color | Developability Guidance |
|---|---|---|---|
RPEMHC VL \(\le 0.2000\) |
Good / Low Risk | Green (#88B46C) |
Low light-chain helper T-cell activation probability in the human population. |
RPEMHC VL \(0.2000 – 0.3000\) |
Warning / Moderate | Yellow (#F8D548) |
Moderate light-chain epitope presentation probability. |
RPEMHC VL \(> 0.3000\) |
Severe / High Risk | Red (#CF4A3C) |
High probability of light-chain helper T-cell priming across human HLA allotypes. |
RPEMHC VH \(\le 0.2000\) |
Good / Low Risk | Green (#88B46C) |
Low heavy-chain helper T-cell activation probability in the human population. |
RPEMHC VH \(0.2000 – 0.3000\) |
Warning / Moderate | Yellow (#F8D548) |
Moderate heavy-chain epitope presentation probability. |
RPEMHC VH \(> 0.3000\) |
Severe / High Risk | Red (#CF4A3C) |
High probability of heavy-chain helper T-cell priming across human HLA allotypes. |
RPEMHC Fv \(\le 0.3500\) |
Good / Low Risk | Green (#88B46C) |
Clinically favorable variable domain construct probability. |
RPEMHC Fv \(0.3500 – 0.4500\) |
Warning / Moderate | Yellow (#F8D548) |
Moderate whole-Fv helper T-cell activation risk. |
RPEMHC Fv \(> 0.4500\) |
Severe / High Risk | Red (#CF4A3C) |
Elevated whole-Fv ADA risk (flags Bococizumab at 0.4630 in RED). |
4. Scientific & Biophysical Basis for RPEMHC Probability Model
The RPEMHC score is derived directly from first-principles biophysical pocket binding kinetics and human population genetics:
- Pocket Binding Activation Function: For each 15-mer window \(i\) and HLA allele \(a\), raw neural network affinity \(s_{i,a}\) is converted to an activation probability \(P_{\text{bind}} = \min(1.0, ((s_{i,a} - s_0)/(1.0 - s_0))^{\gamma})\) above the biophysical IC50 threshold (\(s_0 = 0.638\)).
- Discrete Physical Cluster Pooling: Overlapping 15-mers within \(\le 6\) residues are merged into distinct structural epitope clusters \(C\).
- Allele Presentation Probability: The probability that an individual carrying HLA allotype \(a\) mounts an immune response is \(P(\text{Trigger} \mid a) = 1 - \prod_{c \in C} (1 - P_{\text{bind}}(s_{c,a}))\).
-
Global Population Frequency Weighting: Integrated across the global human population using standardized DRB1 frequencies \(f_a\):
\[\text{RPEMHC Probability} = \sum_{a} f_a \cdot P(\text{Trigger} \mid a)\]
5. High-Risk MHC-II Epitopes (Hotspots Card)
The standalone RPEMHC workspace features a dedicated High-Risk MHC-II Epitopes (Hotspots) table displayed prominently beneath the aggregate summary score cards:
- Automated Clustering: Identifies and clusters 15-mer sliding windows with high groove binding affinity (\(\ge 0.638\)) or broad population presentation (\(\ge 4\) presenting alleles).
- Human Germline Context & Self-Tolerance: Distinguishes 100% human germline peptides (which are naturally tolerized by central immune tolerance) from true neo-epitope liabilities containing somatic hypermutations (SHM), point mutations, or CDR3 junctions.
- Germline Status Badges & Tooltips: Displays clear visual origin badges (🟢
100% Germline, 🟡1 Mut: Pos, 🟣CDR3 Junction, 🔴Non-Germline). Hovering over any badge or peptide provides an interactive tooltip showing the query sequence vs. matched human germline sequence, assigned V/J gene family, and exact amino acid mutations. - 1-Click Noise-Reduction Filter: Allows users to filter the table instantly between All Hotspots, Non-Germline Only (focusing de-immunization engineering purely on actionable liabilities), and Germline Only.
- Peak Representation: Reports the peak 15-mer peptide sequence, top presenting HLA allotype, peak binding score badge, region, and population breadth out of 20 alleles.
-
Interactive Jumping: Clicking any row immediately scrolls both the chain's Linear Sequence Track and Residue Diagnostic Table to that position, applying a temporary outline highlight to speed up de-immunization workflows.
-
Population Promiscuity / Total Binders (\(0 – 20\) Alleles):
- \(0\) Alleles (White): \(0\%\) human population presentation risk.
- \(1 – 3\) Alleles (Yellow): Restricted to specific rare HLA haplotypes (\(\sim 10\%–25\%\) global population).
- \(4 – 7\) Alleles (Orange): Broad presentation across multiple HLA lineages (\(\sim 30\%–65\%\) global population).
- \(\ge 8\) Alleles (Red): Promiscuous helper T-cell hotspot presented across \(> 80\%\) of diverse global human populations.
Understanding HLA Class II Alleles & Nomenclature
1. WHO / IMGT-HLA Nomenclature (e.g., HLA-DRB1*04:01)
Allele designations follow standard World Health Organization (WHO) and IMGT-HLA nomenclature:
- Gene Locus: Prefix (e.g.,
HLA-DRB1*,HLA-DQA1*,HLA-DPA1*) designates the Class II \(\alpha\) or \(\beta\) chain gene. - Field 1 (e.g.,
04): The broad serological antigen group / allele family (e.g., DR4). - Field 2 (e.g.,
:01,:05): The specific protein allotype / subtype. Specific amino acid polymorphisms between subtypes alter the chemical topology and charge distribution of peptide-binding pockets (P1, P4, P6, P7, P9), resulting in distinct epitope repertoires.
2. Tier 1 vs. Tier 2 Allele Panels
- Tier 1 (Core Global Reference Panel): 11 high-prevalence
HLA-DRB1alleles (01:01,03:01,04:01,04:05,07:01,08:02,09:01,11:01,12:01,13:02,15:01) representing over 95% cumulative coverage across diverse global human populations. - Tier 2 (Extended Panel): 9 additional high-impact alleles across secondary
DRB3/4/5,DQA1/DQB1, andDPA1/DPB1loci for complete HLA Class II immunogenicity risk profiling.
3. HLA Class II Reference Panel
| Allele (WHO / IMGT) | Panel / Tier | Estimated Population Frequency | Binding Pocket Motif Preferences | Clinical / Therapeutic Risk Associations |
|---|---|---|---|---|
| HLA-DRB1*01:01 | Tier 1 Core | ~10–15% Global (EUR ~15–20%, ASN ~5–10%) | P1: Large hydrophobic (F, Y, W, L, I); P4/P6/P9: Aliphatic | Broad immunogenic peptide presentation; common ADA epitope presenter |
| HLA-DRB1*03:01 | Tier 1 Core | ~10–15% Global (EUR ~12–18%, AFR ~10–15%) | P1: Hydrophobic (L, I, M, F); P4: Acidic (D, E); P6: Basic/polar | Associated with SLE, Type 1 diabetes, autoimmune thyroiditis, biological drug ADA |
| HLA-DRB1*04:01 | Tier 1 Core | ~8–14% Global (EUR ~15–20%, AMR ~10–15%) | P1: Aromatic/aliphatic (F, Y, W, L); P4: Basic/polar; P9: Hydrophobic | Shared epitope for Rheumatoid Arthritis; high clinical ADA risk in TNFα biologics |
| HLA-DRB1*04:05 | Tier 1 Core | ~5–10% Global (East Asian ~15–25%, EUR ~2–4%) | P1: Aromatic/aliphatic; P4: Acidic/polar (pocket 4 subtype difference) | Key Asian-predominant RA & therapeutic immunogenicity risk allele |
| HLA-DRB1*07:01 | Tier 1 Core | ~15–22% Global (EUR ~25%, Middle Eastern ~25%) | P1: Large aromatic (F, Y, W); P4: Large hydrophobic; P9: Aliphatic | High frequency globally; co-expressed with DRB4*01:01; frequent mAb ADA epitope |
| HLA-DRB1*08:02 | Tier 1 Core | ~3–8% Global (Hispanic/Amerindian ~10–20%, ASN ~5–10%) | P1: Aliphatic/aromatic; P4: Acidic/polar; P6: Polar; P9: Hydrophobic | Important representation for Hispanic/Latin American and Indigenous cohorts |
| HLA-DRB1*09:01 | Tier 1 Core | ~5–10% Global (East Asian ~15–20%, EUR ~2–5%) | P1: Large hydrophobic; P4: Polar/aliphatic; P9: Small neutral | Prominent in East Asian populations; associated with autoimmune hepatitis & ADA |
| HLA-DRB1*11:01 | Tier 1 Core | ~10–18% Global (Mediterranean ~20–25%, AFR ~10–15%) | P1: Aliphatic/aromatic; P4: Small/polar; P6: Basic/polar; P9: Aliphatic | High Mediterranean and Hispanic frequency; co-expressed with DRB3*02:02 |
| HLA-DRB1*12:01 | Tier 1 Core | ~3–8% Global (East Asian ~8–15%, AFR ~5–10%) | P1: Hydrophobic; P4: Polar/basic; P9: Aliphatic | Asian and African cohort diversity representation |
| HLA-DRB1*13:02 | Tier 1 Core | ~6–12% Global (AFR ~8–15%, EUR ~6–10%) | P1: Hydrophobic; P4: Acidic/polar; P6: Polar; P9: Aliphatic | Broad peptide binding capacity; protective in certain viral/autoimmune contexts |
| HLA-DRB1*15:01 | Tier 1 Core | ~12–20% Global (EUR ~20%, ASN ~15–20%) | P1: Hydrophobic/aromatic; P4: Aliphatic; P7: Polar; P9: Small neutral | Strong association with Multiple Sclerosis; major driver of therapeutic protein ADA |
| HLA-DRB3*01:01 | Tier 2 (DRB3) | ~15–30% (Haplotype-linked with DR3, DR11, DR13, DR14) | P1: Hydrophobic; P4: Basic; P6/P9: Aliphatic | Co-expressed secondary β-chain; contributes additive Class II presentation |
| HLA-DRB3*02:02 | Tier 2 (DRB3) | ~20–35% (Co-expressed with DR11, DR13, DR14) | P1: Aliphatic/aromatic; P4: Polar; P9: Aliphatic | High expression in individuals carrying DRB1*11:01 / DRB1*13:01 haplotypes |
| HLA-DRB4*01:01 | Tier 2 (DRB4) | ~30–45% (Haplotype-linked with DR4, DR7, DR9) | P1: Hydrophobic; P4: Neutral/aliphatic; P7/P9: Polar | Highly prevalent secondary DRB molecule in DR4/DR7 carriers; prominent ADA mediator |
| HLA-DRB5*01:01 | Tier 2 (DRB5) | ~15–25% (Haplotype-linked with DR15, DR16) | P1: Large aromatic; P4: Polar/aliphatic; P9: Hydrophobic | Co-expressed with DRB1*15:01; implicated in MS and myelin peptide presentation |
| HLA-DQA1*05:01 / DQB1*02:01 | Tier 2 (DQ2.5) | ~10–20% Global (EUR ~20%, Hispanic ~15%) | P1/P9: Hydrophobic; P4/P6/P7: Negative charge (Glu, Asp, deamidated Gln) | Canonical DQ2.5 heterodimer; strong celiac disease & deamidated therapeutic epitope risk |
| HLA-DQA1*03:01 / DQB1*03:02 | Tier 2 (DQ8) | ~10–18% Global (AMR ~20–30%, EUR ~15%) | P1/P9: Hydrophobic/acidic; P4: Negative/polar | Canonical DQ8 heterodimer; associated with autoimmune diabetes and specific ADA |
| HLA-DQA1*01:02 / DQB1*06:02 | Tier 2 (DQ6.2) | ~15–25% Global (EUR ~20%, ASN ~15%) | P1: Aliphatic; P4: Hydrophobic; P9: Aliphatic | Strongest known genetic association with narcolepsy; high population prevalence |
| HLA-DPA1*01:03 / DPB1*04:01 | Tier 2 (DP401) | ~35–50% Global (Highest frequency DP allele globally) | P1: Large aromatic/hydrophobic; P6: Acidic/polar; P9: Hydrophobic | Predominant DP allele globally (>40% of humans); broad peptide presentation capacity |
| HLA-DPA1*01:03 / DPB1*04:02 | Tier 2 (DP402) | ~15–25% Global | P1: Aromatic/hydrophobic; P6: Polar; P9: Hydrophobic | Secondary major DP allele; essential for complete HLA Class II risk coverage |
Constant Domains: Biophysical Binding Affinity vs. Immunological Self-Tolerance
When evaluating full-length antibody constructs (containing \(\text{CH}_1\), Hinge, \(\text{CH}_2\), \(\text{CH}_3\), or \(\text{C}_\kappa\)/\(\text{C}_\lambda\) domains), users may observe regions in the constant domains with high predicted MHC-II binding scores:
-
Biophysical Groove Affinity: Neural network MHC-II predictors (including RPEMHC and NetMHCIIpan) evaluate the physical chemical affinity of 15-mer peptide windows into the polymorphic binding pockets (P1, P4, P6, P9) of HLA Class II heterodimers. Many native human constant sequences naturally contain hydrophobic or aromatic core residues (e.g. Leu, Val, Phe, Tyr) that satisfy MHC-II pocket geometry, resulting in high predicted binding scores.
-
Central & Peripheral Immune Tolerance: In humans, the immune system is centrally tolerized to native germline constant sequences during thymic development. Negative selection in the thymus eliminates autoreactive CD4+ T-cell clones with high affinity for self-peptides. Consequently, physical MHC-II groove binding in native human constant regions does not translate into clinical helper T-cell activation or anti-drug antibody (ADA) responses.
-
Variable Domain (Fv) Focus: Clinical immunogenicity risk and de-immunization engineering are focused on the Variable Domain (\(\text{Fv}\))—where somatic hypermutations, non-human CDR grafts, framework germline deviations, and V/J junctional sequences create foreign neo-epitopes that can trigger ADA responses in patients.
References
- RPEMHC Model & Publication: Wang et al. (2024), RPEMHC: a relative position encoding-based deep learning framework for MHC-peptide binding prediction, Bioinformatics, Volume 40, Issue 1, btad785. DOI: https://doi.org/10.1093/bioinformatics/btad785.
- RPEMHC Source Code: Official GitHub repository available at https://github.com/lennylv/RPEMHC.