From an MS GWAS Locus to the Gene It Acts Through
A genome-wide association study hands you a locus, not a gene. Most of the more than 200 variants on the multiple sclerosis (MS) susceptibility map are non-coding and regulatory, and deciding which gene each one acts through is a separate inference with its own failure modes. Upcott and colleagues ran that inference at scale, pairing the discovery phase of the largest MS susceptibility GWAS, 14,802 patients against 26,703 unaffected individuals, with publicly available expression quantitative trait loci data from 7,466 blood donors and 5,494 brain donors. The output is a list of 240 genes whose expression the risk alleles appear to change. The scores they later built from that list are a separate question; this post is about how the list was made and what survived the filters.
What Summary Data-Based Mendelian Randomization Tests
The variant-to-gene step ran on summary data-based Mendelian randomization, which uses a GWAS variant as an instrument to ask whether its effect on disease runs through its effect on a gene's expression. The appeal is that it needs only summary statistics from each side, the variant-to-expression effect and the variant-to-disease effect, rather than individual-level data linking both in the same people. The thresholds are specific: eQTLs were selected at p < 5×10⁻⁸ with an R² between 0.05 and 0.90 against the top GWAS SNP at each locus, and gene expression differences associated with MS risk alleles were called at a false discovery rate-corrected p below 0.05. The R² window matters more than it looks, because it excludes both variants too weakly linked to the GWAS signal to be informative and variants so tightly linked that they carry no independent information.
The Problem That HEIDI Exists to Rule Out
A significant result from the first step has two possible readings, and they lead to opposite conclusions. Either the variant affects disease through the gene's expression, which is the finding the method is designed to produce, or two distinct causal variants happen to sit near each other in linkage disequilibrium, one driving expression and the other driving disease, which produces the same statistical signature and means nothing mechanistically. The Heterogeneity In Dependent Instruments test separates them by checking whether the estimated effect stays consistent across the multiple variants at a locus: under true mediation it should, while two separate causal variants produce heterogeneity. Anything with a HEIDI p at or below 0.1 was excluded. The threshold is deliberately permissive in the direction of caution, since here a low p-value is evidence against the gene rather than for it.
Colocalization, Run at Cell-Type Resolution
Surviving variants went to Bayesian colocalization, and the choice of reference data is what makes this step do double duty. Rather than bulk tissue, the authors used single-cell atlases: 1.27 million peripheral blood mononuclear cells from 927 donors covering B lymphocytes, CD4+ and CD8+ T cells, dendritic cells, monocytes, natural killer cells and plasma cells, and 750,614 single-nucleus central nervous system cells from 192 individuals covering astrocytes, endothelial cells, excitatory and inhibitory neurons, microglia, oligodendrocytes, oligodendrocyte precursor cells and pericytes. A posterior probability above 0.80 was taken as evidence that a variant is associated with an eQTL. Beyond adding a further test that one variant drives both signals, this assigns each gene to the cell or tissue type where its eQTL operates, which is what allowed the genes to be classified as immune, CNS or both, and what makes a cell-type-specific score possible at all.
Where the Pipeline Loses Variants
One practical detail is worth stating because papers often leave it in a supplement. SMR and HEIDI identified 387 variants. Carrying those into the Welsh scoring cohort required each one to be present in that cohort's genotype data, which came from Illumina Infinium CoreExome-24 v2 or v3 arrays with imputation. Where a variant was absent, the authors looked for a proxy in perfect linkage disequilibrium, at R² = 1 in the CEU population. Five such proxies were used. For 47 of the 387 variants, no proxy could be found, so roughly one in eight of the discovered instruments never made it into the score that was eventually tested. The discovery set and the scoring set are therefore not identical, and the gap runs in the direction of a weaker score rather than an inflated one.
240 Genes, Most of Them New to MS
SMR and HEIDI filtering returned 240 significant eGenes associated with MS susceptibility. Of these, 43 overlapped between immune and CNS tissues, leaving 88 CNS-specific and 109 immune-specific genes. Most of them were new: 77 of the 88 CNS genes and 85 of the 109 immune genes had not previously been associated with MS, so integrating expression data with an existing GWAS turned up far more new candidates than it confirmed old ones. Gene set enrichment split along the same immune and CNS lines, with the CNS genes linked to kinase activity and immune signalling, which the authors read as pointing to the importance of neuro-immune interactions, and the immune genes linked to lymphocyte biology.
Three External Datasets Asked to Check the List
A gene list derived entirely from statistical integration invites the question of whether these genes do anything in real MS tissue, and the authors put it to three independent datasets. In non-inflammatory post-mortem brain, the highest percentage of cells expressing the identified genes were inhibitory and excitatory neurons. In MS post-mortem lesions, the majority of the CNS genes are differentially expressed, in both white and grey matter and in a cell-type specific manner. In MS peripheral blood and CSF, a high number of the immune genes are differentially expressed cell-type specifically in one compartment or the other, some across several immune subsets. What the authors draw from this is that MS susceptibility genetics, classically considered mainly immune driven, reaches CNS tissue and CNS cells more than previously assumed, with susceptibility variants in this analysis associated with both excitatory and inhibitory neurons, matching what a recent GWAS meta-analysis using different eQTL datasets reported. Two genes carry that reading concretely. IL-12AS1 is an antisense RNA regulating IL-12 production, set against work in the EAE model showing that the IL-12 receptor beta on neurons regulates anti-inflammatory properties and prevents neurodegeneration. IL-7 came out as a CNS-specific susceptibility variant, and IL-7 protein is raised in pre-active MS lesions, where CD8+ cytotoxic T cells co-expressing the IL-7 receptor alpha chain are also found. On scope the authors are direct: they applied this pipeline only to the susceptibility GWAS, because the progression GWAS yields a single genome-wide significant hit, and they say this makes them more likely to underestimate than overestimate the number of CNS associations in MS. The work is a preprint and has not been through peer review.
Disclaimer: This blog post is based on the cited preprint and is intended for informational purposes only. The preprint has not been certified by peer review and should not be used to guide clinical practice. It is not intended to provide medical advice. Please consult with a healthcare professional for any health concerns.
Reference:
Upcott, M., Shao, B., Bray, N., Loveless, S., Tallantyre, E. C., Robertson, N. P., & Kreft, K. L. (2026). Cell-type specific transcriptional risk scores and longitudinal multiple sclerosis outcomes. medRxiv. https://doi.org/10.64898/2026.09.21.26363545
