Loading icon

Multiancestry MS GWAS Meets Single-Cell and Spatial Data

Multiancestry MS GWAS Meets Single-Cell and Spatial Data
Share:

Multiple sclerosis (MS) is more than ten times as prevalent among people of European ancestry as among those of East Asian ancestry, and the genetics record carries that asymmetry forward. More than 200 susceptibility associations have been reported, almost all from European cohorts; MS genome-wide association studies in non-European populations have been limited to a few biobanks, and there had been none among East Asian individuals, partly because the low prevalence makes recruitment hard. Until now the only clearly identified MS risk variants in East Asian individuals sat in the major histocompatibility complex. Fujimoto and colleagues ran the first Japanese MS GWAS, 688 cases against 205,199 controls, folded it into a four-population meta-analysis totalling 29,374 cases and 1,843,563 controls, and then took the result into single-cell and spatial data to ask which cells and which parts of a lesion the risk actually lands on.

Two Loci in the Japanese Data, One Seen Nowhere Else
The Japanese analysis combined two datasets and tested 8,858,013 variants passing post-imputation criteria of minor allele frequency above 0.5% and imputation quality R² above 0.7, with a genomic inflation factor of 1.026. Two loci cleared genome-wide significance. The lead variant in the MHC was rs80079921, nearest to BTNL2, with an odds ratio of 2.71 (95% CI 2.28–3.23) at P = 4.9×10⁻²⁹, and it had not been reported in that region before. The second, rs2199759 at 11q24, carried an odds ratio of 1.58 at P = 2.1×10⁻⁸ and had not been reported in any population. It sits in an intergenic region 670 kb upstream of KIRREL3 and 914 kb downstream of ETS1, which encodes a transcription factor essential to the differentiation of lymphoid lineages and whose downregulation in brain endothelial cells has been reported to impair the blood-brain barrier. On that basis the authors propose the locus may act as a distal enhancer-like regulator of ETS1. The signal did not appear in the other populations, and it is distinct from the MS-associated locus inside ETS1 itself that earlier studies found.

What Scaling to Four Populations Bought
The European meta-analysis, 27,572 cases and 1,436,801 controls, returned 153 significant loci including 17 novel ones. The African analysis, 819 cases and 155,904 controls, returned two loci, both novel. The cross-population analysis returned 156 loci including three that no single-population analysis had reached, giving 22 novel loci in total. Two sit near genes with existing experimental support in MS: rs41287361 near SEMA4D, which encodes an axonal guidance factor that also has a role in immune responses and whose knockout mice resist experimental autoimmune encephalomyelitis, and rs142239370 near PRF1. For the 17 novel European variants, effect sizes were concordant with the 2019 consortium analysis, which the authors read as detection resting on their larger sample rather than on anything new in the underlying biology. Of the two African hits, one is African-specific at an allele frequency of 1.9% against 0.1% in Europeans, while the other has ample European frequency at 33.9% but was not nominally significant in the earlier consortium data.

Shared and Population-Specific, Measured in Both Directions
Of 125 loci significant in Europeans and testable in the Japanese data, 24 were nominally significant and 11 showed nominally significant heterogeneity, with effect sizes correlating across the two populations at Spearman's ρ = 0.67 (P = 6.3×10⁻¹⁸). The novel Japanese variant ran the other way, absent from every other population. The authors set out five reasons an association can stay hidden in a single cohort: limited sample size, population-specific variants with low frequency elsewhere, population-specific causality through gene-environment or gene-gene interaction, sex-specific effects, and variant filtering issues such as imputation quality. They then tested the fourth of those directly. Sex-stratified GWAS using individual-level genotypes found that rs142239370, nominally significant only in the male-heavy veteran cohorts, reached significance in neither sex (P = 0.768 in men, 0.652 in women), and across 145 lead variants 11 showed nominal sex heterogeneity while none survived Bonferroni correction.

Four Independent HLA Signals, Three New to the Field
HLA imputation on the Japanese genotype data, using a Japanese reference panel, produced 53 two-digit and 87 four-digit HLA alleles and 603 amino acid polymorphisms. The strongest association was phenylalanine at HLA-DQβ1 amino acid position 9, at an odds ratio of 2.24 and P = 4.9×10⁻³⁷. Stepwise conditioning then pulled out three further independent signals: HLA-DQA1*03:01 (OR 2.45, P = 5.6×10⁻¹⁷), HLA-DRB1*15:01 (OR 1.83, P = 3.6×10⁻¹⁰) and glycine at HLA-DQβ1 position 125 (OR 0.37, P = 2.7×10⁻⁸). Because stepwise conditioning assumes additivity, they repeated it incorporating non-additive effects and recovered the same four. Only HLA-DRB1*15:01 is the familiar one, the most common MS risk allele in the MHC across ancestries; the other three have not been identified in European populations. HLA-DRB1*04:05, previously reported as an East Asian-specific MS risk allele, was significant in the initial analysis but did not hold after conditioning on position 9, which the authors read as its association being largely explained by that amino acid.

Where the Risk Lands in Cells
Single-cell disease relevance scoring was applied to the cross-population meta-analysis against scRNA-seq of 163,541 peripheral blood mononuclear cells from 20 patients with MS. Among the eight major cell types, CD4+ T cells were significant (P = 0.0020), and within them naive (P = 0.015), central memory (P = 0.0020), effector memory (P = 0.010) and memory regulatory T cells (P = 0.0040), with memory regulatory T cells showing a stronger association than naive ones. The CD4+ finding was confirmed in an independent dataset from the Asian Immune Diversity Atlas. Turning to the brain, in single-nucleus data from MS subcortical lesions covering 74,686 cells, T cells, myeloid cells, endothelial cells and B cells were all significant. Removing the immune populations to look at what remained left endothelial cells significant (P = 0.0010) and brought astrocytes (P = 0.0090) and stromal cells (P = 0.013) through as well. Endothelial cells are the main component of the blood-brain barrier and astrocytes have recently been identified as early and highly active players in lesion formation, so the authors read this as those two cell types having substantial roles alongside the immune cells and microglia that are usually named.

Spatial Niches, a Specificity Check, and What Is Still Out of Reach
The spatial step used gsMap on transcriptomic data from eight active and four inactive MS subcortical lesions, against six niches defined in the original study by gene expression and deconvoluted cell type proportions. In the active lesion, the vascular infiltrating, lesion rim, lesion core and periplaque white matter niches were all significant (P = 2.9×10⁻¹², 5.0×10⁻¹¹, 5.2×10⁻⁷ and 3.9×10⁻⁶). The same four reached significance in the inactive lesion but more weakly (2.0×10⁻⁹, 4.5×10⁻⁵, 0.0010 and 0.019). The vascular infiltrating niche sits primarily within the lesion core, carries endothelial, stromal and immune cells and resembles lesion-associated perivascular compartments, while the rim is characterized by myeloid cells. To test whether this is MS-specific rather than generic inflammation, they repeated the analysis with GWAS data for rheumatoid arthritis, systemic lupus erythematosus, ulcerative colitis, Crohn's disease and allergy, plus Alzheimer's disease, depression and low-density lipoprotein as controls, and the vascular infiltrating and rim associations in active lesions stood out for MS. Comparing gsMap strength against deconvoluted cell type proportions gave positive correlations for that niche with stromal cells (ρ = 0.43) and the rim with myeloid cells (ρ = 0.69), though no single cell type drove the associations on its own. Four limitations close the paper. The meta-analyses relied primarily on biobank data, and part of the sample collection was based on ICD-10 codes, which may have introduced uncertainty in the medical diagnoses. Cases were included regardless of disease stage or clinical subtype to raise sample size, so within-case GWAS will need more clinical data than exists. Brain spatial transcriptomic data are scarce and the analysis methods have not been standardized. And both scoring methods are polygenic, which leaves the specific genes and pathways driving the observed associations unidentified.

Disclaimer: This blog post is based on the cited study and is intended for informational purposes only. It is not intended to provide medical advice. Please consult with a healthcare professional for any health concerns.

Reference:
Fujimoto, R., Ogawa, K., Namba, S., Ogawa, Y., Edahiro, R., Sonehara, K., Tagawa, S., Watanabe, M., Yata, T., Shirai, Y., Yamamoto, Y., Sato, G., Kai, C., Naito, T., Hosokawa, A., Yamamoto, M., Matsuda, K., Shimizu, F., Kinoshita, M., Mihara, M., Nakamori, M., Shimizu, Y., Kawachi, I., Miyamoto, K., Niino, M., Kumanogoh, A., Nakatsuji, Y., Matsushita, T., Mochizuki, H., Kira, J., Okuno, T., Isobe, N., & Okada, Y. (2026). Multiancestry genome-wide association and multiomics analyses elucidate spatiocellular features of multiple sclerosis genetics. Nature Genetics, 58, 2192–2200. https://doi.org/10.1038/s41588-026-02741-5