A method, apparatus, device, medium, and product for generating a systemic lupus erythematosus risk index

CN122822331APending Publication Date: 2026-09-25PEKING UNION MEDICAL COLLEGE HOSPITAL
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610971305.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-01
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

(1)孤立分析:大多数研究仅单独分析常见变异(如构建多基因风险评分)或单独分析罕见变异,缺乏将二者进行整合分析以评估其对疾病联合贡献的系统性框架

Benefits of technology

[0012]根据本申请提供的具体实施例,本申请具有了以下技术效果:通过整合受试者的全基因组单核苷酸多态性分型数据及全外显子组测序数据,实现了对系统性红斑狼疮风险的全面、精准评估。首先,利用荟萃分析构建目标人群的多基因风险评分计算模型,并结合质量控制与群体分层校正,提高了常见变异位点风险评分的可靠性和人群特异性。其次,通过全外显子组测序数据识别罕见有害变异,弥补了仅关注常见变异的不足,增强了遗传风险的覆盖广度。最后,将标准化多基因风险评分、罕见有害变异携带状态等多维度特征融合至逻辑回归预测模型中,输出系统性红斑狼疮风险指数。该方法不仅提升了风险预测的准确性和个体化程度,还能有效指导科研数据分析与健康管理。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122822331A_ABST
    Figure CN122822331A_ABST
Patent Text Reader

Abstract

The application discloses a systemic lupus erythematosus risk index generation method, device, equipment, medium and product, relates to the medical data processing field, and the method comprises the following steps: quality control and population stratification correction are carried out on the whole genome single nucleotide polymorphism typing data of a subject; a multi-gene risk score calculation model of a target population is constructed, and the standardized multi-gene risk score of the subject is calculated according to the quality-controlled single nucleotide polymorphism typing data; the rare harmful variation carrying state of the subject is determined according to whole exome sequencing data and a risk gene set; and the systemic lupus erythematosus risk index of the subject is determined by using a logistic regression prediction model according to the standardized multi-gene risk score and the rare harmful variation carrying state of the subject. The application improves the accuracy and individualization degree of systemic lupus erythematosus risk prediction, and effectively guides scientific research data analysis and health management.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of medical data processing, and in particular to a method, apparatus, device, medium, and product for generating a systemic lupus erythematosus risk index. Background Technology

[0002] Systemic lupus erythematosus (SLE) is a highly heterogeneous autoimmune disease. Current research indicates that its genetic structure is complex, comprising two types of variants: common variants (frequency >1%) identified through genome-wide association studies, and rare variants (frequency <0.1%) identified through gene / exon sequencing. These two types of variants contribute to the genetic risk of SLE. Currently, the main limitations of existing technologies are: (1) Isolated analysis: Most studies only analyze common variants (such as constructing polygenic risk scores) or rare variants in isolation, lacking a systematic framework to integrate the two to assess their combined contribution to the disease.

[0003] (2) Population bias: Existing genetic models are mostly based on European populations, and their predictive performance is significantly reduced in East Asian populations, while East Asian populations have a higher SLE prevalence and different genetic backgrounds.

[0004] (3) Insufficient explanation of clinical heterogeneity: Existing genetic scores can predict disease risk, but cannot effectively explain the huge differences in clinical phenotypes among SLE patients (such as different organ involvement and differences in autoantibody profiles), and cannot achieve refined clinical stratification.

[0005] (4) The mechanism is unclear: there is a lack of analytical strategies to directly link specific genetic variation loads with specific immune pathway dysregulation and final clinical manifestations. Summary of the Invention

[0006] The purpose of this application is to provide a method, device, equipment, medium, and product for generating a systemic lupus erythematosus (SLE) risk index, which can improve the accuracy and individualization of SLE risk prediction and effectively guide scientific research data analysis and health management.

[0007] To achieve the above objectives, this application provides the following solution: In a first aspect, this application provides a method for generating a systemic lupus erythematosus risk index, including: Obtain whole-genome single nucleotide polymorphism (SNP) typing data and whole-exome sequencing data from the subjects; The whole-genome single nucleotide polymorphism (SNP) genotyping data were subjected to quality control and population stratification correction to obtain quality-controlled SNP genotyping data. A meta-analysis was conducted on the genome-wide association study dataset of systemic lupus erythematosus in the target population, and a multi-gene risk score calculation model for the target population was constructed. Based on the quality-controlled single nucleotide polymorphism typing data, and using a multi-gene risk scoring calculation model for the target population, the standardized multi-gene risk score of the subjects was calculated. Based on the whole exome sequencing data and the pre-determined systemic lupus erythematosus risk gene set, the rare harmful variant carrier status of the subjects was determined; The logistic regression prediction model is trained based on the standardized polygenic risk score and rare harmful variant carrier status of the subjects. The trained logistic regression prediction model receives the standardized polygenic risk score and rare harmful variant carrier status of new subjects and outputs the systemic lupus erythematosus risk index of new subjects, which is used to guide the prediction of disease incidence and the prediction of possible future clinical phenotypes.

[0008] Secondly, this application provides a systemic lupus erythematosus risk index generation device, comprising: The data acquisition module is used to acquire whole-genome single nucleotide polymorphism typing data and whole-exome sequencing data of the subjects; The data quality control module is used to perform quality control and population stratification correction on the whole genome single nucleotide polymorphism (SNP) typing data to obtain quality-controlled SNP typing data. The model building module is used to perform meta-analysis of genome-wide association study datasets of systemic lupus erythematosus in the target population and to build a multi-gene risk score calculation model for the target population. The scoring calculation module is used to calculate the standardized multigene risk score of the subjects based on the quality-controlled single nucleotide polymorphism typing data and the multigene risk scoring calculation model of the target population. The status determination module is used to determine the rare harmful variant carrier status of the subject based on the whole exome sequencing data and a pre-determined systemic lupus erythematosus risk gene set; The index calculation module is used to train a logistic regression prediction model based on the standardized polygenic risk score and rare harmful variant carrier status of the subjects. The trained logistic regression prediction model receives the standardized polygenic risk score and rare harmful variant carrier status of new subjects and outputs the systemic lupus erythematosus risk index of the new subjects to guide scientific research data analysis or health management.

[0009] Thirdly, this application provides a computer device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the above-described method for generating a systemic lupus erythematosus risk index.

[0010] Fourthly, this application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the above-described method for generating a systemic lupus erythematosus risk index.

[0011] Fifthly, this application provides a computer program product, including a computer program that, when executed by a processor, implements the above-described method for generating a systemic lupus erythematosus risk index.

[0012] According to the specific embodiments provided in this application, this application achieves the following technical effects: By integrating whole-genome single nucleotide polymorphism (SNP) genotyping data and whole-exome sequencing data of subjects, a comprehensive and accurate assessment of systemic lupus erythematosus (SLE) risk is realized. First, a multi-gene risk score calculation model for the target population is constructed using meta-analysis, and combined with quality control and population stratification correction, the reliability and population specificity of risk scores for common variant sites are improved. Second, rare harmful variants are identified through whole-exome sequencing data, compensating for the shortcomings of focusing only on common variants and enhancing the breadth of genetic risk coverage. Finally, standardized multi-gene risk scores, rare harmful variant carrier status, and other multi-dimensional features are integrated into a logistic regression prediction model to output a SLE risk index. This method not only improves the accuracy and individualization of risk prediction but also effectively guides scientific research data analysis and health management. Attached Figure Description

[0013] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0014] Figure 1 This is an application environment diagram of a method for generating a systemic lupus erythematosus risk index according to an embodiment of this application.

[0015] Figure 2 This is a flowchart illustrating a method for generating a systemic lupus erythematosus risk index, provided as an embodiment of this application.

[0016] Figure 3 This is a schematic diagram of the functional modules of a systemic lupus erythematosus risk index generation device provided in an embodiment of this application.

[0017] Figure 4 This is a schematic diagram of the structure of a computer device provided in an embodiment of this application.

[0018] Figure 5 This is a schematic diagram of the genetic structure of systemic lupus erythematosus in East Asian populations in one embodiment of this application.

[0019] Figure 6 This is a schematic diagram illustrating the combined effect of common and rare variants on SLE susceptibility in one embodiment of this application.

[0020] Figure 7 This is a schematic diagram illustrating the association between polygenic risk and clinical characteristics, and the identification results of three genetically distinct clinical subtypes of SLE in one embodiment of this application. Detailed Implementation

[0021] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0022] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, the application will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0023] The systemic lupus erythematosus risk index generation method provided in this application embodiment can be applied to, for example... Figure 1 In the application environment shown, terminal 101 communicates with server 102 via a network. A data storage system can store the data that server 102 needs to process. The data storage system can be set up independently, integrated into server 102, or placed in the cloud or on another server. Terminal 101 can send the subject's whole-genome single nucleotide polymorphism (SNP) typing data and whole-exome sequencing data to server 102. Server 102 determines the subject's SLE risk index based on the received data. Server 102 can then feed back the obtained SLE risk index to terminal 101.

[0024] The terminal 101 can be, but is not limited to, various desktop computers, laptops, smartphones, tablets, IoT devices, and portable wearable devices. IoT devices can include smart speakers, smart TVs, smart air conditioners, and smart in-vehicle devices. Portable wearable devices can include smartwatches, smart bracelets, and head-mounted devices. The server 102 can be implemented using a standalone server or a server cluster composed of multiple servers, or it can be a cloud server.

[0025] In one exemplary embodiment, such as Figure 2As shown, a method for generating a systemic lupus erythematosus risk index is provided. This method is executed by a computer device, specifically by a terminal or server alone, or by both a terminal and a server. In this embodiment, the method is applied to... Figure 1 Taking server 102 as an example, the explanation includes the following steps 201 to 207.

[0026] Step 201: Obtain whole-genome single nucleotide polymorphism typing data and whole-exome sequencing data of the subjects.

[0027] In a specific application example, step 201 includes steps 11 to 13.

[0028] Step 11: Obtain the subject's genomic DNA.

[0029] Step 12: Perform whole-genome single nucleotide polymorphism (SNP) typing on the genomic DNA to obtain the subject's whole-genome SNP typing data, i.e., common SNP genotype data across the whole genome.

[0030] Step 13: Perform whole exome sequencing (WES) on the genomic DNA to obtain the subject's whole exome sequencing data.

[0031] Step 202: Perform quality control and population stratification correction on the whole genome single nucleotide polymorphism (SNP) typing data to obtain quality-controlled SNP typing data.

[0032] In a specific application example, step 202 specifically includes: performing variation-level and sample-level quality control on the whole-genome single nucleotide polymorphism (SNP) genotyping data to obtain quality-controlled SNP genotyping data. The quality control process specifically includes: removing samples with low detection rates, removing SNP sites with low detection rates, removing SNP sites that deviate from expectations in the Hardy-Weinberg equilibrium test, and removing SNP sites with a minor allele frequency (MAF) of less than 1% (i.e., retaining only common variants).

[0033] This application further performs principal component analysis (PCA) on the quality-controlled SNP genotyping data, using the first five principal components (PC1-PC5) as common SNP principal components. Please use the common SNP principal components and sex of the subjects as covariates to correct for possible genetic stratification in the population.

[0034] Step 203: A meta-analysis was performed on the Genome-Wide Association Study (GWAS) dataset of the target population to construct a polygenic risk score (PRS) calculation model for the target population. The PRS calculation model includes the weight of each SNP.

[0035] In a specific application example, the target population is East Asian, which includes two types of data: external population data and self-generated data. Step 203 includes steps 31 to 33.

[0036] Step 31: Collect SLE GWAS datasets for the target population. Specifically, collect summary statistics for SLE GWAS from at least three independent East Asian populations, with a total sample size of 190,413 (including 4,222 cases and 8,431 controls; 512 cases and 994 controls; and 317 cases and 175,937 controls).

[0037] Step 32: Perform a fixed-effects meta-analysis on the SLE GWAS dataset of the target population. Specifically, a fixed-effects inverse variance weighted model is used to perform a meta-analysis on the SLE GWAS dataset of the target population, calculating the pooled effect size (β value or OR value) and pooled p-value for each SNP to obtain the meta-analysis effect size.

[0038] Step 33: Using the meta-analysis effect size as prior information, the posterior effect size of each SNP is inferred using the Polygenic Risk Score via Continuous Shrinkage (PRS-CS) algorithm based on Bayesian regression and continuous shrinkage prior, and a PRS calculation model (i.e., the weight of each SNP) suitable for the target population is constructed.

[0039] Step 204: Based on the quality-controlled single nucleotide polymorphism (SNP) typing data, calculate the standardized multigene risk score of the subjects using the multigene risk scoring calculation model for the target population.

[0040] In a specific application example, step 204 includes steps 41 to 43.

[0041] Step 41: Extract the genotypes of the SNPs included in the PRS calculation model from the quality-controlled SNP genotyping data.

[0042] Step 42: The genotypes of each SNP locus are weighted and summed with their corresponding weights to obtain the original standardized PRS of the subjects.

[0043] Step 43: Perform Z-score standardization on the raw standardized PRS of the subjects so that the mean PRS of the entire analysis population is 0 and the standard deviation is 1, thus obtaining the standardized PRS of the subjects.

[0044] Step 205: Based on the whole exome sequencing data and the pre-determined systemic lupus erythematosus risk gene set, determine the subject's rare harmful variant carrier status.

[0045] In a specific application example, step 205 includes steps 51 to 52.

[0046] Step 51: Select variants located in the SLE risk gene set from the whole exome sequencing data to obtain the target gene set.

[0047] Specifically, a set of risk genes clearly associated with SLE was first selected, comprising the following 36 genes: FAS, FASL, IKZF1, PTEN, RELA, TNFAIP3, P2RY8, PRKCD, RAG1, RAG2, ACP5, ADAR, DDX58, DNASE1, DNASE1L3, IFIH1, RNASEH2A, RNASEH2B, RNASEH2C, SAMHD1, SAT1, STING1, TLR7, TREX1, UNC93B1, C1QA, C1QB, C1QC, C1R, C1S, C2, C3, C4A, C4B, CYBB, NCF1 Then, variants located in the aforementioned 36 genes were screened from the subjects' whole exome sequencing data to obtain the target gene set.

[0048] Step 52: Detect whether there are rare harmful variants in the target gene set to determine the subject's rare harmful variant carrier status.

[0049] Specifically, the definition of rare harmful mutations: Rare: In public databases (gnomAD Exome East Asian v4.1, gnomAD Genome v4.1, ExAC East Asian v0.3.1), the MAF of this variant is <0.1%.

[0050] Harmful (meets any of the following conditions): (1) Loss-of-Function (LoF) variants: including nonsense mutations, frameshift mutations, and classical splice site variants, and marked as "high confidence" by the LOFTEE plugin; (2) Harmful missense variants: meet at least two of the following three prediction tools: Combined Annotation Dependent Depletion (CADD) score ≥20, Rare Exome Variant Ensemble Learner (REVEL) score >0.5, and AlphaMissense predicts "likely pathogenic".

[0051] If at least one rare harmful variant that meets the above criteria is detected in the target gene set, it is marked as a "carrier"; otherwise, it is marked as a "non-carrier".

[0052] Step 206: Train a logistic regression prediction model based on the subject's standardized polygenic risk score and rare harmful variant carrier status. The trained logistic regression prediction model receives the standardized polygenic risk score and rare harmful variant carrier status of new subjects and outputs the new subject's systemic lupus erythematosus (SLE) risk index to guide research data analysis or health management. The SLE risk index can be used to characterize the multidimensional combination level of genetic information.

[0053] Specifically, the following variables will be used as predictors: Independent variables (primary predictors): Standardized PRS, rare harmful variant carrier status.

[0054] Covariates (correction factors): gender, SNP principal components (PC1-PC5).

[0055] Using SLE disease status (case=1, control=0) as the dependent variable, logistic regression was used to fit the above features to obtain a logistic regression prediction model.

[0056] For new subjects, their corresponding standardized PRS and rare harmful variant carrier status are input, and the logistic regression prediction model outputs the predicted probability of the subject having SLE.

[0057] Step 207: Genetic risk stratification is performed based on the standardized polygenic risk score, the polygenic risk score of multiple immune pathways, and the rare harmful variant carrier status to predict the possible clinical phenotype of the subject's future disease.

[0058] Genetic risk stratification is performed based on the standardized polygenic risk score and the carrier status of the rare harmful variant.

[0059] PRS stratification: Standardized PRS is divided into three tertile groups according to population distribution: low PRS group (<33.3% quantile), medium PRS group (33.3%-66.7% quantile), and high PRS group (≥66.7% quantile).

[0060] Carrier status stratification: Subjects are divided into two categories: carriers and non-carriers.

[0061] This resulted in six subgroups, with the "low PRS and non-carrier" group as the reference (OR=1). The odds ratio (OR) and its 95% confidence interval were calculated for each subgroup. Examples of the results are as follows: Low PRS / Carrier OR=1.46, Medium PRS / Non-Carrier OR=2.76, Medium PRS / Carrier OR=3.85, High PRS / Non-Carrier OR=6.86, High PRS / Carrier OR=11.62.

[0062] Pathway analysis was performed on the standardized polygenic risk scores of the subjects to determine the polygenic risk scores of multiple immune pathways and the total polygenic risk score. Based on the polygenic risk scores of multiple immune pathways, the total polygenic risk score, and the carrier status of rare harmful variants, the possible clinical phenotypes of the subjects in the future were predicted.

[0063] Specifically, the association between immune pathways and clinical phenotypes is first analyzed to determine the pathway-phenotype relationship. In practical applications, this pathway-phenotype relationship can be used to partially predict the possible clinical phenotype of the subject in the future.

[0064] In one exemplary embodiment, such as Figure 3 As shown, a systemic lupus erythematosus risk index generation device is provided, which includes the following functional modules.

[0065] The data acquisition module 301 is used to acquire whole-genome single nucleotide polymorphism typing data and whole-exome sequencing data of the subjects.

[0066] Specifically, the genomic DNA of the subjects is first obtained, then whole-genome SNP typing is performed on the genomic DNA to obtain the subjects' whole-genome SNP typing data, and finally WES is performed on the genomic DNA to obtain the subjects' whole-exome sequencing data.

[0067] The data quality control module 302 is used to perform quality control and population stratification correction on the whole genome single nucleotide polymorphism (SNP) typing data to obtain quality-controlled SNP typing data.

[0068] Specifically, firstly, the whole-genome SNP genotyping data were subjected to quality control at both the variation level and sample level to obtain quality-controlled SNP genotyping data. Then, PCA was performed on the quality-controlled SNP genotyping data, and the first five principal components (PC1-PC5) were used as common SNP principal components to correct for possible genetic stratification in the population.

[0069] Model building module 303 is used to perform meta-analysis of the genome-wide association study dataset of systemic lupus erythematosus in the target population and to build a multi-gene risk score calculation model for the target population.

[0070] Specifically, firstly, SLE GWAS datasets for the target population were collected. Then, a fixed-effects meta-analysis was performed on the SLE GWAS datasets for the target population. Finally, using the meta-analysis effect size as prior information, the PRS-CS algorithm was used to infer the posterior effect size of each SNP, and a PRS calculation model suitable for the target population was constructed.

[0071] The scoring calculation module 304 is used to calculate the standardized multigene risk score of the subject based on the single nucleotide polymorphism typing data after quality control and the multigene risk scoring calculation model of the target population.

[0072] Specifically, firstly, the genotypes of SNPs included in the PRS calculation model are extracted from the quality-controlled SNP genotyping data. Then, the genotypes of each SNP locus are weighted and summed with their corresponding weights to obtain the subject's raw standardized PRS. Finally, the subject's raw standardized PRS is Z-score standardized to obtain the subject's standardized PRS.

[0073] The status determination module 305 is used to determine the rare harmful variant carrier status of the subject based on the whole exome sequencing data and a pre-determined set of systemic lupus erythematosus risk genes.

[0074] Specifically, firstly, variants located within the SLE risk gene set are screened from whole-exome sequencing data to obtain the target gene set. Then, the presence of rare harmful variants in the target gene set is detected to determine the subject's rare harmful variant carrier status.

[0075] The index calculation module 306 is used to train a logistic regression prediction model based on the standardized polygenic risk score, rare harmful variant carrier status, and gender of the subjects. The trained logistic regression prediction model receives the standardized polygenic risk score and rare harmful variant carrier status of new subjects and outputs the systemic lupus erythematosus risk index of the new subjects to guide scientific research data analysis or health management.

[0076] Specifically, the following variables will be used as predictors: Independent variables (primary predictors): Standardized PRS, rare harmful variant carrier status.

[0077] Covariates (correction factors): gender, SNP principal components (PC1-PC5).

[0078] Using SLE disease status (case=1, control=0) as the dependent variable, logistic regression was used to fit the above features to obtain a logistic regression prediction model.

[0079] For new subjects, their corresponding standardized PRS and rare harmful variant carrier status are input, and the logistic regression prediction model outputs the subject's risk score for SLE.

[0080] Phenotypic prediction module 307 is used to perform genetic risk stratification based on the standardized polygenic risk score, the polygenic risk score of multiple immune pathways, and the rare harmful variant carrier status, so as to predict the possible clinical phenotype of the subject's future disease.

[0081] Specifically, the PRS was first stratified: the standardized PRS was divided into three tertile groups according to the population distribution: low PRS group (<33.3% quantile), medium PRS group (33.3%-66.7% quantile), and high PRS group (≥66.7% quantile).

[0082] Then, the carrier status was stratified: the subjects were divided into two categories, namely carriers and non-carriers.

[0083] This resulted in six subgroups, with the "low PRS and non-carrier" group as the reference (OR=1). The odds ratio (OR) and its 95% confidence interval were calculated for each subgroup. Examples of the results are as follows: Low PRS / Carrier OR=1.46, Medium PRS / Non-Carrier OR=2.76, Medium PRS / Carrier OR=3.85, High PRS / Non-Carrier OR=6.86, High PRS / Carrier OR=11.62.

[0084] Pathway analysis was performed on the standardized polygenic risk scores of the subjects to determine the polygenic risk scores of multiple immune pathways and the total polygenic risk score. Based on the polygenic risk scores of multiple immune pathways, the total polygenic risk score, and the carrier status of rare harmful variants, the possible clinical phenotypes of the subjects in the future were predicted.

[0085] In summary, this application achieves a comprehensive and accurate assessment of systemic lupus erythematosus (SLE) risk by integrating subject gender, whole-genome single nucleotide polymorphism (SNP) genotyping data, and whole-exome sequencing data. First, a multi-gene risk scoring model for the target population is constructed using meta-analysis, and quality control and population stratification correction are combined to improve the reliability and population specificity of risk scores for common variant sites. Second, rare and harmful variants are identified using whole-exome sequencing data, compensating for the shortcomings of focusing only on common variants and enhancing the breadth of genetic risk coverage. Finally, standardized multi-gene risk scores, rare and harmful variant carrier status, and other multi-dimensional features are integrated into a logistic regression prediction model to output a SLE risk index. This method not only improves the accuracy and individualization of risk prediction but also effectively guides scientific data analysis and health management.

[0086] In one exemplary embodiment, a computer device is provided, which may be a server or a terminal, and its internal structure diagram may be as follows. Figure 4 As shown, the computer device includes a processor, memory, input / output (I / O) interfaces, and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is also connected to the system bus via the I / O interfaces. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides the environment for the operation of the operating system and computer programs in the non-volatile storage media. The database stores the subject's sex, whole-genome single nucleotide polymorphism (SNP) typing data, and whole-exome sequencing data. The I / O interfaces are used for information exchange between the processor and external devices. The communication interface is used for communication with external terminals via a network connection. When executed by the processor, the computer program implements a method for generating a systemic lupus erythematosus (SLE) risk index.

[0087] Those skilled in the art will understand that Figure 4 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer equipment to which the present application is applied. Specific computer equipment may include, for example, [the following is a list of possible additional structures]. Figure 4 The diagram shows more or fewer components, or combinations of certain components, or different component arrangements.

[0088] In one exemplary embodiment, a computer-readable storage medium is provided storing a computer program that, when executed by a processor, implements the steps in the above-described method embodiments.

[0089] In one exemplary embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above-described method embodiments.

[0090] In one specific application scenario, an East Asian woman with a family history of systemic lupus erythematosus (SLE) visited the rheumatology and immunology department of a hospital for a genetic risk assessment. After signing the informed consent form, the doctor collected her peripheral venous blood, and the clinical laboratory extracted genomic DNA. The samples were then tested using whole-genome SNP genotyping chips and whole-exome sequencing technology. After the testing was completed, the laboratory uploaded the generated SNP genotyping data and whole-exome sequencing data to a bioinformatics analysis server. The server first performed strict quality control on the SNP genotyping data, removing low-quality loci. Subsequently, the system called a multi-gene risk scoring calculation model pre-constructed based on a meta-analysis of genome-wide association studies of SLE in East Asian populations, extracted the genotypes of the corresponding SNP loci of the subject, weighted them, and obtained a standardized multi-gene risk score. At the same time, the system screened rare variants located in the known SLE risk gene set from the whole-exome sequencing data, and determined whether the subject was a carrier of a rare harmful variant based on the criteria for loss-of-function variants and various harmfulness prediction tools. Finally, the genetic characteristics from the two dimensions mentioned above, along with the subject's gender information, are input into a pre-trained logistic regression prediction model. The model outputs a systemic lupus erythematosus (SLE) risk index for the subject and stratifies the genetic risk based on this index. Doctors then develop personalized health management plans and follow-up schedules based on this risk index, the subject's risk subgroup, family history, and clinical signs.

[0091] like Figures 5 to 7 As shown below, a specific embodiment is used to describe the method for generating the systemic lupus erythematosus risk index provided in this application.

[0092] (1) Data source.

[0093] like Figure 5As shown in Part A, this embodiment integrates aggregated statistical data from three independent, publicly available cohorts of 190,413 East Asian individuals to conduct a GWAS meta-analysis: Chinese Cohort 1 (GCST90011866; 4,222 cases, 8,431 controls), Chinese Cohort 2 (GCST90014238; 512 cases, 994 controls), and Japanese Cohort (GCST90018697; 317 cases, 175,937 controls). The independent validation cohort consisted of 2,198 SLE patients recruited by Peking Union Medical College Hospital (as shown in Table 1) and 1,068 healthy controls. SLE patients were enrolled from the Chinese Systemic Lupus Erythematosus Research Group (CSTAR) registry between May 2009 and December 2024. All patients met at least one of the following established classification criteria: the 1997 American College of Rheumatology (ACR) revised criteria, the 2012 Systemic Lupus Erythematosus International Collaborative Clinical Group criteria, or the 2019 European League Against Rheumatism / ACR criteria. The control group was drawn from the Shunyi study, excluding individuals diagnosed with autoimmune diseases. All participants signed written informed consent forms, and the study was approved by the Ethics Committee of Peking Union Medical College Hospital (I-22PJ026 and I-25PJ0522).

[0094] Table 1. Baseline characteristics of 2,198 patients diagnosed with SLE

[0095] (2) SNP genotyping, quality control and filling.

[0096] Genotyping of genomic DNA from all 3,266 participants was performed using Illumina Infinium Asian ScreeningArray v1.0. Rigorous quality control at the variant and sample level was performed using PLINK (v1.9), followed by haplotype typing using SHAPEIT (v2.r904) and filling with Minimac4 (v1.0.2) using the East Asian population reference panel from the 1000 Genomes Project Phase 3. Common variants with a minor allele frequency (MAF) >1% were retained for analysis. Principal component analysis (PCA) was performed based on common variants.

[0097] (3) Whole exome sequencing and rare variant analysis.

[0098] Whole-exome sequencing (WES) was performed on all participants. Variance detection was performed using GATK (v4.3.0.0), and annotation was performed using VEP (v105) and the LOFTEE plugin. Loss-of-function (LOF) variants were defined as variants marked as high confidence by LOFTEE. Missense variants were classified as potentially harmful if they were marked by at least two of the following: CADD ≥ 20, REVEL > 0.5, or AlphaMissense. Rare variants were defined as having a MAF < 0.1% in the gnomADExome East Asian population (v4.1), gnomAD Genome (v4.1), and ExAC East Asian population (v0.3.1) databases. Gene burden analysis employed a two-sided Fisher exact test, focusing on 36 identified SLE risk genes, including FAS, FASL, IKZF1, PTEN, RELA, TNFAIP3, P2RY8, PRKCD, RAG1, RAG2, ACP5, ADAR, DDX58, DNASE1, DNASE1L3, IFIH1, RNASEH2A, RNASEH2B, RNASEH2C, SAMHD1, SAT1, STING1, TLR7, TREX1, UNC93B1, C1QA, C1QB, C1QC, C1R, C1S, C2, C3, C4A, C4B, CYBB, and NCF1.

[0099] (4) GWAS meta-analysis and functional annotation.

[0100] Fixed-effects inverse variance weighted meta-analysis was performed using METAL software (v2021.03.10). Gene-level analysis was performed using MAGMA (v1.10) with a 1000-genome East Asian population linkage disequilibrium reference panel, followed by gene ontology (GO) enrichment analysis using the R package clusterProfiler (v4.6.2).

[0101] (5) Polygenic Risk Score (PRS) analysis.

[0102] The PRS models were constructed from eight sources: meta-analysis, three component GWAS datasets (GCST90011866, GCST90014238, and GCST90018697), and four published models from large Asian or European cohorts (as shown in Table 2). For datasets with available summary statistics, PRS-CS was used to infer the posterior effect size of SNPs; existing models were scored using PLINK2. Model effectiveness was assessed using AUC with 95% confidence intervals, and DeLong's test was used for intergroup comparisons. Pathway-based PRS maps SNPs to 16 selected immune pathways (including antigen processing and presentation, apoptosis, DNA damage / repair and autophagy pathways, B cell receptor signaling pathway, complement and coagulation cascade, cytoplasmic DNA sensing pathway, cell burial pathway, JAK-STAT signaling pathway, mTOR signaling pathway, NF-κB signaling pathway, PI3K-Akt signaling pathway, RIG-I-like receptor signaling pathway, T cell receptor signaling pathway, Th1 and Th2 cell differentiation, Th17 cell differentiation, and Toll-like receptor signaling pathway) using MAGMA with a ±100kb window, with weights derived from PRS-CS.

[0103] Table 2. PRS models used for comparison in this application

[0104] (6) Unsupervised clustering analysis.

[0105] Unsupervised clustering was performed using the K-prototypes algorithm based on standardized pathway PRS values ​​and rare variant carrier states. The optimal number of clusters was determined using the elbow rule, and cluster stability was evaluated using Jaccard similarity and adjusted RAND index (ARI) through bootstrapping resampling (n=100). The clustering results were visualized using factor analysis of mixed data (FAMD).

[0106] (7) Statistical analysis.

[0107] The association between genetic traits and SLE susceptibility or clinical phenotype was assessed by logistic regression, adjusting for sex and the first five principal components. Multiple test correction was performed using the Benjamini-Hochberg method. For intergroup comparisons, continuous variables were analyzed using ANOVA or the Kruskal-Wallis test, and categorical variables were analyzed using the chi-square test or Fisher's exact test. All analyses were performed in R (v4.5.1).

[0108] The above analysis yields the following results.

[0109] (1) GWAS meta-analysis identified 84 SLE susceptibility sites.

[0110] A fixed-effects meta-analysis of 190,413 East Asian individuals (5,051 cases and 185,362 controls) identified 84 independent genome-wide significant associated loci (P < 5 × 10⁻⁶). -8 This includes known sites and several new sites (such as...). Figure 5 (Part B of the equation). The strongest correlation signal is located in HLA-DQA1 (P=1.2×10). -82 This is consistent with previous reports. Gene-level analysis (MAGMA) identified several significantly enriched genes, and GO enrichment analysis revealed significant enrichment in pathways such as type I IFN signaling, T cell activation, and apoptosis regulation (e.g., Figure 5 (As shown in parts C, D, and E of the diagram).

[0111] (2) The PRS derived from meta-analysis is superior to all comparative models.

[0112] A systematic comparison of eight PRS models was performed in an independent validation cohort (2,198 cases, 1,068 controls). The meta-analysis-derived PRS (metaGWAS-PRS) had an AUC of 0.733, significantly outperforming all other models (DeLong test, all P < 0.05; e.g.) Figure 6 (As shown in Parts A and B of the dataset). The AUC of the PRS model based on the three constituent GWAS datasets ranged from 0.658 to 0.712, while the published model from the European cohort showed a significant decline in performance in the East Asian population (AUC = 0.601 to 0.643), suggesting a significant loss of cross-population transferability.

[0113] (3) PRS is associated with the patient’s clinical phenotype.

[0114] In the case cohort, higher PRS was significantly associated with younger age of onset, higher SLEDAI score, renal involvement, and more widespread positivity for autoantibodies (anti-dsDNA, anti-Sm, anti-nucleosome antibodies; all P < 0.05) (e.g. Figure 7 (As shown in Part A of the diagram). These associations remained robust after adjusting for sex and principal components, suggesting that the burden of common polygenic variants not only affects disease susceptibility but also shapes clinical phenotypes. Building upon the global PRS, pathway-based PRS analysis further revealed the fine correspondence between specific immune pathways and clinical phenotypes. Association analysis of PRS and clinical phenotypes for 16 SLE-related immune pathways identified 19 significant pathway-phenotype associations after FDR correction (e.g., ...). Figure 7(As shown in Part B of the diagram). Adaptive immune pathways—especially antigen processing and presentation, Th17 cell differentiation, and Th1 / Th2 differentiation pathways—are positively correlated with anti-rRNP and anti-SSA antibody positivity, but negatively correlated with anticardiolipin and anti-β2GPI antibodies, with the anticardiolipin antibody pathway association being the most significant and consistent. Regarding organ involvement, chronic rash is positively correlated with antigen processing and presentation pathways, while thrombocytopenia is negatively correlated with this pathway. These results indicate that the burden of common polygenic variants does not uniformly affect the overall phenotype, but rather selectively shapes clinical features through specific immune pathways: increased burden of adaptive immune pathways corresponds to a more immunocompetent phenotype (high disease activity, renal involvement, and broad-spectrum autoantibody positivity), while antiphospholipid antibody-related phenotypes exhibit the opposite genetic structure, suggesting that their pathogenesis may be partially independent of adaptive immune pathways driven by common variants.

[0115] (4) Rare harmful variants were independently enriched in SLE cases.

[0116] Gene burden analysis was performed on rare harmful variants (MAF < 0.1%) among 36 identified SLE risk genes. The carrier rate of rare harmful variants was significantly higher in cases compared to controls (OR = 1.39, 95% CI: 1.01–1.92, P = 0.044). Figure 6 (As shown in section C). The effect of loss of function (LOF) variants was stronger (OR=2.85, 95% CI: 1.19-6.84, P=0.019).

[0117] Cases carrying rare harmful variants showed no significant differences in age of onset, disease activity, or major organ involvement compared to non-carriers (all P>0.05), suggesting that rare variants primarily affect the disease susceptibility threshold rather than clinical severity. This contrasts with the phenotypic associations of polygenic PRS, which are significantly associated with a variety of clinical features.

[0118] (5) Common and rare variants contribute additively to the risk of SLE.

[0119] To examine the combined effect of common and rare variants, this embodiment constructed a joint model integrating metaGWAS-PRS and rare variant carrier status. The joint model achieved an AUC of 0.793 (95% CI: 0.778-0.808), significantly outperforming either PRS alone (AUC=0.733, DeLong P<0.001) or rare variant status alone (AUC=0.556, DeLong P<0.001). Figure 6 (as shown in part D of the diagram).

[0120] Stratified analysis based on PRS quartiles and rare variant carrier status showed that both types of genetic factors had an additive effect on disease risk (e.g., Figure 6 (As shown in sections E and F of the diagram). In the highest PRS quartile, carriers of rare variants had the highest disease risk (OR=12.4, 95% CI: 7.8–19.7), while in the lowest PRS quartile, carriers of rare variants also had a significantly higher risk than non-carriers (OR=3.1, 95% CI: 1.4–6.9), suggesting that the two types of variants jointly promote disease development through independent mechanisms. Additive interaction tests did not reveal a significant statistical interaction (P=0.31), supporting an additive rather than synergistic effect model.

[0121] (6) Unsupervised clustering revealed three genetically significantly different SLE subtypes.

[0122] Based on PRS values ​​and rare variant carrier status across 16 immune pathways, K-prototypes unsupervised clustering was performed on 2,198 SLE patients. The elbow rule determined the optimal number of clusters to be 3. Bootstrap stability analysis showed that all three clusters were highly stable (mean Jaccard similarity = 0.82, ARI = 0.79).

[0123] Cluster 1 (adaptive immune-driven, 30.8%) was characterized by high PRS in adaptive immune pathways (T cell activation, B cell signaling, and immunoglobulin production), with the highest positive rates for anti-dsDNA (78.2%), anti-Sm (45.6%), and anti-nucleosome antibodies (62.3%), and kidney involvement was the most common (52.4%).

[0124] Cluster 2 (innate immune-coagulation type, 36.4%) was characterized by high PRS in the innate immune pathway (type I IFN signaling, complement activation) and coagulation pathway, with the highest positive rate of antiphospholipid antibodies (38.7%) and the most prominent involvement of the hematologic system (thrombocytopenia, hemolytic anemia) (41.2%).

[0125] Cluster 3 (low genetic burden, 32.8%) generally had lower pathway PRS values, the lowest rare variant carrier rates, the latest age of onset, the lowest disease activity, and the fewest organ involvements. This subtype may represent a patient group with a relatively low genetic burden and a relatively large contribution from environmental factors.

[0126] The three subtypes showed significant differences in age of onset (P<0.001), SLEDAI score (P<0.001), renal involvement (P<0.001), hematologic involvement (P=0.003), and autoantibody profile (all P<0.05). Figure 7(See Parts C and D in the table, and Tables 3 and 4). In Tables 3 and 4, Cluster 1 vs. 2 represents the p-values ​​for pairwise comparisons between Cluster 1 and Cluster 2. Cluster 1 vs. 3 represents the p-values ​​for pairwise comparisons between Cluster 1 and Cluster 3. Cluster 2 vs. 3 represents the p-values ​​for pairwise comparisons between Cluster 2 and Cluster 3.

[0127] Table 3. PRS values ​​of 16 pathways and frequency of rare harmful variant carriers in three clusters.

[0128] Table 4. Comparison of clinical characteristics among three clusters based on PRS and rare pathogenic variant carrier status in 16 pathways.

[0129] In summary, this embodiment established a genetic analysis framework for SLE that integrates common and rare variants in East Asian populations, and achieved the following main findings: (1) Based on a GWAS meta-analysis of 190,413 East Asians, 84 SLE susceptibility loci were identified, and the PRS model derived from the meta-analysis was superior to all comparative models; (2) Rare harmful variants were independently enriched in SLE cases, with the effect of LOF variants being particularly significant; (3) Common and rare variants made additive contributions to SLE risk, and the combined model had the highest predictive power; (4) Unsupervised clustering revealed three biologically different SLE subtypes, each with characteristic genetic structures and clinical phenotypes.

[0130] A key finding of this embodiment is that polygenic PRS is significantly associated with a variety of clinical features, while rare variant carrier status is not significantly associated with clinical severity. This difference may reflect the different biological mechanisms of the two types of variants: common variants affect disease phenotype by regulating the overall activity level of immune pathways, while rare variants may primarily lower the genetic threshold for disease development by disrupting key nodes in specific molecular pathways, but once the disease occurs, its clinical manifestations are more determined by the polygenic background. This hypothesis is consistent with the diversity of clinical phenotypes among rare variant carriers in previous studies.

[0131] The identification of the three SLE subtypes has significant clinical implications. The adaptive immune-driven subtype, characterized by a highly active adaptive immune response, may respond better to biologics targeting B cells (e.g., belimumab) or T cells (e.g., abatacept). The innate immune-coagulation subtype, characterized by innate immune activation and coagulation abnormalities, may benefit more from anticoagulation therapy with type I IFN pathway inhibitors (e.g., anilumab) or antiphospholipid syndrome. The low genetic burden subtype, with a relatively low genetic burden, may respond well to standard immunosuppressive therapy and has a relatively better prognosis.

[0132] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.

[0133] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments described above. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM).

[0134] The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, etc., and are not limited to these.

[0135] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0136] This document uses specific examples to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the methods and core ideas of this application. Furthermore, those skilled in the art will recognize that, based on the ideas of this application, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of this application.

Claims

1. A method for generating a systemic lupus erythematosus risk index, characterized in that, include: Obtain whole-genome single nucleotide polymorphism (SNP) typing data and whole-exome sequencing data from the subjects; The whole-genome single nucleotide polymorphism (SNP) genotyping data were subjected to quality control and population stratification correction to obtain quality-controlled SNP genotyping data. A meta-analysis was conducted on the genome-wide association study dataset of systemic lupus erythematosus in the target population, and a multi-gene risk score calculation model for the target population was constructed. Based on the quality-controlled single nucleotide polymorphism typing data, and using a multi-gene risk scoring calculation model for the target population, the standardized multi-gene risk score of the subjects was calculated. Based on the whole exome sequencing data and the pre-determined systemic lupus erythematosus risk gene set, the rare harmful variant carrier status of the subjects was determined; The logistic regression prediction model is trained based on the standardized polygenic risk scores and rare harmful variant carrier status of the subjects. The trained logistic regression prediction model receives the standardized polygenic risk scores and rare harmful variant carrier status of new subjects and outputs the systemic lupus erythematosus risk index of the new subjects to guide scientific research data analysis or health management.

2. The method for generating a systemic lupus erythematosus risk index according to claim 1, characterized in that, Obtain whole-genome single nucleotide polymorphism (SNP) typing data and whole-exome sequencing data from the subjects, including: Obtain the subject's genomic DNA; Whole-genome single nucleotide polymorphism (WNP) typing was performed on the genomic DNA to obtain the subject's whole-genome single nucleotide polymorphism (WNP) typing data; Whole exome sequencing was performed on the genomic DNA to obtain the subject's whole exome sequencing data.

3. The method for generating a systemic lupus erythematosus risk index according to claim 1, characterized in that, The whole-genome single nucleotide polymorphism (SNP) genotyping data were subjected to quality control and population stratification correction to obtain quality-controlled SNP genotyping data, including: The genome-wide single nucleotide polymorphism (SNP) genotyping data were subjected to variation level and sample level quality control to obtain quality-controlled SNP genotyping data.

4. The method for generating a systemic lupus erythematosus risk index according to claim 1, characterized in that, The multigene risk scoring calculation model includes a weight for each single nucleotide polymorphism. Based on the quality-controlled single nucleotide polymorphism (SNP) genotyping data, and using a multigene risk scoring model for the target population, a standardized multigene risk score was calculated for each subject, including: Extract the genotypes of single nucleotide polymorphisms included in the multi-gene risk scoring calculation model from the quality-controlled single nucleotide polymorphism typing data; The genotypes of each single nucleotide polymorphism site and their corresponding weights are weighted and summed to obtain the original polygenic risk score of the subject. The original polygenic risk scores of the subjects were Z-score standardized to obtain the standardized polygenic risk scores of the subjects.

5. The method for generating a systemic lupus erythematosus risk index according to claim 1, characterized in that, Based on the whole-exome sequencing data and a pre-defined set of systemic lupus erythematosus risk genes, the subject's rare harmful variant carrier status was determined, including: The target gene set is obtained by screening variants located within the systemic lupus erythematosus risk gene set from the whole exome sequencing data; The presence of rare harmful variants in the target gene set is detected to determine the subject's rare harmful variant carrier status.

6. The method for generating a systemic lupus erythematosus risk index according to claim 1, characterized in that, Also includes: Genetic risk stratification is performed based on the standardized polygenic risk score, the polygenic risk score of multiple immune pathways, and the carrier status of the rare harmful variants to predict the possible clinical phenotype of the subject's future disease.

7. A systemic lupus erythematosus (SLE) risk index generation device, applied to the SLE risk index generation method according to any one of claims 1-6, characterized in that, The device includes: The data acquisition module is used to acquire whole-genome single nucleotide polymorphism typing data and whole-exome sequencing data of the subjects; The data quality control module is used to perform quality control and population stratification correction on the whole genome single nucleotide polymorphism (SNP) typing data to obtain quality-controlled SNP typing data. The model building module is used to perform meta-analysis of genome-wide association study datasets of systemic lupus erythematosus in the target population and to build a multi-gene risk score calculation model for the target population. The scoring calculation module is used to calculate the standardized multigene risk score of the subjects based on the quality-controlled single nucleotide polymorphism typing data and the multigene risk scoring calculation model of the target population. The status determination module is used to determine the rare harmful variant carrier status of the subject based on the whole exome sequencing data and a pre-determined systemic lupus erythematosus risk gene set; The index calculation module is used to train a logistic regression prediction model based on the standardized polygenic risk score and rare harmful variant carrier status of the subjects. The trained logistic regression prediction model receives the standardized polygenic risk score and rare harmful variant carrier status of new subjects and outputs the systemic lupus erythematosus risk index of the new subjects to guide scientific research data analysis or health management.

8. A computer device, comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that the processor executes the computer program to implement the method for generating a systemic lupus erythematosus risk index according to any one of claims 1-6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the computer program implements the method for generating the systemic lupus erythematosus risk index as described in any one of claims 1-6.

10. A computer program product, comprising a computer program, characterized in that, When executed by a processor, the computer program implements the method for generating the systemic lupus erythematosus risk index as described in any one of claims 1-6.