A primary sjogren's syndrome disease prediction system based on genetic polymorphisms

By constructing a multimodal counterfactual interpretation-Transformer model and a three-valued gene feature construction method, the problems of inconsistent data processing and insufficient model interpretability in the risk prediction of primary Sjögren's syndrome were solved, and reliable risk identification and individualized prevention and control of primary Sjögren's syndrome were realized.

CN122417375APending Publication Date: 2026-07-17CENT SOUTH UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CENT SOUTH UNIV
Filing Date
2026-05-28
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

Existing methods for predicting the risk of primary Sjögren's syndrome suffer from non-standard genetic data processing, limited multimodal fusion capabilities, and insufficient model interpretability, making it difficult to achieve accurate and interpretable risk identification.

Method used

A multimodal counterfactual interpretation-Transformer model is constructed, which combines a three-valued gene feature construction method and an anti-class neighbor retrieval mechanism. Through attention contribution quantification and sliding window perturbation interval localization, biologically reasonable individualized counterfactual instances are generated to achieve reliable risk prediction for primary Sjögren's syndrome.

Benefits of technology

It enhances the representation capabilities of high-dimensional genomic data and the interpretability of models, provides transparent and traceable risk prediction results, and supports the formulation of individualized prevention and control strategies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122417375A_ABST
    Figure CN122417375A_ABST
Patent Text Reader

Abstract

This invention relates to the field of artificial intelligence technology and discloses a disease prediction system for primary Sjögren's syndrome based on gene polymorphism. The system consists of a gene data acquisition module, a feature processing server, an artificial intelligence prediction server, and medical and follow-up terminals. By performing quality control, functional site screening, and three-valued dose encoding on genotype data, structured gene features are constructed, and multimodal input features are generated by combining clinical phenotype and environmental exposure information. The system employs a multimodal counterfactual interpretation-Transformer model to quantify feature contributions, locate perturbation intervals, and perform counterfactual optimization inference, thereby outputting risk prediction values ​​and risk levels. This system effectively improves the accuracy and interpretability of gene polymorphism risk prediction, enables traceable and transparent risk assessment, and supports population risk identification and individualized prevention management.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology, and in particular to a disease prediction system for primary Sjögren's syndrome based on gene polymorphism. Background Technology

[0002] Primary Sjögren's syndrome (SLS) is a complex rheumatic and immunological disease driven by both genetic polymorphism and environmental exposure factors, exhibiting specific clinical phenotypic features. With the development of high-throughput sequencing and genotyping technologies, the application of genotype data in risk research is increasing. However, current genomic data processing lacks unified standards in quality control, site selection, and feature encoding, making it difficult to extract effective features from a large number of SNPs. Simultaneously, significant heterogeneity exists between clinical characteristics and environmental factors. Traditional risk prediction models based on logistic regression and gradient boosting trees are not robust enough to high-dimensional sparse genomic data, have weak multimodal fusion capabilities, and their prediction results are difficult to interpret. While current deep learning models possess strong feature learning capabilities, they generally suffer from strong black-box characteristics, difficulty in quantifying feature contributions, and a lack of individualized interpretation in multimodal medical scenarios. In particular, Transformer-based models lack interpretability and counterfactual analysis capabilities in medical genomic polymorphism tasks, limiting their application effectiveness in risk assessment. Therefore, there is an urgent need for a method that can integrate genetic polymorphism, clinical phenotype and environmental exposure information, and improve the transparency and traceability of intelligent risk prediction, so as to achieve more reliable risk identification and prevention support for people with primary Sjögren's syndrome. Summary of the Invention

[0003] This invention aims to address the problems of non-standard gene data processing, limited multimodal fusion capabilities, and insufficient model interpretability in existing methods for predicting the risk of primary Sjögren's syndrome (PSS). It provides a PSS disease prediction system based on gene polymorphism, achieving intelligent prediction of PSS by constructing a multimodal fusion architecture oriented towards genotype data, clinical phenotype data, and environmental exposure data. The core innovations of this invention are: first, proposing a gene feature construction method based on a three-value system, achieving effective representation of high-dimensional gene data through systematic quality control, site screening, and dose encoding; second, constructing a multimodal counterfactual interpretation-Transformer model, combining attention contribution quantification, sliding window perturbation interval localization, and a counterfactual generation optimization framework to enhance the model's feature expression capabilities and interpretability on high-dimensional heterogeneous medical data; and third, generating biologically plausible individualized counterfactual instances through anti-class neighbor retrieval and a three-objective evolutionary optimization mechanism, ensuring that risk prediction is not only accurate but also traceable and interpretable. This invention overcomes the problems of weak fusion ability, insufficient interpretability, and difficulty in handling high-dimensional gene features in existing models, and achieves reliable identification of the risk of primary Sjögren's syndrome, providing technical support for early prevention and risk management in the population.

[0004] This invention provides a disease prediction system for primary Sjögren's syndrome based on gene polymorphism, which can be applied in hospital rheumatology and immunology outpatient clinics, physical examination centers and regional public health monitoring platforms. The system includes: a gene data acquisition module, a feature extraction module, a risk assessment module, a prediction output module and a data center. The gene data acquisition module acquires genotype sample data, clinical phenotype data, and environmental exposure information of the tested individuals; The feature extraction module is deployed on the feature processing server in the data center, and includes a gene feature extraction submodule and a clinical-environment feature processing submodule. The processor calls a pre-built feature engineering program to perform quality control, site screening, and encoding on genotype sample data through the gene feature extraction submodule, forming a three-valued gene feature system. The clinical-environment feature processing submodule standardizes, imputes missing values, and encodes features on clinical phenotype data and environmental exposure factors to obtain clinical phenotype features and environmental exposure features. The three-valued gene features, clinical phenotype features, and environmental exposure features are integrated to construct multimodal input features. The risk assessment module is deployed on an AI prediction server in the data center. The AI ​​prediction server includes a multi-core central processing unit (CPU), memory, and network interface. A multimodal counterfactual interpretation-Transformer model is constructed. Multimodal input features are input into the multimodal counterfactual interpretation-Transformer model, which outputs the predicted risk value and risk level of primary Sjögren's syndrome. The construction method of the multimodal counterfactual interpretation-Transformer model is as follows: based on the Transformer model, attention contribution quantification and sliding window perturbation interval localization methods are introduced, and a three-objective counterfactual generation evolutionary optimization framework is adopted to enhance the feature expression method and reasoning mechanism of the Transformer model, thereby constructing a multimodal counterfactual interpretation-Transformer model with interpretable counterfactual reasoning capabilities. The prediction output module includes a medical staff workstation terminal and a patient follow-up terminal. The medical staff workstation terminal communicates with the artificial intelligence prediction server through the hospital's internal LAN to receive and display the disease risk prediction value and risk level of primary Sjögren's syndrome, and generate early screening warning prompts based on preset thresholds. The patient follow-up terminal uses a mobile application to push individualized prevention and control suggestions to high-risk groups, realizing the formulation of individualized prevention and control strategies for primary Sjögren's syndrome.

[0005] Furthermore, the gene feature extraction submodule includes a quality control unit, a site screening unit, and a dose encoding unit; The quality control unit performs strict quality control operations on genotype sample data. It removes unacceptable SNP sites from the genotype sample data through missing rate threshold filtering, minor allele frequency filtering (MAF), and Hardy-Weinberg balance test to obtain quality-controlled SNP data. The site selection unit, based on a pre-constructed list of genes related to primary Sjögren's syndrome, selects functionally relevant SNP sites located in the immune response regulation region and the inflammatory response regulation region from the quality-controlled SNP data to form a candidate site set; a machine learning feature selection method is introduced, and L1 regularized logistic regression is used to perform feature compression on the candidate site set to obtain a set of retained SNP features; The dose coding unit encodes the retained SNP feature set according to allele dose, forming a three-valued gene feature system; the three-valued gene feature system is represented by a 0, 1, 2 three-valued system. Furthermore, the multimodal counterfactual interpretation-Transformer model includes a preprocessing unit, an attention contribution calculation unit, a perturbation interval localization unit, an anti-class neighbor retrieval unit, a counterfactual evolution optimization unit, a counterfactual scoring selection unit, and a risk calculation and level output unit. The preprocessing unit normalizes, positions, and maps the multimodal input features to form the original multimodal input samples of the current individual, thus obtaining the multimodal input sequence. The attention contribution calculation unit inputs the multimodal input sequence into the Transformer model, extracts the multi-head self-attention matrix composed of multiple attention heads, and performs a weighted summation of the attention distribution of each attention head according to a preset contribution factor to construct the importance weight vector of the time step. The perturbation interval localization unit takes the importance weight vector of the time step as input and uses a sliding window cumulative weight search algorithm to scan the multimodal input sequence segment by segment; the continuous segment with the largest cumulative weight sum is selected from all candidate windows to determine the optimal perturbation interval; The anti-class neighbor retrieval unit has a pre-set reference sample set. Based on the classification output of the Transformer model, it determines the target class of the current individual. It retrieves a set of samples with the opposite target class from the reference sample set and selects the nearest anti-class neighbor sample according to the distance metric of the multimodal feature space. It maps the multimodal feature fragments of the anti-class neighbor sample within the optimal perturbation interval to the feature dimension region corresponding to the current individual, performs a fragment-level replacement operation, and generates an initial counterfactual instance. The counterfactual evolutionary optimization unit inputs initial counterfactual instances into a three-objective counterfactual generation evolutionary optimization framework. It introduces a three-objective optimization mechanism—maximizing confidence, constraining perturbation sparsity, and preserving instance similarity—to construct a three-objective optimization function for counterfactual search. Based on this function, a target category confidence threshold is set as a hard constraint. A reference point generation strategy based on the NSGA-III framework, an initial population construction method using binary random sampling, and genetic operators involving two-point crossover and bit-flip mutation are employed to iteratively evolve the initial counterfactual instances across multiple generations, forming a set of instances that satisfy the confidence threshold. The goal is to find a counterfactual instance set that achieves Pareto optimality in both sparsity and similarity. The three objective optimization functions include: 1. Maximizing prediction confidence: maximizing the predicted probability that a counterfactual instance is classified as a high-risk category for primary Sjögren's syndrome; 2. Minimizing perturbation sparsity: minimizing the Hamming distance between the counterfactual instance and the original sample to achieve minimal perturbation; 3. Preserving instance similarity: minimizing the vector distance (L2 norm) between the counterfactual instance and the original sample to ensure the biological feasibility and rationality of counterfacts within the joint feature space of genetic, clinical, and environmental exposures. The counterfact scoring selection unit calculates a weighted score for each counterfact in the counterfact instance set and selects the counterfact with the highest weighted score as the optimal counterfact interpretation sample. The risk calculation and grade output unit inputs the original multimodal input sample and the optimal counterfactual interpretation sample of the current individual into the Transformer model to obtain the original prediction result and the counterfactual prediction result, respectively. Based on the difference between the two in the prediction result of the target category, a risk metric is constructed to calculate the disease risk prediction value of primary Sjögren's syndrome. Based on the disease risk prediction value of primary Sjögren's syndrome, the risk grade is output to realize interpretable risk prediction based on gene polymorphism and clinical phenotype.

[0006] By adopting the above solution, the beneficial effects achieved by the present invention are as follows: First, this invention achieves systematic modeling and unified expression of the risk of primary Sjögren's syndrome by constructing a multimodal fusion architecture that integrates genotype data, clinical phenotypic data, and environmental exposure information. This effectively solves the problems of inconsistent gene data processing, fragmented feature sources, and difficulty in coordinating different modalities in existing methods. In particular, the introduction of a gene feature construction method based on a three-value system, through rigorous quality control, functional site screening, and dosage-based encoding, enables large-scale SNP sites to be stably and accurately represented in a structured space. This improves the reliability of gene-level risk modeling and provides a high-quality input foundation for robust predictions by subsequent deep learning models.

[0007] Secondly, the multimodal counterfactual interpretation-Transformer model proposed in this invention significantly enhances the feature representation capability and interpretability of heterogeneous medical data. Through attention contribution quantification, cross-channel weight aggregation, and sliding window perturbation interval localization, this invention not only achieves contribution analysis of key gene loci, clinical features, and environmental factors, but also solves the problem that traditional Transformer models struggle to locate important feature fragments in medical scenarios. Simultaneously, by combining counterfactual generation with an evolutionary optimization framework, the model can generate counterfactual samples with clearly defined feature differences, providing a transparent and traceable basis for distinguishing different risk levels. This mechanism improves the interpretability of prediction results, enabling healthcare professionals to understand the sources of risk, thereby enhancing the model's credibility in practical applications.

[0008] Furthermore, this invention combines an anti-class neighbor retrieval mechanism with a three-objective evolutionary optimization strategy to generate counterfactual examples for individuals, providing biologically plausible explanations for risk prediction. This mechanism maintains the similarity and sparsity of the original samples while satisfying confidence constraints, ensuring that the generated counterfactual instances not only reflect key factors in changes in the risk of primary Sjögren's syndrome but can also be directly applied to the development of individualized prevention recommendations. Therefore, this invention effectively solves the application dilemma of traditional prediction systems that "only provide results, not causes," enhancing the practicality, interpretability, and population management value of risk prediction, and providing data support and decision-making basis for hospitals, public health platforms, and follow-up management. Attached Figure Description

[0009] Figure 1 This is a schematic diagram of the overall structure of a primary Sjögren's syndrome disease prediction system based on gene polymorphism proposed in this invention; Figure 2 This is a cluster plot of the SNP genotypes proposed in Example 2. Detailed Implementation

[0010] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.

[0011] Example 1, according to Figure 1 This invention provides a disease prediction system for primary Sjögren's syndrome based on gene polymorphism, which can be applied in hospital rheumatology and immunology clinics, physical examination centers and regional public health monitoring platforms. The system includes: a gene data acquisition module, a feature extraction module, a risk assessment module, a prediction output module and a data center. The gene data acquisition module acquires genotype sample data, clinical phenotype data, and environmental exposure information of the tested individuals; The gene data acquisition module specifically includes: The biological sample collection subunit uses intravenous blood collection tubes, EDTA anticoagulant tubes, and saliva collection devices to collect samples from individuals and bind them to their unique personal identification numbers via sample barcodes. The gene testing subunit includes a fully automated nucleic acid extractor, a genotyping chip scanner, and a high-throughput sequencer. The gene testing subunit performs DNA extraction, library construction, and SNP site detection on the tested individuals according to a preset testing procedure, generating genotype sample data covering the entire genome. The Clinical and Environmental Information Collection Subunit is used for structured collection and unified modeling of clinical phenotypic data and environmental exposure data of the examined individuals. Clinical phenotypic data is retrieved through interfaces to the Hospital Information System (HIS) and Electronic Medical Record System (EMR), including chief complaint information, present illness history, past medical history, personal history, marital and reproductive history, menstrual history, family history, physical examination results, laboratory test results, and pathological biopsy results. Environmental exposure data is obtained through a standardized questionnaire system and an external environmental database interface, including annual average PM2.5 exposure concentration, regional carbon dioxide emission levels, ultraviolet radiation intensity, smoking frequency, green space coverage, climate characteristic parameters, and chemical and occupational exposure information. Chemical and occupational exposure information includes organic solvent exposure, silica dust exposure, particulate matter exposure, and heavy metal exposure. The feature extraction module is deployed on the feature processing server in the data center, and includes a gene feature extraction submodule and a clinical-environment feature processing submodule. The processor calls a pre-built feature engineering program to perform quality control, site screening, and encoding on genotype sample data through the gene feature extraction submodule, forming a three-valued gene feature system. The clinical-environment feature processing submodule standardizes, imputes missing values, and encodes features on clinical phenotype data and environmental exposure factors to obtain clinical phenotype features and environmental exposure features. The three-valued gene features, clinical phenotype features, and environmental exposure features are integrated to construct multimodal input features. The risk assessment module is deployed on an AI prediction server in the data center. The AI ​​prediction server includes a multi-core central processing unit (CPU), memory, and network interface. A multimodal counterfactual interpretation-Transformer model is constructed. Multimodal input features are input into the multimodal counterfactual interpretation-Transformer model, which outputs the predicted risk value and risk level of primary Sjögren's syndrome. The construction method of the multimodal counterfactual interpretation-Transformer model is as follows: based on the Transformer model, attention contribution quantification and sliding window perturbation interval localization methods are introduced, and a three-objective counterfactual generation evolutionary optimization framework is adopted to enhance the feature expression method and reasoning mechanism of the Transformer model, thereby constructing a multimodal counterfactual interpretation-Transformer model with interpretable counterfactual reasoning capabilities. The prediction output module includes a medical staff workstation terminal and a patient follow-up terminal. The medical staff workstation terminal communicates with the artificial intelligence prediction server through the hospital's internal LAN to receive and display the disease risk prediction value and risk level of primary Sjögren's syndrome, and generate early screening warning prompts based on preset thresholds. The patient follow-up terminal uses a mobile application to push individualized prevention and control suggestions to high-risk groups, realizing the formulation of individualized prevention and control strategies for primary Sjögren's syndrome.

[0012] Example 2, according to Figure 2 This embodiment is based on Embodiment 1. In this embodiment, the gene feature extraction submodule includes a quality control unit, a site screening unit, and a dose encoding unit. The quality control unit performs strict quality control operations on genotype sample data. Through deletion rate threshold filtering, minor allele frequency filtering (MAF), and Hardy-Weinberg equilibrium test, it removes SNP sites that do not meet the requirements in the genotype sample data to obtain quality-controlled SNP data. Among them, the removed sites include: sites with a sample deletion rate higher than a preset threshold; sites with a minor allele frequency lower than a preset threshold; and sites that deviate significantly from Hardy-Weinberg equilibrium at the population level. The site selection unit, based on a pre-constructed list of genes related to primary Sjögren's syndrome, selects functionally relevant SNP sites located in the immune response regulation region and the inflammatory response regulation region from the quality-controlled SNP data to form a candidate site set; a machine learning feature selection method is introduced, and L1 regularized logistic regression is used to perform feature compression on the candidate site set to obtain a set of retained SNP features; The dose-coding unit encodes the retained SNP feature set according to allele dose, forming a three-valued gene feature system. This system is represented by a 0, 1, 2 ternary system: 0 corresponds to a genotype without the target risk allele (homozygous wild-type), 1 corresponds to a heterozygous genotype carrying a single risk allele (heterozygous genotype), and 2 corresponds to a homozygous genotype carrying two risk alleles (homozygous risk type). Specifically, for each SNP locus, a reference allele (Ref) and a substitute allele (Alt) are pre-determined, and one of them is selected as the target risk allele based on the risk allele annotation information. The risk factor (Risk) is determined by analyzing the biallelic combinations of an individual at that locus and counting the copy number k of the risk allele in the genotype. When k=0, it is encoded as 0; when k=1, it is encoded as 1; and when k=2, it is encoded as 2. Specifically, when the risk allele is Alt, Ref / Ref, Ref / Alt, and Alt / Alt correspond to k=0, 1, and 2, respectively; when the risk allele is Ref, Alt / Alt, Ref / Alt, and Ref / Ref correspond to k=0, 1, and 2, respectively. The encoding process is achieved by counting the copy number of the target risk allele, ensuring a consistent correspondence between the three-valued encoding and the risk allele dosage. Figure 2The cluster plot representing SNP genotypes shows the signal intensity of Allele A on the horizontal axis and the signal intensity of Allele B on the vertical axis. Three distinct cluster regions are visible, corresponding to the homozygous wild-type AA, the heterozygous genotype AB, and the homozygous risk genotype BB, respectively. Based on the clustering results, the dose encoding unit of this invention maps the three genotypes to allele dose codes of 0, 1, and 2, respectively, to construct a three-valued gene feature vector.

[0013] Example 3, based on Example 2, in which the multimodal counterfactual interpretation-Transformer model includes a preprocessing unit, an attention contribution calculation unit, a perturbation interval localization unit, a counter-class neighbor retrieval unit, a counterfactual evolution optimization unit, a counterfactual scoring selection unit, and a risk calculation and level output unit; The preprocessing unit normalizes, positions, and maps the multimodal input features to form the original multimodal input samples of the current individual, thus obtaining the multimodal input sequence. The attention contribution calculation unit inputs the multimodal input sequence into the Transformer model and extracts a multi-head self-attention matrix composed of multiple attention heads. It then performs a weighted summation of the attention distribution of each attention head based on a preset contribution factor to construct an importance weight vector for each time step. Finally, it employs a cross-channel aggregation and normalization strategy for the weights of different modal channels to obtain a unified feature contribution sequence, which is used for subsequent counterfactual subsequence localization. The perturbation interval localization unit takes the importance weight vector of the time step as input and uses a sliding window cumulative weight search algorithm to scan the multimodal input sequence segment by segment. The continuous segment with the largest cumulative weight sum is selected from all candidate windows to determine the optimal perturbation interval. The perturbation interval is used to limit the modification range of the counterfactual generation process, realize the local constraint on feature perturbation, and avoid irrelevant or excessive changes to the overall sequence. The sliding window cumulative weight search algorithm refers to the algorithm that uses a fixed-length window to gradually slide across the input sequence based on the time step importance vector after multi-head self-attention weighting, accumulates the weights in each window, and selects the one with the largest cumulative weight as the optimal perturbation interval to limit the modification range of counterfactual generation. The anti-class neighbor retrieval unit has a pre-set reference sample set. Based on the classification output of the Transformer model, it determines the target class of the current individual. It retrieves a set of samples with the opposite target class from the reference sample set and selects the nearest anti-class neighbor sample according to the distance metric of the multimodal feature space. It maps the multimodal feature fragments of the anti-class neighbor sample within the optimal perturbation interval to the feature dimension region corresponding to the current individual, performs a fragment-level replacement operation, and generates an initial counterfactual instance. The counterfactual evolutionary optimization unit inputs initial counterfactual instances into a three-objective counterfactual generation evolutionary optimization framework. It introduces a three-objective optimization mechanism—maximizing confidence, constraining perturbation sparsity, and preserving instance similarity—to construct a three-objective optimization function for counterfactual search. Based on this function, a target category confidence threshold is set as a hard constraint. A reference point generation strategy based on the NSGA-III framework, an initial population construction method using binary random sampling, and genetic operators involving two-point crossover and bit-flip mutation are employed to iteratively evolve the initial counterfactual instances across multiple generations, forming a set of instances that satisfy the confidence threshold. The goal is to find a counterfactual instance set that achieves Pareto optimality in both sparsity and similarity. The three objective optimization functions include: 1. Maximizing prediction confidence: maximizing the predicted probability that a counterfactual instance is classified as a high-risk category for primary Sjögren's syndrome; 2. Minimizing perturbation sparsity: minimizing the Hamming distance between the counterfactual instance and the original sample to achieve minimal perturbation; 3. Preserving instance similarity: minimizing the vector distance (L2 norm) between the counterfactual instance and the original sample to ensure the biological feasibility and rationality of counterfacts within the joint feature space of genetic, clinical, and environmental exposures. The three-objective optimization function specifically includes: The formula for maximizing prediction confidence is as follows: ; in, This represents the target value for the prediction confidence level. Indicates the first One candidate counterfactual instance; Indicates the original sample The set of all generated candidate counterfactual instances; The model represents the counterfactual samples. The output category prediction; Indicates the target category; This represents the samples given by the model. Category The predicted probability; The minimum perturbation sparsity objective is expressed by the following formula: ; in, This represents the objective value for perturbation sparsity. Indicates the length of the time step. Indicates the number of feature channels. Indicates the normalization factor; Represents the Hamming distance, i.e., the distance between the original samples. Counterfactual How many elements changed across all time steps and all channels; The instance similarity preservation objective is expressed by the following formula: ; in, This indicates that instance similarity maintains the target value. Represents the original sample Counterfactual samples Distance in a continuous feature space; The counterfact scoring selection unit calculates a weighted score for each counterfact in the counterfact instance set and selects the counterfact with the highest weighted score as the optimal counterfact interpretation sample. ; in, Indicates counterfactual Overall score; This represents the weight hyperparameters, which are preset by the user or the system. The risk calculation and grade output unit inputs the original multimodal input sample and the optimal counterfactual interpretation sample of the current individual into the Transformer model to obtain the original prediction result and the counterfactual prediction result, respectively. Based on the difference between the two in the prediction result of the target category, a risk metric is constructed to calculate the disease risk prediction value of primary Sjögren's syndrome. Based on the disease risk prediction value of primary Sjögren's syndrome, the risk grade is output to realize interpretable risk prediction based on gene polymorphism and clinical phenotype.

[0014] In the conventional technical field, the Transformer model includes a preprocessing unit, a sequence construction and encoding unit, and a risk reasoning and rating output unit; The preprocessing unit preprocesses the multimodal input features, including missing value imputation, normalization, and categorical feature encoding, and converts genotype data, clinical features, and environmental exposure information into a unified input format acceptable to the model. The sequence construction and encoding unit constructs an input sequence from the preprocessed multimodal input features according to fixed rules, and directly inputs the input sequence into the Transformer model; the model encodes the features across time and modes through a multi-head self-attention mechanism to obtain a sequence-level latent space representation; The risk reasoning and level output unit estimates the risk probability of primary Sjögren's syndrome by using the sequence-level latent space representation output by the Transformer model through a fully connected layer. It then classifies the predicted value according to a preset risk range or threshold and outputs the corresponding risk level.

[0015] Example 4, this example is based on Example 3, in this example, The risk assessment module is deployed on the AI ​​prediction server in the data center. The AI ​​prediction server includes a multi-core central processing unit (CPU), memory, and network interface. A multimodal counterfactual interpretation-Transformer model is constructed. Multimodal input features are input into the multimodal counterfactual interpretation-Transformer model, which outputs the disease risk prediction value and risk level of primary Sjögren's syndrome. Original predicted probability: 0.78; Counterfactual prediction probability: 0.31; Risk predictor for primary Sjögren's syndrome: 0.656; The risk levels are output based on the disease risk prediction values ​​for primary Sjögren's syndrome, as shown in Table 1: Table 1 ; Final risk level: Medium to high risk (Level III).

[0016] The prediction output module includes a medical staff workstation terminal and a patient follow-up terminal. The medical staff workstation terminal communicates with the artificial intelligence prediction server through the hospital's internal LAN to receive and display the disease risk prediction value and risk level of primary Sjögren's syndrome, and generate early screening warning prompts based on preset thresholds. The patient follow-up terminal uses a mobile application to push individualized prevention and control suggestions to high-risk groups, realizing the formulation of individualized prevention and control strategies for primary Sjögren's syndrome.

[0017] Healthcare workstation terminal: Individual ID: X_12345; Overall risk prediction for primary Sjögren's syndrome: 0.656; Risk level: Level III (medium to high risk).

[0018] The present invention and its embodiments have been described above. This description is not restrictive. The accompanying drawings are only one embodiment of the present invention, and the actual structure is not limited thereto. In short, if a person skilled in the art is inspired by this description and designs a similar structure and embodiment without departing from the spirit of the present invention, such design should fall within the protection scope of the present invention.

Claims

1. A disease prediction system for primary Sjögren's syndrome based on gene polymorphism, characterized in that, The system includes: a gene data acquisition module, a feature extraction module, a risk assessment module, a prediction output module, and a data center; The gene data acquisition module acquires genotype sample data, clinical phenotype data, and environmental exposure information. The feature extraction module is deployed on the feature processing server in the data center, and includes a gene feature extraction submodule and a clinical-environment feature processing submodule. The gene feature extraction submodule performs quality control, site screening, and encoding on genotype sample data to form a three-valued gene feature system. The clinical-environment feature processing submodule standardizes, imputes missing values, and encodes features from clinical phenotype data and environmental exposure information to obtain clinical phenotype features and environmental exposure features. The three-valued gene features, clinical phenotype features, and environmental exposure features are integrated to construct multimodal input features. The risk assessment module is deployed on the artificial intelligence prediction server in the data center. It constructs a multimodal counterfactual interpretation-Transformer model, inputs multimodal input features into the multimodal counterfactual interpretation-Transformer model, and outputs the disease risk prediction value and risk level of primary Sjögren's syndrome.

2. The disease prediction system for primary Sjögren's syndrome based on gene polymorphism according to claim 1, characterized in that, The construction method of the multimodal counterfactual explanation-Transformer model is as follows: based on the Transformer model, attention contribution quantification and sliding window perturbation interval localization methods are introduced, and a three-objective counterfactual generation evolutionary optimization framework is adopted to enhance the feature expression method and reasoning mechanism of the Transformer model, thus constructing the multimodal counterfactual explanation-Transformer model.

3. The disease prediction system for primary Sjögren's syndrome based on gene polymorphism according to claim 1, characterized in that: The gene feature extraction submodule includes a quality control unit, a site screening unit, and a dose encoding unit; The quality control unit performs strict quality control operations on genotype sample data. It removes unacceptable SNP sites from the genotype sample data through missing rate threshold filtering, minor allele frequency filtering, and Hardy-Weinberg balance test to obtain quality-controlled SNP data. The site selection unit filters functional SNP sites located in the immune response regulation region and the inflammatory response regulation region from the quality-controlled SNP data to form a candidate site set; a machine learning feature selection method is introduced to perform feature compression on the candidate site set to obtain a SNP feature-preserving set. The dose-coding unit encodes the retained SNP feature set according to allele dose, forming a three-valued system of gene features.

4. The disease prediction system for primary Sjögren's syndrome based on gene polymorphism according to claim 3, characterized in that: The gene characteristics of the three-value system are represented by a 0, 1, 2 three-value system, specifically: 0 corresponds to homozygous wild type; 1 corresponds to heterozygous genotype; 2 corresponds to homozygous risk type.

5. The disease prediction system for primary Sjögren's syndrome based on gene polymorphism according to claim 2, characterized in that: The multimodal counterfactual interpretation-Transformer model includes a preprocessing unit, an attention contribution calculation unit, a perturbation interval localization unit, an anti-class neighbor retrieval unit, a counterfactual evolutionary optimization unit, a counterfactual scoring selection unit, and a risk calculation and level output unit. The preprocessing unit normalizes, positions, and maps the multimodal input features to form the original multimodal input samples of the current individual and constructs the multimodal input sequence. The attention contribution calculation unit inputs the multimodal input sequence into the Transformer model, extracts the multi-head self-attention matrix composed of multiple attention heads, and performs a weighted summation of the attention distribution of each attention head to construct the importance weight vector of the time step. The perturbation interval localization unit takes the importance weight vector of the time step as input, and uses a sliding window cumulative weight search algorithm to scan the multimodal input sequence segment by segment, select the continuous segment with the largest cumulative weight, and determine the optimal perturbation interval. The anti-class neighbor retrieval unit uses a pre-set reference sample set to determine the target category of the current individual based on the classification output of the Transformer model. Retrieve a set of samples from the reference sample set that are opposite to the target category of the current individual, and select the opposite class neighbor samples according to the distance metric of the multimodal feature space; map the multimodal feature fragments of the opposite class neighbor samples within the optimal perturbation interval, perform fragment-level replacement operation, and generate the initial counterfactual instance; The counterfactual evolutionary optimization unit introduces a three-objective optimization mechanism of maximizing confidence, perturbation sparsity constraint and instance similarity preservation, and constructs a three-objective optimization function; Based on the three-objective optimization function, the target category confidence threshold is set as a hard constraint. The reference point generation strategy based on the NSGA-III framework, the initial population construction method of binary random sampling, and the genetic operators of double-point crossover and bit flip mutation are used to perform multi-generation iterative evolution on the initial counterfactual instance to form a counterfactual instance set. The counterfact scoring selection unit calculates a weighted score for each counterfact in the counterfact instance set and selects the counterfact with the highest weighted score as the optimal counterfact interpretation sample. The risk calculation and rating output unit inputs the original multimodal input sample and the optimal counterfactual interpretation sample of the current individual into the Transformer model to obtain the original prediction result and the counterfactual prediction result, respectively. A risk metric is constructed based on the difference between the two prediction results for the target category, and the disease risk prediction value for primary Sjögren's syndrome is calculated; the risk level is output based on the disease risk prediction value for primary Sjögren's syndrome.

6. The disease prediction system for primary Sjögren's syndrome based on gene polymorphism according to claim 5, characterized in that: The three objective optimization functions include:

1. Maximizing prediction confidence; 2. Minimizing perturbation sparsity; 3. Preserving instance similarity.