Gene data processing method and system, electronic equipment and storage medium
By constructing a gene feature library through self-supervised learning and integrating genetic features, the method enhances the accuracy of risk prediction in PRS models by capturing complex genetic relationships and biological context, addressing the limitations of SNP-based approaches.
Patent Information
- Application Number
- CN202510805935.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-17
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2045-06-17
AI Technical Summary
Existing PRS models rely on the cumulative effects of a large number of SNPs for risk scores, making it difficult to intuitively connect with biological pathways or functions, resulting in low accuracy in risk score prediction.
By constructing a gene feature library, using the embedding model to self-supervised learning of the gene expression profile, extracting gene embedding features, and combining prediction models for feature integration and risk score prediction, capturing complex relationships between genes.
It improves the accuracy of risk score prediction, reduces the complexity of the model, reduces the demand for sample size, improves training efficiency and generalization ability, and can better reflect biological effects.
Smart Images

Figure CN120319318A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of gene analysis technologies, and particularly to a method, a system, an electronic device, and a storage medium for processing gene data. Background Art
[0002] A polygenic risk score (PRS) model is a tool for evaluating an individual's risk of developing a certain disease or having a certain trait. It integrates genetic variation information at multiple gene loci to evaluate an individual's genetic susceptibility. Currently, this model has been widely applied in scenarios such as disease risk prediction, clinical decision support, and drug research and development.
[0003] The PRS model is based on the results of genome-wide association studies (GWAS), determines numerous single nucleotide polymorphisms (SNPs) in an individual that are related to a specific disease or trait, and then integrates the effect values of these related SNPs to calculate a risk score.
[0004] However, in related technologies, the method of using the PRS model for risk scoring only relies on the cumulative effect of a large number of SNPs for risk scoring, and it is difficult to intuitively link the cumulative effect of SNPs with biological pathways or functions, resulting in a limitation in the in-depth understanding of the risk mechanism, and further leading to low accuracy in predicting the risk score. Summary of the Invention
[0005] The main objective of the embodiments of the present application is to propose a method, a system, an electronic device, and a storage medium for processing gene data, aiming to improve the accuracy of predicting risk scores for target risk tasks.
[0006] To achieve the above objective, in the first aspect of the embodiments of the present application, a method for processing gene data is proposed. The method includes: Obtain the genotype data of a to-be-tested individual, and determine multiple associated genes of the to-be-tested individual according to the genotype data and single nucleotide polymorphism data associated with a target risk task; Query the associated gene features corresponding to each associated gene in a pre-constructed gene feature library, where the gene feature library includes embedding representations of different genes constructed based on an embedding model; the embedding model is obtained by performing self-supervised learning on multiple gene expression profiles of the same species as the to-be-tested individual; Integrate the multiple associated gene features to obtain individual features; The prediction model is called to process the individual features, and a risk score corresponding to the individual to be tested and the target risk task is obtained. The prediction model is trained based on the sample features of multiple sample individuals and the corresponding phenotypic information.
[0007] To achieve the above object, a second aspect of the embodiments of the present application proposes a gene data processing system, which includes: A data acquisition unit, configured to acquire the genotype data of the individual to be tested, and determine multiple associated genes of the individual to be tested according to the genotype data and the single nucleotide polymorphism data associated with the target risk task; A query unit, configured to query the associated gene features corresponding to each associated gene in a pre-constructed feature library. The gene feature library includes the embedding representations of different genes constructed based on an embedding model; the embedding model is obtained by performing self-supervised learning on multiple gene expression profiles of the same species as the individual to be tested; An integration unit, configured to perform feature integration on the multiple associated gene features to obtain individual features; A prediction unit, configured to call a prediction model to process the individual features, and obtain a risk score corresponding to the individual to be tested and the target risk task. The prediction model is trained based on the sample features of multiple sample individuals and the corresponding phenotypic information.
[0008] In some embodiments, the data acquisition unit includes: An acquisition subunit, configured to acquire the first single nucleotide polymorphism data associated with the target risk task; A screening subunit, configured to screen the first single nucleotide polymorphism data based on the statistical test results of the allele frequency differences between the risk individuals and the reference individuals of the target risk task to obtain multiple second single nucleotide polymorphism data; A determination subunit, configured to determine multiple third single nucleotide polymorphism data based on the genotype data, and determine multiple associated genes of the individual to be tested according to the multiple second single nucleotide polymorphism data and the multiple third single nucleotide polymorphism data.
[0009] Optionally, in some embodiments, the determination subunit includes: An acquisition module, configured to acquire the genomic positions and functional association information corresponding to the multiple second single nucleotide polymorphism data and the multiple third single nucleotide polymorphism data; A mapping module, configured to map the multiple second single nucleotide polymorphism data and the multiple third single nucleotide polymorphism data to the corresponding genes according to the genomic positions and the functional association information to obtain multiple associated genes.
[0010] Optionally, in some embodiments, the present application further provides a gene feature library construction device, including: An acquisition unit, configured to acquire a plurality of first training samples, where the first training samples include gene expression profiles of human samples; A training unit, configured to perform self-supervised learning on an embedding model by using the first training samples; A construction unit, configured to, when the embedding model training converges, construct a gene feature library based on the embedding representations of different genes extracted by the trained embedding model.
[0011] Optionally, in some embodiments, the training unit includes: A normalization subunit, configured to perform normalization processing on the gene expression profiles of the first training samples, and determine target training samples according to the normalization processing results; A training subunit, configured to perform self-supervised learning by using random masking based on the target training samples to obtain the embedding model.
[0012] Optionally, in some embodiments, the integration unit includes: A second acquisition subunit, configured to acquire weight coefficients corresponding to the plurality of associated gene features; A calculation subunit, configured to perform weighted calculation on the plurality of associated gene features based on the weight coefficients to obtain individual features.
[0013] Optionally, in some embodiments, the integration unit includes: A convolution subunit, configured to perform convolution processing on a feature sequence formed by the plurality of associated gene features to obtain sequence features; An integration subunit, configured to integrate the plurality of associated gene features based on the sequence features to obtain individual features.
[0014] Optionally, in some embodiments, the present application further provides a model training device, including: A third acquisition subunit, configured to acquire a plurality of second training samples, where the second training samples include sample genotype data of a plurality of sample individuals and corresponding phenotypic information labels; A search subunit, configured to determine sample-associated genes of each sample individual based on the sample genotype data, and search for gene features corresponding to each sample-associated gene in the gene feature library to obtain a plurality of sample-associated gene features; A prediction subunit, configured to integrate the plurality of sample-associated gene features to obtain sample individual features, and input the sample individual features into a prediction model to be trained for disease risk prediction to obtain a predicted risk score; An update subunit, configured to calculate a loss value according to the predicted risk score and the corresponding phenotypic information label, and update the parameters of the prediction model based on the loss value.
[0015] To achieve the above object, a third aspect of the embodiments of the present application provides an electronic device, including a memory and a processor, where the memory stores a computer program, and when the processor executes the computer program, the processing method for gene data described in the first aspect is implemented.
[0016] To achieve the above object, a fourth aspect of the embodiments of the present application provides a storage medium storing a computer program, and when the computer program is executed by a processor, the processing method for gene data described in the first aspect is implemented.
[0017] To achieve the above object, a fifth aspect of the embodiments of the present application provides a computer program product, which includes a computer program, and the computer program is read and executed by a processor of a computer device, so that the computer device executes the processing method for gene data described in the first aspect.
[0018] The processing method for gene data proposed in the embodiments of the present application includes: obtaining genotype data of a to-be-tested individual, and determining multiple associated genes of the to-be-tested individual according to the genotype data and single nucleotide polymorphism data associated with a target risk task; querying, in a pre-constructed gene feature library, an associated gene feature corresponding to each associated gene, where the gene feature library includes embedding representations of different genes constructed based on an embedding model; the embedding model is obtained by performing self-supervised learning on multiple gene expression profiles of the same species as the to-be-tested individual; performing feature integration on the multiple associated gene features to obtain an individual feature; and calling a prediction model to process the individual feature to obtain a risk score corresponding to the to-be-tested individual and the target risk task, where the prediction model is trained based on sample features of multiple sample individuals and corresponding phenotypic information.
[0019] It can be seen from this that the processing method for gene data provided by the embodiments of the present application, through the pre-constructed gene embedding feature library, enables the generation of an individual's embedding feature according to the individual's genotype information and the gene embedding feature library when predicting the risk score of the target risk task for the individual. Further, a risk score is obtained by predicting the individual embedding feature based on a prediction model obtained through supervised training. Since the method based on gene embedding integration features captures the complex relationships between genes and can reflect the true biological effects better than simple SNP accumulation, the accuracy of risk score prediction can be improved. Description of the Drawings
[0020] The accompanying drawings are used to provide a further understanding of the technical solutions of the present application, and constitute a part of the specification. Together with the embodiments of the present application, they are used to explain the technical solutions of the present application, and do not constitute a limitation to the technical solutions of the present application.
[0021] Figure 1 It is a schematic flow chart of the method for processing gene data provided by the present application; Figure 2 It is a schematic flow chart of the method for constructing a gene feature library provided by the present application; Figure 3 It is a schematic flow chart of the model training method provided by the present application; Figure 4 It is a schematic diagram of the area under the ROC curve of the risk score prediction model provided by the present application in an independent test set of simulated data; Figure 5 It is a schematic diagram of the F1 score of the risk score prediction model provided by the present application in an independent test set of simulated data; Figure 6 It is a schematic diagram of the confusion matrix of the risk score prediction model provided by the present application in an independent test set of simulated data; Figure 7 It is a schematic diagram for evaluating the importance of each dimension feature of the gene vector; Figure 8a It is a schematic diagram for comparing the ROC of the solution provided by the present application and the model based on SNP score in the related art in an independent test set; Figure 8b It is a schematic diagram for comparing the F1 scores of the solution provided by the present application and the model based on SNP score in the related art in an independent test set; Figure 9 It is a schematic structural diagram of the gene data processing system provided by the embodiments of the present application; Figure 10 It is a schematic hardware structure diagram of the electronic device provided by the embodiments of the present application. Detailed implementation manners
[0022] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application, and are not used to limit the present application.
[0023] Before further elaborating on the embodiments of the present application, the nouns and terms involved in the embodiments of the present application are described. The nouns and terms involved in the embodiments of the present application are applicable to the following explanations: Transformer model: A deep learning model based on the attention mechanism, originally used for machine translation tasks, and now widely used in fields such as natural language processing and computer vision.
[0024] Embedding vector: An embedding vector is essentially a mapping that maps objects in a high-dimensional discrete space to a low-dimensional continuous vector space. For example, in natural language processing, each word in the vocabulary is originally a discrete symbol, and through the embedding operation, it can be converted into a low-dimensional real number vector. This processing enables the model to better capture the semantic and syntactic relationships between objects, and the embedding vectors corresponding to similar objects are closer in the vector space.
[0025] Support Vector Machine (SVM) model: It is a supervised machine learning model that can be used for classification and regression analysis.
[0026] Genotype: The genotype is the sum total of all the genetic material within the cells of an organism. This genetic material contains specific gene combinations that determine the various genetic traits of the organism. A gene is a DNA fragment with genetic effects, and different genes carry different genetic information. The genotype is the combination form of these genes.
[0027] Cerebrovascular disease (Stroke) is a common neurological disease with high disability and fatality rates, imposing a huge burden on global public health. Genetic factors play an important role in the occurrence and development of cerebrovascular disease. Polygenic risk score (PRS) has become an important tool for predicting the risk of complex diseases by integrating information on multiple genetic variants (usually single nucleotide polymorphisms, SNPs) related to the disease to evaluate an individual's genetic susceptibility. In related technologies, the PRS model usually directly uses a large number of SNPs as features. However, this method has problems such as difficult model training due to high dimensionality, difficult biological significance interpretation, and insufficient information utilization. Specifically, since the number of SNPs involved is extremely large (tens of thousands or even millions), this leads to complex model training, easy overfitting, and requires a large number of samples to obtain robust results. The cumulative effect of a large number of SNPs is difficult to intuitively link to biological pathways or functions, limiting the in-depth understanding of the disease mechanism. In addition, simply relying on the presence or absence of SNPs or genotype information will ignore the complex interactions between genes in terms of function and expression regulation.
[0028] Based on this, to solve the problem of inaccurate prediction caused by directly using SNPs as features to process gene data in related technologies, the embodiments of the present application provide a method, system, electronic device, and storage medium for processing gene data, in order to improve the accuracy of processing gene data. Next, the method for processing gene data provided by the embodiments of the present application will be described.
[0029] Refer to Figure 1 , in some embodiments, the method for processing gene data provided by the embodiments of the present application includes but is not limited to steps S101 to S104.
[0030] Step S101: Obtain the genotype data of the individual to be tested, and determine multiple associated genes of the individual to be tested based on the genotype data and the single nucleotide polymorphism data associated with the target risk task.
[0031] Among them, the method for processing gene data provided in the embodiments of the present disclosure can specifically be a method for predicting a risk score of a disease. In the embodiments of the present disclosure, before performing a risk assessment on an individual to be tested, a gene feature library can be pre-constructed, and a risk score prediction model for a specific target risk task, or a prediction model for short, can be trained based on the pre-constructed gene feature library. Then, the trained prediction model is used to predict the risk score of the target risk task for the individual to be tested. Specifically, when predicting the risk score of the individual to be tested, the genotype information of the individual to be tested can be obtained first, and then the associated genes of the individual to be tested can be determined based on the genotype information of the individual to be tested; afterwards, the gene features corresponding to these associated genes can be further found in the pre-constructed gene feature library, and the gene features corresponding to these associated genes are fused to obtain an individual feature; finally, the trained prediction model is used to predict the risk score of the fused individual feature to obtain an accurate risk score of the individual to be tested.
[0032] Next, the construction process of the gene feature library, the training process of the prediction model, and the process of using the trained prediction model to perform a risk score on the individual to be tested will be introduced in detail from these three aspects respectively.
[0033] As Figure 2 shown, it is a schematic flowchart of the method for constructing a gene feature library provided by the present application. In some embodiments, the construction process of the gene feature library includes: Step S201: Obtain multiple first training samples, where the first training samples include the gene expression profiles of human samples.
[0034] In the embodiments of the present application, a method for extracting gene embedding vectors based on a deep learning model is provided. Among them, the deep learning model can specifically be a model based on the Transformer architecture, such as a bidirectional encoder based on Transformer (Bidirectional Encoder Representations from Transformers, BERT), or an autoregressive model based on Generative Pretrained Transformer (GPT). In the embodiments of the present application, the deep learning model used to construct the gene feature library can be called an embedding model.
[0035] After determining the model architecture of the embedding model for constructing the gene feature library, multiple first training samples for model training can be further obtained, where the multiple first training samples can include the gene expression profiles of human samples. That is, in the embodiments of the present application, multiple gene expression profiles can be determined according to the population to which the individual to be tested belongs, and then the trained embedding model can be used to extract the gene embedding representations of the genes corresponding to the population to which the individual to be tested belongs to construct the gene feature library. Among them, the gene expression profile refers to the set of all gene transcripts of a specific cell, tissue or organism in a certain functional state, which can reflect the gene expression situation. The gene expression process is that DNA is transcribed into RNA and then translated into protein. Gene expression profile analysis focuses on the transcriptional level, detecting the types and abundances of mRNAs in cells to understand the gene expression state. For example, when the individual to be tested is human, large-scale human gene expression profiles can be collected, such as bulk RNA sequencing data from public databases such as ARCHS4, GTEx, Gene Expression Omnibus (GEO), Sequence Read Archive (SRA), etc.
[0036] Step S202, perform self-supervised learning on the embedding model using the first training samples.
[0037] Among them, in the embodiments of the present application, the way to train the embedding model can be self-supervised learning training. Among them, the core of self-supervised learning is to construct a pre-training task to let the model mine the supervision signal from the data. The model is trained on the pre-training task to learn the internal structure and feature representation of the data. These representations can be used as the initial features for downstream tasks to improve the performance of the model in downstream tasks. Self-supervised training does not require a large amount of labeled data, so large-scale data training can be achieved on the basis of low data labeling cost, ensuring the training effect of the model.
[0038] In the process of training the embedding model, the training objective can be designed in one of the following two ways to obtain high-quality gene embedding representations: One way is to predict the true identity (such as gene identifier) of the masked gene based on the gene expression context information. This task aims to capture the "semantic" relationship of genes in the expression network (such as the co-regulation relationship in the co-expression network); Another way is to predict the expression level of the masked gene at the numerical level. This expression level can be predicted after discretization (such as binning) to strengthen the model's learning of gene expression intensity information.
[0039] After constructing the training samples, preprocessing is performed based on the gene expression profiles of each training sample. For example, standardization processing is performed on the gene expression of each training sample. Among them, the standardization processing may include calculating the Z-score or performing sorting and binning, etc. Then, according to the results of the standardization processing (such as the sorting information or binning results based on the Z-score values), the target training samples are determined, so as to perform self-supervised learning on the embedding model based on the target training samples (including the sorted gene identification sequences and the results of the standardization processing).
[0040] It can be understood that during the process of self-supervised learning training of the embedding model, the target training samples can be used as the input of the model, and then the gene expression levels (or the gene identifications of some genes) of some genes are randomly masked, so that the embedding model predicts the true expression levels (or true identities) of the masked genes, and uses the internal structure of the data to generate training targets without manual annotation. After training is completed, the gene expression data of the samples can be further used as the input in the inference stage, and the unmasked gene identification sequences are obtained as the input of the embedding model through preprocessing and sorting, so that the embedding model extracts the embedding representations of each gene. For example, the multi-dimensional vectors of the input layer (or a certain hidden layer) of the embedding model are extracted as the embedding representations of each gene.
[0041] Furthermore, the above two self-supervised learning methods can be jointly trained. For example, the same Transformer encoder can be used to simultaneously learn the context dependence and numerical distribution law of gene expression, set different prediction heads (predicting the expression level or true identity), and combine the losses of the two tasks (which can be weighted, and can be used as model training parameters or dynamically adjusted according to the task importance or convergence speed (such as focusing on gene ID prediction in the early stage and expression level prediction in the later stage)), so that the model can simultaneously predict the two tasks and learn more comprehensive embedding representations.
[0042] Step S203, when the training of the embedding model converges, a gene feature library is constructed based on the embedding representations of different genes extracted by the trained embedding model.
[0043] The process of training the embedding model is a process of iterative training. During the iterative training of the embedding model, the number of iterations of the iterative training can be counted, and the change amount of the model parameters can be detected. When the number of iterations reaches the preset number, or the change amount of the model parameters is less than the preset change threshold, it can be determined that the training of the embedding model converges, so as to obtain the trained embedding model.
[0044] After training the embedding model, the embedding vector corresponding to each gene can be extracted from the trained embedding model as the embedding feature of the gene. Specifically, the embedding vector corresponding to each gene can be extracted from the input layer or a certain hidden layer of the model as the embedding feature of the gene. These embedding vectors can be low-dimensional dense real vectors, encoding the functional and relational information of genes in the expression network. Specifically, the dimension of the embedding vector can be from dozens to hundreds of dimensions. For example, it can be 32 dimensions, or in some embodiments, it can be reduced to a lower dimension through an autoencoder or the like. After obtaining the embedding features corresponding to multiple genes, a gene feature library can be constructed based on the embedding features of these multiple genes.
[0045] In the embodiments of the present application, gene embedding vectors are trained using large-scale gene expression data, which can capture the functional roles and interrelationships of genes in complex biological networks, so that high-dimensional sparse gene information can be compressed into a low-dimensional dense vector space. These embedding vectors contain rich biological information autonomously learned by the model from data-driven gene expression patterns, such as gene functions, biological processes participated in, pathway information, and even associations with diseases.
[0046] After constructing the gene feature library, one or more risk score prediction models can be further trained based on the pre-constructed gene feature library. In the embodiments of the present application, the risk score prediction model can be referred to as a prediction model, and the prediction model can specifically be a machine learning classification model, such as Logistic Regression, Support Vector Machine (SVM), Random Forest, Gradient Boosting Machines, etc.
[0047] As Figure 3 shown, it is a schematic flowchart of the model training method provided by the present application. In the embodiments of the present application, the process of training the prediction model includes: Step S301, obtain a plurality of second training samples.
[0048] In the embodiments of the present application, the prediction model can be trained in a supervised training manner, that is, the above prediction model is supervised and trained using labeled sample data. In this way, labeled second training samples can be obtained first. The labeled second training samples can include sample genotype data of multiple sample individuals and phenotypic information labels corresponding to each sample individual. For example, when the sample individual is a human individual and the target risk task is to predict the cerebrovascular disease score, the genotype data of multiple sample individuals and whether each sample individual is a cerebrovascular disease patient can be obtained. When the sample individual is a cerebrovascular disease patient, the corresponding phenotypic information label is determined to be 1; conversely, when the sample individual is not a cerebrovascular disease patient, the corresponding phenotypic information label is determined to be 0.
[0049] Step S302: Determine the sample associated genes of each sample individual based on the sample genotype data, and look up the gene features corresponding to each sample associated gene in the gene feature library to obtain multiple sample associated gene features.
[0050] When training the prediction model based on the obtained multiple second training samples, the sample associated genes corresponding to the sample individuals can be determined first based on the sample genotype data included in each second training sample. Among them, to determine the sample associated genes corresponding to the sample individuals according to the sample genotype data, specifically, the variant sites in the sample genotype data can be mapped first to determine a part of the associated genes. Then, the genes mapped by the single nucleotide polymorphism variant sites significantly associated with the target risk task can be used as another part of the associated genes. The two parts of the associated genes are integrated to obtain the sample associated genes corresponding to each sample individual.
[0051] After determining the sample associated genes corresponding to each sample individual, for the multiple sample associated genes corresponding to each sample individual, the gene features corresponding to each sample associated gene can be looked up in the above-mentioned pre-constructed preset gene library, so as to obtain multiple sample associated gene features corresponding to each sample individual.
[0052] Step S303: Integrate the multiple sample associated gene features to obtain the sample individual features, and input the sample individual features into the prediction model to be trained for risk prediction to obtain the predicted risk score.
[0053] After determining the multiple sample associated gene features corresponding to each sample individual, the multiple sample associated gene features corresponding to each sample individual can be feature-integrated to obtain the sample individual features corresponding to each sample individual.
[0054] Among them, for feature integration of multiple sample-associated genes, specifically, methods such as simple averaging of corresponding dimensions, weighted averaging, or direct vector splicing can be used for integration. Alternatively, in some embodiments, a pooling-based aggregation method, a convolution operation-based aggregation method, an attention mechanism-based aggregation method, or a linear or non-linear dimensionality reduction aggregation method can be used. These aggregation methods will be introduced in detail below.
[0055] In this way, after obtaining the sample individual features corresponding to the sample individual, the sample individual features can be input into the prediction model to be trained for risk score estimation of the target risk task, and the predicted risk score corresponding to the sample individual predicted and output by the prediction model can be obtained.
[0056] Step S304, calculate the loss value according to the predicted risk score and the corresponding phenotypic information label, and update the parameters of the prediction model based on the loss value.
[0057] Furthermore, the loss value can be calculated based on the predicted risk score corresponding to the sample individual predicted and output by the prediction model and the phenotypic information label corresponding to the sample individual related to the target risk task. Specifically, the cross-entropy method can be used to calculate the loss value. After calculating the loss value, the backpropagation gradient can be determined based on the loss value, and then the gradient backpropagation process can be performed to update the parameters of the prediction model.
[0058] Among them, the process of training the prediction model can be to iteratively update the model parameters of the prediction model. Specifically, the training samples in the second training sample can be divided into multiple batches, and then the model parameters of the prediction model can be updated batch by batch. When the number of updated rounds reaches the preset number of rounds, or when it is detected that the change amplitude of the model parameters is less than the preset value, it can be determined that the training of the model has reached the convergence state, the model training is determined to end, and the finally obtained model parameters are used as the model parameters of the trained prediction model.
[0059] After training the prediction model using this method, the effect of the model can be further evaluated. Specifically, methods such as cross-validation can be used to evaluate the effect of the trained prediction model. When it is determined that the model effect is qualified after evaluating the model effect of the prediction model, the prediction model can be deployed online and used for predicting the risk score of the above-mentioned target risk task.
[0060] When it is necessary to use the prediction model deployed online to estimate the risk score of a target risk task for a certain individual to be tested, the genotype data of the individual to be tested can be obtained first, and then multiple associated genes of the individual to be tested can be determined according to the genotype data of the individual to be tested and the single nucleotide polymorphism data associated with the target risk task.
[0061] Among them, in some embodiments, determining multiple associated genes of an individual to be tested according to genotype data and single nucleotide polymorphism data associated with a target risk task includes: Obtaining first single nucleotide polymorphism data associated with the target risk task; Based on the results of statistical tests for allele frequency differences between risk individuals and reference individuals of the target risk task, screening the first single nucleotide polymorphism data to obtain multiple second single nucleotide polymorphism data; Determining multiple third single nucleotide polymorphism data based on genotype data, and determining multiple associated genes of the individual to be tested according to the multiple second single nucleotide polymorphism data and the multiple third single nucleotide polymorphism data.
[0062] Moreover, in some embodiments, determining multiple associated genes of the individual to be tested according to the multiple second single nucleotide polymorphism data and the multiple third single nucleotide polymorphism data includes: Obtaining the genomic positions and functional association information corresponding to the multiple second single nucleotide polymorphism data and the multiple third single nucleotide polymorphism data; Mapping the multiple second single nucleotide polymorphism data and the multiple third single nucleotide polymorphism data to the corresponding genes according to the genomic positions and functional association information to obtain multiple associated genes.
[0063] In the embodiments of the present disclosure, a method for obtaining complete associated genes from two dimensions is provided. Specifically, SNPs significantly associated with the target risk task (i.e., the first single nucleotide polymorphism data) can be obtained based on GWAS first. For example, when the target risk task is a cerebrovascular disease risk prediction task, SNP information related to cerebrovascular diseases can be obtained from the latest GWAS study first; or, all SNP information located in specific gene regions (such as coding regions, promoters). Then, statistical tests can be performed on the allele frequency differences between the diseased individuals and normal individuals (reference individuals) of the target risk task, so as to screen out the significantly associated SNP loci from them. Among them, the statistical test can be specifically implemented by using the χ² test or Fisher's exact test; when screening SNP loci, SNP loci with a P value less than 0.05 (i.e., the second single nucleotide polymorphism data) can be screened out. Further, bioinformatics tools such as SnpEff and VEP can be used to map these SNP loci to the corresponding genes based on their genomic positions and functional association information, ensuring the consistency of the mapping from the same SNP to genes among different individuals. In this way, the associated genes in the first dimension can be obtained.
[0064] In addition, SNP sites of the genotype (i.e., the third single nucleotide polymorphism data) can be determined based on the genotype data of an individual. Then, bioinformatics tools such as SnpEff and VEP can also be used. By comprehensively considering the physical location of the SNP (such as within a gene, promoter region, adjacent region) and functional association information (such as eQTL data), the SNP variant sites corresponding to the genotype are mapped to related genes, thereby obtaining the associated genes in the second dimension of the individual to be tested. Finally, the associated genes in these two dimensions are combined to obtain the associated genes of the individual to be tested. Among them, this standardized mapping process ensures that the method for determining associated genes remains consistent during the training and inference processes of the risk score prediction model, thereby improving the accuracy and reliability of risk score prediction.
[0065] Step S102, query the associated gene features corresponding to each associated gene in the pre-constructed gene feature library.
[0066] After determining multiple associated genes corresponding to the individual to be tested through the above method, the gene features of each associated gene can be further queried based on the pre-constructed gene feature library introduced in step S101, so as to obtain the gene features corresponding to each associated gene. That is, multiple associated gene features of the individual to be tested can be queried and obtained based on multiple associated genes in the gene feature library.
[0067] Specifically, the gene names and corresponding gene features in the gene feature library can be stored in the form of key-value pairs. Then, when querying, the key data corresponding to the associated genes of the individual to be tested can be input, so that the corresponding value data returned by the query can be obtained.
[0068] Step S103, perform feature integration on multiple associated gene features to obtain individual features.
[0069] After querying and obtaining multiple associated gene features corresponding to multiple associated genes of the individual to be tested, these associated gene features can be further integrated to generate a single feature vector related to the target risk task representing the individual to be tested.
[0070] In some embodiments, performing feature integration on multiple associated gene features to obtain individual features includes: Obtain the weight coefficients corresponding to multiple associated gene features; Perform weighted calculation on multiple associated gene features based on the weight coefficients to obtain individual features.
[0071] In an embodiment of the present application, a feature integration method is provided for obtaining individual features by weighted calculation of multiple associated gene features based on associated gene feature weight coefficients. Among them, for different associated genes, the degree of association with the target risk task is different, and different weights can be set for different associated gene features according to the difference in the degree of association. When the correlation between the associated gene and the target risk task is strong, a higher weight is set for the associated gene, and conversely, when the correlation between the associated gene and the target risk task is weak, a lower weight can be set for the associated gene, so that the correlation between the individual feature and the target risk task can be further improved, thereby improving the accuracy of the disease risk score prediction for the target risk task.
[0072] In some embodiments, multiple associated gene features are integrated to obtain individual features, including: Perform convolution processing on the feature sequence composed of multiple related gene features to obtain sequence features; Based on the sequence features, multiple associated gene features are integrated to obtain individual features.
[0073] In an embodiment of the present application, a feature integration method based on convolution operation is provided. Specifically, a feature sequence composed of multiple associated gene features can be convolved to obtain sequence features. The sequence features can include local associations and dynamic patterns between adjacent genes. Then, multiple associated gene features can be integrated based on the sequence features to obtain individual features.
[0074] In some embodiments, the feature integration of multiple associated gene features may also be performed by using a pooling-based aggregation operation. Specifically, a maximum pooling or mean pooling method may be used to sample and aggregate multiple associated gene features, thereby extracting the most representative local information of the genetic characteristics of the individual to be tested.
[0075] In some embodiments, feature integration of multiple associated gene features may also be performed using a feature integration method based on an attention mechanism. Specifically, weights may be adaptively assigned to different associated gene features based on the attention mechanism, thereby highlighting gene information that is more critical to the prediction result.
[0076] In some embodiments, feature integration of multiple associated gene features can also be performed using a feature integration method based on linear or nonlinear dimensionality reduction. Specifically, dimensionality reduction methods such as principal component analysis, linear transformation, and autoencoder can be combined to compress the high-dimensional feature vector aggregated by the aforementioned method, thereby extracting low-dimensional features that best reflect individual genetic information.
[0077] The above complex aggregation method can fully explore the internal connections between different gene features, enhance the discriminative ability of individual feature vectors, and provide more biologically significant feature support for the subsequent polygenic risk score model of the target risk task based on gene features.
[0078] Among them, it can be understood that the above embodiments provide a variety of feature integration methods. In order to ensure the accuracy of the gene risk score for the target risk task, when training and inferring the risk score prediction model corresponding to the target risk task, it is necessary to ensure that the same feature integration method is adopted. For example, when training the risk score prediction model corresponding to the target risk task, a feature integration method based on convolutional operations is used to integrate multiple sample-related features of the sample individual; then when using the trained risk score prediction model corresponding to the target risk task to predict the risk score of the individual to be tested for the target risk task, it is also necessary to use a feature integration method based on convolutional operations to perform feature integration on multiple associated gene features of the individual to be tested, so as to ensure the accuracy of the risk score prediction of the target risk task obtained by the prediction.
[0079] Step S104: Call the prediction model to process the individual features to obtain the risk score corresponding to the individual to be tested for the target risk task.
[0080] After integrating the individual features of the individual to be tested, the trained prediction model can be called to predict the risk score of the target risk task for the individual features, so as to obtain the risk score corresponding to the individual to be tested for the target risk task.
[0081] The gene data processing method provided by the present application uses low-dimensional dense (e.g., dozens to hundreds of dimensions) gene embedding vectors to replace high-dimensional sparse (tens of thousands) SNP features, greatly reducing the complexity of the risk score prediction model and reducing the demand for the sample size, thereby improving the training efficiency and generalization ability of the model.
[0082] In addition, the gene features (i.e., gene embedding vectors) themselves are learned from large-scale gene expression data and contain rich gene function, interaction, and pathway information. Using these embeddings as features can incorporate this biological context information into the PRS model, helping the model capture deeper genetic risk factors. Gene embeddings capture the complex relationships between genes and can better reflect the true biological effects than simple SNP accumulation. Therefore, the PRS model based on gene embedding features can greatly improve the accuracy of risk score prediction.
[0083] Although the embedding vectors themselves are abstract, potential pathogenic pathways or mechanisms can be explored by analyzing the genes and their functions corresponding to the embedding vectors that contribute greatly to the prediction results, providing clues for subsequent functional research.
[0084] Next, a specific example will be used to introduce the risk score prediction method provided in this application in detail. This example specifically provides a risk score prediction method for human cerebrovascular diseases.
[0085] First, construct a gene embedding vector library.
[0086] Data source selection: Use approximately 500,000 human bulk RNA-seq sample data from the ARCHS4 database.
[0087] Data preprocessing: Standardize the gene expression data of each sample. For example, calculate the expression Z-score of each gene in each sample. Then, within each training sample, sort the genes in descending order according to their Z-scores, and select a certain number of genes ranked at the top as the target training samples.
[0088] Model training: Select a 6-layer Transformer model based on the BERT architecture. The input is the gene identifiers corresponding to the sample data and their sorting information. The training task is set as follows: Randomly mask 15% of the gene identifiers in the input data, and predict the original gene identifiers of the masked genes based on the unmasked genes and their expression sorting information. Among them, the AdamW optimizer is used for model training.
[0089] Embedding vector extraction: After training convergence, extract the 32-dimensional embedding vectors corresponding to each gene in the input layer (or a certain hidden layer) of the model to construct a gene embedding vector library.
[0090] Then, train the risk score prediction model for human cerebrovascular diseases.
[0091] Sample data preparation: Use a synthetic dataset of SNPs significantly associated with the healthy population panel based on 1KG and cerebrovascular disease GWAS analysis: Integrate relevant information (basic variant information, OR value, allele frequency, etc.) from GWAS summary data, and select SNP loci that have been confirmed to be significantly associated with cerebrovascular diseases (a total of 80 SNP loci are extracted). Obtain the genotype information (phase1, v3) of the corresponding SNPs in the 1KG healthy population to construct training samples, and simulate case samples according to the genotype distribution data of SNPs in the GIGASTROKE cohort (for example, if the biallelic distribution of a certain SNP in this cohort is AA 50% \ AC 30% \ CC 20%, then refer to this information to calculate the probability and assign the phase information of this SNP to the simulated case sample, that is, 50% probability of 0|0, 30% probability of 1|0 or 0|1, 20% probability of 1|1).
[0092] Genotype and Gene Association Construction: Using public databases (such as dbSNP, GENCODE) and tools (such as VEP, ANNOVAR), the SNP sites and genetic variant sites of each individual are annotated to the genes where they are located or adjacent genes. At the same time, using the eQTL data of brain tissues in databases such as GTEx, SNPs significantly associated with gene expression can also be associated with the corresponding genes.
[0093] Sample Individual Feature Generation: For each sample individual, first identify all SNPs in its genome that are associated with the known risk of cerebrovascular diseases or all SNPs located in specific gene regions (such as coding regions, promoters). Then, according to the genotype and gene association relationship, determine the gene set corresponding to these SNPs. The genes associated with individual i are denoted as {Gene_A, Gene_B, Gene_C, ...}. Look up the corresponding embedding vectors,,... of Gene_A, Gene_B, Gene_C... from the gene embedding vector library. Then calculate the weighted average of these embedding vectors as the feature vector of individual i. The weights can be based on the genotype of the SNP (for example, homozygous risk allele is 2, heterozygous is 1, homozygous non-risk is 0) and / or the effect value (beta coefficient or ln(OR) value) of this SNP in GWAS studies. The formula is as follows: .
[0094] Model Training: Divide the training samples into a training set and a test set at a ratio of 80% / 20%. Then use the individual feature vectors of the training set and the corresponding cerebrovascular disease labels (1 indicates patient, 0 indicates control) to train a support vector machine (SVM) model.
[0095] Model Testing and Evaluation: For the individuals in the test set, generate their individual feature vectors, and then input them into the trained SVM model to obtain the predicted probability of cerebrovascular disease risk. Then calculate the performance metrics based on the predicted probability and the labels of the test set. The performance metrics include AUC (area under the ROC curve), confusion matrix, and F1 score. As Figure 4 shown, it is a schematic diagram of the area under the ROC curve of the risk score prediction model provided by this application in the independent test set of simulated data. As Figure 5 shown, it is a schematic diagram of the F1 score of the risk score prediction model provided by this application in the independent test set of simulated data. As Figure 6 shown, it is a schematic diagram of the confusion matrix of the risk score prediction model provided by this application in the independent test set of simulated data. As Figure 7As shown in the figure, it is a schematic diagram for evaluating the importance of each dimension feature of the gene vector. Generally, through five-fold cross-validation, the optimal hyperparameter model (kernel: linear, gamma: 1.5, C: 10) achieved good classification performance (AUC = 0.87, Max F1Score = 0.85); the confusion matrix (c) indicates the potential false positive risk of the model, and more caution is needed for the positive results of the model; the feature contribution degree (d) is obtained by randomly shuffling each feature value and then calculating the change degree of the classification performance of the model on the test set (the box plot represents the result distribution of randomly shuffled feature values), indicating that feature information such as the 16th dimension (Feature_17, starting from 0) of the gene vector is relatively important for the classification of the cerebrovascular disease risk. Considering that the dimension can be further reduced in the future to reduce the calculation cost and improve the model performance.
[0096] Finally, the cerebrovascular disease risk score prediction model is used to predict the cerebrovascular disease risk score.
[0097] The trained cerebrovascular disease risk score model can be deployed online. When predicting the cerebrovascular disease risk score of an individual, the genotype data of the individual can be obtained first, then the associated genes of the individual can be determined based on the genotype data, and then the corresponding gene embedding vectors can be queried in the gene embedding vector library based on the associated genes and integrated into an individual feature vector. Finally, the trained cerebrovascular disease risk score model is called to process the individual feature vector to obtain the cerebrovascular disease risk score of the individual output by the model.
[0098] Such as Figure 8a and 8b As shown in the figure, it is a schematic diagram for comparing the ROC and F1 scores of the solution provided by this application and the SNP score-based model in the related technology in the independent test set. Both show that this model has stronger classification performance and robustness. Among them, under the threshold condition of the related technology method (about <0.1), the true positive detection is 0, so the F1 value does not exist.
[0099] In some embodiments, this application can also provide a cerebrovascular disease risk prediction system based on cloud computing.
[0100] Backend implementation: Implement pre-computation of gene embeddings using the Python language and deep learning frameworks (such as TensorFlow or PyTorch), which can be completed on a high-performance computing cluster. Store the generated gene embedding vector library in an efficient key-value database (such as Redis) or vector database. Implement the logic of the feature generation module, model training module, and risk prediction module using Python and machine learning libraries (such as Scikit-learn, XGBoost). Use a relational database (such as PostgreSQL) to store user information, genotype data, phenotype data, and model parameters. Provide service interfaces through RESTful APIs.
[0101] Front-end implementation: Develop a web interface that allows authorized users to upload genotype data files of individuals (such as VCF format). The front-end calls the back-end API to submit data for analysis. The back-end performs feature generation and risk prediction and returns the results (risk scores, probabilities, confidence intervals, etc.) to the front-end. The front-end displays the prediction results in the form of charts or reports.
[0102] System workflow: The user logs in to the system, then uploads the genotype file of the tested individual. Then the system receives the file, then calls the feature generation module (queries the gene embedding library and calculates the individual feature vector), then calls the risk prediction module (loads the trained model and makes predictions), then returns the prediction results to the user interface, and finally the user views the risk report.
[0103] Refer to Figure 9 , in some embodiments, the embodiment of the present application further provides a gene data processing system 900, and the gene data processing system 900 includes: A data acquisition unit 910, configured to acquire genotype data of a tested individual, and determine multiple associated genes of the tested individual according to the genotype data and single nucleotide polymorphism data associated with the target risk task; A query unit 920, configured to query, in a pre-constructed gene feature library, the associated gene features corresponding to each associated gene, where the gene feature library includes embedding representations of different genes constructed based on an embedding model; the embedding model is obtained by performing self-supervised learning on multiple gene expression profiles of the same species as the tested individual; An integration unit 930, configured to perform feature integration on multiple associated gene features to obtain individual features; A prediction unit 940, configured to call a prediction model to process the individual features to obtain a risk score corresponding to the tested individual and the target risk task, where the prediction model is trained based on sample features of multiple sample individuals and corresponding phenotype information.
[0104] In some embodiments, the acquisition unit includes: An acquisition subunit, configured to acquire first single nucleotide polymorphism data associated with a target risk task; A screening subunit, configured to screen the first single nucleotide polymorphism data based on the statistical test result of the allele frequency difference between the risk individuals and the reference individuals of the target risk task to obtain a plurality of second single nucleotide polymorphism data; A determination subunit, configured to determine a plurality of third single nucleotide polymorphism data based on genotype data, and determine a plurality of associated genes of the individual to be tested according to the plurality of second single nucleotide polymorphism data and the plurality of third single nucleotide polymorphism data.
[0105] Optionally, in some embodiments, the determination subunit includes: An acquisition module, configured to acquire the genomic positions and functional association information corresponding to the plurality of second single nucleotide polymorphism data and the plurality of third single nucleotide polymorphism data; A mapping module, configured to map the plurality of second single nucleotide polymorphism data and the plurality of third single nucleotide polymorphism data to the corresponding genes according to the genomic positions and functional association information to obtain a plurality of associated genes.
[0106] Optionally, in some embodiments, the present application further provides a gene feature library construction device, including: An acquisition unit, configured to acquire a plurality of first training samples, where the first training samples include the gene expression profiles of human samples; A training unit, configured to perform self-supervised learning on the embedding model using the first training samples; A construction unit, configured to, when the embedding model training converges, construct a gene feature library based on the embedding representations of different genes extracted by the trained embedding model.
[0107] Optionally, in some embodiments, the training unit includes: A normalization subunit, configured to perform normalization processing on the gene expression profiles of the first training samples, and determine target training samples according to the normalization processing results; A training subunit, configured to perform self-supervised learning using random masking based on the target training samples to obtain an embedding model.
[0108] Optionally, in some embodiments, the integration unit includes: A second acquisition subunit, configured to acquire the weight coefficients corresponding to the plurality of associated gene features; A calculation subunit, configured to perform weighted calculation on the plurality of associated gene features based on the weight coefficients to obtain individual features.
[0109] Optionally, in some embodiments, the integration unit includes: A convolutional subunit, configured to perform convolutional processing on a feature sequence composed of multiple associated gene features to obtain sequence features; An integration subunit, configured to integrate multiple associated gene features based on the sequence features to obtain individual features.
[0110] Optionally, in some embodiments, the present application further provides a model training device, including: A third acquisition subunit, configured to acquire multiple second training samples, where the second training samples include sample genotype data of multiple sample individuals and corresponding phenotypic information labels; A lookup subunit, configured to determine the sample associated genes of each sample individual based on the sample genotype data, and look up the gene features corresponding to each sample associated gene in a gene feature library to obtain multiple sample associated gene features; A prediction subunit, configured to integrate multiple sample associated gene features to obtain sample individual features, and input the sample individual features into a prediction model to be trained for disease risk prediction to obtain a predicted risk score; An update subunit, configured to calculate a loss value according to the predicted risk score and the corresponding phenotypic information label, and update the parameters of the prediction model based on the loss value.
[0111] Refer to Figure 10 , Figure 10 illustrates the hardware structure of an electronic device in another embodiment. The electronic device includes: A processor 1001, which can be implemented in a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, etc., and is configured to execute relevant programs to implement the technical solutions provided in the embodiments of the present application; A memory 1002, which can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM), etc. The memory 1002 can store an operating system and other application programs. When implementing the technical solutions provided in the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 1002 and are called by the processor 1001 to execute the method for processing gene data in the embodiments of the present application; An input / output interface 1003, configured to implement information input and output; A communication interface 1004 is used to implement communication and interaction between this device and other devices. Communication can be achieved through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.); A bus 1005 transmits information between various components of the device (such as a processor 1001, a memory 1002, an input / output interface 1003, and a communication interface 1004); Among them, the processor 1001, the memory 1002, the input / output interface 1003, and the communication interface 1004 are communicatively connected to each other inside the device through the bus 1005.
[0112] The embodiment of the present application also provides a computer program product, which includes a computer program. The processor of the computer device reads and executes the computer program, so that the computer device executes the method for processing gene data as described above.
[0113] Terms such as "first", "second", "third", "fourth", etc. (if any) in the specification of the present disclosure and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances, so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "include" and "comprise" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.
[0114] It should be understood that in the present disclosure, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships can exist. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally represents an "or" relationship between the associated objects before and after. "At least one (one) of the following" or a similar expression refers to any combination of these items, including any combination of single item (one) or plural items (ones). For example, at least one (one) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.
[0115] It should be understood that in the description of the embodiments of the present application, the meaning of "a plurality of (or multiple)" is more than two. Understanding such as "greater than", "less than", "exceeding", etc. does not include the base number, and understanding such as "above", "below", "within", etc. includes the base number.
[0116] In several embodiments provided by the present disclosure, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the system embodiments described above are only illustrative. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, indirect couplings or communication connections of devices or units, and can be in electrical, mechanical, or other forms.
[0117] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place, or can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0118] In addition, each functional unit in various embodiments of the present disclosure can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.
[0119] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a storage medium. Based on this understanding, the technical solution of the present disclosure, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in various embodiments of the present disclosure. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs that can store program codes.
[0120] It should also be understood that the various implementation manners provided in the embodiments of the present application can be combined arbitrarily to achieve different technical effects.
[0121] The above is a specific description of the embodiments of the present disclosure. However, the present disclosure is not limited to the above embodiments. Those skilled in the art can also make various equivalent deformations or substitutions without departing from the spirit of the present disclosure, and these equivalent deformations or substitutions are all included within the scope defined by the claims of the present disclosure.
Claims
1. A method for processing gene data, characterized in that, The method includes: Obtaining genotype data of an individual to be tested, and determining multiple associated genes of the individual to be tested according to the genotype data and single nucleotide polymorphism data associated with a target risk task; Querying, in a pre-constructed gene feature library, the associated gene features corresponding to each of the associated genes, where the gene feature library includes embedding representations of different genes constructed based on an embedding model; the embedding model is obtained by performing self-supervised learning on multiple gene expression profiles of the same species as the individual to be tested; Performing feature integration on the multiple associated gene features to obtain an individual feature; Invoking a prediction model to process the individual feature to obtain a risk score corresponding to the individual to be tested and the target risk task, where the prediction model is trained based on sample features of multiple sample individuals and corresponding phenotypic information.
2. The method according to claim 1, characterized in that, The determining of the multiple associated genes of the individual to be tested according to the genotype data and the single nucleotide polymorphism data associated with the target risk task includes: Obtaining first single nucleotide polymorphism data associated with the target risk task; Based on the statistical test results of the allele frequency differences between risk individuals and reference individuals of the target risk task, screening the first single nucleotide polymorphism data to obtain multiple second single nucleotide polymorphism data; Determining multiple third single nucleotide polymorphism data based on the genotype data, and determining the multiple associated genes of the individual to be tested according to the multiple second single nucleotide polymorphism data and the multiple third single nucleotide polymorphism data.
3. The method according to claim 2, wherein The determining of the multiple associated genes of the individual to be tested according to the multiple second single nucleotide polymorphism data and the multiple third single nucleotide polymorphism data includes: Obtaining the genomic positions and functional association information corresponding to the multiple second single nucleotide polymorphism data and the multiple third single nucleotide polymorphism data; Mapping the multiple second single nucleotide polymorphism data and the multiple third single nucleotide polymorphism data to corresponding genes according to the genomic positions and the functional association information to obtain multiple associated genes.
4. The method according to claim 1, wherein The process of pre-constructing the gene feature library includes: Obtaining multiple first training samples, where the first training samples include gene expression profiles of human samples; Performing self-supervised learning on the embedding model using the first training samples; When the embedding model converges in training, constructing a gene feature library based on the embedding representations of different genes extracted from the trained embedding model.
5. The method according to claim 4, characterized in that, The performing of self-supervised learning on the embedding model using the first training samples includes: Performing standardization processing on the gene expression profiles of the first training samples, and determining target training samples according to the standardization processing results; Based on the target training samples, performing self-supervised learning using random masking to obtain the embedding model.
6. The method according to claim 1, wherein The performing of feature integration on the multiple associated gene features to obtain an individual feature includes: Obtaining the weight coefficients corresponding to the multiple associated gene features; Performing weighted calculation on the multiple associated gene features based on the weight coefficients to obtain an individual feature.
7. The method according to claim 1, wherein The performing of feature integration on the multiple associated gene features to obtain an individual feature includes: Perform convolution processing on the feature sequence composed of the multiple associated gene features to obtain sequence features; Integrate the multiple associated gene features based on the sequence features to obtain individual features.
8. The method according to claim 1, characterized in that The training process of the prediction model includes: Obtain a plurality of second training samples, where the second training samples include the sample genotype data of a plurality of sample individuals and the corresponding phenotypic information labels; Determine the sample associated genes of each sample individual based on the sample genotype data, and look up the gene features corresponding to each sample associated gene in the gene feature library to obtain a plurality of sample associated gene features; Integrate the multiple sample associated gene features to obtain sample individual features, and input the sample individual features into the prediction model to be trained for risk prediction to obtain a predicted risk score; Calculate a loss value according to the predicted risk score and the corresponding phenotypic information label, and update the parameters of the prediction model based on the loss value.
9. A processing system for gene data, characterized in that, The gene data processing system includes: A data acquisition unit, configured to acquire the genotype data of the individual to be tested, and determine multiple associated genes of the individual to be tested according to the genotype data and the single nucleotide polymorphism data associated with the target risk task; A query unit, configured to query the associated gene features corresponding to each associated gene in a pre-constructed feature library, where the gene feature library includes the embedding representations of different genes constructed based on an embedding model; the embedding model is obtained by performing self-supervised learning on the gene expression profiles of multiple individuals of the same species as the individual to be tested; An integration unit, configured to perform feature integration on the multiple associated gene features to obtain individual features; A prediction unit, configured to call a prediction model to process the individual features to obtain a risk score corresponding to the individual to be tested and the target risk task, where the prediction model is trained based on the sample features of multiple sample individuals and the corresponding phenotypic information.
10. An electronic device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the gene data processing method according to any one of claims 1 to 8.
11. A storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the gene data processing method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Method and system for predicting disease risk
CN114373547A
Methods and compositions for estimating or predicting genotypes and phenotypes
CN116895334A
Disease-associated gene identification method based on multi-brain-region multi-level gene regulatory network
CN118155718A
Heat recovering ventilation apparatus
KR1020250047428A
Methods and systems for mutation signature attribution
WO2024118594A1
Cited By
Gene data processing method and device, storage medium and electronic equipment
CN120526857A