Omics Data Processing Method, Device, Electronic Device, and Storage Medium

By determining the association relationship between genes in the omics data and extracting features using graph neural network models, the high-dimensional and noise problems of omics data are solved, improving the accuracy of cancer typing and survival analysis.

CN114429786BActive Publication Date: 2025-06-27BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111649901.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-30
Publication Date
2025-06-27
Estimated Expiration
2041-12-30

AI Technical Summary

Technical Problem

The high-dimensional, high noise and batch effects of omics data lead to difficulties in practical applications, especially in improving the accuracy of cancer typing and survival analysis.

Method used

By obtaining the omics data, the association relationship between multiple genes is determined, the graph data is determined based on the gene expression level and association relationship, the graph neural network model is used to encode the graph data, and the characteristics of the omics data are extracted, so as to perform the target task of the omics data.

Benefits of technology

It achieves low-dimensional representations of accurate omics data, improving the accuracy of downstream applications such as cancer typing and survival analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114429786B_ABST
    Figure CN114429786B_ABST
Patent Text Reader

Abstract

The present disclosure provides an omics data processing method, apparatus, electronic device, and storage medium, which relate to the field of artificial intelligence technologies, specifically to the fields of deep learning and intelligent medical technologies. The specific implementation solution is as follows: obtain omics data, where the omics data includes multiple genes; determine the association relationships between the multiple genes; determine graph data according to the expression levels of the multiple genes in the omics data and the association relationships; and based on the graph data, determine the features of the omics data, so as to execute the target task of the omics data according to the features of the omics data. Thus, an accurate low-dimensional representation of the omics data can be obtained, and further, the accuracy of downstream target tasks such as cancer typing classification tasks and individual survival analysis tasks can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technologies, specifically to the fields of deep learning and intelligent medical technologies, and particularly relates to an omics data processing method, apparatus, electronic device, and storage medium. Background Art

[0002] With the development of high-throughput sequencing technology, omics data is increasingly used in modern medicine. Since omics data can comprehensively depict the health status of patients, it is widely applied in disease diagnosis, medication, and other aspects.

[0003] However, characteristics such as the high-dimensionality, high-noise, and batch effects of omics data bring many difficulties to practical applications. Obtaining a low-dimensional representation of accurate omics data is of great significance for improving the accuracy of downstream applications such as cancer subtype classification and individual survival analysis. Summary of the Invention

[0004] The present disclosure provides an omics data processing method, apparatus, electronic device, and storage medium.

[0005] According to one aspect of the present disclosure, an omics data processing method is provided. The method includes: obtaining omics data, where the omics data includes multiple genes; determining the association relationship between the multiple genes; determining graph data according to the expression levels of the multiple genes in the omics data and the association relationship; and determining the features of the omics data based on the graph data, so as to perform the target task of the omics data according to the features of the omics data.

[0006] According to another aspect of the present disclosure, a model training method for omics data processing is provided. The method includes: obtaining training omics data, where the training omics data includes multiple genes, and there is an association relationship between at least two of the genes; determining the association relationship between the multiple genes; determining training graph data according to the expression levels of the multiple genes in the training omics data and the association relationship; adjusting the training graph data using at least two data augmentation strategies to obtain at least two augmented graph data; encoding the at least two augmented graph data using a graph neural network model to obtain corresponding features; and adjusting the model parameters of the neural network model according to the differences between the features of the at least two augmented graph data to minimize the differences.

[0007] According to another aspect of the present disclosure, there is provided an omics data processing device, the device including: a first acquisition module for acquiring omics data, the omics data including a plurality of genes; a first determination module for determining the association relationship between the plurality of genes; a second determination module for determining graph data according to the expression levels of the plurality of genes in the omics data and the association relationship; and a third determination module for determining the characteristics of the omics data based on the graph data to perform the target task of the omics data according to the characteristics of the omics data.

[0008] According to another aspect of the present disclosure, there is provided a model training device for omics data processing, including: a second acquisition module for acquiring training omics data, the training omics data including a plurality of genes; a fourth determination module for determining the association relationship between the plurality of genes; a fifth determination module for determining training graph data according to the expression levels of the plurality of genes in the training omics data and the association relationship; a first adjustment module for adjusting the training graph data by using at least two data augmentation strategies to obtain at least two augmented graph data; an encoding module for encoding the at least two augmented graph data by using a graph neural network model to obtain corresponding features; and a second adjustment module for adjusting the model parameters of the neural network model according to the differences between the features of the at least two augmented graph data to minimize the differences.

[0009] According to another aspect of the present disclosure, there is provided an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the omics data processing method of the present disclosure or execute the model training method for omics data processing of the present disclosure.

[0010] According to another aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, the computer instructions being used to cause the computer to execute the omics data processing method disclosed in the embodiments of the present disclosure or execute the model training method for omics data processing disclosed in the embodiments of the present disclosure.

[0011] According to another aspect of the present disclosure, there is provided a computer program product, including a computer program, the computer program realizing the steps of the omics data processing method of the present disclosure or realizing the steps of the model training method for omics data processing of the present disclosure when executed by a processor.

[0012] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. Description of the Drawings

[0013] The drawings are used to better understand the present solution and do not constitute a limitation to the present disclosure. Among them:

[0014] Figure 1 is a schematic flowchart of a omics data processing method according to the first embodiment of the present disclosure;

[0015] Figure 2 is a schematic flowchart of a omics data processing method according to the second embodiment of the present disclosure;

[0016] Figure 3 is a schematic flowchart of a model training method for omics data processing according to the third embodiment of the present disclosure;

[0017] Figure 4 is a schematic architecture diagram of a model training method for omics data processing according to the third embodiment of the present disclosure;

[0018] Figure 5 is a schematic flowchart of a model training method for omics data processing according to the fourth embodiment of the present disclosure;

[0019] Figure 6 is a schematic structural diagram of an omics data processing device according to the fifth embodiment of the present disclosure;

[0020] Figure 7 is a schematic structural diagram of a model training device for omics data processing according to the sixth embodiment of the present disclosure;

[0021] Figure 8 is a block diagram of an electronic device for implementing the omics data processing method or the model training method for omics data processing of the embodiments of the present disclosure. Detailed Embodiments

[0022] The following describes exemplary embodiments of the present disclosure in conjunction with the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding and should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, descriptions of well-known functions and structures are omitted below for clarity and conciseness.

[0023] The omics data processing method, model training method for omics data processing, device, electronic device, non-transitory computer-readable storage medium, and computer program product provided by the present disclosure relate to the field of artificial intelligence technology, specifically to the fields of deep learning and intelligent healthcare technology.

[0024] Among them, artificial intelligence is a discipline that studies how to make a computer simulate certain human thinking processes and intelligent behaviors (such as learning, reasoning, thinking, planning, etc.), and it has both hardware-level technologies and software-level technologies. Artificial intelligence hardware technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, and big data processing; artificial intelligence software technologies mainly include several major directions such as computer vision, speech recognition technology, natural language processing technology, and machine learning / deep learning, big data processing technology, and knowledge graph technology.

[0025] Intelligent healthcare is to build a regional medical information platform for health records, and use the most advanced Internet of Things technology to achieve the interaction between patients, medical staff, medical institutions, and medical devices, and gradually achieve informatization. In the near future, the medical industry will integrate more high-tech such as artificial intelligence and sensing technology, making medical services truly intelligent and promoting the prosperous development of the medical cause.

[0026] Currently, the characteristics of high-dimensionality, high-noise, batch effects, etc. of omics data bring many difficulties to practical applications. Obtaining a low-dimensional representation of accurate omics data is of great significance for improving the accuracy of downstream applications such as cancer type classification and survival analysis.

[0027] The present disclosure provides an omics data processing method. By obtaining omics data, which includes multiple genes, determining the association relationships between the multiple genes, determining graph data based on the expression levels of the multiple genes and the association relationships in the omics data, and determining the features of the omics data based on the graph data, so as to perform the target task of the omics data according to the features of the omics data, it is possible to obtain a low-dimensional representation of accurate omics data, and further improve the accuracy of downstream target tasks such as cancer typing and survival analysis of patients based on omics data.

[0028] The following describes the omics data processing method, model training method for omics data processing, device, electronic device, non-transitory computer-readable storage medium, and computer program product of the embodiments of the present disclosure with reference to the accompanying drawings.

[0029] Figure 1It is a schematic flowchart of an omics data processing method according to the first embodiment of the present disclosure. It should be noted that for the omics data processing method in this embodiment, the execution subject is an omics data processing device, which can be implemented by software and / or hardware. The omics data processing device can be configured in an electronic device, which can include but is not limited to a terminal device, a server, etc. This embodiment does not make specific limitations on the electronic device.

[0030] As Figure 1 shown, the omics data processing method may include:

[0031] Step 101, obtain omics data, where the omics data includes multiple genes.

[0032] Among them, the omics data is the omics data to be processed, and the omics data includes multiple genes.

[0033] Step 102, determine the association relationship between multiple genes.

[0034] Step 103, determine graph data according to the expression levels of multiple genes in the omics data and the association relationship.

[0035] Among them, gene expression refers to the process in which cells, during the life process, transform the genetic information stored in the DNA (deoxyribonucleic acid) sequence through transcription and translation into biologically active protein molecules. The simultaneous expression of two or more genes is gene co-expression. Correspondingly, there is a co-expression relationship between these two or more genes.

[0036] Among them, the expression level of a gene is a quantitative value of gene expression.

[0037] In the embodiment of the present disclosure, when it is determined that there is a co-expression relationship between two genes, it can be determined that there is an association relationship between these two genes. Furthermore, graph data can be determined according to the expression levels of multiple genes in the omics data and the association relationship.

[0038] Step 104, determine the characteristics of the omics data based on the graph data, so as to execute the target task of the omics data according to the characteristics of the omics data.

[0039] Among them, the characteristics of the omics data are the low-dimensional omics representations of the omics data, that is, low-dimensional representations.

[0040] The target task can be any downstream task such as an individual survival analysis task, a disease diagnosis task, a medication recommendation task, a cancer typing classification task, etc. The present disclosure does not limit this.

[0041] In the embodiments of the present disclosure, after determining graph data based on the expression levels of multiple genes in omics data and the association relationships, feature extraction can be performed on the graph data to determine the features of the omics data, so as to perform the target task of the omics data according to the features of the omics data.

[0042] Since the graph data is determined based on the expression levels of multiple genes in the omics data and the association relationships between at least two genes, by performing feature extraction on the graph data, it is possible to take into account the expression levels of multiple genes in the omics data and the correlation features between multiple genes, so as to fully extract features from the omics data and obtain an accurate low-dimensional representation of the omics data. Thereby, the expression ability of the low-dimensional representation of the omics data is improved, so that the state of the patient can be better described. Furthermore, based on the obtained low-dimensional representation of the omics data, the target task of the omics data is performed, and the accuracy of the target task can be improved.

[0043] The omics data processing method of the embodiments of the present disclosure can obtain an accurate low-dimensional representation of the omics data by acquiring omics data, where the omics data includes multiple genes, determining the association relationships between the multiple genes, determining graph data based on the expression levels of the multiple genes in the omics data and the association relationships, and determining the features of the omics data based on the graph data, so as to perform the target task of the omics data according to the features of the omics data, and further improve the accuracy of downstream target tasks such as cancer typing classification tasks and individual survival analysis tasks.

[0044] Through the above analysis, it can be seen that in the embodiments of the present disclosure, graph data can be determined based on the expression levels of multiple genes in the omics data and the association relationships, and then based on the graph data, the features of the omics data can be determined, so as to perform the target task of the omics data according to the features of the omics data. Among them, the graph data may include the attributes of multiple nodes in the graph and the edges connecting the nodes. The following combines Figure 2 to further illustrate the process of determining the attributes of multiple nodes included in the graph data and the edges connecting the nodes, and determining the features of the omics data based on the graph data in the omics data processing method provided by the present disclosure according to the expression levels of multiple genes in the omics data and the association relationships.

[0045] Figure 2 is a schematic flowchart of the omics data processing method according to the second embodiment of the present disclosure. As Figure 2 shown, the omics data processing method may include the following steps:

[0046] Step 201: Acquire omics data, where the omics data includes multiple genes.

[0047] Step 202: Determine the association relationships between the multiple genes.

[0048] As a possible implementation, the following method can be used to determine the association relationship between at least two genes in omics data: query the PPI network (protein-protein interaction network) based on the proteins synthesized by multiple genes to obtain the interaction relationship between the proteins synthesized by at least two genes; determine the association relationship between at least two genes according to the interaction relationship between the proteins synthesized by at least two genes.

[0049] It can be understood that in the PPI network, each node corresponds to a protein, and there is a connection edge between two nodes, indicating that there is an interaction relationship between the proteins corresponding to the two nodes. In the embodiments of the present disclosure, various formats such as the names and sequences of the proteins synthesized by multiple genes in the omics data can be used to query the PPI network, so as to determine the interaction relationship between the proteins corresponding to at least two nodes according to the connection edges between at least two nodes in the PPI network, and then determine the association relationship between the genes synthesizing the at least two proteins according to the interaction relationship between the proteins corresponding to the at least two nodes.

[0050] For example, assume that based on the proteins synthesized by multiple genes in the omics data, the PPI network is queried, and it is determined that there is a connection edge between node a corresponding to the protein synthesized by gene A and node b corresponding to the protein synthesized by gene B, and there is a connection edge between node c corresponding to the protein synthesized by gene C and node d corresponding to the protein synthesized by gene D, that is, there is an interaction relationship between the protein synthesized by gene A and the protein synthesized by gene B, and there is an interaction relationship between the protein synthesized by gene C and the protein synthesized by gene D. Then, according to the interaction relationship between the protein synthesized by gene A and the protein synthesized by gene B, and the interaction relationship between the protein synthesized by gene C and the protein synthesized by gene D, it can be determined that there is an association relationship between gene A and gene B, and there is an association relationship between gene C and gene D.

[0051] By determining the association relationship between at least two genes based on the PPI network, the association relationship between each gene in the omics data can be accurately determined.

[0052] As another possible implementation, the following method can be used to determine the association relationship between at least two genes in omics data: count the number of times that the change trends of the expression levels of the first gene and the second gene are the same in N comparisons; in the case where the number of times is greater than a preset threshold, determine that there is an association relationship between the first gene and the second gene, where the preset threshold is less than N.

[0053] Among them, the first gene and the second gene are any two genes included in the omics data. N and the preset threshold can be set arbitrarily as needed, and the preset threshold is less than N. For example, N can be set to 100, and the preset threshold can be set to 60, 70, etc. The present disclosure does not limit this.

[0054] In the embodiments of the present disclosure, for any two genes in the omics data, such as the first gene and the second gene, the change trend of the expression level of the first gene can be compared with the change trend of the second gene N times. Among the N comparisons, when the number of times that the change trend of the expression level of the first gene is the same as the change trend of the second gene exceeds the preset threshold, it can be determined that there is a co-expression relationship between the first gene and the second gene. Correspondingly, it can be determined that there is an association relationship between the first gene and the second gene.

[0055] For example, assume that N is 100 and the preset threshold is 60. Then, among 100 comparisons, if the number of times that the expression level of the first gene increases and the expression level of the second gene also increases is 70 times, it can be determined that there is a co-expression relationship between the first gene and the second gene. Correspondingly, it can be determined that there is an association relationship between the first gene and the second gene.

[0056] By determining the association relationship between at least two genes based on the expression levels of the at least two genes, the association relationship between each gene in the omics data is accurately determined.

[0057] In another possible implementation form, the above two methods can also be combined to jointly determine the association relationship between at least two genes in the omics data, so as to improve the accuracy of the determined association relationship between each gene in the omics data. That is, for any two genes in the omics data, the above two methods can be used to respectively determine whether there is an association relationship between the two genes. When it is determined that there is an association relationship between the two genes by any of the above methods, it can be determined that there is an association relationship between the two genes.

[0058] Step 203: Determine the attributes of the nodes corresponding to the genes in the figure according to the expression levels of multiple genes in the omics data.

[0059] In the embodiments of the present disclosure, the graph data can be determined according to the expression levels and association relationships of multiple genes in the omics data. Among them, the graph data includes the attributes of multiple nodes in the graph and the edges connecting the nodes. Each node in the graph corresponds to a gene, and the structure of the graph can be determined according to the connecting edges between the nodes. The connecting edge between two nodes is the edge connecting the two nodes.

[0060] In the embodiments of the present disclosure, the expression levels of each gene in the omics data can be used as the attributes of the corresponding nodes in the graph.

[0061] Step 204: Determine the edges connecting the corresponding nodes in the graph according to the association relationships between at least two genes.

[0062] In the embodiments of the present disclosure, when there is an association relationship between any two genes, it can be determined that there is a connection relationship between the nodes corresponding to these two genes, and thus the edges connecting these two nodes can be determined. According to the connecting edges between the nodes in the graph, the structure of the graph can be determined.

[0063] Step 205: Determine the characteristics of the omics data based on the attributes of multiple nodes in the graph and the edges connecting the nodes, so as to perform the target task of the omics data according to the characteristics of the omics data.

[0064] In the embodiments of the present disclosure, after determining the attributes of the nodes corresponding to the genes in the graph according to the expression levels of multiple genes in the omics data, and determining the edges connecting the corresponding nodes in the graph according to the association relationships between at least two genes, the graph data including the attributes of multiple nodes in the graph and the edges connecting the nodes can be determined. Furthermore, feature extraction can be performed on the graph data to determine the characteristics of the omics data, so as to perform the target task of the omics data according to the characteristics of the omics data.

[0065] Since the attributes of multiple nodes in the graph data are determined according to the expression levels of the corresponding genes in the omics data, and the edges connecting the nodes are determined according to the association relationships between the corresponding genes, by performing feature extraction on the graph data, the expression levels of multiple genes and the correlation characteristics between multiple genes in the omics data can be taken into account, so as to extract more effective characteristics of the omics data from the omics data and obtain a precise low-dimensional representation of the omics data.

[0066] In the embodiments of the present disclosure, to determine the characteristics of the omics data based on the graph data, it can be specifically implemented in the following manner: Use a graph neural network model to encode the graph data to obtain the characteristics of the omics data.

[0067] Among them, the graph neural network model can be any graph neural network model capable of realizing feature extraction, such as neural network models such as GCN (Graph Convolutional Network) and GAT (Graph Attention Network). The present disclosure places no restrictions on this.

[0068] In the embodiments of the present disclosure, a graph neural network model can be pre-trained by means of contrastive learning. The input of the graph neural network model is graph data determined according to the expression levels and association relationships of multiple genes in omics data, and the output is the features of the omics data. Therefore, after obtaining the graph data, the graph data can be input into the trained graph neural network model, and the graph neural network model can be used to extract the features of the omics data to obtain the features of the omics data.

[0069] Among them, the training process of the graph neural network model can refer to the following embodiments and will not be elaborated here.

[0070] After determining the graph data according to the expression levels and association relationships of multiple genes in the omics data, the graph neural network model is used to encode the graph data, realizing the extraction of more effective omics data features from the omics data by using the graph neural network model and obtaining a precise low-dimensional representation of the omics data.

[0071] The omics data processing method of the embodiments of the present disclosure includes obtaining omics data, which includes multiple genes, determining the association relationships between the multiple genes, determining the attributes of the nodes corresponding to the genes in the graph according to the expression levels of the multiple genes in the omics data, determining the edges connecting the corresponding nodes in the graph according to the association relationships between at least two genes, and determining the features of the omics data based on the attributes of the multiple nodes in the graph and the edges connecting the nodes, so as to execute the target task of the omics data according to the features of the omics data, and can obtain a precise low-dimensional representation of the omics data, thereby improving the accuracy of downstream target tasks such as cancer typing classification tasks and individual survival analysis tasks.

[0072] According to an embodiment of the present disclosure, there is also provided a model training method for omics data processing.

[0073] Figure 3 is a schematic flowchart of a model training method for omics data processing according to the third embodiment of the present disclosure.

[0074] Among them, it should be noted that the model training method for omics data processing provided by the embodiments of the present disclosure is executed by a model training device for omics data processing, hereinafter referred to as the model training device for short. The model training device can be implemented by software and / or hardware. The model training device can be configured in an electronic device, and the electronic device can include but is not limited to a terminal device, a server, etc. The embodiment does not make a specific limitation on the electronic device.

[0075] As Figure 3 shown, the model training method for omics data processing may include the following steps:

[0076] Step 301, obtain training omics data, where the training omics data includes multiple genes.

[0077] Step 302, determine the association relationship between multiple genes.

[0078] Step 303, determine the training graph data according to the expression levels of multiple genes in the training omics data and the association relationship.

[0079] Among them, when two or more genes are expressed simultaneously, it is gene co-expression. Correspondingly, there is a co-expression relationship between the two or more genes.

[0080] Among them, the expression level of a gene is a quantitative value of gene expression.

[0081] In the embodiments of the present disclosure, when it is determined that there is a co-expression relationship between two genes, it can be determined that there is an association relationship between the two genes. Furthermore, the training graph data can be determined according to the expression levels of multiple genes in the training omics data and the association relationship.

[0082] Step 304, adjust the training graph data by using at least two data augmentation strategies to obtain at least two augmented graph data.

[0083] The data augmentation strategy is a strategy for data augmentation of the training graph data. The data augmentation strategy can be set as needed, and the present disclosure does not limit this.

[0084] In the embodiments of the present disclosure, at least two data augmentation strategies can be used to perform data augmentation on the training graph data to obtain at least two augmented graph data. Among them, when each data augmentation strategy is used to perform data augmentation on the training graph data, one corresponding augmented graph data is obtained.

[0085] Step 305, encode at least two augmented graph data by using a graph neural network model to obtain corresponding features.

[0086] Among them, the graph neural network model can be any graph neural network model capable of feature extraction, such as neural network models like GCN and GAT. The present disclosure does not limit this.

[0087] In the embodiments of the present disclosure, for each augmented graph data, a graph neural network model can be used to encode the augmented graph data to obtain the features corresponding to the augmented graph data. Among them, the features corresponding to the augmented graph data are the low-dimensional omics representations of the augmented graph data, that is, the low-dimensional representations. Encoding the augmented graph data means extracting features from the augmented graph data.

[0088] Step 306, adjust the model parameters of the neural network model according to the differences between the features of at least two augmented graph data to minimize the differences.

[0089] In the embodiments of the present disclosure, contrastive learning can be used to train the model. Specifically, after obtaining the features corresponding to at least two augmented graph data, the contrastive learning loss function can be jointly calculated based on the differences between the features corresponding to the at least two augmented graph data, and the model parameters of the graph neural network model can be adjusted to minimize the contrastive learning loss function, so as to optimize the model parameters of the graph neural network model. By repeatedly performing the steps of 304-306 multiple times to optimize the model parameters of the graph neural network model multiple times, the trained graph neural network model can be obtained.

[0090] Among them, the model parameters of the graph neural network model can be optimized by any optimization method such as SGD (Stochastic Gradient Descent) and BGD (Batch Gradient Descent) methods, and the present disclosure does not limit this.

[0091] Among them, the contrastive learning loss function can be an InfoNCE (Info Noise-contrastive estimation) loss function, a BARLOW TWINS (a self-supervised learning method) loss function, or other loss functions, and can be selected according to the contrastive learning method used, and the present disclosure does not limit this.

[0092] Taking the BARLOW TWINS loss function as an example, by increasing the batch size, a loss function can be constructed based on the correlation matrix, so as to learn the differences and correlations between omics data samples in the latent space. Among them, an omics data sample can be understood as a kind of augmented graph data.

[0093] Taking the example of using two data augmentation strategies to adjust the training graph data to obtain two augmented graph data and training the model based on the two augmented graph data, refer to Figure 4The architecture diagram shown. In the embodiments of the present disclosure, a two-tower model can be adopted. The model includes a graph data construction module, a data augmentation module, a graph data encoding module, and a loss function calculation module. Among them, the graph data construction module can determine the association relationships of multiple genes in the training omics data 401, and determine the training graph data 402 according to the expression levels of multiple genes in the training omics data 401 and the association relationships between at least two genes. The data augmentation module can adopt two data augmentation strategies to perform data augmentation on the input training graph data 402 to obtain two augmented graph data 403 and 404. The graph data encoding module is implemented by the graph neural network model 405. Input the augmented graph data 403 into the graph neural network model 405 to extract features from the augmented graph data 403 by using the graph neural network model 405, and low-dimensional features 406 corresponding to the augmented graph data 403 can be obtained. Similarly, input the augmented graph data 404 into the graph neural network model 405 to extract features from the augmented graph data 404 by using the graph neural network model 405, and low-dimensional features 407 corresponding to the augmented graph data 404 can be obtained. The loss function calculation module, after obtaining the features 406 and 407 output by the graph data encoding module, can calculate the contrastive learning loss function according to the difference between the features 406 of the augmented graph data 403 and the features 407 of the augmented graph data 404, and optimize the model parameters of the graph neural network model 405 by adjusting the model parameters of the graph neural network model 405 to minimize the contrastive learning loss function. By repeatedly adopting the two data augmentation strategies to perform data augmentation on the training graph data 402 and optimizing the model parameters of the graph neural network model 405 multiple times based on the differences between the features of the two augmented graph data, the optimal parameters of the graph neural network model 405 can be obtained, and the training of the graph neural network model 405 can be completed.

[0094] It should be noted that the trained graph neural network model in the embodiments of the present disclosure can be used to encode the obtained graph data to obtain the features of the corresponding omics data. The process of using the trained graph neural network model to execute the above steps can refer to the description of the embodiments of the above omics data processing method, which will not be elaborated here.

[0095] By simultaneously learning in a manner combining contrastive learning and a graph neural network model, the differences between at least two augmented graph data and the correlations between genes can be learned, enabling the trained graph neural network model to fully extract features from the graph data, obtain an accurate low-dimensional representation of the omics data, improve the expression ability of the low-dimensional representation of the omics data, and thus be able to better describe the patient's state. Furthermore, based on the obtained low-dimensional representation of the omics data, the target task of the omics data can be executed, and the accuracy of the target task can be improved.

[0096] In summary, the model training method for omics data processing provided by the embodiments of the present disclosure obtains training omics data, where the training omics data includes multiple genes, determines the association relationships between the multiple genes, determines training graph data according to the expression levels of the multiple genes in the training omics data and the association relationships, adjusts the training graph data using at least two data augmentation strategies to obtain at least two augmented graph data, encodes the at least two augmented graph data using a graph neural network model to obtain corresponding features, and adjusts the model parameters of the neural network model according to the differences between the features of the at least two augmented graph data to minimize the differences. Based on the training omics data, the graph neural network model is trained to obtain a graph neural network model for omics data processing. Using the trained graph neural network model to process the graph data determined based on the omics data can fully extract features from the omics data, obtain a precise low-dimensional representation of the omics data, and further improve the accuracy of downstream target tasks such as cancer subtyping classification tasks and individual survival analysis tasks.

[0097] The following further describes the model training device for omics data processing provided by the present disclosure in conjunction with Figure 5 ,

[0098] Figure 5 is a schematic flowchart of the model training method for omics data processing according to the fourth embodiment of the present disclosure.

[0099] As Figure 5 shown, the model training method for omics data processing may include the following steps:

[0100] Step 501, obtain training omics data, where the training omics data includes multiple genes.

[0101] Step 502, determine the association relationships between the multiple genes.

[0102] Among them, the method for determining the association relationships between the multiple genes in the training omics data may refer to the method for determining the association relationships between the multiple genes in the omics data in the above embodiments, which will not be elaborated here.

[0103] Step 503, determine the attributes of the nodes corresponding to the genes in the training graph according to the expression levels of the multiple genes in the training omics data.

[0104] In the embodiments of the present disclosure, the training graph data may be determined according to the expression levels of the multiple genes in the training omics data and the association relationships. Among them, the training graph data includes the attributes of multiple nodes in the training graph and the edges connecting the nodes. Each node in the training graph corresponds to a gene, and the structure of the training graph can be determined according to the connecting edges between the nodes. The connecting edge between two nodes is the edge connecting these two nodes.

[0105] In the embodiments of the present disclosure, the expression levels of each gene in the training omics data can be used as the attributes of the corresponding nodes in the training graph.

[0106] Step 504: Determine the edges connecting the corresponding nodes in the training graph according to the association relationship between at least two genes.

[0107] In the embodiments of the present disclosure, when there is an association relationship between any two genes, it can be determined that there is a connection relationship between the corresponding nodes of these two genes, so that the edges connecting these two nodes can be determined. According to the connecting edges between the nodes in the training graph, the structure of the training graph can be determined.

[0108] Step 505: Adjust the training graph data by using at least two data augmentation strategies to obtain at least two augmented graph data, where the training graph data includes the attributes of multiple nodes in the training graph and the edges connecting the nodes.

[0109] In the embodiments of the present disclosure, after determining the attributes of the nodes corresponding to the genes in the training graph according to the expression levels of multiple genes in the training omics data, and determining the edges connecting the corresponding nodes in the training graph according to the association relationship between at least two genes, the training graph data including the attributes of multiple nodes in the training graph and the edges connecting the nodes can be determined. Furthermore, at least two data augmentation strategies can be used to adjust the training graph data to obtain at least two augmented graph data.

[0110] Since the attributes of multiple nodes in the training graph data are determined according to the expression levels of the corresponding genes in the training omics data, and the edges connecting the nodes are determined according to the association relationship between the corresponding genes, by adjusting the training graph data to obtain at least two augmented graph data, and then training the graph neural network model based on the at least two augmented graph data, the graph neural network model can better learn the correlation between genes in the omics data. Therefore, when encoding the graph data by using the trained graph data network model, more effective omics data features can be extracted from the graph data, and a precise low-dimensional representation of the omics data can be obtained.

[0111] As a possible implementation, data augmentation of the training graph data by using at least two data augmentation strategies can be achieved in the following way: Mask the expression levels of at least one node in the training graph data by using at least two data augmentation strategies to obtain at least two augmented graph data.

[0112] Among them, at least two data augmentation strategies can both be masking the expression levels of at least one node in the training graph data, but the masking positions corresponding to different data augmentation strategies are different, that is, different data augmentation strategies mask the expression levels of different nodes. Thus, by adopting at least two data augmentation strategies to mask the expression levels of at least one node in the training graph data, since different data augmentation strategies mask the expression levels of different nodes, at least two augmented graph data can be obtained.

[0113] As another possible implementation, adopting at least two data augmentation strategies to perform data augmentation on the training graph data can be achieved in the following way: adopting at least two data augmentation strategies to add noise to the expression levels of at least one node in the training graph data to obtain at least two augmented graph data.

[0114] Among them, at least two data augmentation strategies can both be adding noise to the expression levels of at least one node in the training graph data, but the noise addition methods corresponding to different data augmentation strategies are different. For example, different data augmentation strategies add noise to the expression levels of different nodes, or different data augmentation strategies add noise with different amplitudes to the expression levels of the same node, or different data augmentation strategies add noise to the expression levels of different nodes and the amplitudes of the added noise are different. Thus, by adopting at least two data augmentation strategies to add noise to the expression levels of at least one node in the training graph data, since there are differences between different data augmentation strategies, at least two augmented graph data can be obtained.

[0115] By adopting at least two data augmentation strategies to mask the expression levels of at least one node in the training graph data, or adopting at least two data augmentation strategies to add noise to the expression levels of at least one node in the training graph data, data augmentation of the training graph data is realized, and at least two augmented graph data are obtained. Furthermore, the anti-interference ability of the graph neural network model trained using the augmented graph data after data augmentation is enhanced when encoding the graph data corresponding to the omics data. Thus, the trained graph neural network model can obtain a more accurate low-dimensional representation of the omics data.

[0116] Step 506: Use the graph neural network model to encode at least two augmented graph data to obtain corresponding features.

[0117] Step 507: According to the differences between the features of at least two augmented graph data, adjust the model parameters of the neural network model to minimize the differences.

[0118] Among them, the specific implementation process and principle of steps 506-507 can refer to the description of the above embodiments and will not be elaborated here.

[0119] In summary, the model training method for omics data processing provided by the embodiments of the present disclosure obtains training omics data, where the training omics data includes multiple genes, determines the association relationships between the multiple genes, determines the attributes of the nodes corresponding to the genes in the training graph according to the expression levels of the multiple genes in the training omics data, determines the edges connecting the corresponding nodes in the training graph according to the association relationships between at least two genes, adjusts the training graph data using at least two data augmentation strategies to obtain at least two augmented graph data, where the training graph data includes the attributes of multiple nodes in the training graph and the edges connecting the nodes, encodes the at least two augmented graph data using a graph neural network model to obtain corresponding features, encodes the at least two augmented graph data using a graph neural network model to obtain corresponding features, realizes training of the graph neural network model based on the training graph data to obtain a graph neural network model for omics data processing, and uses the trained graph neural network model to process the graph data determined based on the omics data, which can fully extract features from the omics data to obtain an accurate low-dimensional representation of the omics data, thereby improving the accuracy of downstream target tasks such as cancer subtyping classification tasks and individual survival analysis tasks.

[0120] Next, in combination with Figure 6 , the omics data processing device provided by the present disclosure will be described.

[0121] Figure 6 FIG. is a schematic structural diagram of an omics data processing device according to the fifth embodiment of the present disclosure.

[0122] As Figure 6 shown, the omics data processing device 600 provided by the present disclosure includes: a first acquisition module 601, a first determination module 602, a second determination module 603, and a third determination module 604.

[0123] Among them, the first acquisition module 601 is configured to acquire omics data, where the omics data includes multiple genes;

[0124] The first determination module 602 is configured to determine the association relationships between the multiple genes;

[0125] The second determination module 603 is configured to determine graph data according to the expression levels of the multiple genes in the omics data and the association relationships;

[0126] The third determination module 604 is configured to determine the features of the omics data based on the graph data, so as to execute the target task of the omics data according to the features of the omics data.

[0127] It should be noted that the omics data processing device 600 provided in this embodiment can execute the omics data processing method of the foregoing embodiment. Among them, the omics data processing device 600 can be implemented by software and / or hardware. The omics data processing device 600 can be configured in an electronic device, which can include but is not limited to a terminal device, a server, etc. This embodiment does not make specific limitations on the electronic device.

[0128] As a possible implementation manner of the embodiment of the present disclosure, the graph data, including the attributes of multiple nodes in the graph and the edges connecting the nodes, and the second determination module 603 includes:

[0129] A first determination unit, configured to determine the attributes of the nodes corresponding to the genes in the graph according to the expression levels of multiple genes in the omics data;

[0130] A second determination unit, configured to determine the edges connecting the corresponding nodes in the graph according to the association relationship between at least two genes.

[0131] As a possible implementation manner of the embodiment of the present disclosure, the first determination module 602 includes:

[0132] A query unit, configured to query a protein-protein interaction (PPI) network according to the proteins synthesized by multiple genes to obtain the interaction relationship between the proteins synthesized by at least two genes;

[0133] A third determination unit, configured to determine the association relationship between at least two genes according to the interaction relationship between the proteins synthesized by at least two genes.

[0134] As a possible implementation manner of the embodiment of the present disclosure, the multiple genes include a first gene and a second gene; the first determination module 602 includes:

[0135] A statistics unit, configured to count the number of times that the change trends of the expression levels of the first gene and the second gene are the same in N comparisons;

[0136] A fourth determination unit, configured to determine that there is an association relationship between the first gene and the second gene when the number of times is greater than a preset threshold, where the preset threshold is less than N.

[0137] As a possible implementation manner of the embodiment of the present disclosure, the third determination module 604 includes:

[0138] An encoding unit, configured to encode the graph data by using a graph neural network model to obtain the features of the omics data.

[0139] It should be noted that the foregoing description of the embodiment of the omics data processing method also applies to the omics data processing device provided in the present disclosure, and details are not described herein again.

[0140] The omics data processing device provided by the embodiments of the present disclosure can obtain omics data which includes multiple genes, determine the association relationships between the multiple genes, determine graph data based on the expression levels of the multiple genes in the omics data and the association relationships, and determine the characteristics of the omics data based on the graph data, so as to perform the target task of the omics data according to the characteristics of the omics data. It can obtain an accurate low-dimensional representation of the omics data, thereby improving the accuracy of downstream target tasks such as cancer subtyping classification tasks and individual survival analysis tasks.

[0141] According to an embodiment of the present disclosure, there is also provided a model training device for omics data processing.

[0142] Next, in conjunction with Figure 7 , the model training device for omics data processing provided by the present disclosure will be described.

[0143] Figure 7 FIG. is a schematic structural diagram of a model training device for omics data processing according to the sixth embodiment of the present disclosure.

[0144] As Figure 7 shown, the model training device 700 for omics data processing provided by the present disclosure includes: a second acquisition module 701, a fourth determination module 702, a fifth determination module 703, a first adjustment module 704, an encoding module 705, and a second adjustment module 706.

[0145] Among them, the second acquisition module 701 is configured to acquire training omics data, and the training omics data includes multiple genes;

[0146] The fourth determination module 702 is configured to determine the association relationships between the multiple genes;

[0147] The fifth determination module 703 is configured to determine training graph data according to the expression levels of the multiple genes in the training omics data and the association relationships;

[0148] The first adjustment module 704 is configured to adjust the training graph data by using at least two data augmentation strategies to obtain at least two augmented graph data;

[0149] The encoding module 705 is configured to encode at least two augmented graph data by using a graph neural network model to obtain corresponding features;

[0150] The second adjustment module 706 is configured to adjust the model parameters of the neural network model according to the differences between the features of at least two augmented graph data to minimize the differences.

[0151] It should be noted that the model training device 700 for omics data processing provided in this embodiment, hereinafter referred to as the model training device, can execute the model training method for omics data processing in the foregoing embodiment. Among them, the model training device can be implemented by software and / or hardware, and the model training device can be configured in an electronic device, which can include but is not limited to a terminal device, a server, etc. This embodiment does not make a specific limitation on the electronic device.

[0152] As a possible implementation manner of the embodiment of the present disclosure, the training graph data includes the attributes of multiple nodes in the graph and the edges connecting the nodes. The fifth determination module 703 includes:

[0153] A fifth determination unit, configured to determine the attributes of the nodes corresponding to the genes in the training graph according to the expression levels of multiple genes in the training omics data;

[0154] A sixth determination unit, configured to determine the edges connecting the corresponding nodes in the training graph according to the association relationship between at least two genes.

[0155] As a possible implementation manner of the embodiment of the present disclosure, the first adjustment module 704 includes:

[0156] A masking unit, configured to mask the expression levels of at least one node in the training graph data by using at least two data augmentation strategies to obtain at least two enhanced graph data.

[0157] As a possible implementation manner of the embodiment of the present disclosure, the first adjustment module 704 includes:

[0158] A processing unit, configured to add noise to the expression levels of at least one node in the training graph data by using at least two data augmentation strategies to obtain at least two enhanced graph data.

[0159] It should be noted that the foregoing description of the embodiment of the model training method for omics data processing is also applicable to the model training device for omics data processing provided in the present disclosure, and will not be repeated here.

[0160] The model training device for omics data processing provided by the embodiments of the present disclosure obtains training omics data, where the training omics data includes multiple genes, determines the association relationships between the multiple genes, determines training graph data according to the expression levels of the multiple genes in the training omics data and the association relationships, adjusts the training graph data by using at least two data augmentation strategies to obtain at least two augmented graph data, encodes the at least two augmented graph data by using a graph neural network model to obtain corresponding features, and adjusts the model parameters of the neural network model according to the differences between the features of the at least two augmented graph data to minimize the differences. Based on the training omics data, the graph neural network model is trained to obtain a graph neural network model for omics data processing. Using the trained graph neural network model to process the graph data determined based on the omics data can fully extract features from the omics data, obtain a precise low-dimensional representation of the omics data, and further improve the accuracy of downstream target tasks such as cancer typing classification tasks and individual survival analysis tasks.

[0161] Based on the above embodiments, the present disclosure further provides an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the omics data processing method of the present disclosure, or execute the model training method for omics data processing of the present disclosure.

[0162] Based on the above embodiments, the present disclosure further provides a non-transitory computer-readable storage medium storing computer instructions, where the computer instructions are used to cause the computer to execute the omics data processing method disclosed in the embodiments of the present disclosure, or execute the model training method for omics data processing disclosed in the embodiments of the present disclosure.

[0163] Based on the above embodiments, the present disclosure further provides a computer program product, including a computer program, where when the computer program is executed by a processor, the steps of the omics data processing method of the present disclosure are implemented, or the steps of the model training method for omics data processing of the present disclosure are implemented.

[0164] According to the embodiments of the present disclosure, the present disclosure further provides an electronic device, a readable storage medium, and a computer program product.

[0165] Figure 8FIG. shows a schematic block diagram of an exemplary electronic device 800 that can be used to implement embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, personal digital processors, cellular telephones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely exemplary and are not intended to limit the implementations of the present disclosure described and / or claimed herein.

[0166] As Figure 8 shown, the electronic device 800 may include a computing unit 801 that can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 802 or a computer program loaded from a storage unit 808 into a random access memory (RAM) 803. In the RAM 803, various programs and data required for the operation of the device 800 can also be stored. The computing unit 801, the ROM 802, and the RAM 803 are connected to each other via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.

[0167] Multiple components in the device 800 are connected to the I / O interface 805, including: an input unit 806, such as a keyboard, a mouse, etc.; an output unit 807, such as various types of displays, speakers, etc.; a storage unit 808, such as a magnetic disk, an optical disk, etc.; and a communication unit 809, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 809 allows the device 800 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0168] The computing unit 801 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 801 executes the various methods and processes described above, such as the omics data processing method or the model training method for omics data processing. For example, in some embodiments, the omics data processing method or the model training method for omics data processing can be implemented as a computer software program, which is tangibly included in a machine-readable medium, such as the storage unit 808. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 800 via the ROM 802 and / or the communication unit 809. When the computer program is loaded into the RAM 803 and executed by the computing unit 801, one or more steps of the omics data processing method or the model training method for omics data processing described above can be executed. Alternatively, in other embodiments, the computing unit 801 can be configured to execute the omics data processing method or the model training method for omics data processing in any other suitable manner (e.g., by means of firmware).

[0169] Various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGA), application-specific integrated circuits (ASIC), application-specific standard products (ASSP), systems-on-a-chip (SOC), complex programmable logic devices (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs, which can be executed and / or interpreted on a programmable system including at least one programmable processor, the programmable processor can be a special or general programmable processor, can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0170] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general purpose computer, a special purpose computer, or other programmable data processing device, such that the program codes, when executed by the processor or controller, cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0171] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0172] In order to provide interaction with a user, the systems and techniques described herein can be implemented on a computer that has: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, speech input, or tactile input).

[0173] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), the Internet, and blockchain networks.

[0174] A computer system can include a client and a server. The client and the server are generally remote from each other and typically interact through a communication network. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, and solves the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services ("Virtual Private Server", or simply "VPS" for short). The server can be a cloud server, or a server of a distributed system, or a server combined with blockchain.

[0175] It should be understood that various forms of the processes shown above can be used, steps can be reordered, added, or deleted. For example, the steps recited in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and no limitation is imposed herein.

[0176] The above specific embodiments do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the protection scope of this disclosure.

Claims

1. An omics data processing method, comprising: Obtaining omics data, where the omics data includes multiple genes; Determining the association relationship between the multiple genes; Determining graph data according to the expression levels of the multiple genes in the omics data and the association relationship; Based on the graph data, determining the characteristics of the omics data to perform the target task of the omics data according to the characteristics of the omics data; Wherein, the multiple genes include a first gene and a second gene; the determining the association relationship between the multiple genes includes: Counting the number of times that the change trends of the expression levels of the first gene and the second gene are the same in N comparisons; When the number of times is greater than a preset threshold, determining that the first gene and the second gene have the association relationship, where the preset threshold is less than N.

2. The method according to claim 1, wherein, The graph data includes the attributes of multiple nodes in the graph and the edges connecting the nodes; The determining graph data according to the expression levels of the multiple genes in the omics data and the association relationship includes: Determining the attributes of the nodes corresponding to the genes in the graph according to the expression levels of the multiple genes in the omics data; Determining the edges connecting the corresponding nodes in the graph according to the association relationship between at least two of the genes.

3. The method according to claim 1, wherein The determining the association relationship between the multiple genes includes: Querying a protein-protein interaction (PPI) network according to the proteins synthesized by the multiple genes to obtain the interaction relationship between the proteins synthesized by at least two genes; Determining the association relationship between the at least two genes according to the interaction relationship between the proteins synthesized by the at least two genes.

4. The method according to any one of claims 1-3, wherein, The determining the characteristics of the omics data based on the graph data includes: Encoding the graph data using a graph neural network model to obtain the characteristics of the omics data.

5. A model training method for omics data processing, comprising: Obtaining training omics data, where the training omics data includes multiple genes, and there is an association relationship between at least two of the genes; Determining the association relationship between the multiple genes; Determining training graph data according to the expression levels of the multiple genes in the training omics data and the association relationship; Adjusting the training graph data using at least two data augmentation strategies to obtain at least two augmented graph data; Encoding the at least two augmented graph data using a graph neural network model to obtain corresponding characteristics; Adjusting the model parameters of the neural network model according to the differences between the characteristics of the at least two augmented graph data to minimize the differences; Wherein, the multiple genes include a first gene and a second gene; the determining the association relationship between the multiple genes includes: Counting the number of times that the change trends of the expression levels of the first gene and the second gene are the same in N comparisons; When the number of times is greater than a preset threshold, determining that the first gene and the second gene have the association relationship, where the preset threshold is less than N.

6. The method according to claim 5, wherein The training graph data includes the attributes of multiple nodes in the training graph and the edges connecting the nodes; Determining training graph data according to the expression levels of multiple genes in the training omics data and the association relationship includes: Determining the attributes of the nodes corresponding to the genes in the training graph according to the expression levels of multiple genes in the training omics data; Determining the edges connecting the corresponding nodes in the training graph according to the association relationship between at least two genes.

7. The method according to claim 6, wherein Adjusting the training graph data by using at least two data augmentation strategies to obtain at least two enhanced graph data, including: Using at least two data augmentation strategies to mask the expression levels of at least one node in the training graph data to obtain at least two enhanced graph data.

8. The method according to claim 6, wherein Adjusting the training graph data by using at least two data augmentation strategies to obtain at least two enhanced graph data, including: Using at least two data augmentation strategies to add noise to the expression levels of at least one node in the training graph data to obtain at least two enhanced graph data.

9. An omics data processing device, including: A first acquisition module for acquiring omics data, where the omics data includes multiple genes; A first determination module for determining the association relationship between multiple genes; A second determination module for determining graph data according to the expression levels of multiple genes in the omics data and the association relationship; A third determination module for determining the characteristics of the omics data based on the graph data to perform the target task of the omics data according to the characteristics of the omics data; Wherein, the multiple genes include a first gene and a second gene; the first determination module includes: A statistics unit for counting the number of times that the change trends of the expression levels of the first gene and the second gene are the same in N comparisons; A fourth determination unit for determining that there is the association relationship between the first gene and the second gene when the number of times is greater than a preset threshold, where the preset threshold is less than N.

10. The apparatus according to claim 9, wherein The graph data includes the attributes of multiple nodes in the graph and the edges connecting the nodes; The second determination module includes: A first determination unit for determining the attributes of the nodes corresponding to the genes in the graph according to the expression levels of multiple genes in the omics data; A second determination unit for determining the edges connecting the corresponding nodes in the graph according to the association relationship between at least two genes.

11. The device according to claim 9, wherein, The first determination module includes: A query unit for querying a protein-protein interaction (PPI) network according to the proteins synthesized by multiple genes to obtain the interaction relationship between the proteins synthesized by at least two genes; A third determination unit for determining the association relationship between the at least two genes according to the interaction relationship between the proteins synthesized by the at least two genes.

12. The device according to any one of claims 9 to 11, wherein, The third determination module includes: An encoding unit for encoding the graph data by using a graph neural network model to obtain the characteristics of the omics data.

13. A model training device for omics data processing, including: A second acquisition module for acquiring training omics data, where the training omics data includes multiple genes; A fourth determination module, configured to determine the association relationship between the multiple genes; A fifth determination module, configured to determine training graph data according to the expression levels of the multiple genes in the training omics data and the association relationship; A first adjustment module, configured to adjust the training graph data by using at least two data augmentation strategies to obtain at least two augmented graph data; An encoding module, configured to encode the at least two augmented graph data by using a graph neural network model to obtain corresponding features; A second adjustment module, configured to adjust the model parameters of the neural network model according to the differences between the features of the at least two augmented graph data to minimize the differences; Wherein, the multiple genes include a first gene and a second gene; determining the association relationship between the multiple genes includes: Counting the number of times that the change trends of the expression levels of the first gene and the second gene are the same in N comparisons; When the number is greater than a preset threshold, determining that the first gene and the second gene have the association relationship, where the preset threshold is less than N.

14. The apparatus according to claim 13, wherein, The training graph data includes the attributes of multiple nodes in the training graph and the edges connecting the nodes; The fifth determination module includes: A fifth determination unit, configured to determine the attributes of the nodes corresponding to the genes in the training graph according to the expression levels of the multiple genes in the training omics data; A sixth determination unit, configured to determine the edges connecting the corresponding nodes in the training graph according to the association relationship between at least two of the genes.

15. The apparatus according to claim 14, wherein, The first adjustment module includes: A masking unit, configured to mask the expression levels of at least one node in the training graph data by using at least two data augmentation strategies to obtain at least two augmented graph data.

16. The apparatus according to claim 14, wherein The first adjustment module includes: A processing unit, configured to add noise to the expression levels of at least one node in the training graph data by using at least two data augmentation strategies to obtain at least two augmented graph data.

17. An electronic device, comprising: At least one processor; And A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method according to any one of claims 1-4, or execute the method according to any one of claims 5-8.

18. A non-transitory computer-readable storage medium storing computer instructions, where the computer instructions are used to cause the computer to execute the method according to any one of claims 1-4, or execute the method according to any one of claims 5-8.

19. A computer program product, comprising a computer program, where the computer program, when executed by a processor, implements the steps of the method according to any one of claims 1-4, or implements the steps of the method according to any one of claims 5-8.

Citation Information

Patent Citations

  • Omics data processing method and device based on a graph neural network, equipment and medium

    CN112364880A