Training of feature fusion models, classification methods and devices for cancer users, and media.

CN116451172BActive Publication Date: 2026-08-14BOE TECHNOLOGY GROUP CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-14
Publication Date
2026-08-14

AI Technical Summary

Technical Problem

[0004]本公开的目的在于提供一种特征融合模型的训练方法、特征融合模型的训练装置、癌症用户的分类方法、计算机可读存储介质以及电子设备,进而至少在一定程度上克服由于相关技术的限制和缺陷而导致的无法对组学数据以及领域知识进行融合的问题

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116451172B_ABST
    Figure CN116451172B_ABST
Patent Text Reader

Abstract

This disclosure relates to a method, apparatus, and medium for training a feature fusion model and classifying cancer users, belonging to the field of machine learning technology. The method includes: acquiring historical patient sample data of historical users, and extracting first omics data of a first gene point of the historical user from the historical patient sample data; acquiring domain knowledge, and constructing a heterogeneous network based on the first gene point, the first omics data, and the domain knowledge; training a network model to be trained according to the heterogeneous network to obtain a feature fusion model; wherein the feature fusion model is used to perform feature fusion on the first omics data based on domain knowledge and to perform data prediction based on the fused features. This disclosure realizes the fusion of omics data and domain knowledge.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of machine learning technology, and more specifically, to a training method for a feature fusion model, a training device for a feature fusion model, a classification method for cancer users, a computer-readable storage medium, and an electronic device. Background Technology

[0002] Existing methods can combine user-defined prior biological information to achieve nonlinear combinations of variables from different omics datasets. However, they cannot fuse omics data with domain knowledge.

[0003] It should be noted that the information in the background section above is only used to enhance the understanding of the background of this disclosure, and therefore may include information that does not constitute prior art known to those skilled in the art. Summary of the Invention

[0004] The purpose of this disclosure is to provide a training method for a feature fusion model, a training device for a feature fusion model, a classification method for cancer users, a computer-readable storage medium, and an electronic device, thereby overcoming, to at least some extent, the problem of the inability to fuse omics data and domain knowledge due to the limitations and defects of related technologies.

[0005] According to one aspect of this disclosure, a method for training a feature fusion model is provided, comprising:

[0006] Acquire historical patient sample data of historical users, and extract the first omics data of the first gene point of the historical users from the historical patient sample data;

[0007] Acquire domain knowledge and construct a heterogeneous network based on the first gene point, the first omics data, and the domain knowledge;

[0008] Based on the heterogeneous network, the network model to be trained is trained to obtain a feature fusion model; wherein, the feature fusion model is used to perform feature fusion on the first omics data based on domain knowledge and to perform data prediction based on the fused features.

[0009] In one exemplary embodiment of this disclosure, the domain knowledge includes biological signaling pathways and current gene points included in the biological signaling pathways;

[0010] The first set of omics data includes one or more of the following: DNA methylation data, gene mutation SNV data, copy number variation CNV data, and gene expression data.

[0011] In one exemplary embodiment of this disclosure, the heterogeneous network includes a first heterogeneous network and / or a second heterogeneous network;

[0012] The first heterogeneous network is a heterogeneous network with a first gene point as the first child node, the first omics data of the first gene point as the first child node as the node feature, the first connection relationship between current gene points in the biological signaling pathway as the first connection edge, and the second connection relationship between the first gene pairs composed of the first gene points as the second connection edge.

[0013] The second heterogeneous network is a heterogeneous network with a first gene point as the first child node, the first omics data of the first gene point as the second child node, the first connection relationship between current gene points in the biological signaling pathway as the first connection edge, the second connection relationship between the first gene pairs composed of the first gene points as the second connection edge, and the third connection relationship between the first child node and the second child node as the third connection edge.

[0014] In one exemplary embodiment of this disclosure, the first heterogeneous network is used to reflect the correlation between data within the same omics dataset, as well as the correlation between multiple omics datasets;

[0015] The second heterogeneous network is used to reflect the correlation between data within the same omics dataset, as well as the correlation between multiple omics datasets.

[0016] In one exemplary embodiment of this disclosure, the network model to be trained includes a first network model to be trained and / or a second network model to be trained;

[0017] The first network model to be trained includes a graph network model and a first classifier, and the second network model to be trained includes an autoencoder model and a second classifier.

[0018] The graph network model includes a relational graph convolutional network model and / or a graph attention network model.

[0019] In one exemplary embodiment of this disclosure, a heterogeneous network is constructed based on the first gene point, the first omics data, and the domain knowledge, including:

[0020] Construct a first set of child nodes based on the first gene point, and construct a first set of node features based on the first omics data of the first gene point;

[0021] Based on the biological signaling pathways included in the domain knowledge and the current gene points included in the biological signaling pathways, determine the first connection relationship between the first gene points, and construct a first set of connection edges based on the first connection relationship;

[0022] Calculate the first correlation coefficient between the first gene pairs consisting of the first gene points, and determine the second connection relationship between the first gene points based on the first correlation coefficient;

[0023] A second set of connecting edges is constructed based on the second connection relationship, and a first heterogeneous network is constructed based on the first set of child nodes, the first set of node features, the first set of connecting edges, and the second set of connecting edges.

[0024] In one exemplary embodiment of this disclosure, constructing a heterogeneous network based on the first gene point, the first omics data, and the domain knowledge further includes:

[0025] Construct a first set of child nodes based on the first gene point, and construct a second set of child nodes based on the first omics data of the first gene point;

[0026] Based on the biological signaling pathways included in the domain knowledge and the current gene points included in the biological signaling pathways, determine the first connection relationship between the first gene points, and construct a first set of connection edges based on the first connection relationship;

[0027] Calculate the first correlation coefficient between the first gene pairs consisting of the first gene points, and determine the second connection relationship between the first gene points based on the first correlation coefficient;

[0028] A second set of connecting edges is constructed based on the second connectivity relationship, and a third set of connecting edges is constructed based on the third connectivity relationship between the first gene point and the first omics data.

[0029] Construct a second heterogeneous network based on the first set of child nodes, the second set of child nodes, the first set of connecting edges, the second set of connecting edges, and the third set of connecting edges.

[0030] In one exemplary embodiment of this disclosure, the first gene pair consisting of the first gene point includes a first sub-gene point and a second sub-gene point;

[0031] The calculation of the first correlation coefficient between the first gene pairs formed by the first gene points includes:

[0032] Obtain the first sub-omics data of the first sub-gene point in the first gene pair consisting of the first gene point, and the second sub-omics data of the second sub-gene point;

[0033] First sub-gene expression data are extracted from the first sub-omics data, and second sub-gene expression data are extracted from the second sub-omics data.

[0034] The first correlation coefficient is calculated based on the expression data of the first daughter gene and the expression data of the second daughter gene.

[0035] In one exemplary embodiment of this disclosure, a feature fusion model is trained based on the heterogeneous network to obtain the network model to be trained, including:

[0036] Based on the first heterogeneous network, train the first network model to be trained to obtain the first feature fusion model; and / or

[0037] Based on the second heterogeneous network, the second network model to be trained is trained to obtain the second feature fusion model.

[0038] In one exemplary embodiment of this disclosure, a first network model to be trained is trained based on a first heterogeneous network to obtain a first feature fusion model, including:

[0039] The first heterogeneous network is input into the graph network model in the first network model to be trained to obtain the first feature representation corresponding to the first heterogeneous network, and the first feature representation is input into the first classifier in the first network model to be trained to obtain the first predicted label.

[0040] Based on the real user tags of the historical users and the first predicted tags, a first target loss function is constructed, and the first network model to be trained is trained based on the first target loss function to obtain a first feature fusion model.

[0041] In one exemplary embodiment of this disclosure, a second network model to be trained is trained according to a second heterogeneous network to obtain a second feature fusion model, including:

[0042] The second heterogeneous network is subjected to representation learning to obtain a second feature representation corresponding to the second heterogeneous network, and the first omics data of the first gene point is input into the encoding module of the second network model to be trained to obtain the first reconstructed data corresponding to the first omics data.

[0043] Based on the first set of data, the first reconstructed data, and the second feature representation, a second objective loss function is constructed, and the autoencoder model in the second network model to be trained is trained according to the second objective function to obtain the trained encoding model;

[0044] The first omics data is input into the encoding module of the trained encoding model to obtain the second reconstructed data, and the second reconstructed data is input into the second classifier of the second network model to be trained to obtain the second predicted label;

[0045] Based on the second predicted label and the real user labels of the historical users, the second classifier is trained to obtain the second feature fusion model.

[0046] In one exemplary embodiment of this disclosure, the first omics data of the first gene point is input into the encoding module of the second network model to be trained to obtain the first reconstructed data corresponding to the first omics data, including:

[0047] The first omics data is mapped using the first nonlinear function in the encoding module of the second network model to be trained, to obtain intermediate variables;

[0048] The intermediate variables are mapped by the second nonlinear function in the encoding module of the second network model to be trained, and the first reconstructed data corresponding to the first omics data is obtained; wherein the data expression form of the first reconstructed data is consistent with the data expression form of the first omics data.

[0049] In one exemplary embodiment of this disclosure, a second objective loss function is constructed based on first omics data, first reconstructed data, and a second feature representation, including:

[0050] The first reconstructed data is input into the decoding module of the second network model to be trained to obtain the third feature representation corresponding to the second heterogeneous network, and the first sub-loss function is constructed based on the first omics data and the first reconstructed data.

[0051] A second sub-loss function is constructed based on the second and third feature representations, and a second objective function is constructed based on the first and second sub-loss functions.

[0052] In one exemplary embodiment of this disclosure, constructing a first sub-loss function based on first omics data and first reconstructed data includes:

[0053] Calculate the absolute value of the first difference between the first omics data and the first reconstructed data, and square the absolute value of the first difference to obtain the first square calculation result;

[0054] The first sub-loss function is obtained by summing the results of the first squared calculation.

[0055] In one exemplary embodiment of this disclosure, constructing a second sub-loss function based on a second feature representation and a third feature representation includes:

[0056] Obtain the first sub-vector representation of the first sub-gene point and the second sub-vector representation of the second sub-gene point in the first gene pair included in the second feature representation, and calculate the first similarity of the first gene pair based on the first sub-vector representation and the second sub-vector representation;

[0057] Obtain the third sub-vector representation of the first sub-gene point and the fourth sub-vector representation of the second sub-gene point in the first gene pair included in the third feature representation, and calculate the second similarity of the first gene pair based on the third sub-vector representation and the fourth sub-vector representation.

[0058] Calculate the absolute value of the second difference between the first similarity and the second similarity, and sum the absolute values ​​of the second difference to obtain the second sub-loss function.

[0059] According to one aspect of this disclosure, a method for classifying cancer patients is provided, comprising:

[0060] Obtain the current sample data of the current user, and extract the second omics data of the second gene point of the current user from the current sample data;

[0061] The second omics data of the second gene point is input into the feature fusion model to obtain the data prediction result; wherein, the feature fusion model is trained based on the feature fusion model training method described above.

[0062] Based on the data prediction results, the user category to which the current user belongs is determined.

[0063] In one exemplary embodiment of this disclosure, the second omics data of the second gene locus is input into a feature fusion model to obtain data prediction results, including:

[0064] The second omics data of the second gene point is input into the graph network model in the first feature fusion model to obtain a first user feature representation including the dimensionality-reduced second omics data and domain knowledge. The first user feature representation is then input into the first classifier of the first feature fusion model to obtain the data prediction result; and / or

[0065] The second omics data of the second gene point is input into the encoding module of the autoencoding model in the second feature fusion model to obtain the second user feature representation, which includes the dimensionality-reduced second omics data and domain knowledge. The second user feature representation is then input into the second classifier of the second feature fusion model to obtain the data prediction result.

[0066] According to one aspect of this disclosure, a training apparatus for a feature fusion model is provided, comprising:

[0067] The first omics data extraction module is used to obtain historical patient sample data of historical users and extract the first omics data of the first gene point of the historical users from the historical patient sample data.

[0068] A heterogeneous network construction module is used to acquire domain knowledge and construct a heterogeneous network based on the first gene point, the first omics data, and the domain knowledge.

[0069] The model training module is used to train the network model to be trained based on the heterogeneous network to obtain a feature fusion model; wherein, the feature fusion model is used to perform feature fusion on the first omics data based on domain knowledge and to perform data prediction based on the fused features.

[0070] According to one aspect of this disclosure, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the training method for the feature fusion model described in any one of the preceding claims, and the classification method for cancer users described in any one of the preceding claims.

[0071] According to one aspect of this disclosure, an electronic device is provided, comprising:

[0072] Processor; and

[0073] Memory for storing the executable instructions of the processor;

[0074] The processor is configured to execute the training method of the feature fusion model described in any one of the preceding claims, and the classification method of cancer users described in any one of the preceding claims, by executing the executable instructions.

[0075] This disclosure provides a training method for a feature fusion model. On one hand, it acquires historical patient sample data of historical users and extracts first omics data of the first gene point of the historical users from the historical patient sample data. Then, it acquires domain knowledge and constructs a heterogeneous network based on the first gene point, the first omics data, and the domain knowledge. Subsequently, it trains the network model to be trained according to the heterogeneous network to obtain a feature fusion model. At the same time, the feature fusion model is used to fuse domain knowledge and the first omics data and make data prediction based on the fused features. That is, the feature fusion model realizes the fusion of omics data and domain knowledge, solving the problem that the prior art cannot fuse omics data and domain knowledge. On the other hand, since it can fuse omics data and domain knowledge and then make data prediction based on the fused features, the accuracy of the data prediction results is improved.

[0076] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description

[0077] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure. It is obvious that the drawings described below are merely some embodiments of this disclosure, and those skilled in the art can obtain other drawings based on these drawings without any inventive effort.

[0078] Figure 1 The diagram illustrates an example of the structure of DeepOmix, a survival prediction algorithm based on integrating multi-omics data.

[0079] Figure 2 The diagram illustrates a flowchart of a training method for a feature fusion model according to an exemplary embodiment of the present disclosure.

[0080] Figure 3 The diagram schematically illustrates an example structure of a first heterogeneous network according to an exemplary embodiment of the present disclosure.

[0081] Figure 4 The diagram schematically illustrates an example structure of a second heterogeneous network according to an exemplary embodiment of the present disclosure.

[0082] Figure 5 The diagram schematically illustrates the structure of a first network model to be trained according to an example embodiment of the present disclosure.

[0083] Figure 6 The diagram schematically illustrates the structure of a second network model to be trained according to an example embodiment of the present disclosure.

[0084] Figure 7 The illustration shows an example scenario of training a first network model to be trained according to an example embodiment of the present disclosure.

[0085] Figure 8 The flowchart illustrates a method for training a second network model to be trained based on a second heterogeneous network, according to an exemplary embodiment of the present disclosure, to obtain a second feature fusion model.

[0086] Figure 9 The illustration shows an example scenario of training a second network model to be trained according to an exemplary embodiment of the present disclosure.

[0087] Figure 10 A flowchart illustrating a method for classifying cancer users according to an exemplary embodiment of this disclosure is shown schematically.

[0088] Figure 11 The diagram schematically illustrates a block diagram of a training apparatus for a feature fusion model according to an exemplary embodiment of the present disclosure.

[0089] Figure 12 An electronic device is illustrated, according to an example embodiment of the present disclosure, for implementing the above-described feature fusion model and / or a method for classifying cancer users. Detailed Implementation

[0090] Example embodiments will now be described more fully with reference to the accompanying drawings. However, example embodiments can be implemented in many forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided to make this disclosure more comprehensive and complete, and to fully convey the concept of the example embodiments to those skilled in the art. The described features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. In the following description, numerous specific details are provided to give a full understanding of embodiments of this disclosure. However, those skilled in the art will recognize that the technical solutions of this disclosure can be practiced with one or more of the specific details omitted, or other methods, components, apparatus, steps, etc., can be employed. In other instances, well-known technical solutions are not shown or described in detail to avoid obscuring various aspects of this disclosure.

[0091] Furthermore, the accompanying drawings are merely illustrative of this disclosure and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and therefore repeated descriptions of them will be omitted. Some block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities may be implemented in software, in one or more hardware modules or integrated circuits, or in different network and / or processor devices and / or microcontroller devices.

[0092] In some medical event analysis methods, the following approach can be used: First, analyze multi-omics data and interaction data to obtain the analysis results of medical events; then, in the data processing stage, the multi-omics data can be transformed into the same multi-dimensional data space, and then the multi-omics data in the transformed multi-dimensional space can be updated based on the interaction data; finally, the updated multi-omics data and the multi-omics data transformed into the multi-dimensional space are fused for features; finally, the fused features are used for medical event analysis.

[0093] In the aforementioned scheme, the interaction data specifically refers to intermolecular interaction networks, such as protein-protein interaction networks, gene regulation networks, gene co-expression networks, and metabolic networks. Furthermore, in this scheme, the fusion of domain knowledge and multi-omics data is achieved as follows: First, the multi-omics data is updated based on the interaction data. This update method involves fusing the interaction data with the multi-omics data transformed into a multi-dimensional data space using a graph convolutional neural network, thus updating the multi-omics data. Then, the updated multi-omics data features are combined with the original multi-omics data features through a weighted summation to obtain the fused features. However, this scheme simultaneously extracts features from both the original and updated multi-omics data, using the weighted summation of these two features as input for downstream tasks. This is equivalent to introducing more features beyond the multi-omics data features, failing to address the overfitting problem caused by the high dimensionality of multi-omics data.

[0094] In another approach, DeepOmix (a survival prediction algorithm based on integrating multi-omics data) builds a framework that combines user-defined prior biological information to achieve non-linear combinations of variables from different omics datasets; for details, please refer to [link to DeepOmix framework]. Figure 1 As shown, it is designed as a feedforward neural network, consisting of five layers: an omics data input layer 101, a first functional module layer 102, a first hidden layer 103, a second hidden layer 104, and a survival time output layer 105. The first input layer consists of normalized data from four different omics networks (protein interaction network, gene regulation network, gene co-expression network, and metabolic network). Further, the second layer represents gene functional modules (the first functional module layer), where the number of nodes corresponds to the number of functional modules (i.e., signaling pathways). The connection between the gene layer (the first input layer) and the functional layer (the first functional module layer) is constructed based on domain knowledge of pathway gene sets. In practical applications, if a gene belongs to a pathway, an edge is added between the g-th gene and the p-th pathway. Furthermore, the encoder of the gene layer, which is a non-fully connected network, constructs the features of the pathway layer. Finally, the pathway features are transformed into the next two hidden layers (the first and second hidden layers), ultimately reaching the survival data output layer (the survival time output layer). The core principle is to learn the representation of modules by integrating multi-omics data and user-defined functional modules. Each module is represented by a nonlinear function of the multiple omics gene values ​​it contains. However, this approach has the following problems: on the one hand, it does not learn the correlation between multi-omics data, nor does it learn the correlation between gene points in the same omics data.

[0095] Based on this, this exemplary embodiment first provides a training method for a feature fusion model, which can run on terminal devices, servers, server clusters, or cloud servers, etc. Of course, those skilled in the art can also run the method disclosed herein on other platforms as needed, and this exemplary embodiment does not impose any special limitations on this. Specifically, refer to... Figure 2 As shown, the training method for this feature fusion model may include the following steps:

[0096] Step S210. Obtain historical patient sample data of historical users, and extract the first omics data of the first gene point of the historical users from the historical patient sample data;

[0097] Step S220. Obtain domain knowledge and construct a heterogeneous network based on the first gene point, the first omics data, and the domain knowledge;

[0098] Step S230. Train the network model to be trained according to the heterogeneous network to obtain a feature fusion model; wherein, the feature fusion model is used to perform feature fusion on the first omics data based on domain knowledge and to perform data prediction based on the fused features.

[0099] In the training method of the aforementioned feature fusion model, on the one hand, historical patient sample data of historical users is obtained, and the first omics data of the first gene point of historical users is extracted from the historical patient sample data; then, domain knowledge is acquired, and a heterogeneous network is constructed based on the first gene point, the first omics data, and the domain knowledge; then, the network model to be trained is trained according to the heterogeneous network to obtain the feature fusion model. At the same time, the feature fusion model is used to perform feature fusion on the first omics data based on domain knowledge and to make data prediction based on the fused features; that is, the feature fusion model realizes the fusion of omics data and domain knowledge, solving the problem that existing technologies cannot fuse omics data and domain knowledge; on the other hand, since omics data and domain knowledge can be fused and data prediction can be made based on the fused features, the accuracy of the data prediction results is improved.

[0100] The training method of the feature fusion model described in the exemplary embodiments of this disclosure will be further explained and illustrated below with reference to the accompanying drawings.

[0101] First, the application scenarios and inventive objectives of the exemplary embodiments of this disclosure will be explained and described. Specifically, the training method for a feature fusion model provided by the exemplary embodiments of this disclosure integrates multi-omics data by fusing domain knowledge and multi-omics data. In practical applications, this can be achieved as follows: First, multi-omics data is acquired, organized, and preprocessed; then, domain knowledge is acquired, and a heterogeneous network with omics data as node features is constructed based on the domain knowledge; next, multi-omics data is integrated by combining domain knowledge to obtain node representations (i.e., fused features) that integrate domain knowledge and multi-omics data; furthermore, after obtaining the fused features, the multi-omics data features (fused features) that integrate domain knowledge can be applied to downstream tasks (i.e., data prediction); based on this method, the problem of difficulty in integrating domain knowledge during multi-omics data integration can be solved, as well as the overfitting problem caused by the high dimensionality of multi-omics data can be addressed.

[0102] Secondly, this disclosure describes two methods for fusing domain knowledge with multi-omics data. The first method involves constructing a heterogeneous network with multi-omics data as node features based on domain knowledge, and then performing heterogeneous network representation learning through a graph neural network, thereby fusing domain knowledge and multi-omics data with related information together. The second method involves constructing a heterogeneous network containing omics data nodes based on domain knowledge, obtaining omics data node representations with related information through representation learning, and then using this representation with related information as supervision information in the training process of the autoencoder, thereby integrating the related information into the features of the multi-omics data. Furthermore, by designing two schemes to fuse domain knowledge with related information with omics data together and applying the features of multi-omics data with related information to downstream tasks, the overfitting problem caused by the high-dimensionality of multi-omics data without related information can be improved.

[0103] Furthermore, the training method for the feature fusion model described in the exemplary embodiments of this disclosure can also solve the following problems: On the one hand, multi-omics data generally suffer from high dimensionality and small sample size; on the other hand, data between multiple omics and data within the same omics are not isolated but have certain correlations, but these correlations are relatively abstract and difficult to extract and use directly by humans; furthermore, existing studies rarely consider the correlations between omics, which can easily lead to the curse of dimensionality and thus overfitting; further still, since genes in cells work on a signaling pathway basis, and genes in the same pathway are responsible for the same functional module, the scheme described in the exemplary embodiments of this disclosure can utilize signaling pathways to mine the correlation information between multi-omics data, thereby reducing the redundancy of multi-omics data.

[0104] Furthermore, the domain knowledge involved in the exemplary embodiments of this disclosure will be explained and described. Specifically, in daily applications, domain knowledge can be divided into the following forms based on its structural characteristics: one is relational domain knowledge, represented by knowledge bases, knowledge graphs, word node embedding vectors, etc.; another is logical domain knowledge, represented by first-order logic, Markov logic networks, Bayesian networks, etc.; and yet another is scientific domain knowledge, represented by partial differential equations. Among these, relational domain knowledge provides information on the relationships between things. It should be noted here that the knowledge in fields such as biological signaling pathways involved in the exemplary embodiments of this disclosure is relational domain knowledge, so it is considered to be transformed into graph networks for representation learning and then fused with multi-omics data.

[0105] The following will explain and describe the heterogeneous network described in the exemplary embodiments of this disclosure. Specifically, the heterogeneous network involved in the exemplary embodiments of this disclosure may include a first heterogeneous network and a second heterogeneous network; wherein, the first heterogeneous network is a heterogeneous network with a first gene point as the first child node, a first omics data of the first gene point as the first child node as the node feature, a first connection relationship between current gene points in a biological signaling pathway as the first connection edge, and a second connection relationship between first gene pairs composed of the first gene points as the second connection edge; the first heterogeneous network is used to reflect the correlation relationship between data within the same omics data, as well as the correlation relationship between multiple omics data; meanwhile, a specific structural example diagram of the first heterogeneous network can be referred to Figure 3 As shown; further, the second heterogeneous network is a heterogeneous network with a first gene point as the first child node, the first omics data of the first gene point as the second child node, the first connection relationship between current gene points in the biological signaling pathway as the first connection edge, the second connection relationship between the first gene pairs composed of the first gene points as the second connection edge, and the third connection relationship between the first child node and the second child node as the third connection edge; this second heterogeneous network is used to reflect the correlation between data within the same omics data, as well as the correlation between multiple omics data; meanwhile, a specific structural example diagram of this second heterogeneous network can be found in [reference]. Figure 4 As shown.

[0106] It is important to clarify that the difference between the first and second heterogeneous networks is as follows: In the first heterogeneous network, nodes only include gene points, and edges include biological signaling pathways between gene points and connections between gene pairs when the Pearson coefficient is greater than a preset threshold; simultaneously, the omics data of each gene point is configured as node features in the form of encoding at the location of the gene point. Furthermore, in the second heterogeneous network, nodes not only include gene points but also the omics data of each gene point; that is, the omics data of each gene point is connected to each gene point as a separate node; simultaneously, the edges included in the second heterogeneous network not only include biological signaling pathways between gene points and connections between gene pairs when the Pearson coefficient is greater than a preset threshold, but also connections between each gene point and the omics data corresponding to the gene point.

[0107] The following will explain and describe the network model to be trained as described in this example embodiment. Specifically, the network model to be trained as described in this example embodiment may include a first network model to be trained and a second network model to be trained, wherein the first network model to be trained may include a graph network model and a first classifier; wherein, a specific structural example diagram can be referred to Figure 5 As shown; the second network model to be trained may include an autoencoder model and a second classifier; a specific structural example diagram can be found in [reference needed]. Figure 6 As shown. In practical applications, the graph network models described above can include relational graph convolutional network models or graph attention network models, etc.

[0108] Specifically, in Figure 5 The first network model to be trained shown may include a first input layer 501, a graph network model 502, a first classifier 503, and a first output layer 504; wherein the first input layer, graph network model, first classifier, and first output layer are connected sequentially; the specific functions of each module and / or each model will be listed later and will not be elaborated further here; furthermore, in Figure 6 The second network model to be trained shown may include a second input layer 601, an autoencoder model 602, a second classifier 603, and a second output layer 604; wherein the second input layer, the autoencoder model, the second classifier, and the second output layer are connected sequentially; and the specific functions of each module and / or each model will be listed in detail later, and will not be elaborated further here. It should be noted that... Figure 5The graph network model shown can also be called an encoder; that is, the first network model to be trained and the second network model to be trained are structurally similar, but the types of encoders used are different and the specific training methods are also different, but the final output data are similar.

[0109] The following will combine Figures 2-6 right Figure 2 The training method of the feature fusion model shown will be further explained and illustrated. Specifically:

[0110] In step S210, historical patient sample data of historical users are obtained, and first omics data of the first gene point of the historical user are extracted from the historical patient sample data.

[0111] Specifically, the first omics data recorded here may include DNA (DeoxyriboNucleic Acid) methylation data, gene mutation SNV (Single Nucleotide Variation) data, copy number variation (CNV) data, and gene expression data. In practical applications, firstly, historical patient sample data of historical users can be downloaded from the TCGA database using a corresponding download tool (e.g., GDC-Client). The historical patient sample data recorded here may include the corresponding first omics data and a metadata.json file. The metadata.json file described here is descriptive data, describing the attributes of the omics data and is mainly used to organize the downloaded omics data. For example, for a certain cancer (historical patient sample data), the information of all N samples is recorded in this metadata.json file, with two samples separated by a right curly brace "}" and a left curly brace "{".

[0112] Secondly, after downloading the data, it can be organized. Since different types of patient sample data (historical patient sample data) in the TCGA database can be stored independently, typically with a separate file for each type of data within a sample, this data must be organized before use. Therefore, firstly, all data files are extracted from the folder, and compressed files are decompressed. Then, key attribute information of the data files is extracted from metadata.json, including the filename (file_name), data category (data_category), and the corresponding TCGA-barcode (entity_submitter_id). Further, using the TCGA-barcode for each data file, multiple omics data files and real sample labels for each sample are obtained.

[0113] Furthermore, the data is preprocessed, wherein the preprocessing described herein may include missing value handling and data standardization. Specifically, this can be achieved in the following way: First, delete samples with fewer than 4 omics data, and use Python to extract the multi-omics data corresponding to the samples from the corresponding data files and organize them into matrices; the specific generation process is as follows: For example, assuming that a certain cancer has N samples (historical patient sample data), assuming that the data dimension of the DNA methylation group is D1, the data dimension of the gene mutation group is D2, the data dimension of the copy number variation group is D3, and the data dimension of the gene expression group is D4, then each omics data can be organized into N*D1, N*D2, N*D3, and N*D4 matrices in turn; then, standardize the data of multiple omics to obtain the first omics data; the standardization process is to convert the value range of multi-omics data of different magnitudes into the same magnitude, so as to make the omics data of different dimensions numerically comparable; in the actual application process, the example embodiment of this disclosure uses the zero mean standardization method (Z-score) to process the data, and its calculation formula is shown in the following formula (1):

[0114]

[0115] Where x is the initial value of the sample data, x' is the standardized value of the sample data, μ is the mean of the sample data, and σ is the standard deviation of the sample data; finally, the first omics data of the first gene point can be obtained.

[0116] In step S220, domain knowledge is acquired, and a heterogeneous network is constructed based on the first gene point, the first omics data, and the domain knowledge.

[0117] Specifically, the domain knowledge described herein may include biological signaling pathways and the current gene points included in those pathways; it may also include correlation coefficients between gene pairs formed by any two gene points from the current gene points, as well as intermolecular interaction networks, etc. Furthermore, in practical applications, domain knowledge containing related information, such as biological signaling pathways, can be obtained from databases; the biological signaling pathways described in the example embodiments of this disclosure can be obtained from the KEGG and Reactome databases; alternatively, the corresponding intermolecular interaction networks can be obtained by querying databases, and then the molecular interaction relationships can be added as edges to the heterogeneous network.

[0118] Secondly, once domain knowledge is acquired, a heterogeneous network can be constructed based on the first gene point, the first omics data, and the domain knowledge. Since the heterogeneous network can include a first heterogeneous network and a second heterogeneous network, the specific construction process of the heterogeneous network in practical applications can be implemented as follows:

[0119] The first implementation method is as follows: Based on the first gene point, the first omics data, and the domain knowledge, a first heterogeneous network is constructed, which can be achieved as follows: First, a first set of child nodes is constructed based on the first gene point, and a first set of node features is constructed based on the first omics data of the first gene point; Second, based on the biological signaling pathways included in the domain knowledge and the current gene points included in the biological signaling pathways, a first connection relationship between the first gene points is determined, and a first set of connection edges is constructed based on the first connection relationship; Then, a first correlation coefficient is calculated between the first gene pairs composed of the first gene points, and a second connection relationship between the first gene points is determined based on the first correlation coefficient; Finally, a second set of connection edges is constructed based on the second connection relationship, and the first heterogeneous network is constructed based on the first set of child nodes, the first set of node features, the first set of connection edges, and the second set of connection edges.

[0120] In one example embodiment, the first gene pair composed of the first gene point includes a first sub-gene point and a second sub-gene point; wherein, the calculation of the first correlation coefficient between the first gene pair composed of the first gene point described above can be implemented in the following manner: First, obtain the first sub-omics data of the first sub-gene point and the second sub-omics data of the second sub-gene point in the first gene pair composed of the first gene point; second, extract the expression data of the first sub-gene from the first sub-omics data and extract the expression data of the second sub-gene from the second sub-omics data; then, calculate the first correlation coefficient based on the expression data of the first sub-gene and the expression data of the second sub-gene.

[0121] The following section will further explain and illustrate the specific construction process of the first heterogeneous network. Specifically, firstly, the first gene point is used as a child node to construct the first child node binding V; secondly, the first omics data (gene expression data, SNV data, methylation data, CNV data) corresponding to the first gene point are encoded to obtain the node feature representation of the first gene point; then, based on the biological signaling pathway and the current gene points included in the biological signaling pathway, it is determined whether there is a first connection relationship between the first gene points; if there is a first connection relationship, a first set of connection edges is constructed based on the first connection relationship; wherein, the first set of connection edges includes the first gene points with the first connection relationship; next, the first correlation coefficient between the first gene points is calculated; wherein, the first correlation coefficient mentioned here can be the Pearson correlation coefficient, or... The first correlation coefficient is the Pierceman correlation coefficient, which is not specifically limited in this example. Further, it is determined whether the first correlation coefficient is greater than a preset threshold (e.g., the preset threshold could be Pth, and its specific value can be determined according to actual needs; this example does not impose any special restrictions on this). If the first correlation coefficient is greater than or equal to the preset threshold, it is determined that a second connection relationship exists between the first gene pairs. If the first correlation coefficient is less than the preset threshold, it is determined that no connection relationship exists between the first gene pairs. Finally, a second set of connection edges is constructed based on the second connection relationship. Then, based on the first set of child nodes, the first set of node features, the first set of connection edges, and the second set of connection edges, a first heterogeneous network is constructed. The second set of connection edges includes first gene points with the second connection relationship.

[0122] In one example embodiment, the specific calculation method of the Pearson correlation coefficient (first correlation coefficient) of the gene expression corresponding to the first gene is as follows: Taking the first sub-gene point included in the first gene pair as gene point A and the second sub-gene point as gene point B as an example, the specific calculation process of the first correlation coefficient is explained and described. Specifically, firstly, obtain the N gene expression data of gene A in N samples, represented as A_e=(m1,m2,m3,…,mN), and similarly, the gene expression data of gene B in N samples is B_e=(n1,n2,n3,…,nN); then, substitute the two data sequences A_e and B_e into the Pearson correlation coefficient calculation formula to obtain the first correlation coefficient of the two data sequences A_e and B_e; wherein, the specific Pearson correlation coefficient calculation formula can be shown in the following formula (2):

[0123]

[0124] Where r is the first correlation coefficient, X i For sequence data of gene point A, Y i This is the sequence data for gene point B. This represents the average value of gene point A. Here, n represents the average value of gene point B, and n is the number of samples. It should be noted that the Pearson correlation coefficient is chosen as the first correlation coefficient in this exemplary embodiment because the Spearman correlation coefficient is not concerned with whether the two datasets are linearly correlated, but rather with whether they are monotonically correlated; it is based on the rank value of each variable, not the original data. Simply put, the Pearson correlation coefficient processes the original values ​​of the variables, while the Spearman correlation coefficient processes the rank values. Therefore, the statistical power of the Pearson correlation coefficient is higher than that of the Spearman correlation coefficient. In practical applications, the specific correlation coefficient to be selected as the first correlation coefficient can be determined in the following way: one method... The formula is as follows: depending on the purpose, for example, if you only want to analyze the regulatory relationship between two genes, you only need to assume that the expression levels of the two genes are monotonically correlated, and Spearman's correlation coefficient can be used. If you want to analyze more specifically, you can use Pearson's correlation coefficient. Another approach is that, due to the complexity of the regulatory mechanisms and interrelationships between genes, coupled with the interference of experimental errors, detection errors, and other factors, it is not intuitive to determine which correlation calculation method is better. Furthermore, according to actual research, it can be found that Pearson's correlation coefficient is frequently used in the literature to calculate the correlation between genes. Therefore, the example embodiment of this disclosure selects Pearson's correlation coefficient as the first correlation coefficient.

[0125] The second implementation method is as follows: Based on the first gene point, the first omics data, and the domain knowledge, a second heterogeneous network is constructed, which can be achieved as follows: First, a first set of child nodes is constructed based on the first gene point, and a second set of child nodes is constructed based on the first omics data of the first gene point; second, based on the biological signaling pathways included in the domain knowledge and the current gene points included in the biological signaling pathways, a first connection relationship between the first gene points is determined, and a first set of connection edges is constructed based on the first connection relationship; then, a first correlation coefficient is calculated between the first gene pairs composed of the first gene points, and a second connection relationship between the first gene points is determined based on the first correlation coefficient; further, a second set of connection edges is constructed based on the second connection relationship, and a third set of connection edges is constructed based on the third connection relationship between the first gene point and the first omics data; finally, a second heterogeneous network is constructed based on the first set of child nodes, the second set of child nodes, the first set of connection edges, the second set of connection edges, and the third set of connection edges.

[0126] The following section will further explain and illustrate the specific construction process of the second heterogeneous network. Specifically, firstly, the first gene point is used as a child node to construct a first child node binding V; secondly, the first omics data (gene expression data, SNV data, methylation data, CNV data) corresponding to the first gene point is abstracted to obtain the second child node binding; then, based on the biological signaling pathway and the current gene points included in the biological signaling pathway, it is determined whether there is a first connection relationship between the first gene points; if there is a first connection relationship, a first set of connection edges is constructed based on this first connection relationship; wherein, the first set of connection edges includes the first gene points with the first connection relationship; next, the first correlation coefficient between the first gene points is calculated; wherein, the first correlation coefficient recorded here can be the Pearson correlation coefficient or the Pearsman correlation coefficient, and this example does not impose any special restrictions on it; further, it is determined whether the first correlation coefficient is greater than a preset threshold. (For example, the preset threshold can be Pth, and the specific value can be determined according to actual needs. This example does not impose any special restrictions on this.) If the first correlation coefficient is greater than or equal to the preset threshold, it is determined that there is a second connection relationship between the first gene pairs. If the first correlation coefficient is less than the preset threshold, it is determined that there is no connection relationship between the first gene pairs. Further, a second set of connection edges is constructed based on the second connection relationship, wherein the second set of connection edges includes the first gene point with the second connection relationship. Further still, a third connection relationship is established between the first gene point and the first omics data corresponding to the first gene point, and a third set of connection edges is constructed based on the third connection relationship. Finally, a second heterogeneous network reflecting the topological relationship between multi-omics data is constructed based on the first set of child nodes, the second set of child nodes, the first set of connection edges, the second set of connection edges, and the third set of connection edges.

[0127] In step S230, the network model to be trained is trained according to the heterogeneous network to obtain a feature fusion model; wherein, the feature fusion model is used to perform feature fusion on the first omics data based on domain knowledge and to perform data prediction based on the fused features.

[0128] Specifically, since the network model to be trained described here may include a first network model to be trained and a second network model to be trained, the feature fusion model can be trained based on the heterogeneous network in two ways: the first method is to train the first network model to be trained based on the first heterogeneous network to obtain the first feature fusion model; the second method is to train the second network model to be trained based on the second heterogeneous network to obtain the second feature fusion model.

[0129] In one example embodiment, training a first network model to be trained based on a first heterogeneous network to obtain a first feature fusion model can be achieved as follows: First, the first heterogeneous network is input into the graph network model in the first network model to be trained to obtain a first feature representation corresponding to the first heterogeneous network, and the first feature representation is input into the first classifier in the first network model to be trained to obtain a first predicted label; Second, based on the real user labels of the historical users and the first predicted label, a first target loss function is constructed, and the first network model to be trained is trained based on the first target loss function to obtain the first feature fusion model.

[0130] The following will combine Figure 7 The specific training process of the first feature fusion model described in the exemplary embodiments of this disclosure will be further explained and illustrated. Specifically, in practical applications, the exemplary embodiments of this disclosure obtain the vector representation of heterogeneous networks through a heterogeneous network classification task based on graph neural networks (graph network models). The heterogeneous network classification task based on graph neural networks involves training the model using heterogeneous networks with known category labels, and then using the trained model to predict the category of unknown heterogeneous networks. Simultaneously, the representations of each heterogeneous network can be obtained during the training process. The node representations in this vector space represent an effective fusion of the correlation information between multi-omics data in the domain knowledge and the multi-omics data itself.

[0131] In one example embodiment, the Relational Graph Convolutional Network (RGCN) model is used as an example. Specifically, RGCN is a model for studying heterogeneous graphs. It considers the influence of different edges on nodes based on the Graph Convolutional Network (GCN). In practical applications, there are two main methods for constructing graph convolution operators: spectral methods and spatial methods. RGCN uses the spatial method, that is, it defines graph convolution from the perspective of the node's neighborhood. In RGCN, the update method of node i is as shown in the following formula (3):

[0132]

[0133] Among them, h i Let i be the vector representation of node i, and let l be the superscript number of the graph neural network in which it is located. This indicates that node i is self-connected, meaning that each update of a node considers not only the information of its neighbors but also its own information. Let R represent the set of neighboring nodes of type r among the neighboring nodes of node i; let R represent the set of current edge type r; c i,r It is a normalization term, representing the number of nodes of type r among the neighboring nodes of node i; Different edge types have their own parameters; σ is the activation function, which can be ReLU. Specifically, in this example embodiment, firstly, for each sample, heterogeneous networks G1, G2, ..., G... are obtained, each with the first gene point as a child node and the corresponding multi-omics data as node features. N And the corresponding real user tags y1, y2, ..., y N These are the training data for the model; then, the heterogeneous networks are input into the RGCN model to obtain the representation h of each heterogeneous network. G (First feature representation); Finally, the representation of the heterogeneous network (first feature representation) is input into the first classifier C to obtain the output label (first predicted label) y of each heterogeneous network. i For specific scenario examples, please refer to... Figure 7 As shown.

[0134] In one example embodiment, the first objective loss function used in this example embodiment can be the cross-entropy loss function. That is, the first network model to be trained can be trained based on the cross-entropy loss function. After training, the representations of known heterogeneous networks (graph network models) can be obtained, and the trained RGCN model and classifier C (first classifier) ​​can be used to represent and classify unknown heterogeneous networks. The first objective loss function can be specifically referred to as shown in formula (4):

[0135]

[0136] Among them, y i ' is the first predicted label for the i-th sample, y i Let be the real user label for the i-th sample.

[0137] At this point, the entire training process for the first feature fusion model has been completed. The following will combine... Figure 8 as well as Figure 9 The specific training process of the second feature fusion model is explained and illustrated.

[0138] In one example embodiment, reference is made to... Figure 8 As shown, the second feature fusion model is obtained by training the second network model to be trained based on the second heterogeneous network, which can be achieved in the following way:

[0139] Step S810: Perform representation learning on the second heterogeneous network to obtain the second feature representation corresponding to the second heterogeneous network, and input the first omics data of the first gene point into the encoding module in the second network model to be trained to obtain the first reconstructed data corresponding to the first omics data.

[0140] In one example embodiment, the first omics data of the first gene point is input into the encoding module of the second network model to be trained to obtain the first reconstructed data corresponding to the first omics data. This can be achieved as follows: First, the first omics data is mapped using a first nonlinear function in the encoding module of the second network model to be trained to obtain intermediate variables; second, the intermediate variables are mapped using a second nonlinear function in the encoding module of the second network model to be trained to obtain the first reconstructed data corresponding to the first omics data; wherein the data representation form of the first reconstructed data is consistent with the data representation form of the first omics data.

[0141] Step S820: Construct a second objective loss function based on the first omics data, the first reconstructed data, and the second feature representation, and train the autoencoder model in the second network model to be trained according to the second objective function to obtain the trained encoding model.

[0142] In one example embodiment, the second target loss function is constructed based on the first omics data, the first reconstructed data, and the second feature representation. This can be achieved as follows: First, the first reconstructed data is input into the decoding module of the second network model to be trained to obtain the third feature representation corresponding to the second heterogeneous network, and a first sub-loss function is constructed based on the first omics data and the first reconstructed data; second, a second sub-loss function is constructed based on the second feature representation and the third feature representation, and a second target function is constructed based on the first sub-loss function and the second sub-loss function.

[0143] In one example embodiment, the construction of a first sub-loss function based on the first omics data and the first reconstructed data can be achieved as follows: First, the absolute value of the first difference between the first omics data and the first reconstructed data is calculated, and the absolute value of the first difference is squared to obtain the first squared calculation result; second, the first squared calculation result is summed to obtain the first sub-loss function.

[0144] In one example embodiment, the second sub-loss function is constructed based on the second feature representation and the third feature representation, which can be achieved as follows: First, the first sub-vector representation of the first sub-gene point and the second sub-vector representation of the second sub-gene point in the first gene pair included in the second feature representation are obtained, and the first similarity of the first gene pair is calculated based on the first sub-vector representation and the second sub-vector representation; Second, the third sub-vector representation of the first sub-gene point and the fourth sub-vector representation of the second sub-gene point in the first gene pair included in the third feature representation are obtained, and the second similarity of the first gene pair is calculated based on the third sub-vector representation and the fourth sub-vector representation; Then, the absolute value of the second difference between the first similarity and the second similarity is calculated, and the absolute values ​​of the second difference are summed to obtain the second sub-loss function.

[0145] Step S830: Input the first omics data into the encoding module of the trained encoding model to obtain the second reconstructed data, and input the second reconstructed data into the second classifier of the second network model to be trained to obtain the second predicted label.

[0146] Step S840: Based on the second predicted label and the real user labels of the historical users, the second classifier is trained to obtain the second feature fusion model.

[0147] The following will combine Figure 9The specific training process of the second feature fusion model is further explained and illustrated. Specifically, in practical applications, firstly, representation learning is performed on the heterogeneous network (second heterogeneous network) that reflects the topological relationships between multiple omics data to obtain node representations (second feature representations) with omics data association information. In the heterogeneous network representation learning described in the example embodiments of this disclosure, the following models can be used: MetaPath2Vec (heterogeneous network) model, MetaGraph2Vec (heterogeneous graph representation learning) model, HIN2Vec (heterogeneous information network representation learning) model, RGCN (Relational Graph Convolutional Network) model, HAN (Heterogeneous Graph Attention Network) model, and GATNE (General Attributed Multiplex Heterogeneous Network). Embedding, large-scale multi-dimensional heterogeneous attribute network (MLN) model, etc., are not limited in this invention; that is, the above-mentioned second heterogeneous network can be input into the above-listed model to obtain the second feature representation; here, the obtained second feature representation can be denoted as E1; the second feature representation can be used to reflect the association relationship between nodes; of course, the second feature representation can also be obtained through matrix decomposition, principal component analysis, etc., and this example does not impose any special restrictions on this.

[0148] Secondly, the representation reflecting the correlation information between multi-omics data will be used as the supervisory information for the multi-omics data feature extraction process, thereby integrating the implicit correlation information in the domain knowledge into the features of the multi-omics data. In practical application, firstly, an autoencoder model is used for feature extraction to obtain the first reconstructed data; the specific implementation process is as follows: the multi-omics data is input into the autoencoder, the input of which is the first omics data of the first gene point, and the output of which is the extracted features of the multi-omics data (features of the first omics data, i.e., intermediate variables); at the same time, the input of the decoding module is the winning variable, and the output is the first reconstructed data; where each column in the parameter matrix of the first reconstructed data represents the representation of an original feature, and the representation space of the first reconstructed data is denoted as E2; at the same time, the representation in this space reflects the information of the data itself.

[0149] Then, the difference between the similarity of vector pairs in E1 space and the similarity of corresponding vector pairs in E2 space is used, and the sum of their absolute values ​​is obtained to achieve the purpose of using domain knowledge with related information to supervise the feature extraction process of multi-omics data; thus, the second objective loss function of the autoencoder with supervised information can be shown in the following formula (5):

[0150]

[0151] In this diagram, the function before the plus sign represents the loss function for the autoencoder's reconstruction task, while the function after the plus sign represents the loss function for the task of incorporating domain knowledge, i.e., the similarity loss function of all corresponding pairs in E1 and E2, thus realizing the supervisory role of domain knowledge on the autoencoder. Simultaneously, the loss function for the autoencoder's reconstruction task is the sum of the squared difference loss functions of the M omics data, specifically M = 4. Here, x represents the input omics data (the first omics data), z represents the input reconstructed by the autoencoder with the same shape as x (the first reconstructed data), y = f(Wx + b) indicates that the autoencoder maps the input x to y through a nonlinear function f, and z = g(Wx + b) represents the input x being mapped to y by the autoencoder. T y+b') represents the autoencoder mapping the embedded y back to the reconstructed z (first reconstructed input) with the same shape as x through another nonlinear function g; further, the loss function for the domain knowledge task is the sum of the absolute values ​​of the differences in the similarities of all corresponding vector pairs in the E1 and E2 spaces, where P is the total number of vector pairs; the specific implementation process can be found in [reference needed]. Figure 9 As shown; it should be added that when using the second objective loss function to train the autoencoder model, in addition to the reconstruction task loss function of the autoencoder model, there is also domain knowledge with omics data association information as the supervision information loss of the autoencoder model. The model uses these losses to update the parameters of the autoencoder model through backpropagation, thereby integrating domain knowledge into the features of multi-omics data.

[0152] Furthermore, after training the autoencoder model, the second classifier also needs to be trained. Specifically, the first omics data can be input into the encoding module of the trained encoding model to obtain the second reconstructed data, and then the second reconstructed data can be input into the second classifier of the second network model to be trained to obtain the second predicted label. Finally, based on the second predicted label and the real user labels of historical users, the second classifier is trained to obtain the second feature fusion model. Here, the second reconstructed data can include domain knowledge and the first omics data because the parameters of the autoencoder model can be updated through backpropagation during the training process, thereby integrating domain knowledge into the features of multi-omics data. Thus, feature fusion can be achieved without increasing the feature dimension, ultimately achieving the goal of dimensionality reduction.

[0153] At this point, the specific training process of the second feature fusion model has been completed. Based on the above description, it can be seen that the training method of the feature fusion model described in the exemplary embodiments of this disclosure can achieve the fusion of domain knowledge and omics data; at the same time, it can be seen from the foregoing description that the exemplary embodiments of this disclosure solve the problem of difficulty in introducing domain knowledge in multi-omics data integration tasks.

[0154] The specific applications of the fusion features obtained by this disclosure will be explained and described below. Specifically, the fusion features obtained by the method described in the example embodiments of this disclosure can be applied to downstream tasks, thereby improving the overfitting problem caused by the high dimensionality of multi-omics data when there is no associated information. Meanwhile, specific downstream tasks may include, but are not limited to, cancer classification, survival prediction, and drug response prediction; furthermore, these features can be fixed when applied to new downstream tasks, or they can be retrained as the task progresses; this example does not impose any special restrictions on this. The specific application process will be explained and described below with reference to specific example embodiments.

[0155] First, this disclosure provides an example embodiment of a method for classifying cancer patients. Specifically, this method can run on terminal devices, servers, server clusters, or cloud servers, etc.; of course, those skilled in the art can also run the method of this disclosure on other platforms as needed, and this exemplary embodiment does not impose any special limitations on this. Further, refer to... Figure 10 As shown, the method for classifying cancer users may include the following steps:

[0156] Step S1010: Obtain the current sample data of the current user, and extract the second omics data of the second gene point of the current user from the current sample data;

[0157] Step S1020: Input the second omics data of the second gene point into the feature fusion model to obtain the data prediction result; wherein, the feature fusion model is trained based on the feature fusion model training method described above.

[0158] Step S1030: Determine the user category to which the current user belongs based on the data prediction results.

[0159] In one example embodiment, inputting the second omics data of the second gene point into a feature fusion model to obtain a data prediction result can be achieved as follows: inputting the second omics data of the second gene point into a graph network model in a first feature fusion model to obtain a first user feature representation including dimensionality-reduced second omics data and domain knowledge, and inputting the first user feature representation into a first classifier in the first feature fusion model to obtain a data prediction result; and / or inputting the second omics data of the second gene point into an encoding module in an autoencoding model in a second feature fusion model to obtain a second user feature representation including dimensionality-reduced second omics data and domain knowledge, and inputting the second user feature representation into a second classifier in the second feature fusion model to obtain a data prediction result. In other words, in practical applications, the second omics data of the second gene points included in the patient's sample data can be input into the first feature fusion model to obtain a first user feature representation that includes the second omics data and domain knowledge. Then, the patient can be classified as a cancer patient based on the first user feature representation. Alternatively, the second omics data of the second gene points included in the patient's sample data can be input into the second feature fusion model to obtain a second user feature representation that includes the dimensionality-reduced second omics data and domain knowledge (which can also be called reconstructed data that includes the second omics data and domain knowledge). The patient can then be classified as a cancer patient based on the second user feature representation.

[0160] The aforementioned classification method for cancer users, on the one hand, achieves the fusion of omics data and domain knowledge through a feature fusion model, solving the problem that existing technologies cannot fuse omics data and domain knowledge; on the other hand, because it can fuse omics data and domain knowledge and then make data predictions based on the fused features, it improves the accuracy of cancer user classification results.

[0161] It should be further noted that the specific prediction process for survival prediction and drug response prediction is largely similar to the classification process for cancer users, and will not be elaborated further here.

[0162] Thus, the methods described in the exemplary embodiments of this disclosure have been fully implemented. Based on the foregoing description, the methods described in the exemplary embodiments of this disclosure have at least the following advantages: Firstly, they creatively propose a method for converting domain knowledge and multi-omics data into graph-structured data. This method uses genes in pathways as nodes, multi-omics data of genes as features of nodes, and domain knowledge as edges between nodes to build a heterogeneous network, laying the foundation for integrating domain knowledge into multi-omics data. Simultaneously, it can also perform representation learning on the heterogeneous network combining domain knowledge and multi-omics data through graph neural networks, thereby achieving the fusion of domain knowledge and multi-omics data. Furthermore, it can also use domain knowledge as supervisory information in the training process of the multi-omics feature extractor, thereby integrating domain knowledge into the multi-omics feature extractor. The present invention integrates the correlation information into the features of multi-omics data. Furthermore, the exemplary embodiments of this invention also design a heterogeneous network that integrates domain knowledge with multi-omics data, transforming the domain knowledge and multi-omics data into graph-structured data and using multi-omics data as the features of nodes, thus solving the problem of difficulty in introducing domain knowledge in multi-omics data integration tasks. Moreover, the exemplary embodiments of this invention, through heterogeneous network representation learning, can mine the correlation information between data within the same omics and between data from different omics. Simultaneously, by integrating domain knowledge with multi-omics data, redundancy in multi-omics data can be reduced, improving the overfitting problem caused by the high dimensionality of multi-omics data in existing studies.

[0163] This disclosure also provides a training apparatus for a special fusion model. Specifically, refer to... Figure 11 As shown, the training device for this feature fusion model may include a first omics data extraction module 1110, a heterogeneous network construction module 1120, and a model training module 1130. Wherein:

[0164] The first omics data extraction module 1110 can be used to obtain historical patient sample data of historical users and extract the first omics data of the first gene point of the historical users from the historical patient sample data.

[0165] The heterogeneous network construction module 1120 can be used to acquire domain knowledge and construct a heterogeneous network based on the first gene point, the first omics data and the domain knowledge.

[0166] The model training module 1130 can be used to train the network model to be trained according to the heterogeneous network to obtain a feature fusion model; wherein, the feature fusion model is used to perform feature fusion on the first omics data based on domain knowledge and to perform data prediction based on the fused features.

[0167] In one exemplary embodiment of this disclosure, the domain knowledge includes biological signaling pathways and current gene points included in the biological signaling pathways; the first omics data includes one or more of DNA methylation data, gene mutation SNV data, copy number variation CNV data, and gene expression data.

[0168] In one exemplary embodiment of this disclosure, the heterogeneous network includes a first heterogeneous network and / or a second heterogeneous network; wherein, the first heterogeneous network is a heterogeneous network with a first gene point as a first child node, a first omics data of the first gene point as a node feature of the first child node, a first connection relationship between current gene points in a biological signaling pathway as a first connection edge, and a second connection relationship between first gene pairs composed of the first gene points as a second connection edge; the second heterogeneous network is a heterogeneous network with a first gene point as a first child node, a first omics data of the first gene point as a second child node, a first connection relationship between current gene points in a biological signaling pathway as a first connection edge, a second connection relationship between first gene pairs composed of the first gene points as a second connection edge, and a third connection relationship between the first child node and the second child node as a third connection edge.

[0169] In one exemplary embodiment of this disclosure, the first heterogeneous network is used to reflect the correlation between data within the same omics dataset and the correlation between multiple omics datasets; the second heterogeneous network is used to reflect the correlation between data within the same omics dataset and the correlation between multiple omics datasets.

[0170] In one exemplary embodiment of this disclosure, the network model to be trained includes a first network model to be trained and / or a second network model to be trained; the first network model to be trained includes a graph network model and a first classifier, and the second network model to be trained includes an autoencoder model and a second classifier; the graph network model includes a relational graph convolutional network model and / or a graph attention network model.

[0171] In one exemplary embodiment of this disclosure, constructing a heterogeneous network based on the first gene point, the first omics data, and the domain knowledge includes: constructing a first set of child nodes based on the first gene point, and constructing a first set of node features based on the first omics data of the first gene point; determining a first connection relationship between the first gene points based on the biological signaling pathways included in the domain knowledge and the current gene points included in the biological signaling pathways, and constructing a first set of connection edges based on the first connection relationship; calculating a first correlation coefficient between first gene pairs composed of the first gene points, and determining a second connection relationship between the first gene points based on the first correlation coefficient; constructing a second set of connection edges based on the second connection relationship, and constructing a first heterogeneous network based on the first set of child nodes, the first set of node features, the first set of connection edges, and the second set of connection edges.

[0172] In one exemplary embodiment of this disclosure, constructing a heterogeneous network based on the first gene point, the first omics data, and the domain knowledge further includes: constructing a first set of child nodes based on the first gene point, and constructing a second set of child nodes based on the first omics data of the first gene point; determining a first connection relationship between the first gene points based on the biological signaling pathways included in the domain knowledge and the current gene points included in the biological signaling pathways, and constructing a first set of connection edges based on the first connection relationship; calculating a first correlation coefficient between first gene pairs composed of the first gene points, and determining a second connection relationship between the first gene points based on the first correlation coefficient; constructing a second set of connection edges based on the second connection relationship, and constructing a third set of connection edges based on the third connection relationship between the first gene point and the first omics data; and constructing a second heterogeneous network based on the first set of child nodes, the second set of child nodes, the first set of connection edges, the second set of connection edges, and the third set of connection edges.

[0173] In one exemplary embodiment of this disclosure, the first gene pair composed of the first gene point includes a first sub-gene point and a second sub-gene point; wherein, calculating the first correlation coefficient between the first gene pair composed of the first gene point includes: acquiring first sub-omics data of the first sub-gene point and second sub-omics data of the second sub-gene point in the first gene pair composed of the first gene point; extracting first sub-gene expression data from the first sub-omics data and extracting second sub-gene expression data from the second sub-omics data; and calculating the first correlation coefficient based on the first sub-gene expression data and the second sub-gene expression data.

[0174] In one exemplary embodiment of this disclosure, training a network model to be trained based on the heterogeneous network to obtain a feature fusion model includes: training a first network model to be trained based on a first heterogeneous network to obtain a first feature fusion model; and / or training a second network model to be trained based on a second heterogeneous network to obtain a second feature fusion model.

[0175] In one exemplary embodiment of this disclosure, training a first network model to be trained based on a first heterogeneous network to obtain a first feature fusion model includes: inputting the first heterogeneous network into a graph network model in the first network model to be trained to obtain a first feature representation corresponding to the first heterogeneous network, and inputting the first feature representation into a first classifier in the first network model to be trained to obtain a first predicted label; constructing a first target loss function based on the real user labels of the historical users and the first predicted label, and training the first network model to be trained based on the first target loss function to obtain the first feature fusion model.

[0176] In one exemplary embodiment of this disclosure, training a second network model to be trained based on a second heterogeneous network to obtain a second feature fusion model includes: performing representation learning on the second heterogeneous network to obtain a second feature representation corresponding to the second heterogeneous network; inputting the first omics data of the first gene point into the encoding module of the second network model to be trained to obtain first reconstructed data corresponding to the first omics data; constructing a second objective loss function based on the first omics data, the first reconstructed data, and the second feature representation; training the autoencoder model in the second network model to be trained based on the second objective function to obtain a trained encoding model; inputting the first omics data into the encoding module of the trained encoding model to obtain second reconstructed data; inputting the second reconstructed data into a second classifier in the second network model to be trained to obtain a second predicted label; and training the second classifier based on the second predicted label and the real user labels possessed by the historical user to obtain the second feature fusion model.

[0177] In one exemplary embodiment of this disclosure, the first omics data of the first gene point is input into the encoding module of the second network model to be trained to obtain the first reconstructed data corresponding to the first omics data. This includes: mapping the first omics data through a first nonlinear function in the encoding module of the second network model to be trained to obtain intermediate variables; mapping the intermediate variables through a second nonlinear function in the encoding module of the second network model to be trained to obtain the first reconstructed data corresponding to the first omics data; wherein the data representation form of the first reconstructed data is consistent with the data representation form of the first omics data.

[0178] In one exemplary embodiment of this disclosure, constructing a second target loss function based on first omics data, first reconstructed data, and second feature representation includes: inputting the first reconstructed data into the decoding module of a second network model to be trained to obtain a third feature representation corresponding to the second heterogeneous network, and constructing a first sub-loss function based on the first omics data and the first reconstructed data; constructing a second sub-loss function based on the second feature representation and the third feature representation, and constructing a second target function based on the first sub-loss function and the second sub-loss function.

[0179] In one exemplary embodiment of this disclosure, constructing a first sub-loss function based on first omics data and first reconstructed data includes: calculating the absolute value of a first difference between the first omics data and the first reconstructed data, and squaring the absolute value of the first difference to obtain a first squared calculation result; summing the first squared calculation result to obtain the first sub-loss function.

[0180] In one exemplary embodiment of this disclosure, constructing a second sub-loss function based on a second feature representation and a third feature representation includes: obtaining a first sub-vector representation of a first sub-gene point and a second sub-vector representation of a second sub-gene point in a first gene pair included in the second feature representation, and calculating a first similarity of the first gene pair based on the first sub-vector representation and the second sub-vector representation; obtaining a third sub-vector representation of a first sub-gene point and a fourth sub-vector representation of a second sub-gene point in the first gene pair included in the third feature representation, and calculating a second similarity of the first gene pair based on the third sub-vector representation and the fourth sub-vector representation; calculating the absolute value of a second difference between the first similarity and the second similarity, and summing the absolute values ​​of the second difference to obtain the second sub-loss function.

[0181] This disclosure also provides an example embodiment of a cancer user classification device. Specifically, the cancer user classification device may include a second omics data extraction module, a data prediction result acquisition module, and a user category determination module. Wherein:

[0182] The second omics data extraction module can be used to obtain the current sample data of the current user and extract the second omics data of the second gene point of the current user from the current sample data.

[0183] The data prediction result acquisition module can be used to input the second omics data of the second gene point into the feature fusion model to obtain the data prediction result; wherein, the feature fusion model is trained based on the training method of the feature fusion model described above.

[0184] The user category determination module can be used to determine the user category to which the current user belongs based on the data prediction results.

[0185] In one exemplary embodiment of this disclosure, inputting the second omics data of the second gene point into a feature fusion model to obtain a data prediction result includes: inputting the second omics data of the second gene point into a graph network model in a first feature fusion model to obtain a first user feature representation including dimensionality-reduced second omics data and domain knowledge, and inputting the first user feature representation into a first classifier in the first feature fusion model to obtain a data prediction result; and / or inputting the second omics data of the second gene point into an encoding module in an autoencoding model in a second feature fusion model to obtain a second user feature representation including dimensionality-reduced second omics data and domain knowledge, and inputting the second user feature representation into a second classifier in the second feature fusion model to obtain a data prediction result.

[0186] The specific details of each module in the training device of the aforementioned special fusion model and the classification device for cancer users have been described in detail in the training method of the corresponding feature fusion model and the classification method for cancer users, so they will not be repeated here.

[0187] It should be noted that although several modules or units for the device used to perform actions have been mentioned in the detailed description above, this division is not mandatory. In fact, according to embodiments of this disclosure, the features and functions of two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.

[0188] Furthermore, although the steps of the method in this disclosure are described in a specific order in the accompanying drawings, this does not require or imply that the steps must be performed in that specific order, or that all the steps shown must be performed to achieve the desired result. Additional or alternative steps may be omitted, multiple steps may be combined into one step, and / or a step may be broken down into multiple steps.

[0189] In an exemplary embodiment of this disclosure, an electronic device capable of implementing the above-described method is also provided.

[0190] Those skilled in the art will understand that various aspects of this disclosure can be implemented as a system, method, or program product. Therefore, various aspects of this disclosure can be specifically implemented in the following forms: a completely hardware implementation, a completely software implementation (including firmware, microcode, etc.), or a combination of hardware and software aspects, collectively referred to herein as a "circuit," "module," or "system."

[0191] The following reference Figure 12 To describe an electronic device 1200 according to such an embodiment of the present disclosure. Figure 12 The electronic device 1200 shown is merely an example and should not be construed as limiting the functionality and scope of use of the embodiments disclosed herein.

[0192] like Figure 12 As shown, the electronic device 1200 is manifested in the form of a general-purpose computing device. The components of the electronic device 1200 may include, but are not limited to: at least one processing unit 1210, at least one storage unit 1220, a bus 1230 connecting different system components (including storage unit 1220 and processing unit 1210), and a display unit 1240.

[0193] The storage unit stores program code that can be executed by the processing unit 1210, causing the processing unit 1210 to perform the steps described in the "Exemplary Methods" section of this specification according to various exemplary embodiments of this disclosure. For example, the processing unit 1210 can perform actions such as... Figure 2 The steps shown are as follows: Step S210: Obtain historical patient sample data of historical users, and extract the first omics data of the first gene point of the historical users from the historical patient sample data; Step S220: Obtain domain knowledge, and construct a heterogeneous network based on the first gene point, the first omics data and the domain knowledge; Step S230: Train the network model to be trained according to the heterogeneous network to obtain a feature fusion model; wherein, the feature fusion model is used to perform feature fusion on the first omics data based on domain knowledge and to perform data prediction based on the fused features.

[0194] For example, the processing unit 1210 can perform actions such as Figure 10 Step S1010: Obtain the current sample data of the current user, and extract the second omics data of the second gene point of the current user from the current sample data; Step S1020: Input the second omics data of the second gene point into the feature fusion model to obtain the data prediction result; wherein, the feature fusion model is trained based on the training method of the feature fusion model described above; Step S1030: Determine the user category to which the current user belongs based on the data prediction result.

[0195] Storage unit 1220 may include readable media in the form of volatile storage units, such as random access memory (RAM) 12201 and / or cache memory 12202, and may further include read-only memory (ROM) 12203. Storage unit 1220 may also include a program / utility 12204 having a set (at least one) of program modules 12205, including but not limited to: an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include an implementation of a network environment.

[0196] Bus 1230 can represent one or more of several types of bus structures, including a memory cell bus or memory cell controller, a peripheral bus, a graphics acceleration port, a processing unit, or a local bus using any of the various bus structures.

[0197] Electronic device 1200 can also communicate with one or more external devices 1300 (e.g., keyboard, pointing device, Bluetooth device, etc.), and with one or more devices that enable a user to interact with electronic device 1200, and / or with any device that enables electronic device 1200 to communicate with one or more other computing devices (e.g., router, modem, etc.). This communication can be performed via input / output (I / O) interface 1250. Furthermore, electronic device 1200 can also communicate with one or more networks (e.g., local area network (LAN), wide area network (WAN), and / or public networks, such as the Internet) via network adapter 1260. As shown, network adapter 1260 communicates with other modules of electronic device 1200 via bus 1230. It should be understood that, although not shown in the figures, other hardware and / or software modules can be used in conjunction with electronic device 1200, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.

[0198] From the above description of the embodiments, those skilled in the art will readily understand that the exemplary embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solutions according to the embodiments of this disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, external hard drive, etc.) or on a network, including several instructions to cause a computing device (such as a personal computer, server, terminal device, or network device, etc.) to execute the methods according to the embodiments of this disclosure.

[0199] In exemplary embodiments of this disclosure, a computer-readable storage medium is also provided, on which a program product capable of implementing the methods described above is stored. In some possible implementations, various aspects of this disclosure may also be implemented as a program product including program code that, when the program product is run on a terminal device, causes the terminal device to perform the steps of the various exemplary embodiments of this disclosure described in the "Exemplary Methods" section above.

[0200] The program product for implementing the above-described method according to embodiments of the present disclosure may employ a portable compact disc read-only memory (CD-ROM) and include program code, and may run on a terminal device, such as a personal computer. However, the program product of the present disclosure is not limited thereto. In this document, the readable storage medium may be any tangible medium containing or storing a program that may be used by or in conjunction with an instruction execution system, apparatus, or device.

[0201] The program product may employ any combination of one or more readable media. A readable medium may be a readable signal medium or a readable storage medium. A readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of readable storage media (a non-exhaustive list) include: an electrical connection having one or more wires, a portable disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0202] Computer-readable signal media may include data signals propagated in baseband or as part of a carrier wave, carrying readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A readable signal medium may also be any readable medium other than a readable storage medium, capable of sending, propagating, or transmitting programs for use by or in conjunction with an instruction execution system, apparatus, or device.

[0203] The program code contained on the readable medium may be transmitted using any suitable medium, including but not limited to wireless, wired, optical fiber, RF, etc., or any suitable combination thereof.

[0204] Program code for performing the operations of this disclosure can be written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Java and C++, and conventional procedural programming languages ​​such as C or similar languages. The program code can execute entirely on the user's computing device, partially on the user's computing device, as a standalone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In cases involving remote computing devices, the remote computing device can be connected to the user's computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (e.g., via the Internet using an Internet service provider).

[0205] Furthermore, the above figures are merely illustrative of the processes included in the method according to exemplary embodiments of this disclosure and are not intended to be limiting. It is readily understood that the processes shown in the above figures do not indicate or limit the temporal order of these processes. Additionally, it is readily understood that these processes may be executed synchronously or asynchronously, for example, in multiple modules.

[0206] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the invention described herein. This application is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not invented by this disclosure. The specification and embodiments are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the claims.

Claims

1. A training method for a feature fusion model, characterized in that, include: Acquire historical patient sample data of historical users, and extract the first omics data of the first gene point of the historical users from the historical patient sample data; Acquire domain knowledge and construct a heterogeneous network based on the first gene point, the first omics data, and the domain knowledge; The heterogeneous network includes a first heterogeneous network and / or a second heterogeneous network; The first heterogeneous network is a heterogeneous network with the first gene point as the first child node, the first omics data of the first gene point as the node feature of the first child node, the first connection relationship between the current gene points in the biological signaling pathway as the first connection edge, and the second connection relationship between the first gene pairs composed of the first gene points as the second connection edge. The second heterogeneous network is a heterogeneous network with a first gene point as the first child node, the first omics data of the first gene point as the second child node, the first connection relationship between current gene points in the biological signaling pathway as the first connection edge, the second connection relationship between the first gene pairs composed of the first gene points as the second connection edge, and the third connection relationship between the first child node and the second child node as the third connection edge. The first heterogeneous network is used to reflect the correlation between data within the same omics dataset and the correlation between multiple omics datasets; the second heterogeneous network is used to reflect the correlation between data within the same omics dataset and the correlation between multiple omics datasets. Based on the heterogeneous network, the network model to be trained is trained to obtain a feature fusion model; wherein, the feature fusion model is used to perform feature fusion on the first omics data based on domain knowledge and to perform data prediction based on the fused features.

2. The training method for the feature fusion model according to claim 1, characterized in that, The domain knowledge includes biological signaling pathways and current gene points included in those pathways; The first set of omics data includes one or more of the following: DNA methylation data, gene mutation SNV data, copy number variation CNV data, and gene expression data.

3. The training method for the feature fusion model according to claim 1, characterized in that, The network model to be trained includes a first network model to be trained and / or a second network model to be trained. The first network model to be trained includes a graph network model and a first classifier, and the second network model to be trained includes an autoencoder model and a second classifier. The graph network model includes a relational graph convolutional network model and / or a graph attention network model.

4. The training method for the feature fusion model according to claim 1, characterized in that, Based on the first gene point, the first omics data, and the domain knowledge, a heterogeneous network is constructed, including: Construct a set of first child nodes based on the first gene point, and construct a set of first node features based on the first omics data of the first gene point; Based on the biological signaling pathways included in the domain knowledge and the current gene points included in the biological signaling pathways, determine the first connection relationship between the first gene points, and construct a first set of connection edges based on the first connection relationship; Calculate the first correlation coefficient between the first gene pairs consisting of the first gene points, and determine the second connection relationship between the first gene points based on the first correlation coefficient; A second set of connecting edges is constructed based on the second connection relationship, and a first heterogeneous network is constructed based on the first set of child nodes, the first set of node features, the first set of connecting edges, and the second set of connecting edges.

5. The training method for the feature fusion model according to claim 1, characterized in that, Based on the first gene point, the first omics data, and the aforementioned domain knowledge, a heterogeneous network is constructed, which further includes: Construct a first set of child nodes based on the first gene point, and construct a second set of child nodes based on the first omics data of the first gene point; Based on the biological signaling pathways included in the domain knowledge and the current gene points included in the biological signaling pathways, determine the first connection relationship between the first gene points, and construct a first set of connection edges based on the first connection relationship; Calculate the first correlation coefficient between the first gene pairs consisting of the first gene points, and determine the second connection relationship between the first gene points based on the first correlation coefficient; A second set of connecting edges is constructed based on the second connectivity relationship, and a third set of connecting edges is constructed based on the third connectivity relationship between the first gene point and the first omics data. Construct a second heterogeneous network based on the first set of child nodes, the second set of child nodes, the first set of connecting edges, the second set of connecting edges, and the third set of connecting edges.

6. The training method for the feature fusion model according to claim 4 or 5, characterized in that, The first gene pair consisting of the first gene point includes a first sub-gene point and a second sub-gene point; The calculation of the first correlation coefficient between the first gene pairs formed by the first gene points includes: Obtain the first sub-omics data of the first sub-gene point in the first gene pair consisting of the first gene point, and the second sub-omics data of the second sub-gene point; First sub-gene expression data are extracted from the first sub-omics data, and second sub-gene expression data are extracted from the second sub-omics data. The first correlation coefficient is calculated based on the expression data of the first daughter gene and the expression data of the second daughter gene.

7. The training method for the feature fusion model according to claim 1, characterized in that, Based on the heterogeneous network, the network model to be trained is trained to obtain a feature fusion model, including: Based on the first heterogeneous network, train the first network model to be trained to obtain the first feature fusion model; and / or Based on the second heterogeneous network, the second network model to be trained is trained to obtain the second feature fusion model.

8. The training method for the feature fusion model according to claim 7, characterized in that, Based on the first heterogeneous network, the first network model to be trained is trained to obtain the first feature fusion model, including: The first heterogeneous network is input into the graph network model in the first network model to be trained to obtain the first feature representation corresponding to the first heterogeneous network, and the first feature representation is input into the first classifier in the first network model to be trained to obtain the first predicted label; Based on the real user tags of the historical users and the first predicted tags, a first target loss function is constructed, and the first network model to be trained is trained based on the first target loss function to obtain a first feature fusion model.

9. The training method for the feature fusion model according to claim 7, characterized in that, Based on the second heterogeneous network, the second network model to be trained is trained to obtain the second feature fusion model, including: The second heterogeneous network is subjected to representation learning to obtain a second feature representation corresponding to the second heterogeneous network, and the first omics data of the first gene point is input into the encoding module of the second network model to be trained to obtain the first reconstructed data corresponding to the first omics data. Based on the first set of data, the first reconstructed data, and the second feature representation, a second objective loss function is constructed, and the autoencoder model in the second network model to be trained is trained according to the second objective function to obtain the trained encoding model; The first omics data is input into the encoding module of the trained encoding model to obtain the second reconstructed data, and the second reconstructed data is input into the second classifier of the second network model to be trained to obtain the second predicted label; Based on the second predicted label and the real user labels of the historical users, the second classifier is trained to obtain the second feature fusion model.

10. The training method for the feature fusion model according to claim 9, characterized in that, The first omics data of the first gene point is input into the encoding module of the second network model to be trained to obtain the first reconstructed data corresponding to the first omics data, including: The first omics data is mapped using the first nonlinear function in the encoding module of the second network model to be trained, to obtain intermediate variables; The intermediate variables are mapped by the second nonlinear function in the encoding module of the second network model to be trained, and the first reconstructed data corresponding to the first omics data is obtained; wherein the data expression form of the first reconstructed data is consistent with the data expression form of the first omics data.

11. The training method for the feature fusion model according to claim 9, characterized in that, Based on the first omics data, the first reconstructed data, and the second feature representation, a second objective loss function is constructed, including: The first reconstructed data is input into the decoding module of the second network model to be trained to obtain the third feature representation corresponding to the second heterogeneous network, and the first sub-loss function is constructed based on the first omics data and the first reconstructed data. A second sub-loss function is constructed based on the second and third feature representations, and a second objective function is constructed based on the first and second sub-loss functions.

12. The training method for the feature fusion model according to claim 11, characterized in that, The first sub-loss function is constructed based on the first omics data and the first reconstructed data, including: Calculate the absolute value of the first difference between the first omics data and the first reconstructed data, and square the absolute value of the first difference to obtain the first square calculation result; The first sub-loss function is obtained by summing the results of the first squared calculation.

13. The training method for the feature fusion model according to claim 11, characterized in that, The second sub-loss function is constructed based on the second and third feature representations, including: Obtain the first sub-vector representation of the first sub-gene point and the second sub-vector representation of the second sub-gene point in the first gene pair included in the second feature representation, and calculate the first similarity of the first gene pair based on the first sub-vector representation and the second sub-vector representation; Obtain the third sub-vector representation of the first sub-gene point and the fourth sub-vector representation of the second sub-gene point in the first gene pair included in the third feature representation, and calculate the second similarity of the first gene pair based on the third sub-vector representation and the fourth sub-vector representation. Calculate the absolute value of the second difference between the first similarity and the second similarity, and sum the absolute values ​​of the second difference to obtain the second sub-loss function.

14. A method for classifying cancer patients, characterized in that, include: Obtain the current sample data of the current user, and extract the second omics data of the second gene point of the current user from the current sample data; The second omics data of the second gene point is input into the feature fusion model to obtain the data prediction result; wherein, the feature fusion model is trained based on the training method of the feature fusion model according to any one of claims 1-13; Based on the data prediction results, the user category to which the current user belongs is determined.

15. The method for classifying cancer patients according to claim 14, characterized in that, The second omics data of the second gene locus is input into the feature fusion model to obtain data prediction results, including: The second omics data of the second gene point is input into the graph network model in the first feature fusion model to obtain a first user feature representation including the dimensionality-reduced second omics data and domain knowledge. The first user feature representation is then input into the first classifier of the first feature fusion model to obtain the data prediction result; and / or The second omics data of the second gene point is input into the encoding module of the autoencoding model in the second feature fusion model to obtain the second user feature representation, which includes the dimensionality-reduced second omics data and domain knowledge. The second user feature representation is then input into the second classifier of the second feature fusion model to obtain the data prediction result.

16. A training device for a feature fusion model, characterized in that, include: The first omics data extraction module is used to obtain historical patient sample data of historical users and extract the first omics data of the first gene point of the historical users from the historical patient sample data. A heterogeneous network construction module is used to acquire domain knowledge and construct a heterogeneous network based on the first gene point, the first omics data, and the domain knowledge. The heterogeneous network includes a first heterogeneous network and / or a second heterogeneous network; The first heterogeneous network is a heterogeneous network with the first gene point as the first child node, the first omics data of the first gene point as the node feature of the first child node, the first connection relationship between the current gene points in the biological signaling pathway as the first connection edge, and the second connection relationship between the first gene pairs composed of the first gene points as the second connection edge. The second heterogeneous network is a heterogeneous network with a first gene point as the first child node, the first omics data of the first gene point as the second child node, the first connection relationship between current gene points in the biological signaling pathway as the first connection edge, the second connection relationship between the first gene pairs composed of the first gene points as the second connection edge, and the third connection relationship between the first child node and the second child node as the third connection edge. The first heterogeneous network is used to reflect the correlation between data within the same omics dataset and the correlation between multiple omics datasets; the second heterogeneous network is used to reflect the correlation between data within the same omics dataset and the correlation between multiple omics datasets. The model training module is used to train the network model to be trained based on the heterogeneous network to obtain a feature fusion model; wherein, the feature fusion model is used to perform feature fusion on the first omics data based on domain knowledge and to perform data prediction based on the fused features.

17. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the training method of the feature fusion model according to any one of claims 1-13, and the classification method of cancer users according to claim 14 or 15.

18. An electronic device, characterized in that, include: processor; as well as Memory for storing the executable instructions of the processor; The processor is configured to execute the training method of the feature fusion model according to any one of claims 1-13 and the classification method of cancer users according to claim 14 or 15 by executing the executable instructions.

Citation Information

Patent Citations

  • Drug sensitivity prediction method and device based on multi-omics data fusion

    CN113782089A

  • Method and system for analyzing biological networks

    WO2016118513A1