Multi-omics data conjoint analysis and model training method and device, equipment and medium

By constructing a graph structure model and training a disease mechanism model with prior biological knowledge, the problem of insufficient biological interpretability of existing pathway analysis methods is solved, and biologically interpretable disease analysis and prediction and accurate characterization of the strength of molecular feature interactions are achieved.

CN121483386APending Publication Date: 2026-02-06THE CHINESE UNIV OF HONG KONG (SHENZHEN)
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511659426.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-13
Publication Date
2026-02-06

AI Technical Summary

Technical Problem

Existing pathway analysis methods are insufficient to provide detailed explanations of biological mechanisms in molecular biology and bioinformatics, and machine learning methods have limited ability to capture the functional dependencies of biological networks.

Method used

A graph structure model is constructed, using prior biological knowledge to use molecular features as vertices and biological interactions as edges. Central and peripheral features are selected, and a disease mechanism model is trained through biological pathway transport layers and fully connected layers to simulate the signal transduction mechanism in organisms.

Benefits of technology

It enables biologically interpretable disease analysis and prediction, enhances the interpretability of the model, accurately performs disease analysis, and provides the strength of biological interactions between molecular features.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121483386A_ABST
    Figure CN121483386A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data processing, in particular to a multi-omics data conjoint analysis and model training method and device, equipment and a medium, and the method comprises the steps: constructing a graph structure based on prior biological knowledge, selecting a center feature from vertexes of the graph structure, and taking other vertexes except the center feature in the graph structure as edge features; constructing a layer network layer based on the graph structure, wherein the first layer is an edge feature with the shortest path from each center feature as a hop in the graph structure; training a preset disease mechanism model by using the training sample to obtain a trained disease mechanism model; a biological pathway transmission layer in the disease mechanism model is used for transmitting the information of the molecular characteristics of the first layer in the molecular characteristics of the sample to the direct connection characteristics in the first layer, gradually transmitting the information of the edge molecular characteristics to the central characteristics of the zero layer, and obtaining the updated central characteristics for disease analysis. According to the technical scheme, path analysis details can be explained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of data processing technology, specifically to a method, apparatus, device, and medium for joint analysis of multi-omics data and model training. Background Technology

[0002] Pathway analysis is a core technology in modern molecular biology and bioinformatics, used to extract meaningful biological insights from high-throughput omics data such as transcriptomics, proteomics, and metabolomics.

[0003] Traditional pathway analysis methods can be broadly categorized into statistical enrichment methods and machine learning methods. Statistical enrichment methods can identify pathways that are overrepresented in differentially expressed features, but they can only provide summary-level biological annotations lacking details about the mechanisms of biological pathways. For example, they can only indicate that a certain pathway is related to a certain disease, but the details of the pathway's internal mechanisms are missing. Machine learning methods can perform high-dimensional modeling of omics data and pathway information, generating powerful predictive models for clinical outcomes; however, such models typically treat molecular features as independent variables, have limited ability to capture the inherent functional dependencies of biological networks, and often operate as a "black box," with their internal representations difficult to interpret in biological terms. Therefore, there is an urgent need for a mechanism-interpretable pathway analysis method. Summary of the Invention

[0004] To address the problems in related technologies, embodiments of this disclosure provide a method, apparatus, device, and medium for joint analysis of multi-omics data and model training.

[0005] In a first aspect, this disclosure provides a method for training a multi-omics data joint analysis model, including: A graph structure is constructed based on prior biological knowledge, wherein each vertex in the graph structure represents a molecular feature, and each edge represents the biological interaction between the two vertices connected by the edge. The molecular features are multi-omics data. Multiple vertices are selected from the vertices of the graph structure as central features, and the other vertices in the graph structure other than the central features are edge features. The central features are biologically relevant to the analysis task. Based on the graph structure, construct Layered network, wherein the 0th layer of the network is composed of the central feature, the 1st layer of the network is composed of the central feature, the 2nd layer of the network is composed of the central feature, the 3rd layer of the network is composed of the central feature, the 4th layer of the network is composed of the central feature, the 5th layer of the network is composed of the central feature, the 6th layer of the network is composed of the central feature, the 7th layer of the network is composed of the central feature, the The layer is the shortest path from each central feature in the graph structure. Edge features of jumps It is an integer greater than 1; Multiple training samples are obtained, the training samples including the molecular features and disease labels of the samples; Using the training samples, a pre-defined disease mechanism model is trained to obtain a trained disease mechanism model; wherein, the disease mechanism model includes a biological pathway transport layer and a fully connected layer, the biological pathway transport layer being used to transfer the molecular features of the samples... Information about the molecular characteristics of the layer is transmitted to its first layer. The direct connection features in the layer gradually transmit the information of the edge molecular features to the center features of the 0th layer to obtain the updated center features, and then transmit the updated center features to the fully connected layer.

[0006] Secondly, this disclosure provides a method for joint analysis of multi-omics data, including: Obtain the molecular characteristics of the individual to be analyzed; The molecular characteristics of the individual to be analyzed are input into a pre-trained disease mechanism model. The biological pathway transport layer in the pre-trained disease mechanism model transmits the molecular characteristics of the individual to be analyzed. Information about the molecular characteristics of the layer is transmitted to its first layer. The direct connection features in the layer gradually transmit the information of the edge features to the center features of the 0th layer to obtain the updated center features, and then transmit the updated center features to the fully connected layer of the pre-trained disease mechanism model to obtain the analysis results output by the pre-trained disease mechanism model. The analysis results and the biological pathways corresponding to the molecular characteristics of the individual to be analyzed are output, and the biological pathways show the interaction strength between the molecular characteristics of the individual to be analyzed.

[0007] Thirdly, this disclosure provides a multi-omics data joint analysis model training device, including: The graph construction module is configured to construct a graph structure based on prior biological knowledge, wherein each vertex in the graph structure represents a molecular feature, each edge represents a biological interaction between the two vertices connected by the edge, and the molecular feature is multi-omics data; The feature selection module is configured to select multiple vertices from the vertices of the graph structure as central features, and the other vertices in the graph structure besides the central features as edge features. The central features are biologically relevant to the analysis task. The network layer construction module is configured to build based on the graph structure. Layered network, wherein the 0th layer of the network is composed of the central feature, the 1st layer of the network is composed of the central feature, the 2nd layer of the network is composed of the central feature, the 3rd layer of the network is composed of the central feature, the 4th layer of the network is composed of the central feature, the 5th layer of the network is composed of the central feature, the 6th layer of the network is composed of the central feature, the 7th layer of the network is composed of the central feature, the The layer is the shortest path from each central feature in the graph structure. Edge features of jumps It is an integer greater than 1; The sample acquisition module is configured to acquire multiple training samples, which include the molecular features and disease labels of the samples; The model training module is configured to use the training samples to train a preset disease mechanism model, thereby obtaining a trained disease mechanism model; wherein, the disease mechanism model includes a biological pathway transport layer and a fully connected layer, the biological pathway transport layer being used to transfer the molecular features of the samples... Information about the molecular characteristics of the layer is transmitted to its first layer. The direct connection features in the layer gradually transmit the information of the edge molecular features to the center features of the 0th layer to obtain the updated center features, and then transmit the updated center features to the fully connected layer.

[0008] Fourthly, this disclosure provides a multi-omics data joint analysis device, comprising: The acquisition module is configured to acquire the molecular characteristics of the individual to be analyzed; The analysis module is configured to input the molecular characteristics of the individual to be analyzed into a pre-trained disease mechanism model, wherein the biological pathway transport layer in the pre-trained disease mechanism model inputs the molecular characteristics of the individual to be analyzed into the molecular pathway transport layer. Information about the molecular characteristics of the layer is transmitted to its first layer. The directly connected features in the layer gradually transmit the information of the edge features to the center features of the 0th layer to obtain the updated center features, and then transmit the updated center features to the fully connected layer of the pre-trained disease mechanism model to obtain the analysis results output by the pre-trained disease mechanism model; the pre-trained disease mechanism model is trained according to the method described in any one of the first aspects; The output module is configured to output the analysis results and the biological pathways corresponding to the molecular characteristics of the individual to be analyzed, wherein the biological pathways show the interaction strength between the molecular characteristics of the individual to be analyzed.

[0009] Fifthly, this disclosure provides an electronic device including a processor and a memory, wherein the memory stores computer program instructions, and when the processor executes the computer program instructions, the processor performs the method described in either the first or second aspect.

[0010] In a sixth aspect, this disclosure provides a computer-readable storage medium storing computer program instructions, which, when executed by a processor, enable the processor to perform the method described in either the first or second aspect.

[0011] The method provided in this embodiment can construct a graph structure based on prior biological knowledge. Each vertex in the graph structure represents a molecular feature, and each edge represents the biological interaction between the two vertices connected by the edge. The molecular features are multi-omics data. Multiple vertices are selected from the vertices of the graph structure as central features, and the other vertices in the graph structure besides the central features are edge features. The central features are biologically relevant to the analysis task. Based on this graph structure, a graph structure is constructed... Layered network, wherein the 0th layer of the network is composed of the central feature, the 1st layer of the network is composed of the central feature, the 2nd layer of the network is composed of the central feature, the 3rd layer of the network is composed of the central feature, the 4th layer of the network is composed of the central feature, the 5th layer of the network is composed of the central feature, the 6th layer of the network is composed of the central feature, the 7th layer of the network is composed of the central feature, the The layer is the shortest path from each central feature in the graph. Edge features of jumps The integer is greater than 1; multiple training samples are obtained, the training samples include the molecular features and disease labels of the samples; using the training samples, a preset disease mechanism model is trained to obtain the trained disease mechanism model; wherein, the disease mechanism model includes a biological pathway transport layer and a fully connected layer, the biological pathway transport layer is used to transfer the molecular features of the samples into the molecular pathway transport layer. Information about the molecular characteristics of the layer is transmitted to its first layer. The direct connections in the layers progressively transmit information from peripheral molecular features to the central features in layer 0, resulting in updated central features, which are then transmitted to the fully connected layers. This biological pathway transport layer of the disease mechanism model simulates the signal transduction mechanisms between pathways in organisms. In this knowledge-driven disease mechanism model neural network, each neuron represents a specific biological feature. Connections between neurons are established only when the corresponding feature participates in a recorded biological response. In each layer, the model simulates how predetermined peripheral features transmit information to selected central features. The cumulative effect of these peripheral features propagates through the network to the central features. Information from the peripheral features flows along pathway connections to biologically important central features. These central features are ultimately used for disease analysis and prediction. This not only enables accurate disease analysis and prediction but also interprets the relationship between these features and the analysis and prediction results through the information transmission from peripheral features to central features in the biological pathway transport layer, providing a biologically intuitive representation and enhancing the model's interpretability. Attached Figure Description

[0012] Figure 1 A flowchart illustrating a multi-omics data joint analysis model training method provided in an embodiment of this disclosure is shown.

[0013] Figure 2 A flowchart illustrating a multi-omics data joint analysis method provided in an embodiment of this disclosure is shown.

[0014] Figure 3This diagram illustrates the structure of a multi-omics data joint analysis model training device provided in an embodiment of the present disclosure.

[0015] Figure 4 This diagram illustrates a structural block diagram of a multi-omics data joint analysis device provided in an embodiment of this disclosure.

[0016] Figure 5 A structural block diagram of an electronic device according to an embodiment of the present disclosure is shown.

[0017] Figure 6 A schematic diagram of the structure of a computer system suitable for implementing the methods of the embodiments of this disclosure is shown. Detailed Implementation

[0018] In the following, exemplary embodiments of the present disclosure will be described in detail with reference to the accompanying drawings to enable those skilled in the art to readily implement them. Furthermore, for clarity, portions unrelated to the description of exemplary embodiments have been omitted from the drawings.

[0019] In this disclosure, it should be understood that terms such as “comprising” or “having” are intended to indicate the presence of features, figures, steps, behaviors, components, parts or combinations thereof disclosed in this specification, and are not intended to exclude the possibility of the presence or addition of one or more other features, figures, steps, behaviors, components, parts or combinations thereof.

[0020] It should also be noted that, unless otherwise specified, the embodiments and features described in this disclosure can be combined with each other. This disclosure will now be described in detail with reference to the accompanying drawings and embodiments.

[0021] Figure 1 This illustration shows a flowchart of a multi-omics data joint analysis model training method provided in an embodiment of this disclosure, such as... Figure 1 As shown, the training method for the multi-omics data joint analysis model includes the following steps S101 to S105.

[0022] In step S101, a graph structure is constructed based on prior biological knowledge, wherein each vertex in the graph structure represents a molecular feature, each edge represents a biological interaction between the two vertices connected by the edge, and the molecular feature is multi-omics data.

[0023] In one possible implementation, the prior biological knowledge refers to existing biological knowledge, such as a pathway database that records the characteristic interactions of molecules within a pathway. A graph structure can be constructed based on this prior biological knowledge. ,in, Let be the set of vertices in the graph. one of the vertices ( ) represents a molecular characteristic. Let be the set of edges in the graph, and each edge ( Connect two vertices and , representing the two vertices whose connection is derived from the pathway database. and Known biological interactions between them.

[0024] In one possible implementation, the graph structure It can be an undirected graph, in which case each edge Represents the two vertices whose connection is derived from the path database. and Known unidirectional regulated biological interactions between them; or, the diagram structure. It can be a directed graph, in which case each edge Represents the two vertices whose connection is derived from the path database. and Biological interactions known to be bidirectional between them.

[0025] In one possible implementation, the molecular feature can be multi-omics data, such as genes (e.g., DNA sequences encoding proteins), proteins (executors of functions, including enzymes, structural proteins, signaling molecules, etc.), metabolites (e.g., small molecule compounds, substrates and products of biochemical reactions), and regulatory RNA molecules (e.g., miRNA / lncRNA).

[0026] Biological interactions can include protein-protein interactions (such as complex formation and functional cooperation), enzyme-substrate metabolic transformation relationships, and regulatory relationships between transcription factors and target genes, as well as signal transduction.

[0027] In step S102, multiple vertices are selected from the vertices of the graph structure as central features, and the other vertices in the graph structure besides the central features are edge features. The central features are biologically relevant to the analysis task.

[0028] In one possible implementation, it can be derived from the vertex set. Select These vertices are used as central features, and these central features can be denoted as: These features were selected due to their potential biological relevance to the analytical task. They can be selected by experts based on experience or automatically based on existing biological knowledge. This vertex set... Except The remaining vertices outside the central feature are considered edge features, and the remaining edge features may affect the central feature through path connections.

[0029] In step S103, a graph structure is constructed based on the graph structure. Layered network, wherein the 0th layer of the network is composed of the central feature, the 1st layer of the network is composed of the central feature, the 2nd layer of the network is composed of the central feature, the 3rd layer of the network is composed of the central feature, the 4th layer of the network is composed of the central feature, the 5th layer of the network is composed of the central feature, the 6th layer of the network is composed of the central feature, the 7th layer of the network is composed of the central feature, the The layer is the shortest path from each central feature in the graph structure. Edge features of jumps It is an integer greater than 1.

[0030] In one possible implementation, for each edge feature This edge feature can be calculated. arrive The shortest path distance among the distances to each central feature can be calculated using the following formula: ; If edge features The shortest path distance is equal to and Then the features Assigned to a set Therefore, layer 0 in this network layer Composed of the central feature, the first layer Including distance The shortest path distance is the edge feature of one hop, the second layer Including distance The shortest path distance is the edge feature of two hops, and so on until the first hop is obtained. layer Thus we can obtain Layered network layer.

[0031] In step S104, multiple training samples are obtained, the training samples including the molecular features and disease labels of the samples.

[0032] In one possible implementation, the training samples can be patients suffering from the disease to be analyzed, or they can include humans who do not suffer from the disease. Assume the number of training samples is... The number of molecular features in the sample is Then the molecular feature matrix of these samples can be denoted as... ,remember This refers to the disease label (i.e., the actual disease diagnosis).

[0033] In step S105, the training samples are used to train a preset disease mechanism model to obtain a trained disease mechanism model.

[0034] In one possible implementation, the disease mechanism model includes a biological pathway transport layer and a fully connected layer, wherein the biological pathway transport layer is used to transfer the molecular characteristics of the sample... Information about the molecular characteristics of the layer is transmitted to its first layer. The direct connection features in the layer progressively transmit information from the edge molecular features to the central features of the 0th layer, obtaining updated central features, and then transmitting the updated central features to the fully connected layer. The specific training process of this disease mechanism model can be as follows: the molecular features of the sample are input into the biological pathway transport layer, and the biological pathway transport layer transmits the information from the molecular features of the sample to the central features of the 0th layer. Information about the molecular characteristics of the layer is transmitted to its first layer. Direct connectivity features in a layer, where direct connectivity refers to connections between the layers in a graph structure via an edge. The molecular features of the layer are directly connected to the first The molecular features of the layer; in this way, the information of the edge molecular features can be gradually transmitted to the central features of the 0th layer to obtain the updated central features, and the updated central features are transmitted to the fully connected layer to obtain the analysis results output by the disease mechanism model. Based on the analysis results of the disease mechanism model and the disease label, the loss function is calculated, and the parameters of the disease mechanism model are updated based on the loss function until the loss function converges or reaches the maximum number of training iterations, thus obtaining the trained disease mechanism model.

[0035] In one possible implementation, the biological pathway transport layer in the disease mechanism model reflects a biological pathway graph, where each node corresponds to a molecular feature and each edge reflects a recorded biological interaction. This ensures that the structure of the disease mechanism model is biologically reasonable and interpretable. Crucially, through learning from training samples, each biological interaction is represented as a trainable mapping, quantifying how one molecular feature influences another along the pathway. This representation, by utilizing residual connections that encode direct and context-dependent effects, naturally adapts to signal propagation, similar to biological processes such as gene regulation, signal transduction, and metabolic control. Therefore, the biological pathway transport layer in the trained disease mechanism model can effectively reflect the strength of biological interactions between various molecular features, as well as the cumulative influence of these interactions on central features layer by layer along the pathway. Consequently, it affects the disease by influencing these key central features, providing strong interpretability for the analysis results.

[0036] This implementation method can construct a graph based on prior biological knowledge, wherein each vertex in the graph structure represents a molecular feature, and each edge represents the biological interaction between the two vertices connected by the edge. The molecular features are multi-omics data. Multiple vertices are selected from the vertices of the graph structure as central features, and the other vertices in the graph structure besides the central features are edge features. The central features are biologically relevant to the analysis task. Based on the graph structure, a graph is constructed... Layered network, wherein the 0th layer of the network is composed of the central feature, the 1st layer of the network is composed of the central feature, the 2nd layer of the network is composed of the central feature, the 3rd layer of the network is composed of the central feature, the 4th layer of the network is composed of the central feature, the 5th layer of the network is composed of the central feature, the 6th layer of the network is composed of the central feature, the 7th layer of the network is composed of the central feature, the The layer is the shortest path from each central feature in the graph structure. Edge features of jumps The integer is greater than 1; multiple training samples are obtained, the training samples include the molecular features and disease labels of the samples; using the training samples, a preset disease mechanism model is trained to obtain the trained disease mechanism model; wherein, the disease mechanism model includes a biological pathway transport layer and a fully connected layer, the biological pathway transport layer is used to transfer the molecular features of the samples into the molecular pathway transport layer. Information about the molecular characteristics of the layer is transmitted to its first layer. The direct connections in the layers progressively transmit information from peripheral molecular features to the central features in layer 0, resulting in updated central features, which are then transmitted to the fully connected layers. This biological pathway transport layer of the disease mechanism model simulates the signal transduction mechanisms between pathways in organisms. In this knowledge-driven disease mechanism model neural network, each neuron represents a specific biological feature. Connections between neurons are established only when the corresponding feature participates in a recorded biological response. In each layer, the model simulates how predetermined peripheral features transmit information to selected central features. The cumulative effect of these peripheral features propagates through the network to the central features. Information from the peripheral features flows along pathway connections to biologically important central features. These central features are ultimately used for disease analysis and prediction. This not only enables accurate disease analysis and prediction but also interprets the relationship between these features and the analysis and prediction results through the information transmission from peripheral features to central features in the biological pathway transport layer, providing a biologically intuitive representation and enhancing the model's interpretability.

[0037] In one possible implementation, the molecular characteristics of the sample are... Information about the molecular characteristics of the layer is transmitted to its first layer. Direct connection features in layers include: The molecular characteristics of the sample are analyzed according to the following formula. Information on the molecular characteristics of the layer Transmitted to the Information is obtained from direct connection features in the layer. : ; In the formula, , , , These are the parameters to be trained. It is an activation function. This represents element-wise multiplication; yes The adjacency matrix, The Middle Line number Column elements for: ; yes The residual connectivity matrix, The Middle row vector for: ; In the formula For position A unit vector with a value of 1 and a value of zero at all other positions; It is the sample Number of molecular features in the layer It is the first of the samples Number of molecular features in the layer This refers to the sample number. The first layer of molecular characteristics Individual molecular characteristics; This refers to the sample number. The first layer of molecular characteristics Molecular characteristics.

[0038] In the above formula, and Biological interactions refer to and They are directly connected together in the graph structure.

[0039] In one possible implementation, the fully connected layer in the above-described disease mechanism model includes a multilayer perceptron (MLP).

[0040] In this implementation, a novel neural network architecture based on multilayer perceptrons (MLPs) is used to constrain network connectivity according to planned biological pathways. This design mitigates overfitting while maintaining strong predictive power and interpretability.

[0041] In one possible implementation, the method further includes: Based on the biological pathway transport layer in the trained disease mechanism model, the connections between molecular features in the biological pathway transport layer with an interaction strength less than a predetermined threshold are removed to obtain the final disease mechanism model.

[0042] In this implementation, to further simplify the biological pathway transport layer in the disease mechanism model, connections where the interaction strength between molecular features is less than a predetermined threshold can be removed. This effectively simplifies signal transmission to an identity mapping, and the source features in the removed connections can be considered informationless. After pruning edges with low absolute weights (i.e., weak interaction strength between molecular features), the remaining connections can be visualized to reflect strong interaction relationships between molecular features. For the completeness of pathway interpretation, the previously removed features are simply no longer used in the disease mechanism model, but can be reintroduced into the biological pathway as context nodes, ensuring that users can further explore biological interpretability.

[0043] In this implementation, parameters from the disease mechanism model can be used. This indicates the strength of interactions between molecular features in the transport layer of biological pathways. for The first in the matrix Okay, number The absolute value of the column position parameter; when When the value is less than the predetermined threshold, it can be removed. and The connection between them.

[0044] Figure 2 This illustration shows a flowchart of a multi-omics data joint analysis method provided in an embodiment of this disclosure, as shown below. Figure 2 As shown, the multi-omics data joint analysis method includes the following steps S201 to S203: In step S201, the molecular characteristics of the individual to be analyzed are obtained; In step S202, the molecular characteristics of the individual to be analyzed are input into a pre-trained disease mechanism model. The biological pathway transport layer in the pre-trained disease mechanism model transmits the molecular characteristics of the individual to be analyzed. Information about the molecular characteristics of the layer is transmitted to its first layer. The direct connection features in the layer gradually transmit the information of the edge features to the center features of the 0th layer to obtain the updated center features, and then transmit the updated center features to the fully connected layer of the pre-trained disease mechanism model to obtain the analysis results output by the pre-trained disease mechanism model. In step S203, the analysis results and the biological pathways corresponding to the molecular characteristics of the individual to be analyzed are output, and the biological pathways show the interaction strength between the molecular characteristics of the individual to be analyzed.

[0045] In one possible implementation, after a disease mechanism model is pre-trained using the above-described multi-omics data joint analysis model training method, the disease mechanism model can be used to perform multi-omics data joint analysis.

[0046] In one possible implementation, the individual to be analyzed can be a patient to be diagnosed. The patient's molecular characteristics can be obtained and input into the biological pathway transport layer of the disease mechanism model. This biological pathway transport layer can then extract the molecular characteristics of the patient... Information about the molecular characteristics of the layer is transmitted to its first layer. The direct connection features in the layer allow the information of the edge features to be progressively transferred to the central features of the 0th layer, resulting in updated central features. These updated central features are then transferred to the fully connected layer of the pre-trained disease mechanism model, thereby obtaining the analysis results output by the pre-trained disease mechanism model. When outputting these analysis results, the biological pathways from each edge feature to the central feature in the patient's molecular features can also be output, showing the interaction strength between these molecular features. This provides a biological interpretation of the analysis results.

[0047] This disclosure also provides a training device for a multi-omics data joint analysis model. Figure 3 This diagram illustrates a structural block diagram of a multi-omics data joint analysis model training device provided in an embodiment of this disclosure. This device can be implemented as part or all of an electronic device through software, hardware, or a combination of both. Figure 3 As shown, the multi-omics data joint analysis model training device includes: Graph construction module 301 is configured to construct a graph structure based on prior biological knowledge, wherein each vertex in the graph structure represents a molecular feature, each edge represents a biological interaction between the two vertices connected by the edge, and the molecular feature is multi-omics data; Feature selection module 302 is configured to select multiple vertices from the vertices of the graph structure as central features, and the other vertices in the graph structure other than the central features are edge features. The central features are biologically relevant to the analysis task. Network layer construction module 303 is configured to construct based on the graph structure. Layered network, wherein the 0th layer of the network is composed of the central feature, the 1st layer of the network is composed of the central feature, the 2nd layer of the network is composed of the central feature, the 3rd layer of the network is composed of the central feature, the 4th layer of the network is composed of the central feature, the 5th layer of the network is composed of the central feature, the 6th layer of the network is composed of the central feature, the 7th layer of the network is composed of the central feature, the The layer is the shortest path from each central feature in the graph structure. Edge features of jumps It is an integer greater than 1; The sample acquisition module 304 is configured to acquire multiple training samples, the training samples including the molecular features and disease labels of the samples; The model training module 305 is configured to use the training samples to train a preset disease mechanism model, thereby obtaining a trained disease mechanism model; wherein, the disease mechanism model includes a biological pathway transport layer and a fully connected layer, the biological pathway transport layer being used to transfer the molecular features of the samples... Information about the molecular characteristics of the layer is transmitted to its first layer. The direct connection features in the layer gradually transmit the information of the edge molecular features to the center features of the 0th layer to obtain the updated center features, and then transmit the updated center features to the fully connected layer.

[0048] This disclosure also provides a device for joint analysis of multi-omics data. Figure 4 This diagram illustrates a structural block diagram of a multi-omics data joint analysis device provided in an embodiment of this disclosure. This device can be implemented as part or all of an electronic device through software, hardware, or a combination of both. Figure 4 As shown, the multi-omics data joint analysis device includes: The acquisition module 401 is configured to acquire the molecular characteristics of the individual to be analyzed; Analysis module 402 is configured to input the molecular characteristics of the individual to be analyzed into a pre-trained disease mechanism model, wherein the biological pathway transport layer in the pre-trained disease mechanism model inputs the molecular characteristics of the individual to be analyzed into the molecular pathway transport layer of the individual to be analyzed. Information about the molecular characteristics of the layer is transmitted to its first layer. The directly connected features in the layer gradually transmit the information of the edge features to the center features of the 0th layer to obtain the updated center features. The updated center features are then transmitted to the fully connected layer of the pre-trained disease mechanism model to obtain the analysis results output by the pre-trained disease mechanism model. The pre-trained disease mechanism model is trained according to the above model training method. The output module 403 is configured to output the analysis results and the biological pathways corresponding to the molecular characteristics of the individual to be analyzed, wherein the biological pathways show the interaction strength between the molecular characteristics of the individual to be analyzed.

[0049] The technical terms and features mentioned in this device embodiment are the same as or similar to those mentioned in the above method embodiment. For explanations and descriptions of the technical terms and features involved in this device, please refer to the explanations of the above method embodiment. They will not be repeated here.

[0050] This disclosure also discloses an electronic device. Figure 5 A structural block diagram of an electronic device according to an embodiment of the present disclosure is shown.

[0051] like Figure 5 As shown, the electronic device 500 includes a memory and a processor, wherein the memory is used to store one or more computer instructions, wherein the one or more computer instructions are executed by the processor to implement the method according to embodiments of the present disclosure.

[0052] Figure 6 A schematic diagram of the structure of a computer system suitable for implementing the methods of the embodiments of this disclosure is shown.

[0053] like Figure 6 As shown, the computer system 600 includes a processing unit that can execute various processes described above, based on a program stored in a read-only memory (ROM) or a program loaded from a storage portion into a random access memory (RAM). The RAM also stores various programs and data required for the operation of the computer system 600. The processing unit, ROM, and RAM are interconnected via a bus. Input / output (I / O) interfaces are also connected to the bus.

[0054] The following components are connected to the I / O interface: input sections including keyboards, mice, etc.; output sections including cathode ray tubes (CRTs), liquid crystal displays (LCDs), and speakers; storage sections including hard disks; and communication sections including network interface cards such as LAN cards and modems. The communication sections perform communication processing via networks such as the Internet. Drives are also connected to the I / O interface as needed. Removable media, such as disks, optical disks, magneto-optical disks, semiconductor memories, etc., are installed on the drive as needed so that computer programs read from them can be installed into the storage section as required. The processing unit can be implemented as a CPU, GPU, TPU, FPGA, NPU, etc.

[0055] In particular, according to embodiments of this disclosure, the methods described above can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising computer instructions that, when executed by a processor, implement the steps of the methods described above. In such embodiments, the computer program product can be downloaded and installed from a network via a communication component, and / or installed from a removable medium.

[0056] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0057] The units or modules described in the embodiments of this disclosure can be implemented in software or programmable hardware. The described units or modules can also be located in a processor, and the names of these units or modules do not necessarily constitute a limitation on the unit or module itself.

[0058] In another aspect, this disclosure also provides a computer-readable storage medium, which may be a computer-readable storage medium included in the electronic device or computer system described above; or it may be a standalone computer-readable storage medium not assembled into a device. The computer-readable storage medium stores one or more programs, which are used by one or more processors to perform the methods described in this disclosure.

[0059] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of the invention involved in this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the inventive concept. For example, technical solutions formed by substituting the above-described features with (but not limited to) technical features disclosed in this disclosure that have similar functions.

Claims

1. A method for training a multi-omics data joint analysis model, characterized in that, The method comprises: constructing a graph structure based on prior biological knowledge, wherein each vertex in the graph structure represents a molecular feature, and each edge represents a biological interaction between two vertices connected by the edge, and the molecular feature is multi-omics data; selecting a plurality of vertices as central features from the vertices of the graph structure, wherein the vertices other than the central features in the graph structure are edge features, and the central features have biological relevance to an analysis task; based on the graph structure construction layer network layer, wherein a 0th layer of the network layer is composed of the center features, a 1st layer of the network layer is composed of the edge features in the first hop from the center features, layer is the shortest path in the graph structure from each center feature, edge feature in the first hop from the center feature, is an integer greater than 1; obtaining a plurality of training samples, wherein each training sample comprises molecular features of a sample and a disease label; Using the training sample, a preset disease mechanism model is trained to obtain a trained disease mechanism model; wherein the disease mechanism model comprises a biological pathway transmission layer and a full connection layer, the biological pathway transmission layer is used to transmit information of a molecular feature of the sample in a first layer to a directly connected feature of the sample in a second layer, and gradually transmit information of an edge molecular feature to a center feature of the first layer to obtain an updated center feature, and transmit the updated center feature to the full connection layer. and gradually transmit information of an edge molecular feature to a center feature of the first layer to obtain an updated center feature, and transmit the updated center feature to the full connection layer.

2. The method of claim 1, wherein, The method further includes transmitting information of the molecular features of the sample at the first layer to their directly connected features at the second layer. The method further includes transmitting information of the molecular features of the sample at the first layer to their directly connected features at the second layer. The method further includes transmitting information of the molecular features of the sample at the first layer to their directly connected features at the second layer. The molecular feature in the sample is divided into a plurality of sub-features according to the following formula The information of the molecular feature in the layer The information of the direct connection feature in the layer The information of the direct connection feature in the layer : ; wherein , , , is a parameter to be trained, is an activation function, denotes element-wise multiplication; is the adjacency matrix of the first row of the first column of the first element ; is the residual connection matrix, the first row vector is: ; wherein is a position is a unit vector with value 1 at the position and zero elsewhere; is the number of molecular features of the sample at the position is the number of molecular features of the sample at the position is the number of molecular features of the sample at the position is the number of molecular features of the sample at the position is the molecular feature of the sample at the position is the molecular feature of the sample at the position is the molecular feature of the sample at the position is the molecular feature of the sample at the position is the molecular feature of the sample at the position is the molecular feature of the sample at the position 3. The method of claim 1, wherein, the full connection layer comprises a multi-layer perceptron (MLP).

4. The method of claim 1, wherein, The method further comprises: based on the biological pathway transmission layer in the trained disease mechanism model, removing pathways in which the interaction intensity between molecular features is less than a predetermined threshold, to obtain a final disease mechanism model.

5. A multi-omics data joint analysis method, characterized in that, The method comprises: obtaining molecular features of an individual to be analyzed; inputting the molecular features of the individual to be analyzed into a pre-trained disease mechanism model, transmitting the information of the molecular features of the individual to be analyzed to the directly connected features of the individual in the biological pathway transmission layer of the pre-trained disease mechanism model layer to the directly connected features of the individual in the layer, and transmitting the information of the edge features to the center features of the 0th layer step by step, obtaining updated center features, and transmitting the updated center features to the fully connected layer of the pre-trained disease mechanism model, and further obtaining the analysis result output by the pre-trained disease mechanism model; outputting the analysis result and a biological pathway corresponding to the molecular features of the individual to be analyzed, wherein the biological pathway displays the interaction intensity between the molecular features of the individual to be analyzed. 6.A device for training a multi-omics data joint analysis model, comprising: The method comprises: a graph construction module configured to construct a graph structure based on prior biological knowledge, wherein each vertex in the graph structure represents a molecular feature, and each edge represents a biological interaction between two vertices connected by the edge, and the molecular feature is multi-omics data; a feature selection module configured to select a plurality of vertices as central features from the vertices of the graph structure, wherein the vertices other than the central features in the graph structure are edge features, and the central features have biological relevance to an analysis task; a network layer construction module configured to construct a network layer based on the graph structure layer network layer, wherein a 0th layer of the network layer consists of the center features, a 1st layer of the network layer consists of the edge features that are one hop away from the center features, layer is the edge feature that is one hop away from the center feature in the graph structure, layer is the edge feature that is two hops away from the center feature in the graph structure, is an integer greater than 1; a sample obtaining module configured to obtain a plurality of training samples, wherein each training sample comprises molecular features of a sample and a disease label; The model training module is configured to use the training samples to train a preset disease mechanism model, thereby obtaining a trained disease mechanism model; wherein, the disease mechanism model includes a biological pathway transport layer and a fully connected layer, the biological pathway transport layer being used to transfer the molecular features of the samples... Information about the molecular characteristics of the layer is transmitted to its first layer. The direct connection features in the layer gradually transmit the information of the edge molecular features to the center features of the 0th layer to obtain the updated center features, and then transmit the updated center features to the fully connected layer. 7.A multi-omics data joint analysis device, characterized by comprising: The method comprises: an obtaining module configured to obtain molecular features of an individual to be analyzed; The analysis module is configured to input the molecular characteristics of the individual to be analyzed into a pre-trained disease mechanism model, a biological pathway transmission layer in the pre-trained disease mechanism model transmits information of the molecular characteristics of the individual to be analyzed to the directly connected characteristics of the individual to be analyzed in the first layer , transmits information of the edge characteristics to the center characteristics of the first layer step by step, obtains updated center characteristics, and transmits the updated center characteristics to a fully connected layer of the pre-trained disease mechanism model, and further obtains an analysis result output by the pre-trained disease mechanism model; the pre-trained disease mechanism model is trained according to the method of any one of claims 1-5. , transmits information of the edge characteristics to the center characteristics of the first layer step by step, obtains updated center characteristics, and transmits the updated center characteristics to a fully connected layer of the pre-trained disease mechanism model, and further obtains an analysis result output by the pre-trained disease mechanism model; the pre-trained disease mechanism model is trained according to the method of any one of claims 1-5. an outputting module configured to output the analysis result and a biological pathway corresponding to the molecular features of the individual to be analyzed, wherein the biological pathway displays the interaction intensity between the molecular features of the individual to be analyzed.

8. An electronic device, comprising: The computer readable storage medium stores computer program instructions, and the computer program instructions are run by a processor to execute the method in any one of claims 1-5.

9. A computer-readable storage medium, characterized in that, The computer readable storage medium stores computer program instructions, and the computer program instructions are run by a processor to execute the method in any one of claims 1-5.