Multi-source data fusion method and device, electronic equipment and storage medium

By preprocessing omics data and phenotypic data and constructing a similarity matrix, a fusion network is generated, which solves the problem of large deviation between the fusion results of omics data and phenotypic data and achieves effective multi-source data fusion.

CN119229962BActive Publication Date: 2025-12-16BGI WUHAN
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310794383.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-06-29
Publication Date
2025-12-16
Estimated Expiration
2043-06-29

AI Technical Summary

Technical Problem

Existing technologies cannot effectively integrate omics data and phenotypic data with significantly different data structures, resulting in large deviations in the fusion results.

Method used

Multiple omics and phenotypic data of the observed objects were collected, and after preprocessing, similarity matrices were constructed. All similarity matrices were merged to generate a fusion network, and attribute maps of the observed objects were built based on the fusion network.

Benefits of technology

While preserving the scale significance of each individual data point, it effectively fused omics data and phenotypic data with significant differences in data structure, reducing the bias of the fusion results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119229962B_ABST
    Figure CN119229962B_ABST
Patent Text Reader

Abstract

The application discloses a multi-source data fusion method and device, electronic equipment and storage medium, wherein the fusion method comprises the following steps: collecting multiple omics data and phenotype data of an observation object, and preprocessing the multiple omics data and the phenotype data to obtain multiple target omics data and target phenotype data; constructing a similarity matrix for the multiple target omics data and the target phenotype data respectively; fusing all the similarity matrices to generate a fusion network; and establishing an attribute graph of the observation object based on the fusion network, the multiple target omics data and the target phenotype data. The application solves the technical problem that in the related art, omics data and phenotype data with large differences in data structure cannot be effectively fused, resulting in large deviations in the fusion result.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of data processing, in particular to a multi-source data fusion method and device, electronic equipment and storage medium. BACKGROUND

[0002] With the continuous development of high-throughput experimental technology, a large amount of omics data can be obtained in a short time. These omics data make it possible to analyze the occurrence and development mechanism of complex system biology from the molecular level. The central dogma of genetics proposes that in living organisms, genetic information in cells is transmitted from DNA (deoxyribonucleic acid) to mRNA (messenger ribonucleic acid), and then to protein. Research on these three types of substances is respectively referred to as genomics, transcriptomics and proteomics. In addition, although metabolomics, lipidomics and glycomics are not part of the central dogma, they still provide valuable information about metabolites, lipids and sugars, which are intermediate products in the flow of cellular information. In the context of precision medicine, any single omics data is insufficient to elucidate the mechanism of complex system biology diseases, and the integration and analysis of multi-omics data is the current research hotspot and trend. Multi-omics data fusion has gradually become a technical basis for biomedical research in recent years, and the exploration of technologies including machine learning analysis is still in its infancy. The development of such data fusion technologies will provide a new technical basis for the study of the occurrence and development of complex biological phenomena.

[0003] Currently, various biological omics data (for example, genomics, transcriptomics and proteomics) contain a large amount of key biological information. These data have characteristics such as redundancy, isomerism, large noise and complex correlation patterns. Traditional biological statistical methods often have limitations such as shallow feature depth, and it is often difficult to effectively process massive multi-dimensional omics data.

[0004] In view of the above problems, no effective solutions have been proposed so far. SUMMARY

[0005] Embodiments of the present application provide a multi-source data fusion method and device, electronic equipment and storage medium, to at least solve the technical problem that in related technologies, omics data and phenotype data with large differences in data structure cannot be effectively fused, resulting in large deviations in the fusion results.

[0006] According to an aspect of an embodiment of the present application, a multi-source data fusion method is provided, comprising: collecting multiple omics data and phenotype data of an observation object, and preprocessing the multiple omics data and the phenotype data to obtain multiple target omics data and target phenotype data; constructing a similarity matrix for each of the multiple target omics data and the target phenotype data; fusing all the similarity matrices to generate a fusion network; and establishing an attribute graph of the observation object based on the fusion network, the multiple target omics data, and the target phenotype data.

[0007] Optionally, the step of preprocessing the multiple omics data and the phenotype data to obtain the multiple target omics data and the target phenotype data comprises: determining omics features and omics feature values corresponding to the omics features of each sample in the omics data; performing missing value processing on the omics data based on the omics features and the omics feature values corresponding to the omics features to obtain initial omics data; performing logarithmic processing on the omics feature values of each sample in the initial omics data to obtain initial omics feature values; performing normalization processing on the initial omics feature values to obtain target omics feature values, wherein the target omics feature values indicate that the samples constitute the target omics data; and performing a preset embedding mapping on the phenotype data to obtain the target phenotype data.

[0008] Optionally, the step of performing missing value processing on the omics data based on the omics features and the omics feature values corresponding to the omics features to obtain initial omics data comprises: determining target omics features whose missing proportions of omics feature values in all samples are greater than a first preset proportion threshold; determining target samples whose missing proportions of all omics feature values are greater than a second preset proportion threshold; deleting the target omics features of each sample in the omics data and the target samples in the omics data; and filling the remaining missing values of each sample in the omics data after deletion using a preset mean value algorithm to obtain the initial omics data, wherein the remaining missing values refer to the missing values of the omics features of each sample in the omics data after deletion.

[0009] Optionally, the step of constructing a similarity matrix for each of the plurality of target omics data and the target phenotype data comprises: determining a first sample diameter of a first sample and a second sample diameter of a second sample in observation data, wherein the observation data is the target omics data or the target phenotype data, the first sample is any sample in the observation data, and the second sample is any sample in the observation data except the first sample; determining a similarity measure indicator value between the first sample and the second sample based on a sample distance between the first sample and the second sample, the first sample diameter, the second sample diameter, a first weighting parameter, and a second weighting parameter; constructing a fully connected similarity matrix based on all the similarity measure indicator values; and constructing the similarity matrix of the observation data based on the fully connected similarity matrix.

[0010] Optionally, the step of determining a first sample diameter of a first sample and a second sample diameter of a second sample in observation data comprises: determining a first neighbor sample set of the first sample; determining a first sample distance between the first sample and each first neighbor sample in the first neighbor sample set, and determining the first sample diameter of the first sample based on all the first sample distances; determining a second neighbor sample set of the second sample; determining a second sample distance between the second sample and each second neighbor sample in the second neighbor sample set, and determining the second sample diameter of the second sample based on all the second sample distances.

[0011] Optionally, the step of fusing all the similarity matrices to generate a fusion network comprises: determining a total information corresponding to each similarity matrix and an information weight corresponding to each similarity matrix; and generating the fusion network based on the total information and the information weight corresponding to all the similarity matrices.

[0012] Optionally, the step of establishing an attribute graph of the observation object based on the fusion network, the plurality of target omics data, and the target phenotype data comprises: determining a data feature value corresponding to each data feature based on the plurality of target omics data and the target phenotype data; splicing all the data feature values corresponding to the data feature to obtain a feature attribute of the data feature; determining a data node corresponding to the data feature based on the fusion network; associating the feature attribute as a node attribute of the data node to generate the attribute graph.

[0013] Optionally, after the attribute graph of the observation object is established based on the fusion network, the plurality of target omics data and the target phenotype data, the method further includes: processing the attribute graph by using a preset learning model to obtain a vector value of each data node, wherein the preset learning model includes an input layer, a hidden layer and an output layer; generating a target attribute graph based on all the vector values; and performing task analysis on the observation object based on the target attribute graph.

[0014] Optionally, the step of processing the attribute graph by using a preset learning model to obtain a vector value of each data node includes: receiving the attribute graph through the input layer; performing preset neighbor node sampling processing on each data node in the attribute graph through the hidden layer to obtain an aggregated node set of each layer field of the data node; performing preset neighbor aggregation processing on the aggregated node set to obtain an aggregated feature value; generating a target aggregated feature value based on all the aggregated features; representing the target aggregated feature value as the vector value, and outputting the vector value through the output layer.

[0015] According to another aspect of the embodiment of the present application, a multi-source data fusion device is also provided, which includes: a processing unit configured to collect a plurality of omics data and phenotype data of an observation object, and to pre-process the plurality of omics data and the phenotype data to obtain a plurality of target omics data and target phenotype data; a construction unit configured to construct a similarity matrix for each of the plurality of target omics data and the target phenotype data; a fusion unit configured to fuse all the similarity matrices to generate a fusion network; and an establishment unit configured to establish an attribute graph of the observation object based on the fusion network, the plurality of target omics data and the target phenotype data.

[0016] Optionally, the processing unit includes: a first determination module configured to determine an omics feature of each sample in the omics data and an omics feature value corresponding to the omics feature; a first processing module configured to perform missing value processing on the omics data based on the omics feature and the omics feature value corresponding to the omics feature to obtain initial omics data; a second processing module configured to perform logarithmic processing on the omics feature value of each sample in the initial omics data to obtain initial omics feature values; a third processing module configured to perform normalization processing on the initial omics feature values to obtain target omics feature values, wherein the target omics feature values indicate that the samples constitute the target omics data; and a first mapping module configured to perform preset embedding mapping on the phenotype data to obtain the target phenotype data.

[0017] Optionally, the first processing module comprises: a first determining submodule, configured to determine target omics features whose missing proportions of omics feature values in all the samples are greater than a first preset proportion threshold; a second determining submodule, configured to determine target samples whose missing proportions of all omics feature values are greater than a second preset proportion threshold; a first deleting submodule, configured to delete the target omics features of each of the samples in the omics data and delete the target samples in the omics data; and a first filling submodule, configured to fill the remaining missing values of each of the samples in the omics data after deletion by using a preset mean value algorithm to obtain the initial omics data, wherein the remaining missing values refer to values of the omics features missing in each of the samples in the omics data after deletion.

[0018] Optionally, the construction unit comprises: a second determining module, configured to determine a first sample diameter of a first sample and a second sample diameter of a second sample in observation data, wherein the observation data is the target omics data or the target phenotype data, the first sample is any sample in the observation data, and the second sample is any sample in the observation data except the first sample; a third determining module, configured to determine a similarity measure indicator value between the first sample and the second sample based on a sample distance between the first sample and the second sample, the first sample diameter, the second sample diameter, a first weighting parameter, and a second weighting parameter; a first construction module, configured to construct a full-connection similarity matrix based on all the similarity measure indicator values; and a second construction module, configured to construct the similarity matrix of the observation data based on the full-connection similarity matrix.

[0019] Optionally, the second determining module comprises: a third determining submodule, configured to determine a first neighbor sample set of the first sample; a fourth determining submodule, configured to determine a first sample distance between the first sample and each first neighbor sample in the first neighbor sample set and determine the first sample diameter of the first sample based on all the first sample distances; a fifth determining submodule, configured to determine a second neighbor sample set of the second sample; and a sixth determining submodule, configured to determine a second sample distance between the second sample and each second neighbor sample in the second neighbor sample set and determine the second sample diameter of the second sample based on all the second sample distances.

[0020] Optionally, the fusion unit comprises: a fourth determining module, configured to determine total information corresponding to each of the similarity matrices and information weights corresponding to each of the similarity matrices; and a first generating module, configured to generate the fusion network based on the total information corresponding to all the similarity matrices and the information weights.

[0021] Optionally, the establishing unit comprises: a fifth determination module configured to determine a data feature value corresponding to each data feature based on the multiple target omics data and the target phenotype data; a first splicing module configured to splice all the data feature values corresponding to the data feature to obtain a feature attribute of the data feature; a sixth determination module configured to determine a data node corresponding to the data feature based on the fusion network; and a second generation module configured to associate the feature attribute as a node attribute of the data node to generate the attribute graph.

[0022] Optionally, the fusion device further comprises: a fourth processing module configured to, after establishing the attribute graph of the observation object based on the fusion network, the multiple target omics data and the target phenotype data, process the attribute graph by using a preset learning model to obtain a vector value of each data node, wherein the preset learning model comprises an input layer, a hidden layer and an output layer; a third generation module configured to generate a target attribute graph based on all the vector values; and a first analysis module configured to perform task analysis on the observation object based on the target attribute graph.

[0023] Optionally, the fourth processing module comprises: a first receiving submodule configured to receive the attribute graph through the input layer; a first processing submodule configured to perform preset neighbor node sampling processing on each data node in the attribute graph through the hidden layer to obtain an aggregated node set of each layer field of the data node; a second processing submodule configured to perform preset neighbor aggregation processing on the aggregated node set to obtain an aggregated feature value; a first generation submodule configured to generate a target aggregated feature value based on all the aggregated features; and a first output submodule configured to represent the target aggregated feature value as the vector value and output the vector value through the output layer.

[0024] According to another aspect of the embodiments of the present application, a computer readable storage medium is also provided, which comprises a stored computer program, wherein the computer readable storage medium controls a device where the computer readable storage medium is located to perform the fusion method of multi-source data when the computer program runs.

[0025] According to another aspect of the embodiments of the present application, an electronic device is also provided, which comprises one or more processors and a memory, and the memory is configured to store one or more programs, wherein the one or more programs make the one or more processors implement the fusion method of multi-source data when the one or more programs are executed by the one or more processors.

[0026] In the present disclosure, multiple omics data and phenotype data of observation objects are collected, and the multiple omics data and phenotype data are preprocessed to obtain multiple target omics data and target phenotype data. Similarity matrices are constructed for the multiple target omics data and target phenotype data, respectively. All similarity matrices are fused to generate a fusion network. Based on the fusion network, the multiple target omics data and target phenotype data, an attribute graph of the observation objects is established. In the present disclosure, the collected multiple omics data and phenotype data can be preprocessed first, and then the processed multiple target omics data and target phenotype data are constructed into similarity matrices, respectively. All similarity matrices are fused to generate a fusion network. Then, based on the fusion network, the multiple target omics data and target phenotype data, an attribute graph of the observation objects is established. The fusion of omics data and phenotype data with large structural differences can be realized while retaining the dimensional significance of each single data, so as to obtain the attribute graph of the observation objects, thereby solving the technical problem that the omics data and phenotype data with large structural differences cannot be effectively fused in the related art, resulting in large deviation of the fusion result. BRIEF DESCRIPTION OF DRAWINGS

[0027] The accompanying drawings, which are included to provide a further understanding of the application and are incorporated in and constitute a part of this application, illustrate embodiments of the application and together with the description serve to explain the application. In the drawings:

[0028] Figure 1 is a flowchart of an optional multi-source data fusion method according to an embodiment of the present application;

[0029] Figure 2 is a flowchart of an optional data collection and data preprocessing according to an embodiment of the present application;

[0030] Figure 3 is a flowchart of an optional similarity matrix calculation and matrix fusion according to an embodiment of the present application;

[0031] Figure 4 is a schematic diagram of an optional multi-source data fusion process based on a graph convolution network according to an embodiment of the present application;

[0032] Figure 5 is a schematic diagram of an optional multi-source data fusion device according to an embodiment of the present application;

[0033] Figure 6 is a hardware structure block diagram of an electronic device (or mobile device) for a multi-source data fusion method according to an embodiment of the present application. DETAILED DESCRIPTION

[0034] In the following, the technical solutions in the embodiments of the present application will be described clearly and completely with reference to the drawings in the embodiments of the present application by persons in the technical field of the present application, obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by persons in the technical field of the present application without creative labor should belong to the scope of protection of the present application.

[0035] It should be noted that the terms "first", "second" and the like in the specification and claims of the present application and the above-described drawings are used to distinguish similar objects, and do not necessarily indicate a specific order or a chronological sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device including a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but can include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0036] It should be noted that the relevant information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present disclosure are all information and data authorized by the user or authorized by all parties, and the collection, use and processing of relevant data need to comply with relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation portal for user to choose authorization or refusal.

[0037] In the present application, many new data integration methods can be introduced from the fields of mathematics, statistics and computing science. These data integration methods are generally divided into three categories: serial-based integration methods, model-based integration methods, and conversion-based integration methods, wherein,

[0038] The concatenation-based integration method considers using joint data matrix to develop models, which is combined by multiple omics datasets. First, the raw data from individual omics and the corresponding phenotype information are directly concatenated without any preprocessing, and multiple omics data are spliced into a single large matrix for supervised and unsupervised learning. Once the concatenation of individual omics is completed, simple machine learning methods can be used for analysis. Such methods mainly include traditional machine learning algorithms such as DT (Decision Tree), ANN (Approximate Nearest Neighbor), SVM (Support Vector Machine), LASSO (Least absolute shrinkage and selection operator), SVN (Support Vector Machine and upgrade algorithm), DNN (Deep Neural Network), Joint NMF (Non-negative Matrix Factorization), iCluster (an integrated classification algorithm based on multi-omics data), JIVE (an open Java source project), MoCluster (a clustering algorithm), and the like. Such methods can equally use all connected features and can select the most discriminative features for a given phenotype.

[0039] The model-based integration method creates multiple intermediate models for different omics data, and then builds a final model based on each intermediate model. The process is as follows: first, multiple single-omics data matrices are constructed for the raw data of each individual omics and the corresponding phenotype information, and a separate model is developed for each omics, and then these models are integrated into a joint model for analysis. Such methods mainly include Majority-based voting, Hierarchical Classifiers, MOLI (a multi-omics late integration method based on deep neural network), ATHENA (a formalized analysis method), PSDF (Power Spectral Density-based Fault Diagnosis), MDI (Mean-Shift Clustering Algorithm), Lemon-Tree (a decision tree algorithm), and the like. Such methods can be used to combine models based on different omics types, and each model can be developed from different sample populations with the same label.

[0040] The transformation-based integration method firstly transforms each omics dataset into a graph or kernel matrix, and then integrates all the data before building a model. The process is as follows: firstly, establish a single omics raw data and phenotype information correspondence for each omics, then develop a single graph or kernel matrix transformation method for each omics, and then integrate these single-omics transformations into a joint transformation, which is analyzed using a machine learning algorithm. This kind of method mainly includes SDP-SVM (i.e. semi-definite programming support vector machine), FSMKL (i.e. feature set multi-kernel learning), RVM (i.e. relevance vector machine), Graphsharpening (i.e. image sharpening), PAMOGK (i.e. incomplete multi-kernel clustering analysis), NEMO (i.e. a moving object detector) and other methods. In this kind of method, if there is some unique information (such as sample feature number) in the data set, it can be used to analyze more dimensional omics information.

[0041] However, the integration method based on concatenation cannot obtain the effective correlation between omics features, and due to the large difference in the number of features and data distribution between different omics, the results often have large deviations. The deep learning-based technology, such as the Autoencoder (i.e. an unsupervised learning algorithm) method, needs a large number of samples in order to better learn the features of multiple omics. In actual multi-omics applications, the sample is not too much, resulting in unsatisfactory results of the Autoencoder method. Subsequently, unsupervised multi-omics integration methods such as I-Cluster (i.e. a hierarchical clustering algorithm), Similarity Network Fusion (SNF for short), MAUI (i.e. a large-scale automated omics information integration analysis tool) and Multi-Omics Factor Analysis (MOFA+ for short) have achieved some precision improvement, but only the multi-omics data of the observation object is learned unsupervisedly to obtain a one-dimensional observation object similarity network and then analyzed, ignoring the multi-dimensional information of the observation object itself, resulting in poor results in the application process. In addition, most of the above data integration methods consider integrating omics data with similar data structures, and cannot integrate phenotype data which is quite different from omics data structure.

[0042] Therefore, the present application provides a multi-source data fusion analysis method based on a graph convolution network, which can integrate multi-omics data and phenotype data of observation objects, and can realize data fusion while preserving the scale significance of each single data. The present application also adopts a representation learning method based on a fusion network, which can capture shared complementary information from different data sources, deeply mine the internal relationship between different omics data, and has good robustness to data noise and heterogeneity between data.

[0043] The present application will be described in detail below in conjunction with various embodiments.

[0044] Embodiment one

[0045] According to an embodiment of the present application, an embodiment of a multi-source data fusion method is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described herein can be executed in an order different from that shown herein.

[0046] Figure 1 is a flowchart of an optional multi-source data fusion method according to an embodiment of the present application, as shown in Figure 1 The method comprises the following steps:

[0047] Step S101, collecting multi-omics data and phenotype data of an observation object, and pre-processing the multi-omics data and phenotype data to obtain target multi-omics data and target phenotype data.

[0048] Step S102, constructing a similarity matrix for each of the target multi-omics data and target phenotype data.

[0049] Step S103, fusing all the similarity matrices to generate a fusion network.

[0050] Step S104, establishing an attribute graph of the observation object based on the fusion network, the target multi-omics data and the target phenotype data.

[0051] Through the above steps, the multi-omics data and phenotype data of the observation object can be collected and pre-processed to obtain target multi-omics data and target phenotype data. A similarity matrix is constructed for each of the target multi-omics data and target phenotype data. All the similarity matrices are fused to generate a fusion network. Based on the fusion network, the target multi-omics data and the target phenotype data, an attribute graph of the observation object is established. In an embodiment of the present application, the collected multi-omics data and phenotype data can be pre-processed first. Then, a similarity matrix is constructed for each of the processed target multi-omics data and target phenotype data. All the similarity matrices are fused to generate a fusion network. Then, based on the fusion network, the target multi-omics data and the target phenotype data, an attribute graph of the observation object is established. The fusion of multi-omics data and phenotype data with large structural differences can be realized while retaining the dimensional significance of each single data, so as to obtain an attribute graph of the observation object. Thus, the technical problem that multi-omics data and phenotype data with large structural differences cannot be effectively fused in the related art, resulting in a large deviation in the fusion result, is solved.

[0052] The embodiments of the present application will be described in detail below in combination with the above steps.

[0053] In step S101, the multiple omics data and the phenotype data of the observation object are collected and preprocessed to obtain the multiple target omics data and the target phenotype data.

[0054] In the embodiments of the present application, the multiple omics data information and the phenotype data of the observation object (i.e., various organisms, such as microorganisms, etc.) can be collected and sorted first (i.e., the multiple omics data and the phenotype data of the observation object are collected). The multiple omics data include but are not limited to the genomics, transcriptomics, proteomics, metabolomics, etc. The phenotype data include the quantitative data (such as weight, height, etc.), the qualitative data (such as behavior), the photo, the clinical feature description, the medical history, the physical examination, the biochemical examination, etc. Then, the collected multiple omics data can be preprocessed and converted, including the standardization, the normalization, the feature selection, the missing value processing, etc. of the omics data, and the data preprocessing, including the encoding conversion, the data normalization, etc. of the phenotype data (i.e., the multiple omics data and the phenotype data are preprocessed to obtain the multiple target omics data and the target phenotype data).

[0055] Optionally, the step of pre-processing the multiple omics data and the phenotype data to obtain the multiple target omics data and the target phenotype data includes: determining the omics feature of each sample in the omics data and the omics feature value corresponding to the omics feature; performing missing value processing on the omics data based on the omics feature and the omics feature value corresponding to the omics feature to obtain initial omics data; performing logarithmic processing on the omics feature value of each sample in the initial omics data to obtain initial omics feature value; performing normalization processing on the initial omics feature value to obtain target omics feature value, wherein the target omics feature value indicates the sample composition of the target omics data; and performing pre-set embedding mapping on the phenotype data to obtain the target phenotype data.

[0056] In the embodiments of the present application, the omics features of each sample in the omics data and the omics feature values corresponding to the omics features can be determined first (that is, each sample in each omics data has multiple omics features (for example, accurate value, mean value and the like), and each omics feature has a corresponding omics feature value in the sample). Then, the omics data is processed for missing values according to the omics features and the omics feature values corresponding to the omics features to obtain initial omics data. Thereafter, the omics feature values of each sample in the initial omics data can be logarithmically processed to obtain initial omics feature values, for example, all numerical values of the omics features are subjected to log2(x+1) conversion, x representing the original omics feature value, which can reduce the range of data, so that the omics feature data is more stable. The initial omics feature values are normalized to obtain target omics feature values, for example, in order to ensure that the omics feature values fall within a certain range, normalization and standardization are performed on all omics feature values, and the Z-score (that is, standardized score) standardization method can be used to normalize the omics feature values, wherein the Z-score standardization formula is: μ is the mean value of x, σ is the variance of x, and z is taken as the new omics feature value (that is, the target omics feature value) after processing all omics feature original values x. Then, the samples indicated by the target omics feature values form target omics data. Moreover, the text data in the phenotype data can be subjected to embedding mapping based on a transformer (a kind of neural network model) (that is, the phenotype data is subjected to preset embedding mapping) to obtain target phenotype data.

[0057] Optionally, the step of processing the omics data for missing values based on the omics features and the omics feature values corresponding to the omics features to obtain the initial omics data comprises: determining target omics features whose missing proportions of omics feature values in all samples are greater than a first preset proportion threshold; determining target samples whose missing proportions of all omics feature values are greater than a second preset proportion threshold; deleting the target omics features of each sample in the omics data and deleting the target samples in the omics data; filling the remaining missing values of each sample in the deleted omics data by using a preset mean algorithm to obtain the initial omics data, wherein the remaining missing values refer to the missing values of the omics features of each sample in the deleted omics data.

[0058] In the embodiments of the present application, since there are some missing values in the omics features, and these missing values have a greater impact on the overall process results, therefore, it is necessary to process the missing values of the omics data, specifically: the target omics features whose missing proportion of omics feature values in all samples is greater than a first preset proportion threshold (for example, 20%) can be determined, the target samples whose missing proportion of all omics feature values is greater than a second preset proportion threshold (for example, 20%) are determined, then the target omics features of each sample in the omics data are deleted, and the target samples in the omics data are deleted, for example, if the missing proportion of a certain omics feature in all sample expression values is greater than 20%, this omics feature is directly deleted, and when the missing proportion of all omics features of a certain sample is greater than 20%, the sample is deleted. Then, for the retained features and samples, the remaining missing values can be filled by k-nearest neighbor mean (i.e., a preset mean algorithm) (i.e., using the preset mean algorithm to fill the remaining missing values of each sample in the deleted omics data to obtain initial omics data, the remaining missing values refer to the missing values of omics features of each sample in the deleted omics data).

[0059] Figure 2 is an optional data acquisition and data preprocessing flowchart according to the embodiments of the present application, as shown in Figure 2 , including the following processes:

[0060] (1) acquiring multiple omics data and phenotype data of observation objects;

[0061] (2) performing missing value processing, logarithmic change, standardization, and feature selection on the acquired data;

[0062] (3) integrating the multiple omics data using a similarity network fusion algorithm;

[0063] (4) constructing an attribute graph of the observation objects using the fused similarity matrix and the multiple omics splicing data.

[0064] Step S102, constructing a similarity matrix for multiple target omics data and target phenotype data respectively.

[0065] Optionally, the step of constructing the similarity matrix for each of the plurality of omics data and the target phenotype data comprises: determining a first sample diameter of a first sample and a second sample diameter of a second sample in the observation data, wherein the observation data is the omics data or the target phenotype data, the first sample is any sample in the observation data, and the second sample is any sample in the observation data except the first sample; determining a similarity measurement index value between the first sample and the second sample based on a sample distance between the first sample and the second sample, the first sample diameter, the second sample diameter, a first weighting parameter, and a second weighting parameter; constructing a full connection similarity matrix based on all the similarity measurement index values; and constructing a similarity matrix of the observation data based on the full connection similarity matrix.

[0066] In the embodiments of the present application, the similarity matrices can be constructed for the multi-omics data and the phenotype data (i.e., the similarity matrices are constructed for each of the plurality of omics data and the target phenotype data), and then the similarity network fusion algorithm can be used to fuse the plurality of networks (i.e., the plurality of similarity matrices).

[0067] In the embodiments, a single similarity matrix t = 1, 2, …, T, T represents the number of omics, i represents sample i, j represents sample j, and N represents the number of samples, and the construction relationship is as follows:

[0068]

[0069] wherein N k (i) represents the k nearest neighbors of sample i; and ε represents a very small value, which can be set to 0 as an optional embodiment of the present application. When ε = 0, the formula (1) can be explained as follows: only when sample j is the nearest neighbor of sample i, the edge between them will be retained, and the remaining edges will be “corrected”, and thus the is more robust.

[0070] is the full connection similarity matrix.

[0071] In the embodiments, the full connection similarity matrix S (t) is obtained by row normalization matrix is obtained, is a similarity measurement index, the smaller the distance between sample i and sample j, the greater the value, and the greater the similarity. is calculated by a Gaussian kernel function based on sample i and sample j, and the calculation relationship is as follows:

[0072]

[0073] wherein, Let i be the local diameter. Let j be the local diameter of sample j. Let α be the distance between sample j and sample i, and let β be the weighting parameters. As an optional implementation of this application, they can both be set to 1 / 6.

[0074] In this embodiment, the local diameter of the sample is obtained through the following relationship:

[0075]

[0076] In this embodiment, the distance between samples It can be calculated in a variety of ways. As an optional implementation method of this application, Euclidean distance calculation can be used.

[0077] In this embodiment, the first sample diameter of the first sample (i.e., any sample in the observation data, such as sample i) and the second sample diameter of the second sample (i.e., any sample in the observation data other than the first sample, such as sample j) in the observation data (i.e., target omics data or target phenotypic data) can be determined first. Then, based on the sample distance between the first sample and the second sample (e.g., ... ), first sample diameter (e.g. ), second sample diameter (e.g. The first weighted parameter (e.g., α) and the second weighted parameter (e.g., β) are used to determine the similarity metric value between the first sample and the second sample (e.g., ...). (That is, the similarity metric values ​​can be calculated using formulas (2) and (3). Then, based on all the similarity metric values, a fully connected similarity matrix (such as S) is constructed. (t) Then, based on the fully connected similarity matrix, a similarity matrix (such as W) is constructed for the observed data. (t) (That is, the similarity matrix can be obtained through formula (1).)

[0078] Optionally, the steps of determining the first sample diameter of the first sample and the second sample diameter of the second sample in the observation data include: determining the first nearest neighbor sample set of the first sample; determining the first sample distance between the first sample and each first nearest neighbor sample in the first nearest neighbor sample set, and determining the first sample diameter of the first sample based on all first sample distances; determining the second nearest neighbor sample set of the second sample; determining the second sample distance between the second sample and each second nearest neighbor sample in the second nearest neighbor sample set, and determining the second sample diameter of the second sample based on all second sample distances.

[0079] In this embodiment of the invention, the local diameter (i.e., the sample diameter) of the sample can be obtained by formula (3), specifically: the set of nearest neighbor samples of the sample can be determined first (e.g., ) (i.e. determining a first neighbor sample set of the first sample, determining a second neighbor sample set of the second sample), then determining the sample distance between the sample and each neighbor sample in the neighbor sample set (e.g. ) (i.e. determining a first sample distance between the first sample and each first neighbor sample in the first neighbor sample set, determining a second sample distance between the second sample and each second neighbor sample in the second neighbor sample set), and then determining the sample diameter of the sample according to all sample distances (e.g. ) (i.e. determining a first sample diameter of the first sample based on all first sample distances, determining a second sample diameter of the second sample based on all second sample distances).

[0080] Step S103: fusing all similarity matrices to generate a fusion network.

[0081] In the embodiment of the present application, after the similarity matrices of all samples in each omics dimension are constructed respectively, the multiple similarity matrices can be further fused by a similarity network fusion algorithm to obtain a sample similarity matrix (i.e. a fusion network) containing omics data information of each omics (i.e. fusing all similarity matrices to generate a fusion network).

[0082] In the embodiment of the present application, the similarity network fusion algorithm uses the method of constructing a matrix by using multi-omics data to transform and integrate the detection information matrices of different omics, which can accelerate the calculation process of the research and provide valuable information.

[0083] Optionally, the step of fusing all similarity matrices to generate a fusion network comprises: determining total information corresponding to each similarity matrix and information weight corresponding to each similarity matrix; and generating the fusion network based on the total information corresponding to all similarity matrices and the information weight.

[0084] In the embodiment of the present application, when the different omics correlation matrices are fused (i.e. all similarity matrices are fused), the total information of each omics is determined. When it represents a full connection or a non-bipartite graph, the information of different omics can be integrated by random walk at this time, and a stable state can be reached when two-step random walk (r=2).

[0085] In the embodiment of the present application, W (t) represents the total information of omics t after integrating the information of other omics, wherein, represents the sum of the information of (T-1) omics other than omics t, k represents any omics other than omics t, and ω k represents the information weight of omics k. W (t) is obtained in the form of random walk by integrating the information of other omics. (t)The multi-omics similarity matrices W (i.e. the fusion network) are obtained by adding the weights, i.e. the total information W of each similarity matrix can be determined first (t) and the information weight ω of each similarity matrix t Then, the fusion network W is generated according to the total information and the information weight of all the similarity matrices, and the formula is as follows:

[0086]

[0087]

[0088] Wherein, α represents a weighting parameter.

[0089] In step S104, the attribute graph of the observed object is established based on the fusion network, the multi-target omics data and the target phenotype data.

[0090] Optionally, the step of establishing the attribute graph of the observed object based on the fusion network, the multi-target omics data and the target phenotype data comprises: determining the data feature value corresponding to each data feature based on the multi-target omics data and the target phenotype data; splicing all the data feature values corresponding to the data feature to obtain the feature attribute of the data feature; determining the data node corresponding to the data feature based on the fusion network; and associating the feature attribute as the node attribute of the data node to generate the attribute graph.

[0091] In the embodiment of the present application, the attribute graph of the observed object can be established according to the obtained fusion network, and the filtered data features are spliced as the attributes of each node in the attribute graph (i.e. the attribute graph of the observed object is established based on the fusion network, the multi-target omics data and the target phenotype data), and specifically, the data feature value corresponding to each data feature can be determined according to the multi-target omics data and the target phenotype data, and then all the data feature values corresponding to the data feature are spliced to obtain the feature attribute of the data feature (i.e. all the data feature values corresponding to the data feature are spliced to obtain the feature attribute of the data feature), and then the data node corresponding to the data feature is determined based on the fusion network (i.e. each data node on the attribute graph represents a data feature), and then the feature attribute is associated as the node attribute of the data node (i.e. the feature attribute is taken as the node attribute of the corresponding data node) to generate the attribute graph.

[0092] Figure 3 is a flowchart of an optional similarity matrix calculation and matrix fusion according to an embodiment of the present application, as shown in Figure 3 comprises the following processes:

[0093] (1) Constructing a similarity matrix for each of the multiple omics data;

[0094] (2) using a similarity network fusion algorithm to integrate multiple similarity matrices into a single matrix;

[0095] (3) constructing an undirected graph by fusing the similarity matrix;

[0096] (4) splicing multiple omics data and phenotype data as the attributes of the graph nodes to obtain an attribute graph of the observation object.

[0097] Optionally, after the attribute graph of the observation object is established based on the fusion network, the multiple target omics data, and the target phenotype data, the method further includes: processing the attribute graph by using a preset learning model to obtain a vector value of each data node, wherein the preset learning model includes an input layer, a hidden layer, and an output layer; generating a target attribute graph based on all the vector values; and performing task analysis on the observation object based on the target attribute graph.

[0098] In the embodiments of the present application, after the attribute graph of the observation object is established, graph neural network representation learning can be performed on the attribute graph to learn the embedding (i.e., the vector value) of each data node (i.e., the preset learning model (including an input layer, a hidden layer, and an output layer) is used to process the attribute graph to obtain a vector value of each data node), and then a target attribute graph is generated based on all the vector values. The attribute graph after representation learning (i.e., the target attribute graph) can be used for subsequent tasks (i.e., the observation object is analyzed based on the target attribute graph), for example, labeled data can be used for classification tasks, and unlabeled data can be used for clustering tasks.

[0099] Alternatively, before the attribute graph is processed by using the preset learning model to obtain a vector value of each data node, the method further includes: determining a preset learning framework; constructing an initial learning model based on the preset learning framework, wherein the initial learning model includes an input layer, a hidden layer, and an output layer; training the initial learning model to determine a target parameter value of a preset parameter; and setting the parameter value of the preset parameter in the initial learning model to the target parameter value to obtain the preset learning model.

[0100] In the embodiment of the present application, the preset learning framework (i.e., the framework of representation learning) can be determined first, and the framework of representation learning can adopt a framework based on a graph autoencoder. The graph algorithm used in the framework can adopt a GraphSAGE model (i.e., a graph neural network model based on a spatial domain). The architecture of the GraphSAGE model mainly includes an input layer, a hidden layer, and an output layer (i.e., based on the preset learning framework, an initial learning model is constructed, and the initial learning model includes an input layer, a hidden layer, and an output layer). In the embodiment, the input of the input layer is the obtained attribute graph. The hidden layer is composed of multiple layers of neighbor node feature aggregation layers. As an optional implementation of the present application, the number of layers can be set to two. The aggregation layer can mainly include two parts: neighbor node random sampling and neighbor aggregation.

[0101] In the embodiment of the present application, the initial learning model can be trained to determine the target parameter value of the preset parameter (i.e., the pre-determined hyperparameter, such as a learning rate lr, a batch sample size batch_size, and an iteration number epoch). Then, the parameter value of the preset parameter in the initial learning model is set to the target parameter value, so as to obtain a trained preset learning model.

[0102] In the embodiment of the present application, the purpose of designing the hyperparameters (including a learning rate lr, a batch sample size batch_size, and an iteration number epoch) of the constructed model is to accelerate the training speed and converge to a global optimal solution. These parameters affect the convergence speed of the algorithm and the performance of the algorithm in the training process. Unreasonable setting will increase the training time. If the learning rate is too small, the convergence speed will be reduced. If the learning rate is too large, the model will not converge. Unreasonable setting of the iteration number will cause overfitting. The batch sample size batch_size also affects the convergence of the algorithm.

[0103] In the embodiment of the present application, GraphSAGE can also be trained by random combination of subgraphs. The full-graph training mode of GCN (i.e., a graph convolutional neural network) can be changed into a small-batch training mode centered on nodes, so that the model can be trained on large-scale graph data and can make predictions on new graph structures, improving the applicability of industrial scenarios. For the neighbor aggregation operation, other centralized extraction of neighbor features can be replaced by convolution operation.

[0104] Optionally, the step of processing the attribute graph by using the preset learning model to obtain the vector value of each data node comprises: receiving the attribute graph through an input layer; performing preset neighbor node sampling processing on each data node in the attribute graph through a hidden layer to obtain an aggregated node set of each layer field of the data node; performing preset neighbor aggregation processing on the aggregated node set to obtain an aggregated feature value; generating a target aggregated feature value based on all aggregated features; and representing the target aggregated feature value as a vector value and outputting the vector value through an output layer.

[0105] In the embodiment of the application, the attribute graph can be received through the input layer, and then each data node in the attribute graph is processed through the hidden layer to obtain an aggregated node set of each layer field of the data node through preset neighbor node sampling (i.e. neighbor node random sampling). The neighbor node random sampling process is as follows:

[0106] For the target node v, the depth of feature aggregation can be assumed to be K, and the feature of the target node v can be represented as B k , 1≤k≤K represents the nodes contained in the kth layer neighborhood of node v, N k (B k ) represents the result of random sampling of the kth layer neighborhood, that is, a certain number of nodes in the kth layer neighborhood are randomly selected as aggregated nodes. Thus, the aggregated node set of each layer neighborhood of the target node v can be obtained, and the number of nodes in each layer neighborhood is different.

[0107] Then, the aggregated node set is processed through preset neighbor aggregation (i.e. neighbor aggregation) to obtain an aggregated feature value, and the target aggregated feature value is generated based on all aggregated features. The neighbor aggregation process is as follows:

[0108] On the basis of the obtained randomly sampled nodes, the final aggregated representation of the target node can be obtained by aggregating their feature representations. The specific learning strategy is as follows:

[0109]

[0110] wherein, and (k=1, 2, 3, …, K) are the kth layer features extracted through the GraphSAGE model; is N(u), N(u) ′ is the feature aggregated through Aggregate at the tth layer; represents the neighbor node set of node u, N(u) ′ represents the neighbor node set of node u'; CONCAT is feature concatenation; σ l and σ2 are activation functions; is a weight parameter that needs to be trained by the model.

[0111] In the embodiment, the Aggregate aggregation method generally includes a mean aggregator, a GCN (i.e., graph convolutional neural network) aggregator, an LSTM (i.e., long short-term memory network) aggregator, and a maximum pooling aggregator. As an optional implementation of the present application, the present application adopts two aggregation functions, namely mean aggregation and maximum pooling aggregation.

[0112] The mean aggregation is based on taking the average of each dimension of the neighbor nodes of the graph, expressing the correlation distribution of the node neighbors through the mean, and the mean aggregation function is expressed as follows:

[0113]

[0114] wherein, represents the kth layer feature of the model, mean is a mean operator, represents the neighbor node set of node u.

[0115] The maximum pooling aggregation can capture the most prominent performance between neighbor nodes in space and time. After the neighbor nodes are nonlinearly converted, the maximum value of each dimension of the neighbor nodes is selected, and the maximum pooling aggregation function is expressed as follows:

[0116]

[0117] wherein, represents the kth layer feature of the model, max is a maximum value operator, W pool represents a pooling weight parameter, b represents a bias parameter, represents the neighbor node set of node u, and sigma represents an activation function.

[0118] In the embodiment, the activation function is generally a nonlinear function, which can make the network have more abundant expression ability. As an optional implementation of the present application, the present application selects a ReLU activation function, which is expressed as follows:

[0119]

[0120] wherein, z represents a variable of the activation function, and when z is negative, the value of the function is 0, which is used to ensure that the neural network satisfies a nonlinear activation condition.

[0121] Figure 4 is a schematic diagram of an optional multi-source data fusion process based on a graph convolutional network according to the embodiment of the present application, as shown in Figure 4As shown, preprocess gene expression data, DNA methylation data, copy number variation data, phenotype data and other multi-source data, then through association network fusion (ANF) and concatenation of multi-source data, obtain graph data, then perform graph representation learning on the graph data, and the learned graph data can be applied to task analysis of downstream applications.

[0122] In the embodiment, the micro-omics data and the macro-phenotype data can be fused, and in order to test the effectiveness of the scheme, the method can be applied to two different fields of supervised learning and unsupervised learning respectively, wherein the accuracy of the graph convolution-based supervised learning method can reach 85.3%, and the clustering result of the graph-based representation learning also has good performance, and the data can be optimally clustered into different categories, and the clustering Land coefficient reaches 0.7653. As shown in Table 1, in the supervised learning application, the framework proposed in the embodiment is proved by experiments to be superior to the related art fusion based on AutoEncoder (self-encoder), fusion based on SNF (similarity network) and fusion based on GBDT (gradient boosting number model) in terms of three indexes of accuracy, F1 (an index for evaluating the performance of a binary classifier) and AUC (Area Under Curve, curve under area).

[0123] Table 1

[0124]

[0125] In the embodiment, in order to make the various multi-source data more effectively serve the lower layer application, a framework for fusing multi-source data into an attribute graph is constructed, and the graph convolutional neural network algorithm used in the embodiment can also be replaced by other similar graph algorithms (for example, a graph algorithm such as a self-attention mechanism graph convolutional neural network).

[0126] In the embodiment, the multi-source data can be fused efficiently and meaningfully, and the constructed observation object attribute graph can be better applied to the analysis and prediction tasks of the downstream observation objects after graph representation learning, and the effectiveness of the framework can also be verified on the experimental results.

[0127] The following will be described in detail in combination with another embodiment.

[0128] Embodiment two

[0129] The multi-source data fusion device provided in the embodiment includes a plurality of implementation units, and each implementation unit corresponds to each implementation step in the above embodiment one.

[0130] Figure 5is a schematic diagram of an optional multi-source data fusion device according to an embodiment of the present application, as shown in the figure, the fusion device can include: a processing unit 50, a construction unit 51, a fusion unit 52, an establishment unit 53, wherein, Figure 5

[0131] The processing unit 50 is configured to collect multi-omics data and phenotype data of the observation object, and pre-process the multi-omics data and the phenotype data to obtain target multi-omics data and target phenotype data.

[0132] The construction unit 51 is configured to construct a similarity matrix for each of the target multi-omics data and the target phenotype data.

[0133] The fusion unit 52 is configured to fuse all the similarity matrices to generate a fusion network.

[0134] The establishment unit 53 is configured to establish an attribute graph of the observation object based on the fusion network, the target multi-omics data, and the target phenotype data.

[0135] The fusion device can collect multi-omics data and phenotype data of the observation object through the processing unit 50, and pre-process the multi-omics data and the phenotype data to obtain target multi-omics data and target phenotype data. The construction unit 51 constructs a similarity matrix for each of the target multi-omics data and the target phenotype data. The fusion unit 52 fuses all the similarity matrices to generate a fusion network. The establishment unit 53 establishes an attribute graph of the observation object based on the fusion network, the target multi-omics data, and the target phenotype data. In the embodiment of the present application, the collected multi-omics data and phenotype data can be pre-processed first, and then a similarity matrix is constructed for each of the processed target multi-omics data and the target phenotype data. All the similarity matrices are fused to generate a fusion network. Then, an attribute graph of the observation object is established based on the fusion network, the target multi-omics data, and the target phenotype data. The fusion of multi-omics data and phenotype data with large structural differences can be realized while preserving the scale of each single data, so as to obtain an attribute graph of the observation object, thereby solving the technical problem that the multi-omics data and the phenotype data with large structural differences cannot be effectively fused in the related art, resulting in a large deviation in the fusion result.

[0136] ​Optionally, the processing unit comprises: a first determination module configured to determine omics features and omics feature values corresponding to the omics features of each sample in the omics data; a first processing module configured to perform missing value processing on the omics data based on the omics features and the omics feature values corresponding to the omics features, to obtain initial omics data; a second processing module configured to perform logarithmic processing on the omics feature values of each sample in the initial omics data, to obtain initial omics feature values; a third processing module configured to perform normalization processing on the initial omics feature values, to obtain target omics feature values, wherein the target omics feature values indicate samples that form target omics data; and a first mapping module configured to perform a preset embedding mapping on the phenotype data, to obtain target phenotype data.

[0137] Optionally, the first processing module comprises: a first determination sub-module configured to determine target omics features whose missing proportions of omics feature values in all samples are greater than a first preset proportion threshold; a second determination sub-module configured to determine target samples whose missing proportions of all omics feature values are greater than a second preset proportion threshold; a first deletion sub-module configured to delete the target omics features of each sample in the omics data, and delete the target samples in the omics data; and a first filling sub-module configured to fill the remaining missing values of each sample in the omics data after deletion by using a preset mean value algorithm, to obtain the initial omics data, wherein the remaining missing values refer to values of omics features missing in each sample in the omics data after deletion.

[0138] Optionally, the construction unit comprises: a second determination module configured to determine a first sample diameter of a first sample and a second sample diameter of a second sample in the observation data, wherein the observation data is the target omics data or the target phenotype data, the first sample is any sample in the observation data, and the second sample is any sample in the observation data except the first sample; a third determination module configured to determine a similarity measurement index value between the first sample and the second sample based on a sample distance between the first sample and the second sample, the first sample diameter, the second sample diameter, a first weighting parameter, and a second weighting parameter; a first construction module configured to construct a full connection similarity matrix based on all similarity measurement index values; and a second construction module configured to construct a similarity matrix of the observation data based on the full connection similarity matrix.

[0139] Optionally, the second determining module comprises: a third determining submodule, configured to determine a first neighbor sample set of the first sample; a fourth determining submodule, configured to determine a first sample distance between the first sample and each first neighbor sample in the first neighbor sample set, and determine a first sample diameter of the first sample based on all the first sample distances; a fifth determining submodule, configured to determine a second neighbor sample set of the second sample; and a sixth determining submodule, configured to determine a second sample distance between the second sample and each second neighbor sample in the second neighbor sample set, and determine a second sample diameter of the second sample based on all the second sample distances.

[0140] Optionally, the fusion unit comprises: a fourth determining module, configured to determine total information corresponding to each similarity matrix and an information weight corresponding to each similarity matrix; and a first generating module, configured to generate the fusion network based on the total information corresponding to all the similarity matrices and the information weights.

[0141] Optionally, the establishing unit comprises: a fifth determining module, configured to determine a data feature value corresponding to each data feature based on the multiple target omics data and the target phenotype data; a first splicing module, configured to splice all the data feature values corresponding to the data features to obtain a feature attribute of the data features; a sixth determining module, configured to determine a data node corresponding to the data features based on the fusion network; and a second generating module, configured to associate the feature attribute as a node attribute of the data node, and generate the attribute graph.

[0142] Optionally, the fusion apparatus further comprises: a fourth processing module, configured to, after establishing the attribute graph of the observation object based on the fusion network, the multiple target omics data and the target phenotype data, process the attribute graph by using a preset learning model to obtain a vector value of each data node, wherein the preset learning model comprises: an input layer, a hidden layer and an output layer; a third generating module, configured to generate a target attribute graph based on all the vector values; and a first analyzing module, configured to perform task analysis on the observation object based on the target attribute graph.

[0143] Optionally, the fourth processing module comprises: a first receiving submodule, configured to receive the attribute graph through the input layer; a first processing submodule, configured to perform preset neighbor node sampling processing on each data node in the attribute graph through the hidden layer to obtain an aggregated node set of each layer field of the data node; a second processing submodule, configured to perform preset neighbor aggregation processing on the aggregated node set to obtain an aggregated feature value; a first generating submodule, configured to generate a target aggregated feature value based on all the aggregated features; and a first output submodule, configured to represent the target aggregated feature value as the vector value, and output the vector value through the output layer.

[0144] The fusion device can further include a processor and a memory, and the processing unit 50, the construction unit 51, the fusion unit 52, the establishment unit 53 and the like are stored in the memory as program units, and the corresponding functions are realized by the processor executing the program units stored in the memory.

[0145] The processor includes a core, and the core calls the corresponding program units in the memory. The core can be one or more, and the attribute graph of the observation object is established based on the fusion network, the multiple target omics data and the target phenotype data by adjusting the core parameters.

[0146] The memory can include a non-permanent memory in a computer readable medium, a random access memory (RAM) and / or a non-volatile memory such as a read-only memory (ROM) or a flash memory (flash RAM), and the memory includes at least one memory chip.

[0147] The application also provides a computer program product adapted to execute the program of the following method steps when executed on a data processing device: collecting multiple omics data and phenotype data of an observation object, and pre-processing the multiple omics data and the phenotype data to obtain multiple target omics data and target phenotype data, constructing a similarity matrix for each of the multiple target omics data and the target phenotype data, fusing all the similarity matrices to generate a fusion network, and establishing an attribute graph of the observation object based on the fusion network, the multiple target omics data and the target phenotype data.

[0148] According to another aspect of the embodiment of the application, a computer readable storage medium is also provided, which includes a stored computer program, wherein the computer readable storage medium controls the device where the computer readable storage medium is located to execute the fusion method of the multiple source data when the computer program is running.

[0149] According to another aspect of the embodiment of the application, an electronic device is also provided, which includes one or more processors and a memory, and the memory is used to store one or more programs, wherein the one or more programs make the one or more processors realize the fusion method of the multiple source data when the one or more programs are executed by the one or more processors.

[0150] Figure 6 It is a hardware structure block diagram of an electronic device (or a mobile device) for a fusion method of multiple source data according to the embodiment of the application. As shown in Figure 6 the electronic device can include one or more Figure 6The processor 602 (the processor 602 can include, but is not limited to, a processing device such as a microprocessor MCU or a programmable logic device FPGA) and a memory 604 for storing data are shown by using 602a, 602b,..., 602n. In addition, a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which can be included as one of the ports of the I / O interface), a network interface, a keyboard, a power supply and / or a camera can be included. Those skilled in the art can understand that Figure 6 The structure shown is only schematic, and does not limit the structure of the above-mentioned electronic device. For example, the electronic device can further include more or fewer components than Figure 6 The structure shown is only schematic, and does not limit the structure of the above-mentioned electronic device. For example, the electronic device can further include more or fewer components than Figure 6 The structure shown is only schematic, and does not limit the structure of the above-mentioned electronic device. For example, the electronic device can further include more or fewer components than

[0151] The above-mentioned embodiment numbers of the application are only for description, and do not represent the advantages and disadvantages of the embodiments.

[0152] In the above-mentioned embodiments of the application, the description of each embodiment has its own emphasis, and the parts not described in detail in a certain embodiment can be referred to the related description of other embodiments.

[0153] In several embodiments provided in the present application, it should be understood that the disclosed technical content can be implemented by other ways. Among them, the above-mentioned device embodiments are only schematic, for example, the division of the units can be a logical function division, and actual implementation can have another division way, for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the units or modules shown or discussed can be indirect coupling or communication connection through some interfaces, units or modules, which can be electrical or other forms.

[0154] The units described as separate components can or can not be physically separated, and the components shown as units can or can not be physical units, that is, they can be located in one place, or they can be distributed on multiple units. Part or all of the units can be selected to achieve the purpose of the embodiment scheme according to actual needs.

[0155] In addition, each functional unit in each embodiment of the application can be integrated in one processing unit, or each unit can exist physically, or two or more units can be integrated in one unit. The above-mentioned integrated unit can be realized in the form of hardware or in the form of software functional unit.

[0156] The integrated unit, if implemented in the form of a software function unit and sold or used as an independent product, can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application, essentially or in other words, the part that contributes to the prior art or the whole or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium, including a number of instructions to make a computer device (which can be a personal computer, a server or a network device, etc.) execute all or part of the steps of the methods described in various embodiments of the present application. The aforementioned storage medium includes: a U disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a mobile hard disk, a magnetic disk or an optical disk, and various media that can store program codes.

[0157] The above is only the preferred embodiment of the present application, it should be pointed out that, for those skilled in the art, without departing from the principles of the present application, can make a number of improvements and refinements, these improvements and refinements should also be considered as the protection scope of the present application.

Claims

1. A method for fusing multi-source data, characterized in that, The method comprises the following steps: Collecting and preprocessing multiple omics data and phenotype data of an observation object to obtain multiple target omics data and target phenotype data; Constructing a similarity matrix for each of the multiple target omics data and the target phenotype data; wherein, a first sample diameter of a first sample and a second sample diameter of a second sample in observation data are determined, wherein the observation data is the target omics data or the target phenotype data, the first sample is any sample in the observation data, and the second sample is any sample in the observation data except the first sample; a similarity measurement index value between the first sample and the second sample is determined by a Gaussian kernel function based on a sample distance between the first sample and the second sample, the first sample diameter, the second sample diameter, a first weighting parameter, and a second weighting parameter; a fully connected similarity matrix is constructed based on all the similarity measurement index values; the similarity matrix of the observation data is constructed based on the fully connected similarity matrix; the first sample diameter or the second sample diameter is determined based on a distance between a current sample and any sample of k nearest neighbor samples of the current sample and k, wherein the current sample is the first sample or the second sample; Fusing all the similarity matrices to generate a fusion network; Establishing an attribute graph of the observation object based on the fusion network, the multiple target omics data, and the target phenotype data.

2. The fusion method of claim 1, wherein, The preprocessing of the multiple omics data and the phenotype data to obtain the multiple target omics data and the target phenotype data comprises the following steps: Determining an omics feature of each sample in the omics data and an omics feature value corresponding to the omics feature; Performing missing value processing on the omics data based on the omics feature and the omics feature value corresponding to the omics feature to obtain initial omics data; Performing logarithmic processing on the omics feature value of each sample in the initial omics data to obtain initial omics feature values; Performing normalization processing on the initial omics feature values to obtain target omics feature values, wherein the target omics feature values indicate the samples that constitute the target omics data; Performing a preset embedding mapping on the phenotype data to obtain the target phenotype data.

3. The fusion method of claim 2, wherein, The step of performing missing value processing on the omics data based on the omics feature and the omics feature value corresponding to the omics feature to obtain initial omics data comprises the following steps: Determining a target omics feature whose missing proportion of omics feature values in all samples is greater than a first preset proportion threshold; Determining a target sample whose missing proportion of omics feature values is greater than a second preset proportion threshold; Deleting the target omics feature of each sample in the omics data and the target sample in the omics data; The remaining missing values of each sample in the omics data after deletion are filled by using a preset mean value algorithm to obtain the initial omics data, wherein the remaining missing values refer to the missing values of the omics characteristics of each sample in the omics data after deletion.

4. The fusion method of claim 1, wherein The steps of determining the first sample diameter of the first sample and the second sample diameter of the second sample in the observation data include: determining a first neighbor sample set of the first sample; determining a first sample distance between the first sample and each first neighbor sample in the first neighbor sample set, and determining a first sample diameter of the first sample based on all the first sample distances; determining a second neighbor sample set of the second sample; determining a second sample distance between the second sample and each second neighbor sample in the second neighbor sample set, and determining a second sample diameter of the second sample based on all the second sample distances.

5. The fusion method of claim 1, wherein The steps of fusing all the similarity matrices to generate a fusion network include: determining the total information corresponding to each similarity matrix and the information weight corresponding to each similarity matrix; generating the fusion network based on the total information corresponding to all the similarity matrices and the information weight.

6. The fusion method of claim 1, wherein The steps of establishing the attribute graph of the observation object based on the fusion network, multiple target omics data and target phenotype data include: determining the data feature value corresponding to each data feature based on multiple target omics data and target phenotype data; splicing all the data feature values corresponding to the data feature to obtain the feature attribute of the data feature; determining the data node corresponding to the data feature based on the fusion network; associating the feature attribute as the node attribute of the data node to generate the attribute graph.

7. The fusion method of claim 1, wherein After establishing the attribute graph of the observation object based on the fusion network, multiple target omics data and target phenotype data, it further includes: processing the attribute graph by using a preset learning model to obtain the vector value of each data node, wherein the preset learning model includes an input layer, a hidden layer and an output layer; generating a target attribute graph based on all the vector values; performing task analysis on the observation object based on the target attribute graph.

8. The fusion method of claim 7, wherein, The steps of processing the attribute graph by using a preset learning model to obtain the vector value of each data node include: receiving the attribute graph through the input layer; performing preset neighbor node sampling processing on each data node in the attribute graph through the hidden layer to obtain the aggregated node set of each layer field of the data node; performing preset neighbor aggregation processing on the aggregated node set to obtain an aggregated feature value; generating a target aggregated feature value based on all the aggregated features; representing the target aggregated feature value as the vector value, and outputting the vector value through the output layer.

9. A fusion device of multi-source data, characterized in that, The processing unit is configured to collect multiple omics data and phenotype data of an observation object, and pre-process the multiple omics data and the phenotype data to obtain multiple target omics data and target phenotype data. ​ The construction unit is configured to construct a similarity matrix for each of the plurality of target omics data and the target phenotype data; wherein a first sample diameter of a first sample and a second sample diameter of a second sample in observation data are determined, the observation data being the target omics data or the target phenotype data, the first sample being any sample in the observation data, and the second sample being any sample in the observation data except the first sample; a similarity measurement index value between the first sample and the second sample is determined by a Gaussian kernel function based on a sample distance between the first sample and the second sample, the first sample diameter, the second sample diameter, a first weighting parameter, and a second weighting parameter; a full connection similarity matrix is constructed based on all the similarity measurement index values; and the similarity matrix of the observation data is constructed based on the full connection similarity matrix; the first sample diameter or the second sample diameter is determined based on a distance between a current sample and any sample of k nearest neighbor samples of the current sample and k, the current sample being the first sample or the second sample; The fusion unit is configured to fuse all the similarity matrices to generate a fusion network. The establishment unit is configured to establish an attribute graph of the observation object based on the fusion network, the plurality of target omics data, and the target phenotype data.

10. A computer-readable storage medium, characterized in that, The computer readable storage medium includes a stored computer program, wherein the computer readable storage medium controls a device where the computer readable storage medium is located to execute the multi-source data fusion method of any one of claims 1 to 8 when the computer program is running.

11. An electronic device, comprising: The device includes one or more processors and a memory, and the memory is configured to store one or more programs, wherein the one or more programs, when executed by the one or more processors, cause the one or more processors to implement the multi-source data fusion method of any one of claims 1 to 8.

Citation Information

Patent Citations

  • Clustering analysis method and system for multi-omics data

    CN113392894A

  • Drug relocation system and method based on heterogeneous association network deep learning

    CN114038574A