Multi-source data fusion lung cancer prediction method and system
Through multi-source data fusion and graph neural network technology, gene expression, clinical and imaging data are integrated to build a high-precision lung cancer prediction model, solving the problem of low prediction accuracy in the existing technology.
Patent Information
- Application Number
- CN202510274149.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-10
- Publication Date
- 2025-06-27
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In the prior art, cancer risk prediction has the problem of low prediction accuracy.
A multi-source data fusion method was used to obtain a data set of multi-source lung cancer patients including gene expression data, clinical data and imaging data. Through preprocessing, feature extraction and graph neural network (GNN) training, a unified feature vector was constructed, and a deep learning model was used to construct a lung cancer prediction model.
By integrating information from different data sources, multi-dimensional characteristic data is extracted, which improves the accuracy of lung cancer prediction and captures potential lung cancer risk factors.
Smart Images

Figure CN120221072A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of lung cancer prediction, and particularly to a lung cancer prediction method and system based on multi-source data fusion. Background Art
[0002] With the continuous progress of medical technology, early screening and risk prediction of lung cancer have become increasingly important. Lung cancer is one of the most lethal malignant tumors globally. Effective prediction of its risk not only helps with early diagnosis but also improves the treatment effect. In the prior art, the patent with publication number CN 111899882 B discloses a method and system for predicting cancer, including: performing differential analysis on the gene expression profile data of cancer patients and normal people to obtain differential genes; analyzing the gene expression profile data of cancer patients and normal people based on weighted gene co-expression network analysis to obtain hub genes; and processing the gene expression profile data of differential genes through a variational autoencoder algorithm to obtain dimensionality-reduced data; using the gene expression profile data of hub genes and the dimensionality-reduced data as classification features of a preset type of cancer classifier, and achieving accurate classification of cancer patients and normal people through the cancer classifier; in the above solution, when using weighted gene co-expression network analysis, it is equivalent to extracting unit data features, and when using a variational autoencoder to process the expression data of differential genes, the dimensionality reduction method may lead to information loss, thus affecting the classification performance.
[0003] In summary, there is a problem of low prediction accuracy in current cancer risk prediction. Summary of the Invention
[0004] In order to overcome the deficiencies of the prior art, the purpose of the present invention is to provide a lung cancer prediction method and system based on multi-source data fusion, and the present invention solves the problem of low prediction accuracy in cancer risk prediction in the prior art.
[0005] To achieve the above objective, the present invention provides the following solutions:
[0006] A lung cancer prediction method based on multi-source data fusion, comprising:
[0007] Obtaining a multi-source lung cancer patient dataset, where the multi-source lung cancer patient dataset includes: gene expression data, clinical data, and imaging data;
[0008] Performing corresponding preprocessing on different data in the multi-source lung cancer patient dataset to obtain a preprocessed multi-source lung cancer patient dataset, where the preprocessed multi-source lung cancer patient dataset includes: gene expression data, preprocessed clinical data, and structured features;
[0009] Performing feature extraction on the preprocessed multi-source lung cancer patient dataset to obtain first feature data, second feature data, and third feature data;
[0010] Convert the first feature data, the second feature data, and the third feature data into a graph structure and train them with a GNN to obtain a unified feature vector;
[0011] Construct a lung cancer prediction model based on the unified feature vector and a deep learning model to perform lung cancer prediction and obtain a prediction result.
[0012] Preferably, corresponding preprocessing is performed on different data in the multi-source lung cancer patient dataset to obtain a preprocessed multi-source lung cancer patient dataset:
[0013] Perform normalization and differential expression analysis on the gene expression data using the limma package in R language to obtain gene expression data;
[0014] Perform data cleaning on the clinical data and process missing values and outliers to obtain preprocessed clinical data;
[0015] Extract features from the imaging data to obtain corresponding structured features.
[0016] Preferably, feature extraction is performed on the preprocessed multi-source lung cancer patient dataset, and the first feature data, the second feature data, and the third feature data include:
[0017] Perform differential gene extraction and correlation analysis on the gene expression data to screen out important genes related to lung cancer to obtain the first feature data;
[0018] Convert the preprocessed clinical data into a hierarchical variable and construct interaction features to extract the second feature data;
[0019] Extract high-dimensional features from the structured features and process them using a dimensionality reduction algorithm to obtain the third feature data.
[0020] Preferably, the step of converting the first feature data, the second feature data, and the third feature data into a graph structure and training them with a GNN to obtain a unified feature vector includes:
[0021] Construct a feature graph structure based on the first feature data, the second feature data, and the third feature data;
[0022] Use a feature clustering method to construct a geometric model on the graph surface of the feature graph structure to obtain an unfolded feature graph;
[0023] Use the GNN network to train the unfolded feature graph to obtain an embedded vector set;
[0024] Aggregate the embedded vector set to obtain the unified feature vector.
[0025] Preferably, constructing the feature map structure according to the first feature data, the second feature data, and the third feature data includes:
[0026] Calculating the similarity of the first feature data, the second feature data, and the third feature data to obtain similarity data;
[0027] Determining the edges of the feature map structure according to the similarity data;
[0028] Determining the first feature data, the second feature data, and the third feature data as the nodes of the feature map structure.
[0029] Preferably, the GNN is set as a multi-layer network structure.
[0030] Preferably, the expression of the unified feature vector is:
[0031]
[0032] where h uniform is the unified feature vector, w i is the adaptive weight, N is a natural number, and h i is the i-th embedding vector.
[0033] Preferably, the expression of the adaptive weight is:
[0034] w i = f(h i , context) = σ(W w ·h i + b w );
[0035] where σ is the activation function, W w is the weight matrix, b w is the bias term, and context represents the feature information from neighboring nodes.
[0036] Preferably, the calculation expression of the similarity data is:
[0037] Sim combined (q i , g j , f k ) = α·Sim cos (q i , g j ) + β·Sim cos (q i , f k ) + γ·Sim cos (g j , f k );
[0038] Among them, q i represents the first feature data embedding vector of the i-th sample, and g j represents the second feature data embedding vector of the j-th sample, and f k represents the third feature data embedding vector of the k-th sample. α, β, and γ are the first weight parameter, the second weight parameter, and the third weight parameter respectively. Sim combined is the comprehensive similarity. Among them, the matrix corresponding to the comprehensive similarity is:
[0039]
[0040] S is the similarity matrix, which is used to determine the edge connections between the nodes in the feature map structure.
[0041] A lung cancer prediction system for multi-source data fusion includes:
[0042] A data acquisition module for acquiring a multi-source lung cancer patient dataset, where the multi-source lung cancer patient dataset includes: gene expression data, clinical data, and imaging data;
[0043] A preprocessing module for performing corresponding preprocessing on different data in the multi-source lung cancer patient dataset to obtain a preprocessed multi-source lung cancer patient dataset, where the preprocessed multi-source lung cancer patient dataset includes: gene expression data, preprocessed clinical data, and structured features;
[0044] A feature extraction module for extracting features from the preprocessed multi-source lung cancer patient dataset, including first feature data, second feature data, and third feature data;
[0045] A feature unification module for converting the first feature data, the second feature data, and the third feature data into a graph structure and training them by a GNN to obtain a unified feature vector;
[0046] A prediction module for constructing a lung cancer prediction model based on the unified feature vector and a deep learning model to perform lung cancer prediction and obtain a prediction result.
[0047] The present invention discloses the following technical effects:
[0048] The present invention provides a lung cancer prediction method and system for multi-source data fusion. The method includes: obtaining a multi-source lung cancer patient dataset, where the multi-source lung cancer patient dataset includes: gene expression data, clinical data, and imaging data; performing corresponding preprocessing on different data in the multi-source lung cancer patient dataset to obtain a preprocessed multi-source lung cancer patient dataset, where the preprocessed multi-source lung cancer patient dataset includes: gene expression data, preprocessed clinical data, and structured features; performing feature extraction on the preprocessed multi-source lung cancer patient dataset to obtain first feature data, second feature data, and third feature data; converting the first feature data, second feature data, and the third feature data into a graph structure and training them by a GNN to obtain a unified feature vector; constructing a lung cancer prediction model based on the unified feature vector and a deep learning model to perform lung cancer prediction and obtain a prediction result. By effectively integrating multiple data from different sources (gene expression data, clinical data, and imaging data), the present invention forms a comprehensive lung cancer patient dataset. By extracting the first, second, and third feature data, this method not only considers gene correlation but also comprehensively considers clinical data and imaging features. This multi-dimensional feature extraction enables the model to be more comprehensive in learning data relationships, helps capture potential lung cancer risk factors, and improves the accuracy of lung cancer prediction. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the embodiments. Obviously, the following described drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0050] Figure 1 It is a flowchart of a lung cancer prediction method for multi-source data fusion provided by an embodiment of the present invention;
[0051] Figure 2 It is a schematic diagram of the construction process of the unified feature vector provided by an embodiment of the present invention;
[0052] Figure 3 It is a schematic diagram of the structure of a lung cancer prediction system for multi-source data fusion provided by an embodiment of the present invention.
[0053] Reference numerals:
[0054] 1 - Data acquisition module, 2 - Preprocessing module, 3 - Feature extraction module, 4 - Feature unification module, 5 - Prediction module. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0055] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0056] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0057] As Figure 1 shown, the present invention provides a lung cancer prediction method for multi-source data fusion, including:
[0058] Step 100: Obtain a multi-source lung cancer patient dataset, where the multi-source lung cancer patient dataset includes: gene expression data, clinical data, and imaging data;
[0059] Specifically, for gene expression data: Obtain the gene expression data of lung cancer patients using public databases (such as TCGA, GEO, ArrayExpress); select eligible studies, such as relevant datasets for non-small cell lung cancer (NSCLC) or small cell lung cancer (SCLC). Use R language or Python scripts to download the required gene expression matrix and corresponding sample information from the public database. Conduct a preliminary screening of the downloaded data and retain relevant samples (such as samples diagnosed with lung cancer).
[0060] For clinical data: Obtain the clinical information of lung cancer patients from hospitals or cancer registries, including medical history, treatment records, physical examination results, tumor stage, typing, etc. Ensure the privacy protection and legality of the data, and follow the review and approval procedures of the ethics committee. Extract the clinical diagnosis information of patients through the electronic medical record system (EMR). Relevant characteristics of each patient, such as age, gender, smoking history, family history, and past medical history, are classified and sorted.
[0061] For imaging data: Collect lung imaging data, such as CT scans or X-rays, which can usually be obtained through the hospital's imaging database. The imaging data needs to be labeled and archived to ensure its diagnostic value and usability. Use the hospital information system (HIS) or the radiology picture archiving and communication system (PACS) to obtain the imaging data. Digitally process the images to ensure that the data meets the storage and analysis standards (such as DICOM format).
[0062] Specifically, ensure the uniform format of gene expression data, clinical data, and imaging data, and convert them into a compatible data structure (such as CSV, JSON, or database format). Integrate the gene expression data, clinical data, and imaging data into a unified database based on the patient ID. Create a data dictionary to explain the meaning and source of each variable in the dataset for convenient subsequent analysis.
[0063] More specifically, gene expression information: the gene expression matrix of each sample.
[0064] Clinical information: age, gender, smoking and alcohol history, treatment plan, follow-up status, etc.
[0065] Imaging information: CT scan image files and descriptions.
[0066] Step 200: Perform corresponding preprocessing on different data in the multi-source lung cancer patient dataset to obtain a preprocessed multi-source lung cancer patient dataset, where the preprocessed multi-source lung cancer patient dataset includes: gene expression data, preprocessed clinical data, and structured features;
[0067] Step 300: Extract features from the preprocessed multi-source lung cancer patient dataset, including first feature data, second feature data, and third feature data;
[0068] Step 400: Convert the first feature data, second feature data, and third feature data into a graph structure and train them with a GNN to obtain a unified feature vector;
[0069] Step 500: Construct a lung cancer prediction model based on the unified feature vector and a deep learning model to perform lung cancer prediction and obtain a prediction result.
[0070] Furthermore, the corresponding preprocessing of different data in the multi-source lung cancer patient dataset to obtain a preprocessed multi-source lung cancer patient dataset:
[0071] Perform standardization and differential expression analysis on the gene expression data using the limma package in R language to obtain gene expression data;
[0072] Specifically, data loading and installation of dependent packages - reading gene expression data - data standardization - setting the comparison design matrix - differential expression analysis - obtaining the final result.
[0073] Perform data cleaning on the clinical data and handle missing values and outliers to obtain preprocessed clinical data;
[0074] The specific step process is: reading clinical data - checking for missing values - handling missing values - identifying and handling outliers - saving the preprocessed clinical data.
[0075] Feature extraction is performed on the image data to obtain corresponding structured features.
[0076] The specific step process is: reading image data - image processing and feature extraction - image preprocessing - extracting structured features - saving the extracted features.
[0077] Furthermore, feature extraction is performed on the preprocessed multi-source lung cancer patient dataset, including first feature data, second feature data, and third feature data:
[0078] Differentially expressed gene extraction and correlation analysis are performed on the gene expression data to screen out important genes related to lung cancer, obtaining the first feature data;
[0079] Specifically, advanced statistical methods are used to compare gene expression levels in different groups (such as the tumor group and the control group); multiple hypothesis testing correction (such as FDR) is introduced to reduce false positive results, thereby improving the reliability of gene screening; the correlation between the screened significant genes and multiple clinical variables (such as pathological stage, patient survival rate) is calculated, and the correlation coefficient is used to evaluate the association strength between genes and disease phenotypes. A scientific threshold selection method is adopted to retain only genes with strong correlation with clinical characteristics, forming the first feature dataset.
[0080] The preprocessed clinical data is converted into categorical variables and interaction features are constructed to extract the second feature data;
[0081] Specifically, continuous clinical variables (such as age, tumor size) are converted into categorical variables, and cut using quantiles or clinical significance (such as age groups) to simplify the complexity of the model and improve interpretability; potential very important combinations of clinical variables (such as smoking status and tumor stage) are identified and interaction features are generated, aiming to capture the non-linear relationship and potential interaction effects between variables; new clinical features are generated through feature combination techniques to enhance the performance of the model; feature selection algorithms (such as recursive feature elimination, tree-based algorithms) are used to evaluate and select the clinical features that contribute most to model prediction.
[0082] High-dimensional features are extracted from the structured features and processed using a dimensionality reduction algorithm to obtain the third feature data.
[0083] Specifically, a series of high-dimensional features are extracted from the image data. These features can comprehensively describe the important information in the image (such as shape, texture, intensity, etc.) to ensure the capture of biomedical information that is not visible to the naked eye. An appropriate dimensionality reduction method, such as principal component analysis (PCA), t-SNE, or UMAP, is selected to process the extracted high-dimensional features. The dimensionality reduction algorithm should fully consider retaining most of the information while reducing the data complexity. Visualization tools (such as scatter plots and heat maps) are used to display the features after dimensionality reduction to facilitate the identification of potential population structures or data signals. The linear separability and clustering performance of the dimensionality reduction results are evaluated.
[0084] Further, as Figure 2 shown, the conversion of the first feature data, the second feature data, and the third feature data into a graph structure and training by the GNN to obtain a unified feature vector includes:
[0085] Step 401: Construct a feature graph structure according to the first feature data, the second feature data, and the third feature data;
[0086] Step 402: Use a feature clustering method to construct a geometric model on the graph surface of the feature graph structure to obtain an unfolded feature graph;
[0087] Step 403: Use the GNN network to train the unfolded feature graph to obtain an embedded vector set;
[0088] Step 404: Aggregate the embedded vector set to obtain the unified feature vector.
[0089] Specifically, integrate the features from gene expression data, clinical data, and image data to ensure that these features can be seamlessly combined to form a comprehensive feature set; convert each type of feature data into nodes of a graph structure. The first feature data (genes), the second feature data (clinical), and the third feature data (image features) are used as different node types in the graph respectively; clarify the attributes of each node. For example, gene nodes can contain expression levels, clinical nodes can include patient age, gender, etc., and image nodes can include the extracted high-dimensional features; construct edges based on the similarity between features; which genes show similarity under the same clinical background to form the connections of the graph; ensure that the weights of the edges in the graph can reflect the similarity or relationship strength between nodes to construct a feature graph structure.
[0090] Further, the construction of the feature graph structure according to the first feature data, the second feature data, and the third feature data includes:
[0091] Calculate the similarity of the first feature data, the second feature data, and the third feature data to obtain similarity data;
[0092] Specifically, to eliminate the influence between different feature magnitudes (such as the difference in the magnitude of gene expression and the numerical value of imaging features), all features are normalized (such as Z-score normalization or Min-Max scaling).
[0093] Determine the edges of the feature map structure according to the similarity data;
[0094] Determine the first feature data, the second feature data, and the third feature data as the nodes of the feature map structure.
[0095] Furthermore, the GNN is set to a multi-layer network structure.
[0096] The overall architecture of the GNN usually consists of multiple graph convolutional layers and fully connected layers, forming a deep network structure. The multi-layer network structure aims to capture the complex relationships between nodes and the non-linear mapping between features.
[0097] Furthermore, the expression of the unified feature vector is:
[0098]
[0099] where h uniform is the unified feature vector, w i is the adaptive weight, N is a natural number, and h i is the i-th embedding vector.
[0100] Reducing the weight of noise features reduces the interference of invalid information to the model, enhances the stability and robustness of the model. Integrating the properties of different features (such as gene, clinical, and imaging data) can generate richer feature representations, providing a more comprehensive basis for subsequent model learning and prediction tasks. This aggregation method allows for corresponding feature fusion when facing different types or scales of data, improving the flexibility of model construction and being applicable to a variety of application scenarios.
[0101] Furthermore, the expression of the adaptive weight is:
[0102] w i = f(h i , context) = σ(W w ·h i + b w );
[0103] where σ is the activation function, W w is the weight matrix, b w is the bias term, and context represents the feature information from neighboring nodes.
[0104] Specifically, the adaptive weights are dynamically adjusted according to the specific values of the features and the context information, and fixed weights are no longer used. This flexibility allows the model to better adapt to the variations between different datasets and features. The adaptive weights can automatically learn the relationships between features, which helps to reveal the potential hidden features in the data, and this is crucial for multi-source data fusion.
[0105] Furthermore, the calculation expression of the similarity data is as follows:
[0106] Sim combined (q i ,g j ,f k ) = α·Sim cos (q i ,g j ) + β·Sim cos (q i ,f k ) + γ·Sim cos (g j ,f k );
[0107] where q i represents the first feature data embedding vector of the i-th sample, g j represents the second feature data embedding vector of the j-th sample, f k represents the third feature data embedding vector of the k-th sample, α, β, and γ are the first weight parameter, the second weight parameter, and the third weight parameter respectively, and Sim combined is the comprehensive similarity. Among them, the matrix corresponding to the comprehensive similarity is:
[0108]
[0109] S is the similarity matrix, which is used to determine the edge connections between the nodes in the feature map structure.
[0110] Specifically, the comprehensive similarity calculation formula is used to evaluate the similarity between different features (such as gene expression, clinical data, and imaging data). By introducing adaptive weights, the formula can dynamically adapt to the correlations between features. Combining the similarities of different features can effectively extract information from multi-dimensional data, enabling the model to comprehensively consider multiple data sources and improve the comprehensiveness of the analysis. By weighted combination of different similarity metrics (such as cosine similarity), the comprehensive similarity formula can more comprehensively reflect the relationships between data features and enhance the effect of traditional similarity calculation.
[0111] Furthermore, for the similarity matrix: The final similarity matrix provides the similarity results between samples, which can intuitively display the positional relationship of different samples in the feature space, helping to identify samples with similar features. This matrix provides the basic data for the subsequent establishment of the feature map structure. Using the information in the similarity matrix, graph nodes and edges can be constructed more effectively, and then graph neural network (GNN) training can be carried out.
[0112] The combined similarity calculation formula and the role of the final similarity matrix in lung cancer prediction complement each other. The former extracts and integrates information from different data sources in a dynamic and weighted manner, while the latter integrates this information into an intuitive matrix form, facilitating data understanding and detection.
[0113] As Figure 3 shown, this embodiment also provides a lung cancer prediction system for multi-source data fusion, including:
[0114] A data acquisition module 1 for acquiring a multi-source lung cancer patient dataset, where the multi-source lung cancer patient dataset includes: gene expression data, clinical data, and imaging data;
[0115] A preprocessing module 2 for performing corresponding preprocessing on different data in the multi-source lung cancer patient dataset to obtain a preprocessed multi-source lung cancer patient dataset, where the preprocessed multi-source lung cancer patient dataset includes: gene expression data, preprocessed clinical data, and structured features;
[0116] A feature extraction module 3 for extracting features from the preprocessed multi-source lung cancer patient dataset, including first feature data, second feature data, and third feature data;
[0117] A feature unification module 4 for converting the first feature data, second feature data, and third feature data into a graph structure and training them by GNN to obtain a unified feature vector;
[0118] A prediction module 5 for constructing a lung cancer prediction model based on the unified feature vector and a deep learning model to perform lung cancer prediction and obtain a prediction result.
[0119] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The same or similar parts among the various embodiments can be referred to each other.
[0120] In this article, specific examples are used to elaborate on the principles and implementation modes of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation modes and application scopes. To sum up, the content of this specification should not be construed as a limitation to the present invention.
Claims
1. A lung cancer prediction method based on multi-source data fusion, characterized in that: include: Acquire a multi-source lung cancer patient dataset, wherein the multi-source lung cancer patient dataset includes: gene expression data, clinical data, and imaging data; Preprocessing the different data in the multi-source lung cancer patient data set accordingly to obtain a preprocessed multi-source lung cancer patient data set, wherein the preprocessed multi-source lung cancer patient data set includes: gene expression data, preprocessed clinical data and structured features; Performing feature extraction on the preprocessed multi-source lung cancer patient data set to obtain first feature data, second feature data, and third feature data; Convert the first feature data, the second feature data, and the third feature data into a graph structure and train them using a GNN to obtain a unified feature vector; A lung cancer prediction model is constructed based on the unified feature vector and the deep learning model to perform lung cancer prediction and obtain a prediction result.
2. The lung cancer prediction method based on multi-source data fusion according to claim 1, characterized in that: The different data in the multi-source lung cancer patient data set are preprocessed accordingly to obtain the preprocessed multi-source lung cancer patient data set: The gene expression data are standardized and differentially expressed using the limma package of the R language to obtain gene expression data; Performing data cleaning on the clinical data and processing missing values and abnormal values to obtain preprocessed clinical data; Feature extraction is performed on the image data to obtain corresponding structured features.
3. The lung cancer prediction method based on multi-source data fusion according to claim 2, characterized in that: The preprocessed multi-source lung cancer patient data set is subjected to feature extraction, and the first feature data, the second feature data and the third feature data include: Extracting differential genes from the gene expression data and performing correlation analysis to screen out important genes related to lung cancer to obtain first characteristic data; converting the preprocessed clinical data into hierarchical variables and constructing interactive features to extract second feature data; A high-dimensional feature is extracted from the structured feature and processed using a dimensionality reduction algorithm to obtain third feature data.
4. The lung cancer prediction method based on multi-source data fusion according to claim 2, characterized in that: The converting the first feature data, the second feature data, and the third feature data into a graph structure and training them by GNN to obtain a unified feature vector includes: Constructing a feature graph structure according to the first feature data, the second feature data and the third feature data; A geometric model is constructed on the surface of the feature graph structure by using a feature clustering method to obtain an expanded feature graph; Using the GNN network to train the expanded feature map to obtain an embedding vector set; The embedding vector set is aggregated to obtain the unified feature vector.
5. The lung cancer prediction method based on multi-source data fusion according to claim 4, characterized in that: The constructing a feature graph structure according to the first feature data, the second feature data and the third feature data comprises: Calculating similarities among the first feature data, the second feature data and the third feature data to obtain similarity data; Determining the edge of the feature graph structure according to the similarity data; The first feature data, the second feature data and the third feature data are determined as nodes of the feature graph structure.
6. The lung cancer prediction method based on multi-source data fusion according to claim 5, characterized in that: The GNN is set as a multi-layer network structure.
7. The lung cancer prediction method based on multi-source data fusion according to claim 5, characterized in that: The expression of the unified eigenvector is: Among them, h uniform is a unified eigenvector, w i is the adaptive weight, N is a natural number, h i is the i-th embedding vector.
8. The lung cancer prediction method based on multi-source data fusion according to claim 7, characterized in that: The expression of the adaptive weight is: w i =f(h i ,context)=σ(W w ·h i +b w ); Among them, σ is the activation function, W w is the weight matrix, b w is a bias term, and context represents feature information from neighboring nodes.
9. The lung cancer prediction method based on multi-source data fusion according to claim 7, characterized in that: The calculation expression of the similarity data is: Sim combined (q i ,g j ,f k )=α·Sim cos (q i ,g j )+β·Sim cos (q i ,f k )+γ·Sim cos (g j ,f k ); Among them, q i represents the first feature data embedding vector of the i-th sample, g j represents the second feature data embedding vector of the jth sample, f k represents the third feature data embedding vector of the kth sample, α, β and γ are the first weight parameter, the second weight parameter and the third weight parameter respectively, Sim combined is the comprehensive similarity, wherein the matrix corresponding to the comprehensive similarity is: S is a similarity matrix, which is used to determine the edge connections between nodes in the feature graph structure.
10. A lung cancer prediction system based on multi-source data fusion, characterized in that: include: A data acquisition module, used to acquire a multi-source lung cancer patient data set, wherein the multi-source lung cancer patient data set includes: gene expression data, clinical data and imaging data; A preprocessing module, used to perform corresponding preprocessing on different data in the multi-source lung cancer patient data set to obtain a preprocessed multi-source lung cancer patient data set, wherein the preprocessed multi-source lung cancer patient data set includes: gene expression data, preprocessed clinical data and structured features; A feature extraction module, used to extract features of the preprocessed multi-source lung cancer patient data set, including first feature data, second feature data and third feature data; A feature unification module, used to convert the first feature data, the second feature data and the third feature data into a graph structure and train the structure by GNN to obtain a unified feature vector; The prediction module is used to construct a lung cancer prediction model according to the unified feature vector and the deep learning model to perform lung cancer prediction and obtain a prediction result.
Citation Information
Patent Citations
A method and system for predicting cancer
CN111899882B