Cancer data synchronous detection and classification method based on model analysis
By integrating multimodal cancer data, performing data preprocessing and feature fusion, screening for sensitive biomarkers, constructing and optimizing an LSTM prediction model, the problems of large errors and long processing times in existing cancer detection technologies have been solved, achieving rapid and accurate cancer detection and classification.
Patent Information
- Application Number
- CN202511354217.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-22
- Publication Date
- 2025-11-07
AI Technical Summary
Existing cancer detection technologies suffer from large cumulative errors, long processing times, waste of resources, and complex subtypes, making it difficult to achieve early and accurate diagnosis and precise classification.
By integrating multimodal cancer data, performing data preprocessing and cross-modal feature fusion, screening for sensitive biomarkers, and combining LSTM prediction models and Monte Carlo simulation analysis algorithms for simultaneous detection and classification, the model is optimized to improve accuracy and efficiency.
It enables rapid and accurate detection of cancer presence and its subtypes in a single analysis, reducing misdiagnosis rate and diagnosis time, and improving resource utilization efficiency.
Smart Images

Figure CN120913822A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of cancer data analysis, in particular to a cancer data synchronous detection and classification method based on model analysis. BACKGROUND
[0002] Cancer is a malignant proliferative disease that occurs when cells in the body accumulate genetic mutations, and its core features include abnormal cell cycle regulation, tumor formation, spread to distant organs through blood vessels / lymphatic vessels, and the presence of genotype / phenotype difference cell subgroups within the same tumor. The early diagnosis rate of cancer is relatively high, the typing complexity is large, and the multi-modal data is fragmented, so it is necessary to integrate multi-modal cancer data to achieve cancer detection. Traditional process for detecting cancer, including image detection of cancer, has the characteristics of large error accumulation, long time consumption, and easy waste of resources. Through multi-modal data fusion and modeling, the purpose of whether there is cancer and what type of cancer can be obtained in a single analysis, and the speed is faster than traditional methods. Synchronous detection and classification of cancer is not only a technical iteration, but also a paradigm shift in diagnosis and treatment - from "segmented diagnosis" to "holistic decision-making", providing a core driving force for precision medicine. While reducing the medical burden, it truly achieves the clinical goal of "early detection, accurate classification, and precise treatment". Therefore, a cancer data synchronous detection and classification method based on model analysis is proposed. SUMMARY
[0003] The present application overcomes the shortcomings of the prior art and provides a cancer data synchronous detection and classification method based on model analysis.
[0004] To achieve the above purpose, the technical scheme adopted by the present application is as follows: The present application provides a cancer data synchronous detection and classification method based on model analysis, comprising the following steps: Integrate the multi-modal data of the cancer data, and perform data preprocessing on the multi-modal data of the cancer data to obtain preprocessed cancer integration data; Perform cross-modal feature fusion on all preprocessed cancer integration data, and select sensitive biomarkers of the preprocessed cancer integration data; Combine the sensitive biomarkers to perform joint optimization modeling of synchronous detection and synchronous classification of the multi-modal cancer data, and obtain a target LSTM prediction model; Combine the Monte Carlo simulation analysis algorithm to infer the synchronous detection and classification data of the target LSTM prediction model, and perform model iteration optimization on the target LSTM prediction model according to the comprehensive confidence index obtained by inference.
[0005] Further, in a preferred embodiment of the present application, the multi-modal data of the integrated cancer data is processed, and the multi-modal data of the integrated cancer data is preprocessed to obtain preprocessed integrated cancer data, specifically: Obtain multi-modal data of cancer data, and label it as multi-modal cancer data; The multi-modal cancer data includes clinical indicator data and genomic data of different types of cancer data, and the multi-modal cancer data is biomedical data related to cancer; Introduce a data processing terminal and a dynamic time warping method, obtain time series of different types of multi-modal cancer data, and perform time alignment processing on different types of multi-modal cancer data according to the time series of different types of multi-modal cancer data; At the same time, introduce a spatial registration algorithm to perform spatial alignment processing on the different types of multi-modal cancer data after time alignment processing, wherein the spatial alignment processing is to map the different types of multi-modal cancer data after time alignment processing in the same spatial coordinate system; After time alignment and spatial alignment processing, the multi-modal cancer data is labeled as spatiotemporal alignment multi-modal cancer data, and a filling algorithm based on a graph neural network is used to infer the similarity of the spatiotemporal alignment multi-modal cancer data, and based on the inference result, the missing values of the spatiotemporal alignment multi-modal cancer data are filled, and all the spatiotemporal alignment multi-modal cancer data after filling of the missing values are normalized to obtain preprocessed integrated cancer data.
[0006] Further, in a preferred embodiment of the present application, the multi-modal feature fusion of all preprocessed integrated cancer data is performed, and sensitive biomarkers of the preprocessed integrated cancer data are screened, specifically: In the data processing terminal, the preprocessed integrated cancer data is subjected to cross-modal feature embedding representation processing to obtain a feature vector of the preprocessed integrated cancer data, which is labeled as a cancer integration data feature vector; The cross-modal feature embedding representation processing is to extract a local gene mutation feature vector from the genomic data in the preprocessed integrated cancer data by a convolutional neural network, and to extract a dynamic feature vector from the clinical indicator data in the preprocessed integrated cancer data by a gated recurrent unit; In the data terminal, the cross-modal attention weights of different cancer integration data feature vectors are calculated by a multi-modal cross-attention mechanism, and a cross-modal causal chain of the preprocessed integrated cancer data is constructed according to the cross-modal attention weights of the different cancer integration data feature vectors; The gene-protein interaction graph of the preprocessed cancer integrated data is constructed by KEGG pathway, and the cross-modal attention weight of the feature vector of the different cancer integrated data is mapped to the gene-protein interaction graph, the node of the gene-protein interaction graph is updated, and the cross-modal causal chain of the preprocessed cancer integrated data is generated. Based on the big data network, all biomarkers existing in the multi-modal cancer data are retrieved, and the occurrence frequency and the number of occurrences of different biomarkers in the cross-modal causal chain are calculated according to the cross-modal causal chain of the preprocessed cancer integrated data. If the occurrence frequency and the number of occurrences of the biomarker in the cross-modal causal chain are greater than a preset value, the corresponding biomarker is marked as a sensitive biomarker.
[0007] Further, in a preferred embodiment of the present application, the combination of the sensitive biomarker, the joint optimization modeling of the synchronous detection and the synchronous classification of the multi-modal cancer data is performed to obtain a target LSTM prediction model, specifically: An LSTM prediction blank model is introduced into the data terminal, and a full connection layer, an input layer and an output layer are determined in the LSTM prediction blank model; All cancer integrated data feature vectors are input into the input layer of the LSTM prediction blank model, and a general representation vector of all cancer integrated data feature vectors is extracted in the full connection layer of the LSTM prediction blank model and marked as a cancer integrated data general representation vector; After outputting the cancer integrated data general representation vector, the output layer of the LSTM prediction blank model is divided into a detection branch and a classification branch, wherein the detection branch is used for cancer existence probability calculation, and the classification branch is used for cancer subtype classification processing of the multi-modal cancer data; The LSTM prediction blank model is trained, wherein the training method of the LSTM prediction blank model is that, in the detection branch, combined with the sensitive biomarker, if the corresponding vector of the sensitive biomarker is output in the cancer integrated data general representation vector, the cancer existence probability is defined as greater than a preset value, and the cancer existence probability is calculated twice according to the number of sensitive biomarkers to obtain a preliminary trained LSTM prediction model; The number of sensitive biomarkers is positively correlated with the cancer existence probability, and the cancer existence probability corresponding to the number of different sensitive biomarkers is different. The classification branch of the preliminary trained LSTM prediction model is subjected to cancer subtype prediction training, and a target LSTM prediction model is generated based on the training result.
[0008] Further, in a preferred embodiment of the present application, the classification branch of the preliminary training LSTM prediction model is trained for cancer subtype prediction, and a target LSTM prediction model is generated based on the training results, specifically: Through the big data network, different cancer subtype data of multi-modal cancer data are retrieved, and co-occurrence history probabilities of different cancer subtype data are determined in the big data network. Based on the co-occurrence history probabilities of different cancer subtype data, co-occurrence weights of different cancer subtype data are determined; Based on the co-occurrence weights of different cancer subtype data, graph convolution training is performed in the classification branch in combination with sensitive biomarkers, wherein the graph convolution training is to determine all cancer subtype data present according to sensitive biomarkers, and to predict and determine all cancer subtype data present according to the co-occurrence weights of different cancer subtype data. After the graph convolution training, the trained preliminary training LSTM prediction model is output and labeled as the target LSTM prediction model.
[0009] Further, in a preferred embodiment of the present application, the target LSTM prediction model is inferred based on the synchronous detection classification data in combination with the Monte Carlo simulation analysis algorithm, and the target LSTM prediction model is iteratively optimized based on the comprehensive confidence index obtained by inference, specifically: The Monte Carlo simulation analysis algorithm is introduced, and synchronous detection classification data output by the target LSTM prediction model is obtained; Based on the Monte Carlo simulation analysis algorithm, Monte Carlo inference is performed on the synchronous detection classification data, and the probability distribution variance of the synchronous detection classification data is output after the Monte Carlo inference; The probability distribution variance of the synchronous detection classification data obtained each time is analyzed, and the difference value of the probability distribution variance of different synchronous detection classification data is calculated; According to the difference value of the probability distribution variance of different synchronous detection classification data, the comprehensive confidence index of the synchronous detection classification data is calculated, wherein different difference values correspond to different comprehensive confidence indexes; If the comprehensive confidence index of the synchronous detection classification data is greater than a preset value, the target LSTM prediction model is output as a qualified LSTM prediction model; If the comprehensive confidence index of the synchronous detection classification data is not greater than the preset value, a drift detection algorithm is deployed for the target LSTM model, the feature distribution of all cancer integrated data feature vectors input in the input layer of the target LSTM model is monitored in real time, and the distribution position of all cancer integrated data feature vectors in the input layer is determined. If the distribution position of the cancer integration data feature vector in the input layer is not within the preset range, the target LSTM model is automatically triggered for retraining until the distribution position of all cancer integration data feature vectors in the input layer is within the preset range. When the comprehensive confidence index of the synchronous detection classification data output by the target LSTM model is greater than the preset value, the retraining process is stopped, and a qualified LSTM prediction model is obtained.
[0010] The second aspect of the present application also provides a cancer data synchronous detection and classification system based on model analysis, which comprises a memory and a processor, wherein the memory stores a cancer data synchronous detection and classification method, and the cancer data synchronous detection and classification method is executed by the processor to realize the following steps: Integrate the multi-modal data of the cancer data, and perform data preprocessing on the multi-modal data of the cancer data to obtain preprocessed cancer integration data. Perform cross-modal feature fusion on all preprocessed cancer integration data, and screen to obtain sensitive biomarkers of the preprocessed cancer integration data. In combination with the sensitive biomarkers, perform joint optimization modeling of the synchronous detection and synchronous classification of the multi-modal cancer data to obtain a target LSTM prediction model. In combination with the Monte Carlo simulation analysis algorithm, infer the synchronous detection classification data of the target LSTM prediction model, and perform model iteration optimization on the target LSTM prediction model according to the comprehensive confidence index obtained by the inference.
[0011] The present application solves the technical defects in the background art, and has the following advantages: multi-modal cancer data is obtained and integrated and preprocessed, and cross-modal feature fusion is performed on the multi-modal cancer data to extract the features of the cancer data and generate sensitive biomarkers. Through the sensitive biomarkers, an LSTM prediction model is constructed for synchronous detection and classification of the cancer data, and finally the comprehensive confidence index of the model is analyzed, and the model is iteratively optimized according to the analysis result. The present application can obtain whether there is cancer and what type of cancer subtype through a single analysis by combining the modeling method, and the speed is also faster than that of the traditional method. BRIEF DESCRIPTION OF DRAWINGS
[0012] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiments or prior art description. Obviously, the drawings in the following description are only some embodiments of the present application, and those skilled in the art can obtain other drawings of embodiments according to these drawings without creative labor.
[0013] Figure 1 a flow chart of a cancer data synchronous detection and classification method based on model analysis is shown; Figure 2 a flow chart of a method for constructing a target LSTM prediction model is shown; Figure 3 a program view of a cancer data synchronous detection and classification system based on model analysis is shown. DETAILED DESCRIPTION
[0014] In order to enable a more clear understanding of the above-mentioned objects, features and advantages of the present application, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be noted that the embodiments of the present application and the features in the embodiments can be combined with each other without conflict.
[0015] In the following description, many specific details are set forth in order to provide a thorough understanding of the present application, however, the present application can also be implemented in other ways different from those described herein, therefore, the scope of protection of the present application is not limited by the specific embodiments disclosed below.
[0016] Figure 1 a flow chart of a cancer data synchronous detection and classification method based on model analysis is shown, including the following steps: S102: integrate the multi-modal data of the cancer data, and perform data preprocessing on the multi-modal data of the cancer data to obtain preprocessed cancer integrated data; S104: perform cross-modal feature fusion on all preprocessed cancer integrated data, and screen to obtain sensitive biomarkers of the preprocessed cancer integrated data; S106: combined with the sensitive biomarkers, perform joint optimization modeling of synchronous detection and synchronous classification on the multi-modal cancer data to obtain a target LSTM prediction model; S108: combine the Monte Carlo simulation analysis algorithm to reason the synchronous detection and classification data of the target LSTM prediction model, and perform model iterative optimization on the target LSTM prediction model according to the comprehensive confidence index obtained by reasoning.
[0017] Further, in a preferred embodiment of the present application, the multi-modal data of the cancer data is integrated, and the multi-modal data of the cancer data is preprocessed to obtain preprocessed cancer integrated data, specifically: obtain the multi-modal data of the cancer data, and label it as multi-modal cancer data; The multi-modal cancer data includes clinical index data and genomic data of different types of cancer data, and the multi-modal cancer data is biomedical data related to cancer; Introduce a data processing terminal, and introduce a dynamic time warping method, obtain time series of different types of multi-modal cancer data, and perform time alignment processing on different types of multi-modal cancer data according to the time series of different types of multi-modal cancer data; Meanwhile, a spatial registration algorithm is introduced to perform spatial alignment processing on the different types of multi-modal cancer data after time alignment processing, wherein the spatial alignment processing is to map the different types of multi-modal cancer data after time alignment processing in the same spatial coordinate system. The multi-modal cancer data after time alignment and spatial alignment processing is labeled as spatiotemporal alignment multi-modal cancer data, the spatiotemporal alignment multi-modal cancer data is subjected to cancer data similarity inference by a filling algorithm based on a graph neural network, and based on the inference result, the spatiotemporal alignment multi-modal cancer data is subjected to missing value filling, and all the spatiotemporal alignment multi-modal cancer data after missing value filling is subjected to data normalization processing to obtain preprocessed cancer integration data.
[0018] It should be noted that the cancer data of different modalities is integrated to construct condition data for a detection classification model, so the multi-modal cancer data is obtained. After obtaining the data, since the data may be obtained at different times and locations, it is necessary to perform time and space alignment on the obtained multi-modal cancer data to realize the unity of the multi-modal cancer data and facilitate the accuracy of subsequent data analysis and as condition data for modeling. At the same time, there may be missing data when obtaining the data, so it is necessary to perform data normalization and data interpolation filling processing on the data. Through the graph neural network filling algorithm, the similarity of the data can be inferred, the data with high similarity that needs to be filled is inferred, the purpose of missing value filling is achieved, and finally the preprocessed cancer integration data is obtained.
[0019] Further, in one preferred embodiment of the present application, the cross-modal feature fusion is performed on all the preprocessed cancer integration data, and sensitive biomarkers of the preprocessed cancer integration data are screened, specifically: In the data processing terminal, cross-modal feature embedding representation processing is performed on the preprocessed cancer integration data to obtain a feature vector of the preprocessed cancer integration data, which is labeled as a cancer integration data feature vector. Among them, the cross-modal feature embedding representation processing is to respectively extract a local gene mutation feature vector from genomic data in the preprocessed cancer integration data by a convolutional neural network, and extract a dynamic feature vector from clinical index data in the preprocessed cancer integration data by a gated recurrent unit. In the data terminal, the cross-modal attention weights of the different cancer integrated data feature vectors are calculated through the multi-modal cross attention mechanism, and the cross-modal causal chain of the preprocessed cancer integrated data is constructed according to the cross-modal attention weights of the different cancer integrated data feature vectors. Among them, the gene-protein interaction graph of the preprocessed cancer integrated data is constructed through the KEGG pathway, and the cross-modal attention weights of the different cancer integrated data feature vectors are mapped onto the gene-protein interaction graph, the nodes of the gene-protein interaction graph are updated, and the cross-modal causal chain of the preprocessed cancer integrated data is generated. Based on the big data network, all biomarkers existing in the multi-modal cancer data are retrieved, and the occurrence frequency and number of different biomarkers in the cross-modal causal chain are calculated according to the cross-modal causal chain of the preprocessed cancer integrated data. If the occurrence frequency and number of biomarkers in the cross-modal causal chain are greater than the preset value, the corresponding biomarker is marked as a sensitive biomarker.
[0020] It should be noted that feature extraction needs to be performed on the data, and the sensitive biomarkers of the data are determined in combination with the data features. The sensitive biomarkers are different in different cancers, and when the sensitive biomarkers appear, it proves that there is a high probability that the corresponding cancer subtype will appear. Cross-modal feature embedding is used for feature vector extraction of data, which maps heterogeneous data in a unified vector space, eliminates the semantic gap between modalities, and extracts local gene mutation feature vectors from genomic data in the preprocessed cancer integrated data through a convolutional neural network. After obtaining different feature vectors, the cross-modal attention weights of the different cancer integrated data feature vectors are calculated in combination with the multi-modal cross attention mechanism, and the proportion of the integrated feature data is calculated. Because the proportion of cancer data obtained in different ways is different. The cross-modal causal chain of the preprocessed cancer integrated data reflects the relevance between different data, and there is also a correlation between the data, that is, the priority of the data may be different, and the data may also have the possibility of generating other data. The cross-modal causal chain of the preprocessed cancer integrated data can provide conditions for screening sensitive biomarkers. The occurrence frequency and number of biomarkers corresponding to different cancer data exist on the cross-modal causal chain of the preprocessed cancer integrated data, and the largest biomarker is selected as the sensitive biomarker. The KEGG pathway constructs the gene-protein interaction graph of the preprocessed cancer integrated data, the nodes on the graph are biomolecules, that is, biomarkers, and the edges of the graph are regulatory relationships. Mapping the data onto the graph for node updating can obtain the cross-modal causal chain of the preprocessed cancer integrated data.
[0021] Further, in a preferred embodiment of the present application, the combination of the Monte Carlo simulation analysis algorithm infers the synchronous detection classification data of the target LSTM prediction model, and according to the comprehensive confidence index obtained by inference, the target LSTM prediction model is iteratively optimized, specifically: The Monte Carlo simulation analysis algorithm is introduced, and the synchronous detection classification data output by the target LSTM prediction model is obtained; Based on the Monte Carlo simulation analysis algorithm, the synchronous detection classification data is inferred by Monte Carlo, and the probability distribution variance of the synchronous detection classification data is output after Monte Carlo inference; The probability distribution variance of each obtained synchronous detection classification data is analyzed, and the difference value of the probability distribution variance of different synchronous detection classification data is calculated; According to the difference value of the probability distribution variance of different synchronous detection classification data, the comprehensive confidence index of the synchronous detection classification data is calculated, wherein different difference values correspond to different comprehensive confidence indexes; If the comprehensive confidence index of the synchronous detection classification data is greater than the preset value, the target LSTM prediction model is output as a qualified LSTM prediction model; If the comprehensive confidence index of the synchronous detection classification data is not greater than the preset value, the target LSTM model is deployed to detect the drift algorithm, the feature distribution of all cancer integrated data feature vectors input in the input layer of the target LSTM model is monitored in real time, and the distribution position of all cancer integrated data feature vectors in the input layer is judged. If the distribution position of the cancer integrated data feature vector in the input layer is not within the preset range, the target LSTM model is automatically triggered for retraining until the distribution position of all cancer integrated data feature vectors in the input layer is within the preset range; When the comprehensive confidence index of the synchronous detection classification data output by the target LSTM model is greater than the preset value, the retraining process is stopped, and a qualified LSTM prediction model is obtained.
[0022] It should be noted that the obtained synchronous detection classification data output by the target LSTM prediction model may not be accurate, and the obtained data needs to be analyzed to determine its accuracy. If it is not accurate, the model needs to be corrected. The Monte Carlo simulation analysis algorithm performs Monte Carlo reasoning on the synchronous detection classification data to calculate the confidence index. The variance of the forward propagation area of several data is calculated, and the difference is compared after multiple calculations. If the difference is large, it proves that the output data is unstable, and the model accuracy is low. At this time, the target LSTM model needs to deploy a drift detection algorithm to detect whether there is data input abnormality in the training process of the model, which leads to abnormal data training and inaccurate data. If the distribution position of the cancer integrated data feature vector in the input layer is not within the preset range, it proves that there is data shift in the data training process. At this time, retraining is needed until the comprehensive confidence index of the synchronous detection classification data output by the target LSTM model is greater than the preset value.
[0023] Figure 2 A method flow chart for constructing a target LSTM prediction model is shown, including the following steps: S202: Combined with sensitive biomarkers, the synchronous detection and synchronous classification of multi-modal cancer data are jointly optimized to obtain a target LSTM prediction model; S204: The classification branch of the preliminary training LSTM prediction model is trained for cancer subtype prediction, and the target LSTM prediction model is generated based on the training result.
[0024] Further, in a preferred embodiment of the present application, the target LSTM prediction model is obtained by jointly optimizing the synchronous detection and synchronous classification of multi-modal cancer data combined with sensitive biomarkers, specifically: An LSTM prediction blank model is introduced into the data terminal, and a full connection layer, an input layer and an output layer are determined in the LSTM prediction blank model; All cancer integrated data feature vectors are input into the input layer of the LSTM prediction blank model, and the universal representation vectors of all cancer integrated data feature vectors are extracted in the full connection layer of the LSTM prediction blank model and labeled as cancer integrated data universal representation vectors; After outputting the cancer integrated data universal representation vectors, the output layer of the LSTM prediction blank model is divided into a detection branch and a classification branch. The detection branch is used for cancer probability calculation, and the classification branch is used for cancer subtype classification processing of multi-modal cancer data; The LSTM prediction blank model is trained, and a training method of the LSTM prediction blank model is that, in the detection branch, if a corresponding vector of the sensitive biomarker is output in the cancer integration data general representation vector, the cancer presence probability is defined as greater than a preset value in combination with the sensitive biomarker, and the cancer presence probability is calculated twice according to the number of sensitive biomarkers to obtain a preliminary training LSTM prediction model; The number of sensitive biomarkers is positively correlated with the cancer presence probability, and the cancer presence probability corresponding to different numbers of sensitive biomarkers is different. The classification branch of the preliminary training LSTM prediction model is subjected to cancer subtype prediction training, and a target LSTM prediction model is generated based on a training result.
[0025] It should be noted that the LSTM prediction blank model is used for predicting cancer data, detecting cancer data and classifying cancer data, and is a neural network model, which is divided into a full connection layer, an input layer and an output layer. The general representation is extracted in the full connection layer, and the purpose is to ensure that the cancer essence features learned by the detection and classification tasks are consistent, and task conflicts are avoided. A double-branch shared feature extraction form is adopted, which is divided into a detection branch and a classification branch, and is respectively used for cancer presence probability calculation and cancer subtype classification processing. In the detection branch, the cancer presence probability is calculated according to the number of sensitive biomarkers, different numbers correspond to different presence probabilities, and the presence probability is positively correlated.
[0026] Further, in a preferred embodiment of the present application, the classification branch of the preliminary training LSTM prediction model is subjected to cancer subtype prediction training, and a target LSTM prediction model is generated based on a training result, specifically: Different cancer subtype data of multi-modal cancer data is searched through a big data network, and a co-occurrence historical probability of different cancer subtype data is determined in the big data network, and a co-occurrence weight of different cancer subtype data is determined based on the co-occurrence historical probability of different cancer subtype data. Based on the co-occurrence weight of different cancer subtype data, graph convolution training is performed in the classification branch in combination with the sensitive biomarker, wherein the graph convolution training is to determine all cancer subtype data existing according to the sensitive biomarker, and to predict and determine skirt cancer subtype data of all cancer subtype data existing according to the co-occurrence weight of different cancer subtype data. After the graph convolution training, the trained preliminary training LSTM prediction model is output and calibrated as a target LSTM prediction model.
[0027] It needs to be explained that even if there is cancer, the subtypes of cancer are different, and one cancer can have multiple subtypes, and there is relevance between subtypes. For example, different subtypes may be necessarily associated, such as the existence of a subtype a, and the existence of a subtype b; or they may be in conflict, such as the existence of a subtype a, and the non-existence of a subtype b. In a big data network, the co-occurrence history probability of different cancer subtype data is determined, the co-occurrence weight of different cancer subtype data is determined, and the purpose is to realize graph convolution training. The graph convolution training can determine all cancer subtype data that exists according to sensitive biomarkers, and finally determine the output cancer subtype data according to the co-occurrence relationship.
[0028] As shown in Figure 3 The second aspect of the present application also provides a cancer data synchronous detection and classification system based on model analysis, which comprises a memory 31 and a processor 32, the memory 31 stores a cancer data synchronous detection and classification method, and the cancer data synchronous detection and classification method is executed by the processor 32 to realize the following steps: Integrate the multi-modal data of the cancer data, and perform data preprocessing on the multi-modal data of the cancer data to obtain preprocessed cancer integrated data; Cross-modal feature fusion is performed on all preprocessed cancer integrated data, and sensitive biomarkers of the preprocessed cancer integrated data are screened out; Combined with the sensitive biomarkers, the multi-modal cancer data is subjected to joint optimization modeling of synchronous detection and synchronous classification to obtain a target LSTM prediction model; The synchronous detection and classification data of the target LSTM prediction model are inferred by combining the Monte Carlo simulation analysis algorithm, and the target LSTM prediction model is iteratively optimized according to the comprehensive confidence index obtained by inference.
[0029] The above is only a specific embodiment of the present application, but the protection scope of the present application is not limited thereto, any skilled person in the art can easily think of changes or replacements within the technical scope disclosed by the present application, which should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A method for cancer data synchronous detection and classification based on model analysis, characterized in that, The method comprises the following steps: Integrating multi-modal data of cancer data, and data preprocessing of the multi-modal data of cancer data to obtain preprocessed cancer integration data; Cross-modal feature fusion is performed on all preprocessed cancer integration data, and sensitive biomarkers of the preprocessed cancer integration data are screened; Combined with the sensitive biomarkers, a joint optimization modeling of synchronous detection and synchronous classification of the multi-modal cancer data is performed to obtain a target LSTM prediction model; The synchronous detection and classification data of the target LSTM prediction model are reasoned by combining the Monte Carlo simulation analysis algorithm, and the target LSTM prediction model is iteratively optimized according to the comprehensive confidence index obtained by reasoning.
2. The model-based analysis of cancer data synchronization detection and classification method according to claim 1, characterized in that, The method comprises the following steps: Obtain multi-modal cancer data, and calibrate it as multi-modal cancer data; The multi-modal cancer data includes clinical indicator data and genomic data of different types of cancer data, and the multi-modal cancer data is biomedical data related to cancer; Introduce a data processing terminal and a dynamic time warping method, obtain time series of different types of multi-modal cancer data, and perform time alignment processing on different types of multi-modal cancer data according to the time series of different types of multi-modal cancer data; At the same time, introduce a spatial registration algorithm to perform spatial alignment processing on different types of multi-modal cancer data after time alignment processing, wherein the spatial alignment processing is to map different types of multi-modal cancer data after time alignment processing in the same spatial coordinate system; After time alignment and spatial alignment processing, the multi-modal cancer data is calibrated as spatiotemporal alignment multi-modal cancer data, the spatiotemporal alignment multi-modal cancer data is inferred for cancer data similarity based on a filling algorithm based on a graph neural network, and the spatiotemporal alignment multi-modal cancer data is filled with missing values based on the inference result, and all spatiotemporal alignment multi-modal cancer data after missing value filling is subjected to data normalization processing to obtain preprocessed cancer integration data.
3. The model-based analysis of cancer data synchronization detection and classification method according to claim 1, characterized in that, The method comprises the following steps: In the data processing terminal, cross-modal feature embedding representation processing is performed on the preprocessed cancer integration data to obtain a feature vector of the preprocessed cancer integration data, which is calibrated as a cancer integration data feature vector; The cross-modal feature embedding representation processing is to extract a local gene mutation feature vector from the genomic data in the preprocessed cancer integration data by a convolutional neural network, and to extract a dynamic feature vector from the clinical indicator data in the preprocessed cancer integration data by a gated recurrent unit; In the data terminal, the cross-modal attention weights of different cancer integration data feature vectors are calculated by a multi-modal cross-attention mechanism, and a cross-modal causal chain of the preprocessed cancer integration data is constructed according to the cross-modal attention weights of different cancer integration data feature vectors; The gene-protein interaction graph of the preprocessed cancer integrated data is constructed by KEGG pathway, and the cross-modal attention weight of the feature vector of the different cancer integrated data is mapped to the gene-protein interaction graph, the node of the gene-protein interaction graph is updated, and the cross-modal causal chain of the preprocessed cancer integrated data is generated. Based on the big data network, all biomarkers existing in the multi-modal cancer data are retrieved, and the frequency and number of different biomarkers in the cross-modal causal chain are calculated according to the cross-modal causal chain of the preprocessed cancer integrated data. If the frequency and number of biomarkers in the cross-modal causal chain are greater than the preset value, the corresponding biomarker is marked as a sensitive biomarker.
4. The model-based analysis of cancer data synchronization detection and classification method according to claim 1, wherein, The joint optimization modeling of the multi-modal cancer data is performed by combining the sensitive biomarker, and the target LSTM prediction model is obtained, specifically: An LSTM prediction blank model is introduced into the data terminal, and a full connection layer, an input layer and an output layer are determined in the LSTM prediction blank model; All cancer integrated data feature vectors are input into the input layer of the LSTM prediction blank model, and the general representation vector of all cancer integrated data feature vectors is extracted in the full connection layer of the LSTM prediction blank model, which is marked as a cancer integrated data general representation vector; After outputting the cancer integrated data general representation vector, the output layer of the LSTM prediction blank model is divided into a detection branch and a classification branch, wherein the detection branch is used for cancer existence probability calculation, and the classification branch is used for cancer subtype classification processing of the multi-modal cancer data; The LSTM prediction blank model is trained, wherein the training method of the LSTM prediction blank model is that, in the detection branch, combined with the sensitive biomarker, if the corresponding vector of the sensitive biomarker is output in the cancer integrated data general representation vector, the cancer existence probability is defined as greater than a preset value, and the cancer existence probability is calculated twice according to the number of sensitive biomarkers to obtain a preliminary trained LSTM prediction model. The number of sensitive biomarkers is positively correlated with the cancer existence probability, and the cancer existence probability corresponding to the number of different sensitive biomarkers is different. The classification branch of the preliminary trained LSTM prediction model is trained for cancer subtype prediction, and the target LSTM prediction model is generated based on the training result.
5. The model-based analysis of cancer data synchronization detection and classification method according to claim 4, characterized in that, The classification branch of the preliminary trained LSTM prediction model is trained for cancer subtype prediction, and the target LSTM prediction model is generated based on the training result. Different cancer subtype data of the multi-modal cancer data are retrieved through the big data network, and the co-occurrence historical probability of different cancer subtype data is determined in the big data network, and the co-occurrence weight of different cancer subtype data is determined based on the co-occurrence historical probability of different cancer subtype data. Based on the co-occurrence weight of the different cancer subtype data, combined with sensitive biomarkers, graph convolution training is performed within the classification branch, wherein the graph convolution training is to determine all cancer subtype data present according to sensitive biomarkers, and to predict the skirt cancer subtype data of all cancer subtype data present according to the co-occurrence weight of the different cancer subtype data; After the graph convolution training, the trained preliminary training LSTM prediction model is output, which is calibrated as a target LSTM prediction model.
6. The model-based analysis of cancer data synchronization detection and classification method according to claim 1, wherein, The combination of the Monte Carlo simulation analysis algorithm reasons the synchronous detection classification data of the target LSTM prediction model, and the target LSTM prediction model is iteratively optimized according to the comprehensive confidence index obtained by reasoning, specifically: The Monte Carlo simulation analysis algorithm is introduced, and the synchronous detection classification data output by the target LSTM prediction model is obtained; Based on the Monte Carlo simulation analysis algorithm, the synchronous detection classification data is subjected to Monte Carlo reasoning, and the probability distribution variance of the synchronous detection classification data is output after the Monte Carlo reasoning; The probability distribution variance of the synchronous detection classification data obtained each time is analyzed, and the difference value of the probability distribution variance of different synchronous detection classification data is calculated; According to the difference value of the probability distribution variance of different synchronous detection classification data, the comprehensive confidence index of the synchronous detection classification data is calculated, wherein different difference values correspond to different comprehensive confidence indexes; If the comprehensive confidence index of the synchronous detection classification data is greater than the preset value, the target LSTM prediction model is output as a qualified LSTM prediction model; If the comprehensive confidence index of the synchronous detection classification data is not greater than the preset value, the target LSTM model is deployed with a drift detection algorithm, which monitors the feature distribution of all cancer integrated data feature vectors in the input layer of the target LSTM model in real time, and judges the distribution position of all cancer integrated data feature vectors in the input layer; If the distribution position of the cancer integrated data feature vector in the input layer is not within the preset range, the target LSTM model is automatically triggered for retraining until the distribution position of all cancer integrated data feature vectors in the input layer is within the preset range; When the comprehensive confidence index of the synchronous detection classification data output by the target LSTM model is greater than the preset value, the retraining process is stopped, and a qualified LSTM prediction model is obtained.
7. A model-based analysis of cancer data synchronous detection and classification system characterized in that, The cancer data synchronous detection and classification system comprises a memory and a processor, and the memory stores a cancer data synchronous detection and classification method program, which realizes the cancer data synchronous detection and classification method steps of any one of claims 1-6 when executed by the processor.
Citation Information
Patent Citations
A method and system for determining cancer network markers based on probability model
CN109101783A
Cancer risk prediction method based on deep learning
CN118507048A
Multi-modal optimization method for detecting dynamic network biomarkers of individual cancer patients
CN118609658A
Cancer subtype classification method based on KAN network and multi-omics data
CN119252347A