A malignant tumor occurrence probability map acquisition and analysis system
By collecting data from multiple sources and constructing artificial intelligence models, a probability map of malignant tumors is generated and analyzed, which solves the problem of low early diagnosis rate of malignant tumors and realizes comprehensive prevention and treatment of malignant tumors.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SECOND MEDICAL CENT OF CHINESE PLA GENERAL HOSPITAL
- Filing Date
- 2025-11-17
- Publication Date
- 2026-07-24
Smart Images

Figure CN121506531B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of malignant tumor technology, specifically to a system for acquiring and analyzing the probability map of malignant tumor occurrence. Background Technology
[0002] Cancer, medically known as malignant tumors, is a complex and serious disease that seriously threatens human health. It is the second leading cause of death worldwide, after cardiovascular and cerebrovascular diseases. Malignant tumors are characterized by insidious onset, long onset time, rapid progression, and limited screening and treatment methods. By the time patients notice the appearance of corresponding symptoms, the disease has often progressed to the middle or late stages, making clinical intervention and treatment difficult.
[0003] The incidence of malignant tumors is affected by multiple factors such as region, age, gender and lifestyle. The incidence rate increases significantly with age, with people over 60 years old accounting for more than 60%. Lung cancer, breast cancer and colorectal cancer are high-incidence types. Smoking, obesity and environmental pollution are the main risk factors. Early screening can reduce mortality.
[0004] Existing technologies cannot effectively obtain and analyze the probability profiles of malignant tumors, thus reducing the early diagnosis rate of malignant tumors and failing to promote the comprehensive development of malignant tumor prevention, control, and treatment. Summary of the Invention
[0005] The purpose of this invention is to provide a system for acquiring and analyzing the probability map of malignant tumors, which can effectively acquire and analyze the probability map of malignant tumors, improve the early diagnosis rate of malignant tumors, and promote the comprehensive development of malignant tumor prevention, control and treatment, thus solving the problems mentioned in the background art.
[0006] To achieve the above objectives, the present invention provides the following technical solution: A system for acquiring and analyzing the probability map of malignant tumor occurrence, comprising: The data acquisition and integration module is used to collect multi-source patient data from multiple sources and dimensions. The data processing map acquisition module is used to process the collected multi-source patient data, determine the patient characteristic data, and build a malignant tumor occurrence probability prediction model based on artificial intelligence. Based on the malignant tumor occurrence probability prediction model, the module performs pattern recognition on the patient characteristic data, determines the malignant tumor occurrence probability prediction result, and generates a malignant tumor occurrence probability map. The graph analysis and visualization module is used to analyze the probability graph of malignant tumor occurrence and to visualize the analysis report of the probability graph of malignant tumor occurrence.
[0007] Preferably, multi-source patient data is collected from multiple sources and dimensions, and the following operations are performed: Monitor the patient's genome, transcriptome, epigenome, and proteome to obtain patient omics data; Monitor patients' electronic health records, pathology reports, and imaging data to obtain patients' clinical data; Monitor patients' smoking, diet, exercise, and environmental data to obtain patients' lifestyle data; Based on patient omics data, patient clinical data, and patient lifestyle data, multi-source patient data is generated.
[0008] Preferably, the data processing map acquisition module includes: The data processing unit is used to process multi-source patient data and determine patient characteristic data; This includes cleaning the multi-source patient data, removing noise, identifying missing and outlier values, and processing the identified missing and outlier values. Normalize the multi-source patient data to convert it into a unified data format, remove the dimensional differences between the multi-source patient data, and form standardized multi-source patient data. Feature extraction is performed on multi-source patient data to extract features related to the acquisition and analysis of malignant tumor incidence probability maps, thereby determining patient characteristic data.
[0009] Preferably, missing values and outliers in the patient's multi-source data are identified, and the identified missing values and outliers are processed by performing the following operations: Based on the missing matrix diagram, a heatmap is used to visualize the missing values of all values in the multi-source patient data, identify all missing data points in the multi-source patient data, and identify the missing values in the multi-source patient data. Based on the standard deviation method, values exceeding the mean ± 3 standard deviations in multi-source patient data are identified, data points that are significantly different from the behavior patterns of most data are found, and outliers in multi-source patient data are identified. The missing and outlier values in the multi-source patient data were evaluated to determine whether the missing and outlier values in the multi-source patient data were valuable for obtaining and analyzing the probability map of malignant tumor occurrence. When missing values and outliers in the multi-source patient data are valuable for obtaining and analyzing the probability map of malignant tumor occurrence, the missing values in the multi-source patient data are filled and the outliers in the multi-source patient data are replaced. When missing and outlier values in patient multi-source data are of no value for obtaining and analyzing the probability map of malignant tumor occurrence, then the missing and outlier values in the patient multi-source data are removed.
[0010] Preferably, the data processing map acquisition module further includes: The model building unit is used to build a model for predicting the probability of malignant tumor occurrence based on artificial intelligence. Collect patient historical data and divide the collected patient historical data into training set and test set in a 7:3 ratio. The machine learning model is trained using a training set, enabling it to learn autonomously from the training set how to predict the probability of malignant tumors and predict the probability of a patient developing a malignant tumor within a specific timeframe in the future, thus establishing a malignant tumor probability prediction model. The malignant tumor incidence probability prediction model was tested using a test set to evaluate its generalization performance and determine whether it could achieve the expected effect of predicting the probability of a patient developing a malignant tumor within a specific time period in the future. When the malignant tumor incidence probability prediction model fails to achieve the expected effect of predicting the probability of a patient developing a malignant tumor within a specific future time, the parameters of the malignant tumor incidence probability prediction model are adjusted and iteratively optimized until the malignant tumor incidence probability prediction model can achieve the expected effect of predicting the probability of a patient developing a malignant tumor within a specific future time, and the optimal malignant tumor incidence probability prediction model is determined.
[0011] Preferably, the data processing map acquisition module further includes: The probability prediction unit is used to predict the probability of a patient developing a malignant tumor within a specific timeframe in the future. The process involves inputting patient characteristic data into a malignant tumor incidence probability prediction model, analyzing and recognizing patterns in the patient characteristic data based on the model, and predicting the probability of a patient developing a malignant tumor within a specific timeframe in the future, thus determining the prediction result of the malignant tumor incidence probability.
[0012] Preferably, the data processing map acquisition module further includes: The map generation unit is used to generate a map of the probability of malignant tumor occurrence. Based on the prediction results of the incidence probability of malignant tumors and combined with patient characteristic data, a probability map of the incidence of malignant tumors is generated. Specifically, the incidence probability of malignant tumors output by the prediction model is associated with the corresponding biomedical characteristics to construct a probability map of the incidence of malignant tumors that can be queried in a visual form, and to reveal the complex network relationship between patient characteristics and diseases.
[0013] Preferably, the spectral analysis and visualization module includes: The graph analysis unit is used to analyze the probability graph of malignant tumor occurrence, enabling doctors to understand and utilize the probability graph of malignant tumor occurrence. Among them, the system queries based on the probability map of malignant tumor occurrence, inputs patient characteristic data and returns their probability of malignant tumor occurrence, performs population analysis and driving factor analysis based on the probability map of malignant tumor occurrence, compares the probability distribution of malignant tumor occurrence in different subgroups, identifies the features that contribute the most to the probability of malignant tumor occurrence, and develops prevention plans for patients, generating a probability map analysis report of malignant tumor occurrence. The visualization unit is used to visually display the analysis report on the probability of malignant tumors in the form of charts, and to automatically issue early warnings for patients with a high probability of malignant tumors, prompting clinical intervention.
[0014] Preferably, a machine learning model is trained using a training set to determine a predictive model for the probability of malignant tumor occurrence, including: The training data in the training set is input into the multimodal feature extraction layer to extract features from the training data, resulting in multiple sets of modal feature vectors. The training data in the training set includes patients' genomic data, transcriptomic data, epigenomic data, proteomic data, electronic health records, pathology reports, imaging data, and patients' smoking, diet, exercise, and environmental data. The multiple sets of modal feature vectors include genomic feature vectors, molecular expression feature vectors, epigenetic feature vectors, clinical feature vectors, imaging feature vectors, and environmental feature vectors. Intramodal correlation enhancement processing is performed on each of the multiple sets of modal feature vectors to obtain intramodal self-interaction enhanced feature vectors; Causal interaction processing is performed on the combination of environment-clinical and clinical-imaging dual modalities, causal attention weights are calculated and causal interaction feature vectors are generated; association interaction processing is performed on the combination of genome-transcriptome, transcriptome-epigome, and environment-genome dual modalities, association weights are calculated in conjunction with regulatory network priors and association interaction feature vectors are generated. The global weights of each modality are calculated based on the modal quality scoring matrix. The self-interaction enhancement feature vectors, causal interaction feature vectors, and correlation interaction feature vectors within the modality are input into the Transformer encoder for global fusion. After residual connection and layer normalization processing, the cross-modal fusion feature vector is output. The cross-modal fusion feature vector is input into the spatiotemporal probability prediction layer, and combined with the survival analysis model and probability calibration mechanism, the initial prediction results of the probability of malignant tumor occurrence in multiple time windows are output. The model is trained in stages based on a multi-task loss function. When the training results meet the requirements, a prediction model for the probability of malignant tumors is obtained.
[0015] Preferably, a malignant tumor incidence probability map is generated based on the predicted malignant tumor incidence probability and combined with patient characteristic data, including: Constructing a knowledge ontology framework for malignant tumor medicine; Acquire multi-source medical data, and accurately map the multi-source medical data to the medical knowledge ontology framework to obtain the data-ontology mapping result; Relationship prediction is performed on ontology nodes in the medical knowledge ontology framework based on the knowledge graph embedding model. The confidence of the predicted initial relationships is optimized to obtain the target relationships between ontology nodes and their corresponding confidence. Construct a mapping matrix between the probability ontology and multiple ontology nodes, and dynamically fuse the probability and ontology nodes based on the data-ontology mapping results, target relationships and corresponding confidence levels; The medical knowledge ontology framework and mapping matrix are incrementally self-optimized, and a probability map of malignant tumor occurrence is constructed.
[0016] Compared with the prior art, the beneficial effects of the present invention are: This invention collects and processes multi-source patient data to determine patient characteristic data. Based on artificial intelligence, it constructs a malignant tumor incidence probability prediction model. This model analyzes and performs pattern recognition on patient characteristic data, predicting the probability of a patient developing a malignant tumor within a specific future timeframe. The predicted malignant tumor incidence probability is then used to generate a malignant tumor incidence probability map. Analysis of this map identifies the features that contribute most to the probability of malignant tumor incidence, and prevention plans are developed for patients. A malignant tumor incidence probability map analysis report is generated and visualized. Furthermore, it automatically issues warnings and prompts clinical intervention for patients with a high probability of malignant tumor incidence. This invention effectively obtains and analyzes malignant tumor incidence probability maps, improving the early diagnosis rate of malignant tumors. Attached Figure Description
[0017] Figure 1 This is a block diagram of the system for acquiring and analyzing the probability map of malignant tumors according to the present invention. Detailed Implementation
[0018] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0019] To address the current limitations in effectively obtaining and analyzing probabilistic maps of malignant tumor occurrence, which reduces the early diagnosis rate of malignant tumors and hinders the comprehensive development of malignant tumor prevention and treatment, please refer to [link to relevant documentation]. Figure 1 This embodiment provides the following technical solution: A system for acquiring and analyzing probability maps of malignant tumors includes: a data acquisition and integration module, a data processing map acquisition module, and a map analysis and visualization module.
[0020] Specifically, through the interactive communication between the data acquisition and integration module, the data processing map acquisition module, and the map analysis and visualization module, the probability map of malignant tumor occurrence can be effectively acquired and analyzed, thereby improving the early diagnosis rate of malignant tumors and promoting the comprehensive development of malignant tumor prevention, control, and treatment.
[0021] Among them, the data acquisition and integration module is used to collect multi-source patient data from multiple sources and dimensions.
[0022] In this embodiment, multi-source patient data is collected from multiple sources and dimensions, and the following operations are performed: Monitor the patient's genome, transcriptome, epigenome, and proteome to obtain patient omics data; Monitor patients' electronic health records, pathology reports, and imaging data to obtain patients' clinical data; Monitor patients' smoking, diet, exercise, and environmental data to obtain patients' lifestyle data; Based on patient omics data, patient clinical data, and patient lifestyle data, multi-source patient data is generated, which facilitates better prediction of the probability of patients developing malignant tumors at a specific time in the future.
[0023] The data processing map acquisition module is used to process the collected multi-source patient data, determine the patient characteristic data, and construct a malignant tumor occurrence probability prediction model based on artificial intelligence. Based on the malignant tumor occurrence probability prediction model, the module performs pattern recognition on the patient characteristic data, determines the malignant tumor occurrence probability prediction result, and generates a malignant tumor occurrence probability map.
[0024] In this embodiment, the data processing map acquisition module includes: a data processing unit, a model building unit, a probability prediction unit, and a map generation unit.
[0025] The data processing unit is used to process multi-source patient data, including: The patient multi-source data is cleaned to remove noise, and missing and outlier values are identified and processed. Normalize the multi-source patient data to convert it into a unified data format, remove the dimensional differences between the multi-source patient data, and form standardized multi-source patient data. Feature extraction is performed on multi-source patient data to extract features related to the acquisition and analysis of malignant tumor incidence probability maps, thereby determining patient characteristic data.
[0026] In this embodiment, missing values and outliers in the patient's multi-source data are identified, and the identified missing values and outliers are processed by performing the following operations: Based on the missing matrix diagram, a heatmap is used to visualize the missing values of all values in the multi-source patient data, identify all missing data points in the multi-source patient data, and identify the missing values in the multi-source patient data. Based on the standard deviation method, values exceeding the mean ± 3 standard deviations in multi-source patient data are identified, data points that are significantly different from the behavior patterns of most data are found, and outliers in multi-source patient data are identified. The missing and outlier values in the multi-source patient data were evaluated to determine whether the missing and outlier values in the multi-source patient data were valuable for obtaining and analyzing the probability map of malignant tumor occurrence. When missing values and outliers in the multi-source patient data are valuable for obtaining and analyzing the probability map of malignant tumor occurrence, the missing values in the multi-source patient data are filled and the outliers in the multi-source patient data are replaced. When missing and outlier values in patient multi-source data are of no value for obtaining and analyzing the probability map of malignant tumor occurrence, then the missing and outlier values in the patient multi-source data are removed.
[0027] The model building unit, used to construct a predictive model for the probability of malignant tumor occurrence based on artificial intelligence, includes: Collect patient historical data and divide the collected patient historical data into training set and test set in a 7:3 ratio. The machine learning model is trained using a training set, enabling it to learn autonomously from the training set how to predict the probability of malignant tumors and predict the probability of a patient developing a malignant tumor within a specific timeframe in the future, thus establishing a malignant tumor probability prediction model. The malignant tumor incidence probability prediction model was tested using a test set to evaluate its generalization performance and determine whether it could achieve the expected effect of predicting the probability of a patient developing a malignant tumor within a specific time period in the future. When the malignant tumor incidence probability prediction model fails to achieve the expected effect of predicting the probability of a patient developing a malignant tumor within a specific future time, the parameters of the malignant tumor incidence probability prediction model are adjusted and iteratively optimized until the malignant tumor incidence probability prediction model can achieve the expected effect of predicting the probability of a patient developing a malignant tumor within a specific future time, and the optimal malignant tumor incidence probability prediction model is determined.
[0028] The probability prediction unit is used to predict the probability of a patient developing a malignant tumor within a specific future timeframe, including: Patient characteristic data is input into a malignant tumor incidence probability prediction model. The model analyzes and performs pattern recognition on the patient characteristic data, and predicts the probability of the patient developing a malignant tumor within a specific time period in the future, thus determining the prediction result of the malignant tumor incidence probability.
[0029] Among them, the map generation unit is used to generate a map of the probability of malignant tumor occurrence; Based on the prediction results of the incidence probability of malignant tumors and combined with patient characteristic data, a probability map of the incidence of malignant tumors is generated. Specifically, the incidence probability of malignant tumors output by the prediction model is associated with the corresponding biomedical characteristics to construct a probability map of the incidence of malignant tumors that can be queried in a visual form, and to reveal the complex network relationship between patient characteristics and diseases.
[0030] The graph analysis and visualization module is used to analyze the probability graph of malignant tumors and visualize the analysis report of the probability graph of malignant tumors.
[0031] In this embodiment, the spectral analysis and visualization module includes: a spectral analysis unit and a visualization display unit.
[0032] Among them, the map analysis unit is used to analyze the probability map of malignant tumor occurrence, so that doctors can understand and utilize the probability map of malignant tumor occurrence. Among them, the system queries based on the probability map of malignant tumor occurrence, inputs patient characteristic data and returns their probability of malignant tumor occurrence, performs population analysis and driving factor analysis based on the probability map of malignant tumor occurrence, compares the probability distribution of malignant tumor occurrence in different subgroups, identifies the features that contribute the most to the probability of malignant tumor occurrence, and develops prevention plans for patients, generating a probability map analysis report of malignant tumor occurrence. The visualization unit is used to visualize the analysis report on the probability of malignant tumors in the form of charts and graphs, and to automatically issue early warnings for patients with a high probability of malignant tumors, prompting clinical intervention.
[0033] The machine learning model is trained using a training set to determine a predictive model for the probability of malignant tumor occurrence, including: The training data in the training set is input into the multimodal feature extraction layer to extract features from the training data, resulting in multiple sets of modal feature vectors. The training data in the training set includes patients' genomic data, transcriptomic data, epigenomic data, proteomic data, electronic health records, pathology reports, imaging data, and patients' smoking, diet, exercise, and environmental data. The multiple sets of modal feature vectors include genomic feature vectors, molecular expression feature vectors, epigenetic feature vectors, clinical feature vectors, imaging feature vectors, and environmental feature vectors. Intramodal correlation enhancement processing is performed on each of the multiple sets of modal feature vectors to obtain intramodal self-interaction enhanced feature vectors; Causal interaction processing is performed on the combination of environment-clinical and clinical-imaging dual modalities, causal attention weights are calculated and causal interaction feature vectors are generated; association interaction processing is performed on the combination of genome-transcriptome, transcriptome-epigome, and environment-genome dual modalities, association weights are calculated in conjunction with regulatory network priors and association interaction feature vectors are generated. The global weights of each modality are calculated based on the modal quality scoring matrix. The self-interaction enhancement feature vectors, causal interaction feature vectors, and correlation interaction feature vectors within the modality are input into the Transformer encoder for global fusion. After residual connection and layer normalization processing, the cross-modal fusion feature vector is output. The cross-modal fusion feature vector is input into the spatiotemporal probability prediction layer, and combined with the survival analysis model and probability calibration mechanism, the initial prediction results of the probability of malignant tumor occurrence in multiple time windows are output. The model is trained in stages based on a multi-task loss function. When the training results meet the requirements, a prediction model for the probability of malignant tumors is obtained.
[0034] In this embodiment, the methods for obtaining genomic feature vectors, molecular expression feature vectors, epigenetic feature vectors, clinical feature vectors, imaging feature vectors, and environmental feature vectors include: The genomic data in the training set is processed based on the GAT-based genomic variation analysis sub-model, and genomic feature vectors are output. The Transformer-based molecular expression analysis sub-model processes the transcriptome and proteome data (including three time points) in the training set and outputs molecular expression feature vectors. The epigenomic data (including 3 time points) in the training set are processed based on the Attention-LSTM epigenomic parsing sub-model, and the epigenomic feature vector is output. The Time-SeriesGNN clinical data parsing sub-model is used to process electronic health records in the training set of clinical-imaging-environment data and output clinical feature vectors. The 3DResNet-UNet imaging parsing sub-model is used to process the imaging data in the clinical-imaging-environment data of the training set and output the imaging feature vector. The CausalEmbedding environment-lifestyle parsing sub-model processes the environment-lifestyle data in the training set of clinical-imaging-environment data and outputs environment feature vectors.
[0035] In this embodiment, the genomic variation analysis sub-model (GAT-based) has the following inputs: a chromosome variation graph (where the chromosome variation graph is based on chromosomes, nodes are variation sites (features: PathScore, allele frequency, location), and edges are LD coefficients between variations; each chromosome is a subgraph, for a total of 24 subgraphs); layer structure: node embedding layer: maps the variation's PathScore, allele frequency, and other features into a 64-dimensional vector, embedding GRCh38 genomic location information (e.g., "Chromosome 1:1000000-1000100" → location encoding); graph attention layer (GAT): sets 3 attention heads, and calculates attention weights: ;in, For mutation Features For mutation The LD-related neighbors (r²≥0.8) are defined, W is the weight matrix (64×64), and a is the attention vector (128-dimensional). Multi-scale aggregation layer: Local aggregation: Subgraph features are aggregated in 100kb windows to generate "local variation feature vectors"; Chromosome aggregation: Max-pooling is performed on the local features of each chromosome to generate 24 chromosome feature vectors (each 64-dimensional); Global fusion layer: The 24 chromosome features are fused using a Transformer encoder (2 layers, 8 attention heads) to output a 512-dimensional genome feature vector. .
[0036] In this embodiment, the transcriptome-proteome analysis sub-model (Transformer-based) has the following inputs: transcriptome pathway activity matrix (rows: 50 core pathways, columns: 3 time points), proteome expression matrix (rows: 100 tumor-associated proteins, columns: 3 time points); layer structure: molecular function embedding layer: maps pathways / proteins to GO functional annotations (e.g., "cell proliferation" "apoptosis"), generating functional embedding vectors (64-dimensional); self-attention layer: sets up 8-head attention to capture synergistic relationships between pathways (e.g., "CellCycle" and "p53"). The signaling layer incorporates attention weights and KEGG pathway interaction network priors; the temporal fusion layer uses dual LSTMs (2 layers, 256 hidden layers) to process transcriptome and proteome data at three time points respectively; the cross-omics alignment layer calculates the correlation between temporally aligned protein expression and corresponding gene transcripts (e.g., p53 protein and TP53 gene), and uses an attention mechanism to focus on highly correlated (r≥0.6) transcription-protein pairs; the global output layer splices pathway features, temporal features, and protein features, and outputs a molecular expression feature vector through a fully connected layer (512 dimensions). .
[0037] In this embodiment, the epigenome analysis sub-model (Attention-LSTM) inputs: DMR-gene regulation pair tensor (rows: 200 DMRs, columns: 200 target genes, depth: 3 time points, value: regulatory strength); layer structure: DMR embedding layer: mapping the β value, length, and CpG density of DMRs to a 64-dimensional vector, embedding gene promoter / enhancer position information; regulatory attention layer: calculating the regulatory weights of DMRs on target genes, combined with ChIP-seq data from epigenetic databases (such as ENCODE) (overlapping transcription factor binding sites are given high weights); LSTM temporal layer: processing methylation data at 3 time points to capture the dynamic changes of DMRs (e.g., "promoter methylation increases year by year → gene expression decreases"); output layer: max-pooling regulatory features, outputting a 512-dimensional epigenetic feature vector through a fully connected layer. .
[0038] In this embodiment, the clinical data parsing sub-model (Time-Series GNN) has the following inputs: 5-year EHR time-series data (each time point contains 30 clinical indicators), and disease history feature vectors; Layer structure: Time-series node construction: each time point is a node (features: clinical indicators + disease history), and the edges between nodes are "time differences" (e.g., 1 year); Time-series GNN layer: GCN is used to capture the cross-time correlation of clinical indicators (e.g., "elevated blood glucose → increased glycated hemoglobin the following year"), and the edge weights are the reciprocals of the time difference (the closer the time, the higher the weight); Trend extraction layer: CNN (3×3 convolutional kernels) is used to extract the short-term trend of clinical indicators (e.g., blood pressure fluctuations within 6 months), and a fully connected layer is used to extract the long-term trend (5-year changes); Output layer: The time-series trend features and disease history features are concatenated to output a 512-dimensional clinical feature vector. .
[0039] In this embodiment, the imaging analysis sub-model (3DResNet-UNet) has the following inputs: ROI images from CT / MRI / PET-CT (3D tensor: 64×64×64) and WSI pathological sections (2D tensor: 256×256); Layer structure: 3D feature extraction layer: 3D features of the ROI (such as nodule density and edge smoothness) are extracted using 3DResNet50 (pre-trained on MedicalNet), and the output is 256-dimensional; Pathology-image alignment layer: The pathological features of WSI (such as cellular atypia) are spatially matched with the ROI features of CT, and an attention mechanism is used to focus on the "image abnormality-pathological positive" region; Multimodal image fusion layer: The structural features of CT and the metabolic features (SUV value) of PET-CT are stitched together, and a 512-dimensional image feature vector is output through a fully connected layer. .
[0040] In this embodiment, the environment-lifestyle analysis sub-model (CausalEmbedding) takes the following inputs: quantified lifestyle data (such as smoking index, BMI), and regional environmental data (PM2.5, ...). Layer structure: Causal embedding layer: Constructs a causal graph of "environment-lifestyle-clinical indicators" using a causal graph neural network (CGNN), where nodes are variables and edges are causal effect values (initialized based on the UK Biobank cohort); Effect screening layer: Retains edges with causal effect values ≥ 0.3 (e.g., "smoking → decreased lung function" "PM2.5 → increased inflammatory markers"); Output layer: Maps causal features to a 512-dimensional environmental feature vector. .
[0041] In this embodiment, intra-modal correlation enhancement processing is performed on each group of modal feature vectors in multiple groups to obtain intra-modal self-interaction enhanced feature vectors, including: A three-layer interaction mechanism is constructed, consisting of "intramodal self-interaction → bimodal causal / associative interaction → multimodal global interaction," as detailed below: Intramodal self-interaction (enhancing intramodal relationships): Genome: Based on the LD network, the association weights between variant features are updated using GAT; Molecular expression: Enhance pathway-protein synergy using Transformer self-attention layers (e.g., "p53 pathway activity-protein expression"). Epigenetics: Enhancing the DMR-gene temporal regulatory association using LSTM+attention mechanism; Clinically: Based on EHR time series, Time-SeriesGNN is used to update the correlation of indicators across time points; Imagery: Based on the spatial location of the ROI, the feature associations of neighboring regions are updated using 3D-CNN; Environment: Use CGNN to strengthen the causal chain between environmental variables (e.g., "PM2.5 → inflammation → decrease in FEV1"). Output: "Self-interactive enhancement feature vectors" for each modality (all maintaining 512 dimensions).
[0042] In this embodiment, causal interaction processing is performed on the environment-clinical and clinical-imaging dual-modal combinations, calculating causal attention weights and generating causal interaction feature vectors, including: Environment-Clinical Interaction: Input: , Causal effect matrix C (512×512 dimensional, based on UKBiobank); Causal attention calculation: ;in, The environmental modal mass vector (512 dimensions); The clinical modality quality vector (512 dimensions); For Hadamard product; An environmental-clinical causal attention weight matrix; The normalization function is used to make the sum of the weights equal to 1. The causal effect matrix is 512×512 dimensional. The outer product operation expands the two vectors into a matrix; interactive feature generation: ; The feature vector for environment-clinical interaction is 512-dimensional. It is a multilayer perceptron (a 3-layer fully connected network with ReLU activation); This is a vector concatenation operation that horizontally joins two vectors. This is the environmental feature vector (512 dimensions); The clinical feature vector is 512-dimensional; the output is 512-dimensional. Clinical-Image Interaction: Input: (e.g., "FEV1" lung function) (e.g., "SUV value of lung nodules"); Clinical-image mapping: Generate a mask matrix R based on the NCCN guideline rule base (e.g., "FEV1 < 80% and SUVmax > 5 → high risk"); Interactive feature generation: , The clinical-image interaction feature vector is 512-dimensional. For gated multilayer perceptrons, parameters are dynamically adjusted based on the mask matrix R; For vector concatenation operations, and same; The image feature vector is 512-dimensional. The output is a clinical-image mapping mask matrix, constructed based on the NCCN guidelines, where 1 represents a high-risk combination and 0 represents a low-risk combination; the output is 512-dimensional.
[0043] In this embodiment, association-based interaction processing is performed on the genome-transcriptome and transcriptome-epimetame dual-modality combinations. Association weights are calculated based on regulatory network priors, and association-based interaction feature vectors are generated. Genome-transcriptome interaction: Input: , eQTL regulation matrix; Graph construction: The feature vector is split into 128 biological units (e.g., "TP53 region"), and the edge weights between units are... ; This represents the absolute value of the variant-expression correlation in the eQTL database. , For the corresponding modal quality score; Interactive feature generation: , This is a genome-transcriptome interaction feature vector (512 dimensions); This is a graph attention network that processes a graph consisting of 128 biological units; each unit node represents a feature vector of the 128 biological units; edge weights represent the intensity of inter-unit regulation; the output is 512-dimensional. Transcriptome-Epigenome Interaction: Input: , DMR-gene regulatory pairs; association weight calculation: ; The transcriptome-epigenome association weight (scalar) represents the regulatory strength between the two modalities; For the Sigmoid function; Pearson correlation coefficient calculation; interaction feature generation: , The transcriptome-epimetom interaction feature vector (512 dimensions) It is a multilayer perceptron; output is 512-dimensional; environment-genome interaction: input: , Environment-mutation databases (such as COSMIC); interaction feature generation: ; The feature vector for environment-genome interaction (512 dimensions); For causal attention mechanisms, attention weights are calculated by combining environmental-mutation prior knowledge; the output is 512-dimensional.
[0044] In this embodiment, the global weights of each modality are calculated based on the modal quality scoring matrix. The intra-modal self-interaction enhancement feature vectors, causal interaction feature vectors, and correlational interaction feature vectors are input into the Transformer encoder for global fusion. After residual connection and layer normalization processing, the cross-modal fusion feature vector is output, including: Construct a modal quality score matrix (MQM), with quality scores for each data type ranging from 0 to 1 (constrained by clip(·,0,1)): Genomic data: ; Genomic modality quality score; sequencing depth: depth of genome sequencing coverage; variant detection rate: proportion of successfully detected variants; As an upper limit constraint, depths greater than 150 are calculated as 150. As an upper limit constraint, if the detection rate is >98%, it will be calculated as 98%. To constrain values to the range [0,1]; Molecular expression data: ; Molecular expression modal quality scoring; batch effect Values represent RNA integrity (0-10), with lower values indicating more severe RNA degradation; the percentage of undetected proteins is the proportion of proteins that could not be successfully detected out of the total target proteins; epigenomic data: ; Epigenome modality quality score; Methylation site coverage: the proportion of successfully detected methylation sites; Batch number: the number of batches treated in the experiment, reflecting the size of the batch effect; The quality decreases exponentially with increasing batch numbers; clinical data: ; Clinical modality quality score; Missing field count: the number of missing clinical indicators in the electronic health record; Total field count: the total number of clinical indicators in the electronic health record; Imaging data: ; Image modal quality score; ROI segmentation accuracy: the degree of overlap between automatic segmentation and manual annotation of the region of interest; image signal-to-noise ratio: an image quality index (unit: dB), the higher the value, the lower the noise; As an upper limit constraint, if the accuracy is >95%, it is calculated as 95%; As an upper limit constraint, the signal-to-noise ratio is calculated as 35dB when it is >35dB; Environmental data: ; Environmental modal quality is scored; sensor coverage rate is the proportion of the target area covered by environmental monitoring sensors; questionnaire reliability is a lifestyle questionnaire, measuring internal consistency (≥0.8); modal importance is obtained by pre-training a single-modal model in an independent cohort of TCGA / UK Biobank and calculating the SHAP value as... (such as genome) =0.32, Clinical =0.28); Global weight calculation: ; Let m be the global weight of mode m; The quality score for mode m; The importance weight of modality m is calculated based on the SHAP value, reflecting its predictive contribution; feature redundancy handling: clinically relevant features ( , , Gating is performed. ; This is the fused clinical feature vector; The Sigmoid function maps the gate weights to [0,1]. These are the gating weight parameters; Clinical feature vectors (512-dimensional); Transformer global fusion: input 11 feature vectors (6 self-interactions + 5 bimodal interactions, each 512-dimensional); The environment-clinical interaction feature vector (512 dimensions) The clinical-image interaction feature vector is 512-dimensional. Cross-attention layer: Set 16 attention heads, and combine attention weights with dynamic modality weights. Residual connections and layer normalization: Residual connections and LayerNorm are added after each attention layer; Global feature output: After passing through a 2-layer Transformer encoder, a 1024-dimensional global fused feature vector is output. .
[0045] In this embodiment, the cross-modal fusion feature vector is input into the spatiotemporal probability prediction layer. Combined with the survival analysis model and probability calibration mechanism, the initial prediction results of the probability of malignant tumor occurrence in multiple time windows are output, including: Spatiotemporal probability prediction layer: In view of the "uncertainty of occurrence time" characteristic of malignant tumors, a three-level prediction mechanism of "time series modeling - survival analysis - probability calibration" is constructed, which is specifically constructed as follows: Temporal Feature Embedding: Input: (1024-dimensional), patient age (continuous value), duration of medical history (e.g., "5-year history of diabetes"); Processing: Age embedding: Mapping age to a 64-dimensional vector, embedding "age risk interval" (e.g., "50-60 years old → high-risk interval for lung cancer"); Time-series embedding of medical history: Using a pre-trained decoder from Clinical indicators were decoupled, and LSTM was used to process the association between medical history duration and clinical indicators; spatiotemporal feature splicing: (1024+64+64=1152 dimensions); It is a spatiotemporal feature vector (1152 dimensions), which integrates global features, age, and medical history; This is a globally fused feature vector (1024 dimensions); For the age embedding vector (64-dimensional), continuous ages are mapped to risk intervals; The medical history embedding vector (64-dimensional) encodes the duration and severity of the disease history. Survival analysis model (improved version of DeepSurv): Based on the DeepSurv framework, it introduces multi-time window prediction: Input: Patient's "follow-up time" and "whether tumor occurred" labels (training phase); Network structure: Fully connected layer: 2 layers, hidden layer dimension 2048→1024, activation function LeakyReLU; Survival function calculation layer: outputs survival functions for different time windows (t=1 year, 3 years, 5 years). Probability conversion: Probability of malignant tumor occurrence ; The probability of malignant tumor occurrence within time window t; The conditional survival function is given the spatiotemporal features; the loss function is: ;in, Loss function for survival analysis; This represents the number of training samples; This is an event indicator (1 for tumor occurrence, 0 for loss to follow-up). For model parameters, Let i be the feature vector of sample i; The follow-up time for sample i; Summation of the risk set; λ is the L2 regularization term; λ is the L2 regularization coefficient (1e-5). Transpose sign; probability calibration and uncertainty quantification, probability calibration: independent calibration for each time window: ; in, The probability of occurrence of the calibrated time window t; This is the scaling parameter for the time window t; This is the offset parameter for the time window t; The model's original predicted probabilities; Calibration evaluation: The calibration effect is evaluated using the expected calibration error (ECE), with ECE ≤ 0.05 considered acceptable; Uncertainty quantification: Monte Carlo dropout (MCDO): Dropout (rate 0.3) is added to the fully connected layer, and the prediction is repeated 20 times to obtain 20 probability values; Confidence interval calculation: Uncertainty is represented by a 95% confidence interval (CI), such as "1-year probability of occurrence = 0.25 (95% CI: 0.18-0.32)"; Uncertainty grading: CI width < 0.1 is "low uncertainty", 0.1-0.2 is "medium uncertainty", and > 0.2 is "high uncertainty"; Prediction result output: Final output: Probability of malignant tumor occurrence in 3 time windows: , , (Retain 3 decimal places); 95% confidence interval for each probability; Risk level classification: Low risk: <0.1; Medium risk: High risk: .
[0046] In this embodiment, the model is trained in stages based on a multi-task loss function. When the training results meet the requirements, a malignant tumor occurrence probability prediction model is obtained, including: Multi-dimensional loss function design: First-dimensional loss function: ; The second-dimensional loss function is the modality feature consistency loss, which ensures that different modalities have consistent features for the same biological significance (e.g., the risk signal of "TP53 variant" is consistent in the genome and transcriptome). ;in, For modal feature consistency loss; The size of the biological tag set; Let the mean square error function be used. 'b' is the feature extraction function; 'b' is the biological significance label. Genome feature vector; This represents the molecular expression feature vector; Third-dimensional loss function: Data quality-perceived loss: Penalizing the impact of low-quality data. ;in The quality score of modality m for sample i. The gradient of the feature; Perceived loss for data quality; The total number of samples; StopGrad blocks gradient backpropagation to avoid gradient explosion in low-quality data; Let m be the feature vector of mode i; Fourth-dimensional loss function: Clinical prior constraint loss: Introducing clinical guideline priors (e.g., "Smoking history ≥ 20 years → increased risk of lung cancer"): ;in, Loss due to clinical prior constraints; KL divergence measures the difference between two distributions; To predict the distribution for the model, For clinical prior distribution; Total loss function: ; This is the total loss function; To analyze losses for survival; For modal feature consistency loss; For perceived loss of data quality.
[0047] The working principle and beneficial effects of the above technical solution are as follows: The training data covers patients' data from microscopic levels such as genomics, transcriptomics, epigenomics, and proteomics, to macroscopic levels such as electronic health records, pathology reports, imaging data, and lifestyle-related data on smoking, diet, exercise, and environment. This multi-dimensional data integration enables the model to comprehensively capture various biological information and environmental factors related to the occurrence of malignant tumors, avoiding the limitations of a single data source, and thus more accurately reflecting the complex mechanisms of tumor development. Intramodal correlation enhancement processing is performed on each set of modal feature vectors to obtain intramodal self-interaction enhanced feature vectors. This processing method highlights the interrelationships between features within each modality, enhances the expressive power of the features, and enables the model to better capture key information within the same modality, improving the sensitivity and specificity of tumor-related features. The global weights of each modality are calculated based on the modality quality scoring matrix, and the intramodal self-interaction enhanced feature vectors, causal interaction feature vectors, and correlation interaction feature vectors are input into the Tr... The Ansformer encoder performs global fusion; this fusion method can weight and combine features according to the importance of each modality, giving full play to the advantages of different modal features and avoiding information redundancy and feature conflict problems that may be caused by simple splicing, thus obtaining a more representative and discriminative cross-modal fusion feature vector. The cross-modal fusion feature vector is input into the spatiotemporal probability prediction layer, combined with the survival analysis model and probability calibration mechanism, to output the initial prediction results of the probability of malignant tumor occurrence in multiple time windows. This multi-time window prediction method can provide clinicians with tumor occurrence risk information at different time points, which helps to formulate personalized screening and prevention strategies and achieve early intervention and precise control of tumors. The model is trained in stages based on a multi-task loss function, which can optimize the model in a targeted manner according to the characteristics and objectives of different training stages. This staged training method helps to improve the convergence speed and training efficiency of the model, while avoiding overfitting problems, so that the model has good performance on both the training set and the test set.
[0048] Based on the predicted probability of malignant tumor occurrence and combined with patient characteristic data, a malignant tumor occurrence probability map is generated, including: Constructing a knowledge ontology framework for malignant tumor medicine; Acquire multi-source medical data, and accurately map the multi-source medical data to the medical knowledge ontology framework to obtain the data-ontology mapping result; Relationship prediction is performed on ontology nodes in the medical knowledge ontology framework based on the knowledge graph embedding model. The confidence of the predicted initial relationships is optimized to obtain the target relationships between ontology nodes and their corresponding confidence. Construct a mapping matrix between the probability ontology and multiple ontology nodes, and dynamically fuse the probability and ontology nodes based on the data-ontology mapping results, target relationships and corresponding confidence levels; The medical knowledge ontology framework and mapping matrix are incrementally self-optimized, and a probability map of malignant tumor occurrence is constructed.
[0049] In this embodiment, the construction of the malignant tumor medical knowledge ontology framework includes: A four-layer nested medical knowledge ontology is constructed, comprising a core ontology, a risk ontology, a probability ontology, and an evidence and intervention ontology. The core ontology includes a patient ontology and a tumor type ontology. The risk ontology includes a gene risk ontology, a lifestyle ontology, a medical history ontology, and an environmental risk ontology, and is forcibly associated with the patient ontology in the core ontology. The probability ontology includes an occurrence probability ontology, a risk stratification ontology, and a confidence level ontology, and is forcibly associated with risk ontology nodes. The evidence and intervention ontology includes a literature evidence ontology, a test report ontology, and an intervention plan ontology, and is associated with the probability ontology through an evidence chain. Each ontology node is configured with core attributes and association rules. Core attributes are divided into static attributes (such as "gene name") and dynamic attributes (such as "exposure duration" and "last update time"). Dynamic attributes support time-series verification. Association rules are weighted based on authoritative domain data (such as NCCN guidelines) and the applicable time range is marked (such as "smoking history must be ≥10 years"). Each ontology node is bound to a medical standard code, including ICD-O-3 code and SNOMED CT code. In case of conflict, ICD-O-3 code is used first, and conflict logs are recorded for manual review. All ontology node operations are logged, including node creation, updates, and conflict events.
[0050] In this embodiment, the step of acquiring multi-source medical data and accurately mapping the multi-source medical data to the medical knowledge ontology framework to obtain a data-ontology mapping result includes: Acquire six types of multi-source medical data: patient structured data, patient unstructured data, blockchain-based evidence data, predictive model output data, intervention plan data, and public epidemiological data; The mapping relationship between various types of data and corresponding ontology nodes is established based on the mapping rule engine: structured data is directly matched with ontology attributes; unstructured data is matched with ontology nodes through word vectors; blockchain-stored data is linked to the test report ontology and patient ontology through hash verification; public epidemiological data is mapped to environmental risk ontology and probability ontology; the KNN algorithm is used to calculate the data-ontology mapping similarity, and the mapping results are quality checked by combining template-based verification rules (such as data integrity and timeliness thresholds); mapping results with similarity <85% are marked as low quality, the mapping log is recorded, and manual review is triggered.
[0051] In this embodiment, the step of predicting relationships between ontology nodes in the medical knowledge ontology framework based on a knowledge graph embedding model, optimizing the confidence of the predicted initial relationships, and obtaining the target relationships and corresponding confidence levels between ontology nodes includes: The TransE model is used to embed ontology nodes, transforming them into low-dimensional vectors. An initial candidate set of relations is predicted based on vector similarity, comprising relation type, node pairs, and prediction confidence. The initial candidate set is optimized through three rounds of validation: Round 1: Hard validation using clinical guideline templates (e.g., NCCN guidelines) to filter relations that do not conform to authoritative rules; Round 2: Validation using the KNN algorithm for literature and data support, calculating evidence support based on authoritative databases such as PubMed and the Cochrane Library; Round 3: Temporal association validation using an LSTM model only when dynamic attribute nodes (e.g., lifestyle ontology) are involved; otherwise, this round is skipped. The final confidence score is calculated based on the results of the three rounds of validation: Confidence Score = (Guideline Validation Weight × Guideline Matching Degree) + (Literature Validation Weight × Evidence Support Degree) + (Temporal Validation Weight × ... (Time-series relevance); Weight allocation is based on NCCN evidence levels: guideline verification weight 0.6, literature verification weight 0.3, and time-series verification weight 0.1; Relationships with confidence ≥80% are retained as target relationships, and relationship evolution logs are recorded, including the original relationship, new relationship, triggering conditions, update time, and verification details.
[0052] In this embodiment, the construction of the mapping matrix between the probability ontology and multiple ontology nodes, based on the data-ontology mapping results, target relationships, and corresponding confidence levels, to achieve dynamic fusion of probability and ontology nodes, includes: Construct a two-dimensional mapping matrix, where the row dimension represents the probability ontology nodes and the column dimension represents the associated nodes in the risk ontology and the core ontology. The matrix elements include contribution and confidence. Initialize matrix element values based on the relationship between the prediction model output data and the target; Define dynamic update trigger conditions: Static update: triggered when the prediction model version is upgraded; Dynamic update: triggered when the change rate of key patient data is >5% (such as gene mutation status, exposure duration) or the confidence fluctuation of risk stratification ontology nodes is >10%. The KNN algorithm is used to map patient occurrence probabilities to a risk stratification ontology, and the influence weight of each ontology node on risk stratification is labeled. During the mapping process, the confidence of multiple probabilities associated with one piece of evidence (such as a single literature evidence supporting the probability of multiple tumor types) is weighted according to the evidence level (GRADE system): high-quality evidence (such as RCT studies) has a weight of 1.0, medium-quality evidence has a weight of 0.7, and low-quality evidence has a weight of 0.3, so as to realize the linkage and fusion of probability and risk stratification.
[0053] In this embodiment, incremental self-optimization is performed on the medical knowledge ontology framework and mapping matrix, and a probability map of malignant tumor occurrence is constructed. This includes: ontology optimization only for new nodes not bound to standard codes or temporary nodes generated by mapping; detection of semantically similar ontology nodes based on cosine similarity clustering (similarity threshold > 90%); and for detected similar nodes, priority is given to checking the binding status of medical standard codes: if already bound to ICD-O-3 / SNOMED... CT encoding preserves the original nodes, recording semantic variant attributes and conflict logs; if not bound to an encoding, nodes are merged and manually reviewed, with merged nodes inheriting the highest level of evidence confidence; when merging evidence ontology nodes, a "one piece of evidence, multiple probabilities" association is established, but confidence is weighted according to the GRADE system hierarchy and associated with the confidence ontology; mapping matrix optimization: based on relationship evolution logs and mapping logs, low-quality mappings (similarity <85%) are recalibrated; when ontology nodes are merged or deleted, the row and column dimensions of the mapping matrix are updated synchronously; a malignant tumor incidence probability atlas is constructed: using the optimized ontology framework as the skeleton and the mapping matrix as the connector, patient incidence probabilities are dynamically fused; the atlas output includes risk stratification visualization, evidence tracing paths, and confidence labeling; all optimization operations are logged completely, including node merging / deletion, confidence changes, and manual review records, ensuring historical traceability.
[0054] The working principle and beneficial effects of the above technical solution are as follows: Constructing a medical knowledge ontology framework for malignant tumors provides a standardized and structured foundation for integrating multi-source medical data. This framework can systematically organize and define various concepts, entities, and relationships related to malignant tumors, making previously scattered and heterogeneous medical knowledge orderly and easy to manage, laying a solid foundation for subsequent data mapping and knowledge mining. Based on a knowledge graph embedding model, relationship prediction of ontology nodes in the medical knowledge ontology framework can discover potential relationships between ontology nodes. These relationships may be difficult to discover in traditional data processing and analysis, but are of great significance for understanding the occurrence mechanism, development process, and treatment selection of malignant tumors. For example, relationship prediction may reveal the association between certain genes and specific tumor types, providing a basis for precise diagnosis and personalized treatment of tumors. Constructing a mapping matrix between the probability ontology and multiple ontology nodes, and dynamically fusing probability with ontology nodes, can intuitively display the relationship between the probability of malignant tumor occurrence and various related factors. This fusion method makes probability information no longer an isolated value, but associated with specific medical concepts and entities, providing clinicians and researchers with more comprehensive and in-depth information.
[0055] In summary, by collecting and processing multi-source patient data to determine patient characteristic data, a malignant tumor incidence probability prediction model is constructed based on artificial intelligence. This model analyzes and performs pattern recognition on patient characteristic data, predicting the probability of patients developing malignant tumors within a specific future timeframe. Based on the predicted probability and combined with patient characteristic data, a malignant tumor incidence probability map is generated. Analysis of this map identifies the features contributing most to the incidence probability, and prevention plans are developed for patients. A malignant tumor incidence probability map analysis report is generated and visualized. Furthermore, automatic warnings and clinical intervention prompts for patients with a high probability of malignant tumor incidence. This approach effectively obtains and analyzes malignant tumor incidence probability maps, improves the early diagnosis rate of malignant tumors, and promotes the comprehensive development of malignant tumor prevention and treatment.
[0056] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.
[0057] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A system for acquiring and analyzing the probability map of malignant tumor occurrence, characterized in that, include: The data acquisition and integration module is used to collect multi-source patient data from multiple sources and dimensions. The data processing map acquisition module is used to process the collected multi-source patient data, determine the patient characteristic data, and build a malignant tumor occurrence probability prediction model based on artificial intelligence. Based on the malignant tumor occurrence probability prediction model, the module performs pattern recognition on the patient characteristic data, determines the malignant tumor occurrence probability prediction result, and generates a malignant tumor occurrence probability map. The graph analysis and visualization module is used to analyze the probability graph of malignant tumor occurrence and to visualize the analysis report of the probability graph of malignant tumor occurrence. The data processing map acquisition module further includes: The model building unit is used to build a model for predicting the probability of malignant tumor occurrence based on artificial intelligence. Collect patient historical data and divide the collected patient historical data into training set and test set in a 7:3 ratio. The machine learning model is trained using a training set, enabling it to learn autonomously from the training set how to predict the probability of malignant tumors and predict the probability of a patient developing a malignant tumor within a specific timeframe in the future, thus establishing a malignant tumor probability prediction model. The malignant tumor incidence probability prediction model was tested using a test set to evaluate its generalization performance and determine whether it could achieve the expected effect of predicting the probability of a patient developing a malignant tumor within a specific time period in the future. When the malignant tumor occurrence probability prediction model fails to achieve the expected effect of predicting the probability of a patient developing a malignant tumor within a specific future time, the parameters of the malignant tumor occurrence probability prediction model are adjusted and iteratively optimized until the malignant tumor occurrence probability prediction model can achieve the expected effect of predicting the probability of a patient developing a malignant tumor within a specific future time, and the optimal malignant tumor occurrence probability prediction model is determined. The machine learning model is trained using a training set to determine a predictive model for the probability of malignant tumor occurrence, including: The training data in the training set is input into the multimodal feature extraction layer to extract features from the training data, resulting in multiple sets of modal feature vectors. The training data in the training set includes patients' genomic data, transcriptomic data, epigenomic data, proteomic data, electronic health records, pathology reports, imaging data, and patients' smoking, diet, exercise, and environmental data. The multiple sets of modal feature vectors include genomic feature vectors, molecular expression feature vectors, epigenetic feature vectors, clinical feature vectors, imaging feature vectors, and environmental feature vectors. Intramodal correlation enhancement processing is performed on each of the multiple sets of modal feature vectors to obtain intramodal self-interaction enhanced feature vectors; Causal interaction processing is performed on the combination of environment-clinical and clinical-imaging dual modalities, causal attention weights are calculated and causal interaction feature vectors are generated; association interaction processing is performed on the combination of genome-transcriptome, transcriptome-epigome, and environment-genome dual modalities, association weights are calculated in conjunction with regulatory network priors and association interaction feature vectors are generated. The global weights of each modality are calculated based on the modal quality scoring matrix. The self-interaction enhancement feature vectors, causal interaction feature vectors, and correlation interaction feature vectors within the modality are input into the Transformer encoder for global fusion. After residual connection and layer normalization processing, the cross-modal fusion feature vector is output. The cross-modal fusion feature vector is input into the spatiotemporal probability prediction layer, and combined with the survival analysis model and probability calibration mechanism, the initial prediction results of the probability of malignant tumor occurrence in multiple time windows are output. The model is trained in stages based on a multi-task loss function. When the training results meet the requirements, a prediction model for the probability of malignant tumors is obtained.
2. The system for acquiring and analyzing the probability map of malignant tumor occurrence according to claim 1, characterized in that, Collect multi-source patient data from multiple sources and dimensions, and perform the following operations: Monitor the patient's genome, transcriptome, epigenome, and proteome to obtain patient omics data; Monitor patients' electronic health records, pathology reports, and imaging data to obtain patients' clinical data; Monitor patients' smoking, diet, exercise, and environmental data to obtain patients' lifestyle data; Based on patient omics data, patient clinical data, and patient lifestyle data, multi-source patient data is generated.
3. The system for acquiring and analyzing the probability map of malignant tumor occurrence according to claim 2, characterized in that, The data processing map acquisition module includes: The data processing unit is used to process multi-source patient data and determine patient characteristic data; This includes cleaning the multi-source patient data, removing noise, identifying missing and outlier values, and processing the identified missing and outlier values. Normalize the multi-source patient data to convert it into a unified data format, remove the dimensional differences between the multi-source patient data, and form standardized multi-source patient data. Feature extraction is performed on multi-source patient data to extract features related to the acquisition and analysis of malignant tumor incidence probability maps, thereby determining patient characteristic data.
4. The system for acquiring and analyzing the probability map of malignant tumor occurrence according to claim 3, characterized in that, Identify missing and outlier values in multi-source patient data, and process the identified missing and outlier values by performing the following operations: Based on the missing matrix diagram, a heatmap is used to visualize the missing values of all values in the multi-source patient data, identify all missing data points in the multi-source patient data, and identify the missing values in the multi-source patient data. Based on the standard deviation method, values exceeding the mean ± 3 standard deviations in multi-source patient data are identified, data points that are significantly different from the behavior patterns of most data are found, and outliers in multi-source patient data are identified. The missing and outlier values in the multi-source patient data were evaluated to determine whether the missing and outlier values in the multi-source patient data were valuable for obtaining and analyzing the probability map of malignant tumor occurrence. When missing values and outliers in the multi-source patient data are valuable for obtaining and analyzing the probability map of malignant tumor occurrence, the missing values in the multi-source patient data are filled and the outliers in the multi-source patient data are replaced. When missing and outlier values in patient multi-source data are of no value for obtaining and analyzing the probability map of malignant tumor occurrence, then the missing and outlier values in the patient multi-source data are removed.
5. The system for acquiring and analyzing the probability map of malignant tumor occurrence according to claim 4, characterized in that, The data processing map acquisition module further includes: The probability prediction unit is used to predict the probability of a patient developing a malignant tumor within a specific timeframe in the future. The process involves inputting patient characteristic data into a malignant tumor incidence probability prediction model, analyzing and recognizing patterns in the patient characteristic data based on the model, and predicting the probability of a patient developing a malignant tumor within a specific timeframe in the future, thus determining the prediction result of the malignant tumor incidence probability.
6. The system for acquiring and analyzing the probability map of malignant tumor occurrence according to claim 5, characterized in that, The data processing map acquisition module further includes: The map generation unit is used to generate a map of the probability of malignant tumor occurrence. Based on the prediction results of the incidence probability of malignant tumors and combined with patient characteristic data, a probability map of the incidence of malignant tumors is generated. Specifically, the incidence probability of malignant tumors output by the prediction model is associated with the corresponding biomedical characteristics to construct a probability map of the incidence of malignant tumors that can be queried in a visual form, and to reveal the complex network relationship between patient characteristics and diseases.
7. The system for acquiring and analyzing the probability map of malignant tumor occurrence according to claim 6, characterized in that, The spectral analysis and visualization module includes: The graph analysis unit is used to analyze the probability graph of malignant tumor occurrence, enabling doctors to understand and utilize the probability graph of malignant tumor occurrence. Among them, the system queries based on the probability map of malignant tumor occurrence, inputs patient characteristic data and returns their probability of malignant tumor occurrence, performs population analysis and driving factor analysis based on the probability map of malignant tumor occurrence, compares the probability distribution of malignant tumor occurrence in different subgroups, identifies the features that contribute the most to the probability of malignant tumor occurrence, and develops prevention plans for patients, generating a probability map analysis report of malignant tumor occurrence. The visualization unit is used to visually display the analysis report on the probability of malignant tumors in the form of charts, and to automatically issue early warnings for patients with a high probability of malignant tumors, prompting clinical intervention.
8. The system for acquiring and analyzing the probability map of malignant tumor occurrence according to claim 7, characterized in that, Based on the predicted probability of malignant tumor occurrence and combined with patient characteristic data, a malignant tumor occurrence probability map is generated, including: Constructing a knowledge ontology framework for malignant tumor medicine; Acquire multi-source medical data, and accurately map the multi-source medical data to the medical knowledge ontology framework to obtain the data-ontology mapping result; Relationship prediction is performed on ontology nodes in the medical knowledge ontology framework based on the knowledge graph embedding model. The confidence of the predicted initial relationships is optimized to obtain the target relationships between ontology nodes and their corresponding confidence. Construct a mapping matrix between the probability ontology and multiple ontology nodes, and dynamically fuse the probability and ontology nodes based on the data-ontology mapping results, target relationships and corresponding confidence levels; The medical knowledge ontology framework and mapping matrix are incrementally self-optimized, and a probability map of malignant tumor occurrence is constructed.