Multi-mode AI-based early prediction method for multiple cancer species
Through multimodal AI method, combined with the feature extraction of plasma, urine and expiratory samples, the graph neural network is used to construct the tissue association structure between cancer species, solving the accuracy and credibility of the single-modal cancer early screening technology, and achieving fine prediction and clinical decision support for early cancer.
Patent Information
- Application Number
- CN202510461263.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-14
- Publication Date
- 2025-07-11
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing cancer early screening technology relies on biomarkers from a single moist source, making it difficult to maintain high accuracy in complex scenarios where early cancer signals are extremely weak and lesions are unknown. It lacks the ability to effectively integrate multimodal data sources and model dynamic signaling across tissues, resulting in high uncertainty in predicted results and is unable to provide fine-grained clinical decision support.
By using multimodal AI method, by obtaining plasma, urine and exhalation samples, the cfDNA fragments, the exosome proteome characteristics and the volatile organic compound characteristics were extracted respectively, and the tissue association structure between cancer species was constructed using the graph neural network, and the probability of cancer, tissue localization of primary foci and targeted drug sensitivity index were output.
It improves the ability to capture early cancer signals, realizes fine modeling of tumor tissue sources, improves the credibility of predicted results and decision-making reliability, and provides fine-grained clinical support.
Smart Images

Figure CN120299530A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of early cancer screening and artificial intelligence-assisted diagnosis, and specifically to a multi-cancer early prediction method based on multi-modal AI. Background Art
[0002] Current early cancer screening technologies mostly rely on biomarkers from a single body fluid source, such as serum proteins, methylation signals, or cfDNA fragment characteristics. Although such methods have been applied to pre-clinical screening, in the face of complex scenarios with extremely weak early cancer signals and unclear lesions, the accuracy is often difficult to stably maintain at a high level. Especially in asymptomatic populations, the noise interference is strong and the individual differences are significant. The single-modal solutions of the existing technologies lack sufficient structural feature support, resulting in the model output being prone to misjudgment or "insufficient confidence" results, and it is difficult to meet the generalization requirements.
[0003] On the other hand, although in recent years there have been studies attempting to fuse multi-modal data sources for cancer prediction, such as combining proteomic and DNA methylation data, in the specific implementation process, most solutions adopt simple splicing or independent processing of features, and fail to consider the potential biological pathway associations between different modalities. This fragmented modeling method is difficult to capture the cross-tissue dynamic signal transmission during tumor formation, especially lacking the modeling ability for metabolic regulatory pathways and tissue coupling relationships, and there are still significant gaps in tumor tissue localization and treatment target prediction.
[0004] In addition, most existing methods adopt a single target output, such as only outputting the "cancer / non-cancer" classification result, lacking a finer-grained task setting, and unable to provide key clinical decision-making information such as primary focus tracing or treatment sensitivity assessment at the same time. Even if some studies attempt to introduce a multi-task framework, there is generally a lack of evaluation means for the confidence of the output results, resulting in insufficient interpretability of the prediction results in the face of uncertain samples, limiting the practicality and clinical trust of the model in real scenarios. Therefore, the present invention proposes a multi-cancer early prediction method based on multi-modal AI to solve the deficiencies of the existing technologies. Summary of the Invention
[0005] Aiming at the deficiencies of the existing technologies, the present invention provides a multi-cancer early prediction method based on multi-modal AI, which solves the key problems such as limited single-modal information, unclear tissue localization, and insufficient credibility of prediction results.
[0006] To achieve the above purposes, the present invention is realized through the following technical solutions: A multi-cancer early prediction method based on multi-modal AI, including the following steps:
[0007] S1. Obtain multiple body fluid samples from the same subject, and the body fluid samples include plasma samples, urine samples, and exhaled breath samples;
[0008] S2. Sequence the plasma sample to obtain the fragmentomics characteristics of cfDNA, where the fragmentomics characteristics include fragment end sequence frequency information and nucleosome positioning density information;
[0009] S3. Perform proteomic detection on exosomes in the urine sample to obtain the glycosylated protein expression characteristics in the exosomes;
[0010] S4. Perform volatile organic compound detection on the exhaled breath sample to obtain the feature vector composed of ion mobility spectrometry data;
[0011] S5. Preprocess and vectorize the obtained multi-modal features, and input them into a multi-modal deep learning model, where the model constructs an inter-cancer tissue association structure based on a graph neural network;
[0012] S6. Output the prediction results including the canceration probability, primary focus tissue localization information, and targeted drug sensitivity indicators through the deep learning model.
[0013] The present invention also provides a multi-cancer early prediction system based on multi-modal AI, including:
[0014] A cfDNA fragment feature acquisition module for collecting and extracting the end sequence features and nucleosome positioning information of cfDNA in plasma;
[0015] An exosome proteome detection module for obtaining the glycosylated protein expression characteristics of exosomes from urine samples;
[0016] An exhaled breath VOCs analysis module for obtaining the ion mobility spectrometry features in exhaled breath samples;
[0017] A feature preprocessing module for normalizing the three-modal data and constructing feature expressions;
[0018] A graph neural network modeling module for establishing a graph structure based on the metabolic pathway relationship between tissues and performing multi-modal joint prediction;
[0019] An output module for outputting the canceration probability, primary focus tissue localization, and treatment sensitivity results.
[0020] The present invention provides a multi-cancer early prediction method based on multi-modal AI. It has the following beneficial effects:
[0021] 1. The present invention adopts a structured extraction method combining cfDNA end sequences and nucleosome positioning information, realizing fine modeling of the tumor tissue source. Different from the existing technology that only uses the statistical features of cfDNA fragment lengths, it avoids the problem of information loss and greatly improves the ability to capture weak signals of early canceration.
[0022] 2. By introducing ion mobility spectrometry to analyze the migration fingerprint characteristics of exhaled VOCs molecules, the present invention constructs a non-invasive detection pathway that can be used to reflect abnormal whole-body oxidative metabolism. Compared with the traditional detection methods using serum markers or single metabolite quantification, it is not only easy to operate, but also highly sensitive to non-specific tumor states, solving the problem of insufficient sensitivity of current body fluid samples.
[0023] 3. The present invention uses a graph neural network to perform structured joint modeling of multi-modal biomarkers across tissue pathways, and effectively establishes a topological model of functional coupling between tissues. Compared with the common simple splicing or concatenation models in previous modal fusion, this scheme has stronger structural characterization ability and solves the problem of information breakage caused by heterogeneity between data modalities from the mechanism level.
[0024] 4. The present invention refines the prediction results into three types of task outputs: cancer probability, primary focus location, and targeted drug response, and introduces an uncertainty estimation mechanism to make each prediction result have a quantitative reference of "confidence level". Different from the traditional AI diagnosis model with single output and fixed results, this method is closer to the actual clinical needs and significantly improves the decision-making reliability and implementation ability. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] Figure 1 is the flowchart of the method of the present invention;
[0026] Figure 2 is the system architecture diagram of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0027] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0028] Please refer to Figure 1 , the embodiment of the present invention provides a multi-cancer early prediction method based on multi-modal AI, including the following steps:
[0029] S1. Obtain multiple body fluid samples from the same subject, and the body fluid samples include plasma samples, urine samples, and exhaled samples;
[0030] In the construction of a multi-modal AI system, the quality and integrity of data input constitute an important basis for the learning accuracy of the model. The step of obtaining multiple body fluid samples described in the present invention is a precondition for connecting clinical practice and algorithm modeling, directly affecting the stability and interpretability of subsequent feature distributions.
[0031] To achieve the combined prediction goal of multi-cancer screening and tissue localization, the present invention constructs input channels using multiple independent information sources. Generally, it is recommended to synchronously obtain three types of body fluid samples from the same subject, namely plasma samples, urine samples, and exhaled breath samples. The above samples should all be collected on the same day, with an interval of no more than 2 hours, to reduce the impact of temporal deviation on the correlation of multi-omics information.
[0032] Specifically, the collection of plasma samples is usually completed by peripheral vein puncture. It is recommended to use EDTA anticoagulant blood collection tubes to prevent coagulation from interfering with the release of cfDNA. After collection, let it stand at room temperature for 10 minutes, and then perform two-step centrifugation. The first step is at 1600g for 10 minutes to remove blood cells; the second step is at 12000g for 10 minutes to remove cell debris, platelets, and macromolecular complexes. The finally obtained plasma can be used for downstream cfDNA extraction. In a possible implementation, the cfDNA concentration is recommended to reach ≥10 ng / mL to meet the minimum input requirement for WGS sequencing.
[0033] The collection of urine samples usually selects midstream urine to avoid the influence of contaminants. As an option, a vacuum negative pressure sampling bag can be used to ensure closed transmission. After collection, it should be frozen at -80°C within 15 minutes, or stored short-term at 4°C. The part used for extracting exosome signals needs to be processed through the following steps:
[0034] Centrifuge to remove cell residues and large particulate matter;
[0035] Further purify using a 0.22μm filter membrane;
[0036] Use ultracentrifugation (100000g, 90 minutes) to precipitate exosome particles;
[0037] In an optional step, add protease inhibitors to extend storage stability;
[0038] After exosome lysis, perform glycosylated peptide enrichment detection on a proteomics platform (LC-MS / MS).
[0039] The collection of exhaled breath samples is completed relying on a standardized exhaled breath detection device. The subject usually needs to sit still for 5 minutes to avoid exercise-induced metabolic fluctuations. When exhaling, it is necessary to be slow and uniform, continuously exhale no less than 3 times, each time no less than 5 seconds, to ensure sufficient collection of gas at the bottom of the lungs. The sample is analyzed by a GC-IMS (gas chromatography-ion mobility spectrometry) instrument to obtain a two-dimensional spectrum: the horizontal axis is the retention time (RetentionTime), the vertical axis is the drift time (DriftTime), and the signal peak intensity represents the concentration of volatile organic compounds.
[0040] In a specific implementation, a feature vector can be constructed by extracting VOC peaks of fixed types. Suppose a total of k = 100 VOCs are extracted, then a vector with a dimension of 3k is constructed for each sample. The composition of this vector is as follows:
[0041]
[0042] Among them: is the retention time of the i-th VOC (unit: second); is the drift time of the i-th VOC (unit: millisecond); I (i) is the peak intensity of the i-th VOC (unit: millivolt); k is the total number of VOCs, fixed at 100;
[0043] After construction, all numerical values are linearly normalized and mapped to the interval [0, 1].
[0044] In order to unify the data input interfaces of various body fluid samples, a modal set in the following form is established as the starting point of model input:
[0045]
[0046] Among them: is the cfDNA fragmentomics feature vector, including the frequency of 5'-end four-base sequences and the nucleosome positioning density; is the exosome glycoprotein expression vector, with a dimension of approximately 250; is the exhaled breath VOCs feature vector, composed of triples of 100 VOCs; d1 and d2 represent the vector dimensions in the corresponding modalities, and can be mapped to a unified dimension through the downstream preprocessing module.
[0047] In the actual implementation process, the above three types of collection ports can also be integrated through a disposable multimodal sensing system, such as a blood channel, a urine collector, and a replaceable exhaled breath chip, and unified numbering, scanning, and batch processing are combined with the label management system for batch modeling training.
[0048] In some extended implementation modes, if the subject has contraindications or special circumstances and any body fluid sample cannot be collected, the method of the present invention also supports reasoning through the incomplete modality completion mechanism in the case of modality absence, but a higher model confidence level will be obtained with a complete modality input.
[0049] S2. Sequencing the plasma sample to obtain the fragmentomics features of cfDNA, where the fragmentomics features include fragment end sequence frequency information and nucleosome positioning density information;
[0050] In a multi-modal cancer prediction system, cfDNA, as the product of apoptosis in the blood circulation of the subject, has a natural connection with the primary tissue in terms of the fragment structure information it carries. Therefore, extracting cfDNA from plasma and establishing fragmentomics expression is an important component module of the present invention, which follows the sample preparation process of step S1 and directly affects the feature effectiveness of downstream modeling.
[0051] Generally, the extraction of cfDNA can be completed using a column purification kit or the magnetic bead method. In a typical implementation, the silica membrane column method is used to extract cfDNA. By combining treatment with lysis buffer and proteinase K, high-purity free DNA is obtained, and it is recommended that the minimum concentration is not less than 10 ng / mL. To retain the boundary information of cfDNA fragments, an unbiased ligation enzyme protocol is adopted during the library construction process to construct a library without PCR amplification. The sample is sequenced on the Illumina NovaSeq platform, with a paired-end read length of 150 bp, and the sequencing depth is controlled above 30×.
[0052] After obtaining the original fastq file, the following operations need to be completed in sequence: filtering low-quality reads, removing adapter sequences, merging paired ends, mismatch correction, and aligning to the human reference genome (such as hg38). Generally, the BWA-MEM algorithm is selected as the alignment tool, and the parameters are set as follows:
[0053] The shortest matching length: 50 bp;
[0054] The maximum mismatch rate: 5%;
[0055] Secondary alignment is prohibited.
[0056] The requirements for the data quality control link are:
[0057] MappingRate > 98%;
[0058] DuplicateRate < 20%;
[0059] InsertSizePeak ≈ 166 bp, which conforms to the biological characteristics of cfDNA.
[0060] In the implementation path of the present invention, the cfDNA fragmentomics features consist of two parts, namely the fragment end sequence frequency information and the nucleosome positioning density information, forming a complete multi-dimensional expression vector structure.
[0061] First, the 5'-end fragment sequence frequency is counted. Let the total number of cfDNA fragments be N, and for each fragment, the 5'-end consecutive 4-base sequence is extracted. There are 4 4 = 256 combinations for the 4 bases. Each combination is quickly encoded through a hash index as a feature dimension:
[0062]
[0063] Wherein: is the 4-mer frequency vector at the 5' end of the cfDNA fragment; n i is the occurrence times of the i-th 4-mer combination in all cfDNA fragments; is the total number of cfDNA fragments; all frequency vector components satisfy That is, the total frequency is normalized to 1.
[0064] In some extended implementations, the frequency vector can also be further subjected to Logit transformation or Z-score normalization to be compatible with the input space requirements of subsequent deep networks.
[0065] Secondly, calculate the nucleosome positioning density feature of the cfDNA fragment. This feature is based on the mapping relationship between the midpoint of the cfDNA and the chromatin open region. The calculation formula for the midpoint coordinate of each cfDNA fragment is as follows:
[0066]
[0067] Where: x start is the starting coordinate of the cfDNA fragment on the reference genome, in bp; x end is the termination coordinate of the cfDNA fragment on the reference genome, in bp; x mid is the midpoint coordinate of the cfDNA fragment, used for mapping to the chromatin region.
[0068] Map all cfDNA midpoint coordinates to the predefined tissue-specific chromatin open region set, and determine whether it falls into any tissue region. The predefined chromatin region set is constructed as follows:
[0069]
[0070] Wherein: is the chromatin open region set corresponding to the t-th type of tissue; K t is the number of open regions of this tissue; [s t,j , e t,j is the start and end coordinates (in bp) of the j-th region; all region sets come from the DNase-seq annotation data of public databases such as ENCODE or RoadmapEpigenomics, and are classified by tissue and organ, with no less than 10 categories.
[0071] Count the number m t of cfDNA fragment midpoints falling into each tissue region, and form the following density distribution vector:
[0072]
[0073] Wherein: is the tissue nucleosome density vector; T is the number of tissue types; m t is the number of midpoints of cfDNA fragments falling in the chromatin region of the t-th type of tissue; is the total number of midpoints of cfDNA hits in all tissue regions; each term satisfies All terms are normalized and the sum is 1.
[0074] In some technical paths, to prevent deviations caused by the edge region, a window buffer parameter δ can be set to expand the matching interval to [s t,j -δ, e t,j +δ], where δ = 20bp.
[0075] Finally, the integrated feature vector is constructed as follows:
[0076]
[0077] This vector has both fragment generation mode and tissue source pointing ability, and is the main input representation of the plasma channel.
[0078] It should be noted that in the context of some complex tumor tissues, the cfDNA fragment information may also contain supplementary information such as dinucleosome cycle information, 5'-3' digestion deviation maps, CpG methylation island locus, etc. The above content can be used as the subsequent expansion direction of the present invention, but does not affect the feasibility of the core technical solution and the modeling path.
[0079] S3. Perform proteomic detection on the exosomes in the urine sample to obtain the glycosylated protein expression characteristics in the exosomes;
[0080] After completing the extraction of plasma cfDNA fragmentomics features, to further enhance the model's perception ability of cancer-related epigenetic modification changes, the present invention introduces the glycosylated proteome features in urine exosomes as an independent modal input channel. This type of information is at the protein expression level and can capture the abnormalities in the metabolic and signal transduction networks in the tumor microenvironment, forming a complement to the nucleic acid modality.
[0081] Generally, urine samples are widely used in exosome research due to their non-invasive, convenient collection, high storage stability and other characteristics. Exosomes are membranous vesicles with a diameter of about 30 - 150 nanometers, containing a variety of functional proteins, lipids and RNAs, and can stably carry the molecular characteristics of the source cells, especially suitable for the excavation of early tumor-related protein characteristics.
[0082] In this example, urine collection was completed using a sterilized sampler, and 50 mL of mid-stream morning urine was collected. As an option, the collected sample was stored briefly at 4 °C for a maximum of 2 hours and then transferred to -80 °C for long-term storage to maximize the inhibition of protease activity.
[0083] The isolation operation of exosomes involves multiple steps including centrifugation, filtration, and ultracentrifugation, as follows:
[0084] First, centrifuge at 2000 g for 10 minutes to remove large molecular debris;
[0085] Then, centrifuge at 10000 g for 20 minutes to remove non-exosomal particles;
[0086] Subsequently, filter through a 0.22 μm pore size polyethersulfone membrane to further remove large particles;
[0087] Finally, ultracentrifuge at 100000 g for 90 minutes to enrich exosomes and resuspend with sterile PBS.
[0088] The resuspension volume was controlled within 200–500 μL. Immediately after resuspension, lysis buffer containing TritonX-100, protease inhibitor, and reducing agent was added, and the lysis time was controlled at 15 minutes on ice. The lysate was quantified by the BCA method and entered the downstream proteome processing workflow.
[0089] To obtain glycosylation site information, the present invention adopts a combined strategy of affinity enrichment and liquid chromatography-mass spectrometry.
[0090] Specifically:
[0091] Use sugar chain affinity columns such as ConA / WGA / LCA to selectively enrich glycopeptide segments in the lysate;
[0092] After desalting and vacuum drying, it was redissolved in 0.1% formic acid aqueous solution;
[0093] Use an LC-MS / MS system for protein scanning, and the mass spectrometry parameters are set as follows:
[0094] Nano-flow chromatography flow rate: 300 nL / min;
[0095] Precursor ion mass range: 350–1800 m / z;
[0096] Collision-induced dissociation (CID) energy: 35 eV;
[0097] Data-dependent scanning mode, collecting up to 20 product ions;
[0098] Dynamic exclusion time: 30 s.
[0099] MS data is processed on the MaxQuant or ProteomeDiscoverer platform. Peptide identification uses the SEQUEST or Andromeda engine, and the database search parameters include:
[0100] Mass spectrometry tolerance: 10 ppm;
[0101] Targeted modification types: N-glycosylation, O-glycosylation;
[0102] Up to two mismatch sites are allowed.
[0103] In the present invention, only peptide signals with high recognition confidence are retained, and a peak intensity threshold is set to filter background noise. The specific threshold is set as:
[0104] I th = 10 5 ;
[0105] Where: I th Is the lower limit threshold of the signal intensity; the unit is the original ion count (IntensityCount) output by the mass spectrometry platform; all peptides below this threshold will be filtered.
[0106] All peptides that meet the conditions constitute a feature set Where p i Represents the i-th peptide with its original intensity S i . The following expression vector is constructed:
[0107] x Prot = [log(S1), log(S2),..., log(S m )];
[0108] Where: Is the protein expression vector; S i The original signal intensity of the i-th retained peptide; m is the total number of retained peptides, usually between 200 and 300; all values are transformed by the natural logarithm to facilitate the stability of the numerical distribution.
[0109] To eliminate batch bias between samples, further Z-score normalization is performed on each dimension of x Prot To obtain the normalized expression vector
[0110]
[0111] Where: Is the normalized expression value of the i-th peptide; x i Is the original logarithmic expression value of this peptide; μ i Is the mean value of all samples of this peptide in the training set; σ iis the standard deviation of all samples of this peptide segment in the training set.
[0112] After construction will be used as one of the model input modalities and jointly input into the multi-modal deep neural network together with the plasma cfDNA modality and the exhaled breath VOCs modality for subsequent cancer prediction, primary focus tissue localization and other tasks.
[0113] It should be noted that in some scenarios, such as the existence of missing peptide segments or low-quality data, zero-padding, mean imputation or nearest neighbor interpolation algorithms can be selected for processing to ensure that the feature dimensions of all samples are consistent and do not affect the tensor parallel training of the network structure.
[0114] In addition, in some advanced implementations, the site sequence context of glycopeptide segments, the hydrophilicity index of peptide segments or the predicted retention time can also be introduced as auxiliary variables to expand the modality information and enhance the feature space expressiveness of the model.
[0115] S4. Detect volatile organic compounds in the exhaled breath sample to obtain a feature vector composed of ion mobility spectrometry data;
[0116] After constructing the plasma cfDNA features and the urinary exosome glycoprotein expression features, the present invention further introduces the exhaled breath metabolic signal of the subject as the third modality feature source. This step aims to capture the abnormal release of gas molecules caused by tumor-related metabolic reprogramming and provide information supplementation at the end-product level for the model through the quantitative and structured expression of volatile organic compounds (VOCs) in the exhaled breath.
[0117] Generally, exhaled breath samples need to be detected by a high-sensitivity analytical instrument, and a gas chromatography-ion mobility spectrometry (GC-IMS) system is often selected to complete the detection. This system combines chromatographic separation and electric field migration mechanisms, and has the characteristics of fast speed, high throughput, label-free, etc., and is especially suitable for the analysis and identification of complex mixed VOCs in exhaled breath.
[0118] In this embodiment, exhaled breath samples are collected from the subject in the early morning, on an empty stomach and in a resting state. To ensure stable physiological conditions, interfering factors such as eating, strenuous exercise and smoking should be avoided. During the sampling process, three consecutive deep and even exhalations need to be completed, each lasting no less than 5 seconds. The sampling device automatically averages and combines the three samples to improve signal consistency.
[0119] As an implementation method, the core parameter settings of the used GC-IMS device are as follows:
[0120] Column type: MXT-5, 30m × 0.25mm;
[0121] Column temperature: Constant temperature at 60°C;
[0122] IMS drift tube length: 98 mm;
[0123] Electric field strength: 500 V / cm;
[0124] Gas carrier: High-purity nitrogen;
[0125] Ion source temperature: 75 °C;
[0126] Injection volume: 2.0 mL;
[0127] Nebulizer pressure: 2.5 bar.
[0128] The VOCs detection signal is output in the form of a two-dimensional spectrogram. The X-axis is the retention time t r , and the Y-axis is the drift time t d , and the Z-axis signal is the peak intensity I. The three together form a complete single VOCs description vector.
[0129] In a specific implementation, the present invention presets 100 VOCs target peaks as fixed vector templates, and the samples are strictly aligned. Each target peak i corresponds to a unique set of three-dimensional data features Where: is the retention time, representing the elution time of the compound in the GC column, with the unit of second (s); is the drift time, representing the time required for the compound to migrate in the IMS, with the unit of millisecond (ms); I (i) is the peak intensity, which is the maximum response value of the ionization peak of the compound, with the unit of millivolt (mV) or instrument unit (a.u.); i = 1, 2,..., 100 is the fixed VOC number, and the order among samples is fixed.
[0130] For some undetected peaks, or peaks with signals lower than the noise lower limit I noise = 10 -2 mV, the system automatically sets them to 0 to avoid noise pollution. For unified processing structure, the 100 sets of three-dimensional data are constructed into a sample feature vector in the way of "sequential splicing":
[0131]
[0132] Where: is the spliced VOCs feature vector; every three dimensions represent the chromatographic drift characteristics of a VOCs in turn.
[0133] As a commonly used preprocessing method, the above-mentioned original feature vector will be processed by dimension-wise normalization to eliminate the differences in dimension and numerical distribution among different VOCs.
[0134] The normalization operation is defined as follows:
[0135]
[0136] Wherein: is the j-th dimensional feature after normalization; x j is the original j-th dimensional feature; μ j is the mean of the j-th dimensional feature in the training set; σ j is the standard deviation of the j-th dimensional feature in the training set; j = 1, 2,..., 300.
[0137] After normalization, all features are centered around 0 and scaled by a unit standard deviation, which is convenient for the neural network modeling to converge stably.
[0138] To address the problems of chromatographic peak overlap and drift offset, the system introduces a reference standard gas mixture in each analysis to correct the peak position and control the position drift error |Δt r | < 0.3 s, |Δt d | < 0.2 ms. Those that do not meet the error range are automatically set to zero.
[0139] The finally normalized feature vector is used as the exhalation modality input, together with the cfDNA feature x cfDNA and the protein expression feature and input into the multi-modal neural network together.
[0140] In addition, in some edge deployment environments, simplified acquisition devices (such as electronic nose + micro IMS) can be used to capture VOCs changes in real time and dynamically update the model output to improve the response efficiency and adaptability. However, the core feature extraction structure is still based on the input construction strategy of fixed peak + standardized vector.
[0141] S5. Preprocess and vectorize the obtained multi-modal features, and input them into the multi-modal deep learning model, which constructs the tissue association structure among cancer types based on the graph neural network;
[0142] In the multi-modal cancer prediction system of the present invention, the cfDNA fragmentomics features, exosome glycoprotein expression features, and exhaled breath VOCs ion mobility features provide biological information support from the three dimensions of genetics, protein, and metabolism respectively. In order to effectively integrate the three types of high-dimensional heterogeneous data, this step introduces a unified preprocessing and deep fusion mechanism, and completes the tissue-level association modeling through the graph neural network model.
[0143] Generally, there are significant differences in the feature dimensions, numerical scales, and distribution patterns of different modalities. Direct splicing or equal-weight input will lead to information imbalance. The present invention adopts the method of feature remapping + attention fusion + graph-aware propagation to achieve structured and multi-path modeling.
[0144] In a specific implementation, the three types of features obtained in steps S2 - S4 are respectively expressed as:
[0145] is the feature vector of plasma cfDNA;
[0146] is the feature vector of urinary exosome protein expression;
[0147] is the feature vector of exhaled VOCs spectrum.
[0148] First, perform a unified linear mapping on the features of each modality:
[0149] z cfDNA = ReLU(W1x cfDNA + b1);
[0150] z Prot = ReLU(W2x Prot + b2);
[0151] z VOC = ReLU(W3x VOC + b3);
[0152] where: is the unified embedding vector of each modality; is the projection matrix corresponding to the modality; is the bias term; d embed is set to a unified embedding dimension, such as 128 or 256; ReLU(·) is the element-wise activation function.
[0153] As an option, if there is a modality missing (e.g., a sample has no exhalation data), then adopt the modality masking mechanism, that is:
[0154]
[0155] where: is the modality availability mask; is the modality vector finally actually used for fusion.
[0156] Next, concatenate the representation vectors of the three modalities:
[0157]
[0158] To prevent overfitting, add the Dropout random inactivation mechanism:
[0159] Z drop = Dropout(Z concat , p drop );
[0160] where: p dropis the inactivation ratio, generally set to 0.2 - 0.5; Dropout(·) randomly discards some dimensions in the input with a certain probability.
[0161] To guide the network to automatically identify important modal dimensions, the present invention further introduces a modal-level attention mechanism. Let the final fusion representation be:
[0162]
[0163] Where: is the modal attention weight; is the attention score; is the learning parameter.
[0164] The fused vector will be used as the input and fed into the graph neural network module. The graph model constructs a heterogeneous structure with cancer tissue types as nodes, and establishes a coupled graph of metabolic pathways between tissues G=(V, E), where:
[0165] V = {v1, v2,..., v n} is the set of nodes in the graph, representing different cancer tissue types;
[0166] E is the set of edges, connecting any two cancer tissue types v i , v j ;
[0167] n is the number of cancer tissue types, set to 12.
[0168] The edge weights are defined by the overlap degree of metabolic pathways:
[0169]
[0170] Where: T i is the set of metabolic pathways corresponding to the i-th cancer tissue type; |·| is the size of the set; w ij ∈[0, 1] is the edge weight value between nodes v i and v j ; The edge weights form the graph adjacency matrix for GNN propagation.
[0171] The graph neural network adopts a hierarchical design. It is recommended to use the GAT structure, and the node feature propagation form is:
[0172]
[0173] Where: is the representation of node i at the l-th layer; W (l) is the shared weight matrix of this layer; is the attention weight that node i receives from node j in the l-th layer; σ is the activation function, such as ELU; is the adjacency set of node i; l is the number of network layers, generally 2 - 3.
[0174] To prevent the graph structure from being too dense, the present invention introduces an edge weight sparsification mechanism, which only retains the k non-zero edges with the top ij weights. For example, the top 30% of the largest edge weights are retained, and the remaining w
[0175] is set to 0, improving the calculation efficiency and information discrimination. i In the initialization stage, each node v in the graph initializes the vector fused which can be taken as a unit vector, organizational prior features, or obtained from the fused representation h
[0176]
[0177] Finally, all node representations are used to output the multi-cancer type recognition results and tissue probability maps, and enter the next step S6 to perform classification and regression tasks.
[0178] S6. Output the prediction results including the cancer probability, primary site tissue localization information, and targeted drug sensitivity indicators through the deep learning model;
[0179] After completing the multi-modal feature preprocessing, graph structure construction, and deep embedding propagation, the system described in the present invention enters the output stage. This stage relies on the fused representation of the graph neural network, and sets parallel prediction head networks for different task objectives to achieve the three core function outputs of cancer probability assessment, primary site tissue localization, and targeted drug sensitivity prediction. Generally, the output structure of the present invention is based on a shared encoder and adopts a multi-task learning architecture, that is, inferring the results of the three tasks simultaneously from the same graph embedding representation This structure can improve the cooperation between tasks and reduce the risk of overfitting.
[0180] Cancer probability prediction module:
[0181] In this embodiment, the cancer probability is obtained through a simple fully connected + Sigmoid layer, and the form is as follows:
[0182]
[0183] where: p ∈ (0, 1) represents the predicted probability that the subject under examination is a cancer individual; is the embedding feature output by the GNN; is the classifier weight vector; is the bias term; e is the base of the natural logarithm; the Sigmoid function ensures that the output value is a probability-type result.
[0184] The classification output is trained using binary cross-entropy loss and is applicable to the cancer risk modeling scenario.
[0185] Primary tumor tissue localization module:
[0186] Specifically, the output of the primary tumor tissue localization module is a vector where n is the predefined number of tissues (for example, 12 cancer-related tissues such as lung, liver, pancreas, etc.). After normalizing this vector with the Softmax function, a probability distribution is output:
[0187]
[0188] where: P i ∈(0,1) represents the predicted probability that the i-th tissue is the primary tumor; is the original logit value of the i-th dimension in the model output vector; is the total number of all optional tissue categories; e is the base of the natural logarithm; is the output that constitutes the complete probability distribution.
[0189] The Softmax input vector z comes from the following transformation:
[0190] z = W org ·h + b org ;
[0191] where: is the weight of the fully connected layer for the tissue localization task; is the bias term.
[0192] The tissue localization task training uses multi-class cross-entropy loss and supports the output of top-k prediction results (such as top-3 tissue possibilities).
[0193] Targeted drug sensitivity prediction module:
[0194] In a possible implementation, for the sensitivity prediction of each targeted drug k, two values are output:
[0195] The predicted IC50 concentration μ k , representing the concentration required for the drug to effectively inhibit cancer cell proliferation;
[0196] The corresponding predicted uncertainty σ k , used to estimate the confidence of the model in this prediction. The prediction process is as follows:
[0197]
[0198] where: is the predicted IC50 value corresponding to the k-th drug; is the predicted standard deviation corresponding to the k-th drug; is the corresponding weight vector; is the corresponding bias; the Softplus function log(1 + exp(·)) is used to ensure non - negativity of the standard deviation; the unit of IC50 is often nM, and the results can be logarithmically transformed to stabilize the distribution.
[0199] The uncertainty propagation mechanism is used to simulate the stability of the model when extrapolating data distribution. The standard deviation term can be used both to construct confidence intervals and for sample importance ranking. During the training phase, the drug prediction regression task is optimized using the Negative Log - Likelihood (NLL) loss, defined as follows:
[0200]
[0201] where: y k is the true IC50 label of the k-th drug; the definitions of the remaining symbols are the same as above; the loss function penalizes samples with large prediction errors and high confidence.
[0202] Multi - task joint training mechanism:
[0203] During the overall training process, the losses of the three tasks are combined in a weighted manner, defined as:
[0204]
[0205] where: is the binary classification loss (cross - entropy) of the canceration probability; is the multi - classification cross - entropy of the primary tumor tissue location; is the sum of the regression losses for all drugs; is the task loss weight coefficient; the weights can be set empirically or automatically adjusted as learnable parameters.
[0206] In some embodiments, the output module can also be connected to a post - processing system to perform extended functions such as top - k screening, confidence score quantization, and multi - time - point sliding window stability evaluation on the prediction results, for clinical interpretation or deployment to mobile devices.
[0207] Please refer to Figure 2 , the present invention also provides a multi - cancer early prediction system based on multi - modal AI, including:
[0208] The cfDNA fragment feature acquisition module is used to collect and analyze the circulating free DNA (cfDNA) in plasma samples, extract its end - sequence preference, fragment length distribution, and nucleosome positioning features related to chromatin state. The original fragment data is obtained through high - throughput sequencing technology, and combined with reference genome annotation, a structural feature matrix for characterizing apoptosis patterns and tissue sources is generated;
[0209] Exosome proteome detection module, which is used to isolate and enrich exosomes from urine samples, and adopts a method of combining mass spectrometry (such as LC-MS / MS) with glycosylation modification enrichment to extract and quantitatively analyze the expression profiles of glycosylated proteins in exosomes. The obtained proteome data can reflect the changes in tissue microenvironment and metabolic status, and provide biomarker support for early cancer screening and traceability analysis;
[0210] Exhaled breath VOCs analysis module, which is used to detect volatile organic compounds (VOCs) in exhaled breath samples through ion mobility spectrometry (IMS) technology and extract the migration time distribution map under the electric field. Combining pattern recognition algorithms to extract high-dimensional fingerprint features from two-dimensional spectra to reflect metabolic abnormal states such as in vivo oxidative stress and inflammation;
[0211] Feature preprocessing module, which is used to perform normalization, missing value filling, principal component compression and variability analysis on three types of modal data of cfDNA, proteome and VOCs respectively, and construct a unified embedded feature expression vector. This module adopts batch normalization, min-max normalization or Z-score transformation, and automatically aligns the feature dimensions of different modalities according to sample labels to ensure the compatibility of subsequent models;
[0212] Graph neural network modeling module, which is used to construct a graph structure reflecting the connection relationship of metabolic pathways between different tissues, map three types of modal data to graph node attributes, and adopt methods such as graph convolution (GCN), graph attention mechanism (GAT) or message passing neural network (MPNN) for joint modeling. Learn the pathological influence paths between different tissues through the structure information propagation mechanism to improve the modeling ability of complex carcinogenesis processes;
[0213] Output module, which is used to synthesize the graph embedding results and output three key prediction results: the probability of individual carcinogenesis, the possible location of the primary tissue, and the prediction of sensitivity to targeted therapy. Adopt a multi-task deep network structure for classification and regression output respectively, and at the same time provide an uncertainty assessment of the results to support clinical auxiliary diagnosis and drug selection.
[0214] Although the embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A multi-cancer early prediction method based on multi-modal AI, characterized in that, It includes the following steps: S1. Obtain multiple body fluid samples from the same individual to be examined, where the body fluid samples include plasma samples, urine samples, and exhaled breath samples; S2. Sequence the plasma sample to obtain the fragmentomics characteristics of cfDNA, where the fragmentomics characteristics include fragment end sequence frequency information and nucleosome positioning density information; S3. Perform proteomic detection on exosomes in the urine sample to obtain the glycosylated protein expression characteristics in the exosomes; S4. Perform volatile organic compound detection on the exhaled breath sample to obtain a feature vector composed of ion mobility spectrometry data; S5. Preprocess and vectorize the obtained multi-modal characteristics, and input them into a multi-modal deep learning model, where the model constructs an inter-cancer tissue association structure based on a graph neural network; S6. Output prediction results including canceration probability, primary focus tissue localization information, and targeted drug sensitivity indicators through the deep learning model.
2. The multi-cancer early prediction method based on multi-modal AI according to claim 1, wherein The fragment end sequence frequency information is composed of a four-base sequence at the 5' end of the cfDNA fragment, the statistical frequency vector dimension is 256 dimensions, and the frequency satisfies the normalization constraint.
3. The multi-cancer early prediction method based on multi-modal AI according to claim 1, wherein The nucleosome positioning density information is obtained by statistically determining whether the midpoint position of the cfDNA fragment is located in a predefined tissue-specific chromatin open region, and the chromatin open region is from a reference database.
4. The multi-cancer early prediction method based on multi-modal AI according to claim 1, wherein The glycosylated protein expression characteristics are a combination of polypeptide segment expression intensity values with an intensity greater than a preset threshold, the threshold is 10 to the fifth power, and the expression intensity values are logarithmically normalized.
5. The multi-cancer early prediction method based on multi-modal AI according to claim 1, wherein The feature vector of the exhaled breath sample is sequentially spliced by the retention time, drift time, and peak intensity of multiple predefined volatile organic compounds, and the number of volatile organic compounds is fixed at 100.
6. The multi-cancer early prediction method based on multi-modal AI according to claim 1, characterized in that, The graph neural network in the deep learning model is trained by constructing a cancer tissue graph. The nodes in the graph represent different cancer tissue types, and the weight of the edge is determined by the degree of overlap of the corresponding tissue metabolic pathways, and the degree of overlap is calculated in the form of the intersection-to-union ratio of the pathway sets.
7. The multi-cancer early prediction method based on multi-modal AI according to claim 1, characterized in that, The deep learning model includes a feature screening module, which calculates attention weights for each modal feature through an attention mechanism and selects features whose weight values meet the set conditions to participate in modeling.
8. The multi-cancer early prediction method based on multi-modal AI according to claim 1, characterized in that, The primary focus tissue localization information is the probability distribution information output for multiple predefined anatomical tissues, and the probability is calculated through the Softmax function. The Softmax function is: Among them, P i represents the probability that the i-th tissue is the primary focus, and z i is the unnormalized predicted value corresponding to the i-th tissue output by the model, and n is the total number of tissues participating in the prediction.
9. The multi-cancer early prediction method based on multi-modal AI according to claim 1, wherein, The targeted drug sensitivity indicator is the IC50 value and its standard deviation, and the value is output by the trained regression sub-network, and the standard deviation is calculated through the model uncertainty propagation mechanism.
10. A multi-cancer early prediction system based on multi-modal AI, which is applied to the multi-cancer early prediction method based on multi-modal AI according to any one of claims 1-9, characterized in that, It includes: A cfDNA fragment feature acquisition module for collecting and extracting the end sequence features and nucleosome positioning information of cfDNA in plasma; An exosome proteome detection module for obtaining the glycosylated protein expression characteristics of exosomes from urine samples; An exhaled breath VOC analysis module for obtaining the ion mobility spectrometry characteristics of exhaled breath samples; A feature preprocessing module for normalizing the three-modal data and constructing feature expressions; A graph neural network modeling module for establishing a graph structure based on the relationship of inter-tissue metabolic pathways and performing multimodal joint prediction; An output module for outputting the canceration probability, the primary tissue localization, and the treatment sensitivity results.