Global metabolome profiling method, lung cancer prediction model construction method and device
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SUN YAT SEN UNIVERSITY CANCER CENTER (CANCER HOSPITAL AFFILIATED TO SUN YAT SEN UNIVERSITY CANCER RESEARCH INSTITUTE OF SUN YAT SEN UNIVERSITY)
- Filing Date
- 2024-01-03
- Publication Date
- 2026-07-31
AI Technical Summary
The existing deep learning analysis methods have low resolution in lung cancer prediction, ignoring the peak information of metabolic signal, resulting in inaccurate quantitative and qualitative inaccurateness, making it difficult to accurately identify lung cancer-related metabolites.
The chromatography-mass spectrometry raw data of serum samples was obtained by using liquid chromatography-mass spectrometry technology, and the convolution operation was performed in the three-dimensional point cloud space was downsampled into a two-dimensional matrix. Mask gallery was constructed and perturbation prediction was performed. The global metabolic feature spectrum was inferred in convolutional neural network, and the mask ratio and contribution calculation function of the Mask graph were used to improve resolution.
High-resolution global metabolic feature spectrum inference is achieved, which improves the accuracy and interpretation of characteristic importance of lung cancer prediction, especially the ability to distinguish between lung cancer, benign lung nodules and healthy states.
Smart Images

Figure CN122498006A_ABST
Abstract
Description
Global metabolic profile inference method, lung cancer prediction model construction method and device Technical Field
[0001] The present invention relates to the field of metabolomics data analysis, and in particular to a high-resolution global metabolic profile inference method, a lung cancer prediction model construction method and a device. Background Art
[0002] Lung cancer is one of the most common malignant tumors worldwide and has one of the highest mortality rates. The latest data released in my country in 2018 showed that there were 2.1 million new cases of lung cancer in my country, ranking first among malignant tumors, accounting for 18.4% of all cancer deaths (ranking first), and 1.8 million deaths (ranking first), accounting for more than a quarter of all deaths from malignant tumors. Early diagnosis is crucial for improving the treatment, prognosis, and survival rates of cancer patients. Currently, the diagnosis of lung cancer relies on invasive puncture and bronchoscopy to obtain tissue or cells for pathological examination. CT imaging is the primary auxiliary diagnostic tool, but the differential diagnosis of small lung nodules from benign or malignant lesions remains challenging. Serum tests for lung cancer, such as carcinoembryonic antigen, keratin fragments, and squamous cell carcinoma antigen, can be used as auxiliary diagnostic tools or for follow-up monitoring of lung cancer, but their sensitivity and specificity still need to be improved.
[0003] In recent years, thanks to the rapid development of mass spectrometry technology, the application of metabolomics in disease diagnosis has gradually attracted widespread attention. Metabolomics is a new discipline that conducts qualitative and quantitative analysis of small-molecule metabolites with a relative molecular weight of less than 1000 in the body. The metabolome refers to all low-molecular-weight metabolites produced by an organism or cell during a specific physiological period. Many life activities within the cell occur at the metabolite level. Therefore, the detection and identification of the metabolome can determine the pathophysiological state of the body and potentially identify markers related to its pathogenesis. Therefore, metabolomics has broad application prospects in clinical medicine. Metabolites in serum are stable and quantifiable, which provides a possibility for non-invasive diagnosis in clinical applications. Currently, studies have reported the use of serum metabolite markers to diagnose lung cancer or to differentiate between malignant lung tumors and benign nodules.
[0004] Liquid chromatography-mass spectrometry (LC-MS) is a mainstream method for non-targeted metabolomics research. Due to its high retention, high sensitivity, and high resolution, it is widely used in fields such as pharmaceutical research, biology, environmental monitoring, and food safety. Inferring disease-related global metabolic profiles based on LC-MS primarily relies on a data analysis pipeline that combines qualitative and quantitative analysis. However, the complex structure of raw LC-MS data makes analysis very challenging. Deep learning has been proposed to process raw LC-MS data, as described in Chinese patent CN113554176A. However, existing deep learning analysis methods suffer from several major issues. First, during peak extraction and metabolite identification, only retention time (RT), mass-to-charge ratio (m / z), and peak area of metabolic signals are considered as quantitative indicators, ignoring important information such as peak shape, resulting in inaccurate quantitative results. Furthermore, due to database limitations and the presence of adducts, qualitative identification of metabolites is also inaccurate. Secondly, the current deep learning-based inference technology is based on a gradient-based class activation function approach; it calculates the product between the feature map of the network's last convolutional layer and the gradient of the target category to obtain category-specific activation maps. These activation maps can show which locations in the image are important for the network to judge as a specific category. However, the size of the resulting activation map is consistent with the size of the last convolutional layer, which is much smaller than the size of the input data, making it difficult to obtain a high-resolution feature spectrum. In lung cancer prediction analysis based on metabolic signals, a difference of only 1 ppm often represents a different signal. Therefore, to improve the accuracy of lung cancer prediction, a higher-resolution and clearer global metabolic feature spectrum inference method is urgently needed.
[0005] Summary of the Invention
[0006] In order to overcome the problem that existing metabolic profile inference technology has low resolution and is difficult to apply to the prediction of cancer diseases such as lung cancer, the present invention provides a method and device for constructing a lung cancer prediction model based on global metabolic profiles based on deep learning.
[0007] To achieve the above objectives, the present invention provides a method for constructing a lung cancer prediction model based on a global metabolic profile, which comprises:
[0008] Serum samples from people of different genders and age groups were obtained and divided into three groups based on lung cancer, benign lung nodules, and health;
[0009] Metabolites in each serum sample were extracted and detected using liquid chromatography-high-resolution mass spectrometry to obtain chromatography-mass spectrometry raw data;
[0010] The global metabolic profile of metabolites in each serum sample is inferred based on the raw chromatography-mass spectrometry data, which includes using a filter to perform a convolution operation on data converted from the raw chromatography-mass spectrometry data in a three-dimensional point cloud space to downsample the data into a two-dimensional matrix; constructing a mask image with multiple mask ratios and a size consistent with the size of the two-dimensional matrix based on the two-dimensional matrix to form a mask image library; selecting multiple target mask images from the mask image library, multiplying each target mask image by the two-dimensional matrix, and then inputting the result into a trained mass spectrometry inference model based on a convolutional neural network for perturbation prediction, and obtaining the retention information of the sample corresponding to the currently input two-dimensional matrix based on the perturbation prediction to infer the global metabolic profile of metabolites in each serum sample;
[0011] Based on the inferred global metabolic profile of metabolites in each serum sample, the selected classifier was trained to obtain a lung cancer prediction model.
[0012] According to one embodiment of the present invention, the extraction of metabolites in a serum sample includes: extracting and obtaining a metabolite dry sample using a liquid-liquid extraction method; and redissolving and centrifuging the metabolite dry sample to prepare a sample to be tested.
[0013] According to one embodiment of the present invention, the steps for extracting metabolites in serum samples are as follows: after the serum sample is completely thawed on ice, 50 μL is taken into a 1.5 mL EP tube, 225 μL of frozen methanol is added, and vortexing is performed for 30 seconds; then 750 μL of frozen methyl tert-butyl ether is added, vortexing is performed for 30 seconds, and then shaking is performed on ice at 400 rpm for 1 hour; then 188 μL of pure water is added and vortexing is performed for 1 minute; centrifugation is performed at 15,000 × g for 10 minutes at 4°C; after centrifugation, 125 μL of the lower supernatant is taken from two tubes respectively into new EP tubes and dried using a vacuum freeze dryer. All serum metabolite dry samples are stored in a -80°C refrigerator before testing; the serum metabolite dry extract is re-dissolved, and after centrifugation, the supernatant is taken to prepare the test sample for detection by liquid chromatography-high-resolution mass spectrometry.
[0014] According to one embodiment of the present invention, a filter is used to perform a convolution operation on the chromatographic-mass spectrometric raw data of each serum sample in a three-dimensional point cloud space to downsample the three-dimensional chromatographic-mass spectrometric raw data into a two-dimensional matrix, including:
[0015] Convert the original data into .mzml format;
[0016] Set the starting retention time T0, ending retention time Te, starting mass-to-charge ratio R0, and ending mass-to-charge ratio Re;
[0017] Within the retention time range and mass-to-charge ratio range, three-dimensional points are sampled from the .mzml format, and each three-dimensional point includes the retention time t, mass-to-charge ratio r, and ion intensity i of the chromatographic-mass spectrometric raw data;
[0018] Based on the preset convolution kernel size and span size, the maximum pooling convolution function is used to perform convolution operations on the three-dimensional point space cloud using filters. The maximum value in the convolution window is used as the pooling result, thereby downsampling the features of the three-dimensional point cloud space to a two-dimensional matrix.
[0019] According to one embodiment of the present invention, the step of inferring the global metabolic profile of metabolites in each serum sample based on the perturbation prediction of the mass spectrometry inference model of the convolutional neural network includes:
[0020] Multiply each target mask image by a two-dimensional matrix and input it into a trained mass spectrometry inference model based on a convolutional neural network to obtain the perturbed predicted probability corresponding to each target mask image;
[0021] Based on the predicted probability after perturbation of multiple target mask images, the contribution score of each molecular feature in the serum sample corresponding to the two-dimensional matrix is calculated to form a contribution heat map;
[0022] According to the network structure of the mass spectrometry inference model, a mapping function t=map1(x), r=map2(y) is obtained; wherein t is the retention time and r is the mass-to-charge ratio;
[0023] Map the two-dimensional coordinates (x, y) of the feature contribution heat map to retention time and mass-to-charge ratio;
[0024] Molecular features with contribution scores less than the score threshold and ion intensities less than the intensity threshold were filtered out to obtain the retained features;
[0025] Key metabolites were screened based on the retained features, and correlation calculations were performed to infer the metabolic markers and metabolic network patterns of the serum samples corresponding to each two-dimensional matrix, thereby generating a global metabolic profile of the metabolites in each serum sample.
[0026] According to one embodiment of the present invention, after the perturbation prediction probability corresponding to each target Mask image is obtained, the contribution score of each molecular feature is obtained using the following contribution calculation function: S=sum(pred i (x(t,r)*mask i )·mask i ) / (n*mask i _ratio)
[0027] Among them, mask i is the i-th target Mask image, pred i(x(t,r) is the predicted probability of the perturbation of the i-th target Mask image, x(t,r) is the signal strength corresponding to the (t, r) position, n is the total number of target Mask images, mask i _ratio is the mask ratio of the i-th target Mask image. The mask ratio refers to the ratio of 0 values to 1 values in the Mask image.
[0028] According to one embodiment of the present invention, a mask image library is pre-generated and a target mask image is extracted based on boosting for calculating the contribution scores of molecular features in different samples.
[0029] In another aspect, the present invention also provides a device for constructing a lung cancer prediction model based on a global metabolic profile, comprising a sample acquisition module, a metabolite extraction module, a profile inference module, and a model construction module. The sample acquisition module collects serum samples from people of different genders and age groups and categorizes them into three groups based on lung cancer, benign lung nodules, and healthy individuals. The metabolite extraction module extracts metabolites from each serum sample and detects them using liquid chromatography-high-resolution mass spectrometry to generate raw chromatography-mass spectrometry data. The signature spectrum inference module infers the global metabolic signature of metabolites in each serum sample based on the raw chromatography-mass spectrometry data. This includes using a filter to perform a convolution operation on the data converted from the raw chromatography-mass spectrometry data in a three-dimensional point cloud space to downsample the data into a two-dimensional matrix; constructing a mask image with multiple mask ratios and a size consistent with the two-dimensional matrix based on the two-dimensional matrix to form a mask image library; selecting multiple target mask images from the mask image library, multiplying each target mask image with the two-dimensional matrix, and inputting the result into a trained mass spectrometry inference model based on a convolutional neural network for perturbation prediction. Based on the perturbation prediction, the retained information of the sample corresponding to the current input two-dimensional matrix is obtained to infer the global metabolic signature of the metabolites in each serum sample. The model construction module trains the selected classifier based on the inferred global metabolic signature of the metabolites in each serum sample to obtain a lung cancer prediction model.
[0030] On the other hand, the present invention also provides a lung cancer prediction model, which is constructed using the above-mentioned method for constructing a lung cancer prediction model based on a global metabolic profile.
[0031] On the other hand, the present invention also provides a lung cancer prediction device comprising the above lung cancer prediction model.
[0032] In summary, in the method for constructing a lung cancer prediction model based on a global metabolic profile provided by the present invention, metabolites in serum samples are detected using a high-retention, high-resolution liquid chromatography-mass spectrometry technique to obtain three-dimensional chromatographic-mass spectrometry raw data. In the inference of the global profile based on the chromatographic-mass spectrometry raw data, a filter is used to perform a convolution operation on the data converted based on the chromatographic-mass spectrometry raw data in the three-dimensional point cloud space to downsample the data into a two-dimensional matrix; the filter formed by the convolution kernel is a small matrix or vector, which minimizes the sampling spacing to obtain a higher-resolution downsampled two-dimensional matrix, thereby effectively avoiding the false sparsity noise caused by the existing sliding window sampling due to excessive intervals or missing sampling points. Furthermore, after obtaining the two-dimensional matrix, a mask image is set to cover some pixels and the output change of the mass spectrometry inference model after being perturbed by each target mask image is calculated, thereby analyzing the contribution of each molecular feature and obtaining the global metabolic profile of different samples based on the retained features. The perturbation-based feature contribution method can generate a global profile at single-pixel resolution, providing a more nuanced interpretation of feature importance for individual sample predictions. The minimum sampling spacing and the single-pixel resolution of the global profile ensure that the inferred global metabolic profile retains high information and has extremely high resolution, providing the necessary conditions for lung cancer prediction.
[0033] In order to make the above and other objects, features and advantages of the present invention more clearly understood, preferred embodiments are given below with reference to the accompanying drawings for detailed description. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] FIG1 is a schematic flow chart of a method for constructing a lung cancer prediction model based on a global metabolic profile according to a first embodiment of the present invention.
[0035] Figure 2 shows the total ion currents in positive and negative modes when metabolites were detected using liquid chromatography-high-resolution mass spectrometry.
[0036] FIG3 is a schematic diagram showing the process of downsampling the three-dimensional chromatographic-mass spectrometric raw data into a two-dimensional matrix in step S30 in FIG1 .
[0037] FIG4 is a schematic diagram showing the process of obtaining a global metabolic profile based on a mass spectrometry inference model in step S30 in FIG1 .
[0038] Figure 5 shows a schematic diagram of the global metabolic profile inference of lung cancer, benign lung nodules and healthy people.
[0039] FIG6 shows a comparison of subject curves between the global metabolic profile inference method established by the present invention and other inference methods.
[0040] FIG7 is a statistical diagram showing the accuracy of lung cancer prediction using a lung cancer prediction model based on a global metabolic profile.
[0041] FIG8 is a schematic diagram showing the structure of a device for constructing a lung cancer prediction model based on a global metabolic profile according to the first embodiment of the present invention.
[0042] FIG9 is a schematic diagram showing the flow of the global metabolic profile inference method provided in Example 2 of the present invention.
[0043] FIG10 is a schematic flow chart showing step S2 in FIG9 .
[0044] FIG11 is a schematic flow chart showing step S4 in FIG9 .
[0045] FIG12 shows the structural schematic of the global metabolic profile inference device provided in the second embodiment of the present invention. DETAILED DESCRIPTION
[0046] Example 1
[0047] Existing global metabolic profile inference methods suffer from significant information loss during raw data processing due to neglect of metabolic signal peaks. Furthermore, metabolic profiles derived through deep learning using gradient-based class activation functions suffer from low resolution, making it difficult to accurately characterize different global metabolic states. This seriously impacts the accuracy of these prediction models for lung cancer. In light of this, this embodiment provides a method for constructing a lung cancer prediction model based on a global metabolic profile with high information retention and resolution.
[0048] As shown in FIG1 , the method for constructing a lung cancer prediction model based on a global metabolic profile provided in this embodiment includes obtaining serum samples from people of different genders and age groups and dividing the obtained serum samples into three groups based on lung cancer, benign lung nodules, and health (step S10). Metabolites in each serum sample are extracted and detected using liquid chromatography-high-resolution mass spectrometry to obtain chromatography-mass spectrometry raw data (step S20). The global metabolic profile of the metabolites in each serum sample is inferred based on the chromatography-mass spectrometry raw data (step S30). This step includes using a filter to convolve the data converted based on the chromatography-mass spectrometry raw data in a three-dimensional point cloud space to downsample the data into a two-dimensional matrix. Based on the two-dimensional matrix, a mask image with multiple mask ratios and a size consistent with the two-dimensional matrix is constructed to form a mask image library; multiple target mask images are selected from the mask image library, each target mask image is multiplied by the two-dimensional matrix and then input into a trained mass spectrometry inference model based on a convolutional neural network for perturbation prediction. Based on the perturbation prediction, the retention information of the sample corresponding to the currently input two-dimensional matrix is obtained to infer the global metabolic profile of metabolites in each serum sample. Based on the inferred global metabolic profile of metabolites in each serum sample, the selected classifier is trained to obtain a lung cancer prediction model (step S40).
[0049] The following detailed description of the lung cancer prediction model construction method provided by this embodiment is provided in conjunction with Figures 1 to 7. The construction method begins in step S10. In this embodiment, a total of 794 serum samples were collected from men and women of different ages. The serum samples were divided into three groups based on lung cancer, benign lung nodules, and healthy individuals. Within each serum sample group, the samples were matched by gender and age group. Serum samples include serum, plasma, and whole blood.
[0050] After obtaining the serum sample set, step S201 within step S20 is performed to extract metabolites from the serum samples using a three-phase extraction method with methyl tert-butyl ether:methanol:water (10:3:2.5, v / v / v). Metabolites from the serum samples include one or more chemical substances detectable by liquid chromatography-mass spectrometry, such as small molecule organic acids, amino acids, and carnitine. The specific procedure is as follows: After the serum sample is completely thawed on ice, 50 μL is transferred to a 1.5 mL EP tube, 225 μL of chilled methanol is added, and the tube is vortexed for 30 seconds. 750 μL of chilled methyl tert-butyl ether is then added, vortexed for 30 seconds, and shaken on ice at 400 rpm for 1 hour. 188 μL of purified water is then added and vortexed for 1 minute. The tube is centrifuged at 15,000 × g for 10 minutes at 4°C. After centrifugation, 125 μL of the lower layer of the supernatant is transferred from two tubes to new EP tubes and dried using a vacuum freeze dryer. All serum metabolite dry extracts are stored in a -80°C freezer before testing. The serum sample metabolite extraction step also includes re-dissolving the serum metabolite dry extract, centrifuging it, and removing the supernatant to prepare the test sample. Specifically, after obtaining the serum sample metabolic dry extract, 120 μL of redissolution solvent (acetonitrile: water = 4: 1) was added thereto, vortexed for 5 minutes, centrifuged at 15000 × g for 10 minutes at 4 ° C, and 100 μL of supernatant was taken to the inner liner tube to prepare the test sample. In addition, 10 μL of each of the lung cancer, benign pulmonary nodules and healthy human serum test samples was taken, vortexed and mixed, and then QC samples were made. Although this embodiment illustrates the metabolite extraction of serum samples using liquid-liquid extraction as an example, the present invention does not impose any restrictions on the metabolite extraction method.
[0051] After obtaining the serum metabolite sample to be tested, step S202 in step S20 is performed to detect the metabolites using liquid chromatography-high resolution mass spectrometry to obtain chromatography-mass spectrometry raw data. Specifically, the detection method is as follows:
[0052] ①Set up liquid chromatography conditions
[0053] Chromatographic column: BEH Amide (100×2.1mm, 1.7μm).
[0054] Mobile phase: In positive mode, phase A was acetonitrile:water = 95:5 (10 mM ammonium acetate, 0.1% formic acid), and phase B was acetonitrile:water = 50:50 (10 mM ammonium acetate, 0.1% formic acid); in negative mode, phase A was acetonitrile:water = 95:5 (10 mM ammonium acetate, pH adjusted to 9.0 with ammonia), and phase B was acetonitrile:water = 50:50 (10 mM ammonium acetate, pH adjusted to 9.0 with ammonia). The total ion currents in the positive and negative modes are shown in Figure 2.
[0055] The elution gradient is shown in Table 1 below:
[0056] Table 1: LC-HRMS mobile phase elution gradient
[0057] ②Mass spectrometry conditions
[0058] Qualitative analysis was performed using a Q Exactive mass spectrometer (Thermo Fisher Scientific, USA), using an electrospray ionization (ESI) source in both positive and negative full scan modes (Fullscan) and data-dependent scan (ddMS2) modes. The spray voltage was +3800 / -3200 V; the nebulizer temperature was 350°C; high-purity nitrogen was used as the sheath gas and auxiliary gas, with parameters set to 40 arb and 10 arb, respectively; the ion transfer tube temperature was 320°C; the mass scan range was 70–1050 m / z; the primary scan resolution was 70,000 FWHM, and the secondary scan resolution was 35,000 FWHM.
[0059] ③Injection order
[0060] Before each test, six QC samples were injected to stabilize the detection system. Serum samples were randomly injected, with one QC sample inserted for every 10 serum injections. The first and last injections in the test sequence were both QC samples. Finally, the QC samples underwent ddMS2 full and segmented scans to generate the chromatographic-mass spectrometric raw data for compound identification.
[0061] After obtaining the chromatographic-mass spectrometry raw data in step S20, step S30 is executed to infer the global metabolic profile. Specifically, this step includes: using a filter to perform a convolution operation on the data converted from the chromatographic-mass spectrometry raw data in a three-dimensional point cloud space to downsample the data into a two-dimensional matrix (step S301); forming a mask library based on the two-dimensional matrix and selecting multiple target mask images, multiplying each target mask image with the two-dimensional matrix to form input data for the mass spectrometry inference model based on the convolutional neural network (step S302); and obtaining the retention information of the serum sample corresponding to the current input two-dimensional matrix based on the perturbation prediction of the mass spectrometry inference model based on the convolutional neural network to infer the global metabolic profile of the metabolites in each serum sample (step S303).
[0062] In this embodiment, step S301 includes:
[0063] Step S3011: Convert the chromatographic-mass spectrometric raw data into the .mzml format. However, the present invention is not limited to this. In other implementations, the raw data may be converted into a desired format based on the requirements of the programming language.
[0064] Step S3012: setting the starting retention time T0, the ending retention time Te, the starting mass-to-charge ratio R0, and the ending mass-to-charge ratio Re.
[0065] Step S3013: Sample three-dimensional points from the .mzml file within a retention time range and a mass-to-charge ratio range, each of which contains the retention time t, mass-to-charge ratio r, and ion intensity i of the chromatographic-mass spectrometric raw data. The retention time range is the range from the starting retention time T0 to the ending retention time Te, and the mass-to-charge ratio range is the range from the starting mass-to-charge ratio R0 to the ending mass-to-charge ratio R0.
[0066] In step S3014, based on the preset convolution kernel size (c1*c2) and span size (s1*s2), the maximum pooling convolution function is used to perform a convolution operation on the three-dimensional point space cloud using a filter, and the maximum value in the convolution window is used as the pooling result, thereby downsampling the features of the three-dimensional point cloud space to a two-dimensional matrix. Specifically, c1 and c2 refer to the length and width of the convolution kernel, c1*c2 can be 2*2 or 3*3; correspondingly, s1 and s2 are the length and width of the span, s1*s2 can be 3*3 or 5*5. In this embodiment, the size and span of the convolution kernel are set to be different, thereby smoothly removing the data cavity caused by the missing sampling points, which is more conducive to high-resolution sampling. However, the present invention does not impose any limitation on this.
[0067] After obtaining the two-dimensional matrix, step S302 is performed. Specifically, a plurality of mask images with different mask ratios and sizes consistent with the size of the two-dimensional matrix are generated by masking. The plurality of mask images constitute a mask library, and the mask ratio mask_ratio refers to the ratio of 0 values to 1 values in the mask image. Specifically, the two-dimensional matrix is masked one by one based on the set different mask ratios mask_ratio to automatically and randomly generate a mask image. However, the present invention does not impose any restrictions on this. In other embodiments, mask plates with different mask ratios mask_ratio can also be pre-made, and the two-dimensional matrix corresponding to each serum sample is covered based on a plurality of pre-generated mask plates, thereby quickly generating a plurality of mask images of the two-dimensional matrix to form a mask library. After obtaining the mask library, a plurality of target mask images are extracted from the mask library using a boost extraction method and each target mask image is multiplied and fused with the two-dimensional matrix to form the input data of the mass spectrometry inference model. A mask library is randomly pre-generated according to the preset mask ratio, and then the mask image is extracted through boosting for feature contribution calculation of different samples, thereby reducing the generation calculation time of the mask image and effectively improving the speed of global feature spectrum inference.
[0068] Afterwards, step S303 is executed. Specifically, the step includes:
[0069] In step S3031, each target mask image is multiplied by a two-dimensional matrix and then input into a trained mass spectrometry inference model based on a convolutional neural network to obtain the perturbed predicted probability pred(x(t,r)*mask) corresponding to each target mask image.
[0070] In this step, the construction of the trained convolutional neural network-based mass spectrometry inference model includes:
[0071] (1) Construct a training set, validation set, and test set for the mass spectrometry inference model and include data from different sources as an external test set, using sample attributes as classification labels. (2) Construct a convolutional neural network (CNN) model and train it using the training set. (3) Evaluate the model's performance on the validation and test sets. If performance is poor, adjust the model structure and hyperparameters and retrain. (4) Save the model with the highest accuracy and robustness to form a trained CNN-based mass spectrometry inference model.
[0072] In step S3032, based on the predicted probabilities after perturbation of the multiple target mask images, a contribution calculation function is used to calculate the contribution score S of each molecular feature (t, r) in the serum sample corresponding to the two-dimensional matrix to form a contribution heat map. The contribution calculation function used in this embodiment is as follows: S = sum(predi(x(t, r)*maski)·maski) / (n*maski_ratio)
[0073] Among them, maski is the i-th target Mask image, predi(x(t,r)*maski) is the predicted probability of the i-th target Mask image after perturbation, x(t,r) is the signal strength corresponding to the (t, r) position; n is the total number of target Mask images, maski_ratio is the mask ratio of the i-th target Mask image, and the mask ratio refers to the ratio of 0 values to 1 values in the Mask image.
[0074] Step S3033 , according to the network structure of the mass spectrum inference model, obtain the mapping function t=map1(x), r=map2(y); wherein t is the retention time, and r is the mass-to-charge ratio.
[0075] Step S3034, mapping the two-dimensional coordinates (x, y) of the feature contribution heat map to retention time and mass-to-charge ratio;
[0076] Step S3035 , filtering out molecular features (t, r) whose contribution scores are less than the score threshold and whose ion intensities are less than the intensity threshold to obtain retained features [(t1, r1), (t2, r2), (t3, r3), ..., (tn, rn)];
[0077] In step S3036 , key metabolites are screened based on the retained features, and correlation calculation is performed to infer the metabolic markers and metabolic network patterns of the serum samples corresponding to each two-dimensional matrix, thereby generating a global metabolic profile of the metabolites in each serum sample.
[0078] Figure 5 shows a schematic diagram of global metabolic profile inference for lung cancer, benign lung nodules, and healthy subjects; Figure 6 shows the receiver operating characteristic (ROC) curves of the global metabolic profile inference method inferred in step S30 of the present invention and other inference methods. As can be seen from Figure 6, compared with traditional methods such as support vector machines (SVM), random forests (RF), deep learning neural networks (DNN), and densely connected convolutional networks (DenseNet121), the lung cancer prediction model based on global metabolic profiles (DeepMSProfiler) proposed in the present invention has the highest area under the receiver operating characteristic (ROC) curve (AUC = 0.99).
[0079] The global metabolic profile inference provided in this embodiment is performed using the maximum pooling convolution method, which uses the convolution kernel as a filter to perform a convolution operation on the extracted data point cloud, thereby downsampling the features of the three-dimensional point cloud space to the two-dimensional space. The convolution kernel is a small matrix or vector, and the pooling result is the maximum value in the convolution window; the maximum pooling convolution method performs sampling by minimizing the sampling interval to avoid false sparsity noise caused by excessive sampling intervals or missing sampling points, thereby obtaining higher-resolution downsampled input data to retain more original data information. Furthermore, in the sampling process, the convolution kernel size and span are not synchronized, thereby smoothly removing the data cavity caused by the missing sampling points, which is more conducive to high-resolution sampling. The feature contribution method based on Mask image perturbation not only ensures that the input data size and the two-dimensional matrix size are consistent to obtain high resolution, but also the perturbation prediction can provide a more detailed feature importance explanation for the prediction results on individual samples.
[0080] After obtaining a high-resolution global metabolic profile inference of each serum sample metabolite in step S30, step S40 is performed to select a classifier, and a classification training set and a classification test set are constructed based on the three groups of serum samples obtained in step S10. In this embodiment, 686 of the 794 samples obtained in step S10 are divided into a classification training set, and the remaining 108 samples are used as independent classification test sets. The selected classifier is trained with the classification training set and tested with the classification test set to form a lung cancer prediction model. Figure 7 shows the accuracy of the 108 test samples after prediction by the lung cancer prediction model. The accuracy refers to the consistency between the results predicted by the lung cancer prediction model provided by this embodiment for each test sample and the actual pathological state of the test sample. Figure 7 shows the prediction results of lung cancer, benign lung nodules and healthy people based on the global metabolic profile (DeepMSProfiler). The horizontal axis is the predicted sample grouping situation, and the vertical axis is the actual sample grouping situation. The numbers in the boxes are the number of matching samples between the actual sample grouping and the predicted sample grouping. The values in the brackets are the number of samples correctly predicted by the predicted sample grouping divided by the total number of samples in the actual sample grouping. From the statistical data in the figure, it can be seen that the prediction accuracy of the lung cancer prediction model provided in this embodiment is above 83%, the prediction accuracy for benign lung nodules reaches 92%, and the accuracy for lung cancer prediction is as high as 97%; this data shows that the lung cancer prediction model constructed in this embodiment based on the global metabolic feature spectrum has a very high prediction accuracy.
[0081] Corresponding to the above-mentioned method for constructing a lung cancer prediction model based on a global metabolic profile, this embodiment also provides a device for constructing a lung cancer prediction model based on a global metabolic profile. As shown in FIG8 , the device for constructing a lung cancer prediction model based on a global metabolic profile includes a sample acquisition module 10, a metabolite extraction module 20, a profile inference module 30, and a model construction module 40. The sample acquisition module 10 obtains serum samples from people of different genders and age groups and divides the obtained serum samples into three groups based on lung cancer, benign lung nodules, and healthy status. The metabolite extraction module 20 extracts metabolites from each serum sample and detects the metabolites using liquid chromatography-high-resolution mass spectrometry to obtain raw chromatography-mass spectrometry data. The profile inference module 30 infers the global metabolic profile of metabolites in each serum sample based on the raw chromatographic-mass spectrometry data. This includes using a filter to perform a convolution operation on the data converted from the raw chromatographic-mass spectrometry data in a three-dimensional point cloud space to downsample the data into a two-dimensional matrix; constructing a mask image with multiple mask ratios and a size consistent with the two-dimensional matrix based on the two-dimensional matrix to form a mask image library; selecting multiple target mask images from the mask image library, multiplying each target mask image with the two-dimensional matrix, and then inputting the result into a trained mass spectrometry inference model based on a convolutional neural network for perturbation prediction. Based on the perturbation prediction, the retained information of the sample corresponding to the currently input two-dimensional matrix is obtained to infer the global metabolic profile of the metabolites in each serum sample. The model construction module 40 trains the selected classifier based on the inferred global metabolic profile of the metabolites in each serum sample to obtain a lung cancer prediction model.
[0082] In this embodiment, the characteristic spectrum inference module 30 executes S3011 to 3014 in step S301 to convert the chromatographic-mass spectrometric raw data of the three-dimensional cloud space into a two-dimensional matrix, as shown in Figure 3. Furthermore, the characteristic spectrum inference module 30 executes step S302, forms a mask library based on the two-dimensional matrix and selects multiple target mask images, and multiplies each target mask image with the two-dimensional matrix to form the input data of the mass spectrometry inference model based on the convolutional neural network. Finally, the characteristic spectrum inference module 30 executes S3031 to step S3036 in step S303, and infers the global metabolic characteristic spectrum of metabolites in each serum sample based on the perturbation prediction of the mass spectrometry inference model of the convolutional neural network, as shown in Figure 4. The various refinement steps in step S30 have been described in detail in the above method for constructing a lung cancer prediction model based on a global metabolic characteristic spectrum, and will not be repeated here.
[0083] Regarding the specific definition of the device for constructing a lung cancer prediction model based on a global metabolic profile, please refer to the definition of steps S10 to S40 in the construction method above, and will not be repeated here. The various modules in the above-mentioned device for constructing a lung cancer prediction model based on a global metabolic profile can be implemented in whole or in part through software, hardware, or a combination thereof. The above-mentioned modules can be embedded in or independent of the processor in the computer device in hardware form, or can be stored in the memory of the computer device in software form, so that the processor can call and execute the corresponding operations of the above-mentioned modules.
[0084] Corresponding to the above-mentioned method for constructing a lung cancer prediction model based on a global metabolic profile, this embodiment also provides a lung cancer prediction model constructed based on the construction method. The lung cancer prediction model can be stored in a memory in a computer device in the form of software so that the processor can call it to perform lung cancer prediction operations.
[0085] Correspondingly, this embodiment also provides a lung cancer prediction device including the lung cancer prediction model.
[0086] Example 2
[0087] Although Example 1 illustrates the application of a high-resolution global metabolic profile inference method to lung cancer prediction, the present invention is not limited thereto. In other implementations, the high-information-retention and high-resolution global metabolic profile inference method provided by the present invention can also be used to predict other diseases with different pathological or physiological states, particularly cancers with nodular characteristics, such as thyroid cancer and breast cancer.
[0088] To this end, this embodiment provides a method for inferring a high-resolution global metabolic profile based on raw chromatographic-mass spectrometric data from a target sample. The present invention does not limit the application of this method. While this method can be used for predicting pathological or physiological conditions such as lung cancer, as described in Example 1, it can also be used for other research or applications related to global metabolic profiles.
[0089] As shown in FIG9 , the high-resolution global metabolic profile inference method provided in this embodiment includes:
[0090] Step S1, obtaining chromatographic-mass spectrometric raw data of a target sample;
[0091] Step S2, using a filter to perform a convolution operation on the data converted from the chromatography-mass spectrometry raw data in the three-dimensional point cloud space to downsample the data into a two-dimensional matrix;
[0092] In step S3, a mask image with multiple mask ratios and a size consistent with the size of the two-dimensional matrix is constructed based on the two-dimensional matrix to form a mask image library; multiple target mask images are selected from the mask image library, and each target mask image is multiplied by the two-dimensional matrix to form the input data of the mass spectrometry inference model based on the convolutional neural network.
[0093] Step S4, obtaining the retention information of the sample corresponding to the currently input two-dimensional matrix according to the perturbation prediction of the mass spectrometry inference model based on the convolutional neural network, so as to infer the global metabolic characteristic spectrum of the metabolites in the target sample.
[0094] After obtaining the chromatographic-mass spectrometric raw data of the target sample data in step S1, step S2 will be executed. This embodiment provides one possible implementation method for step S2, as shown in FIG10 , and the specific steps are as follows:
[0095] Step S21, converting the chromatographic-mass spectrometric raw data into .mzml format. However, the present invention does not limit this. In other implementations, the raw data may also be converted into a desired format based on the requirements of the programming language.
[0096] Step S22 , setting the starting retention time T0 , the ending retention time Te , the starting mass-to-charge ratio R0 , and the ending mass-to-charge ratio Re .
[0097] Step S23: Sample three-dimensional points from the .mzml format within a retention time range and a mass-to-charge ratio range, each of which includes the retention time t, mass-to-charge ratio r, and ion intensity i of the chromatographic-mass spectrometric raw data. The retention time range is the range between the starting retention time T0 and the ending retention time Te, and the mass-to-charge ratio range is the range between the starting mass-to-charge ratio R0 and the ending mass-to-charge ratio R0.
[0098] Step S24, based on the preset convolution kernel size (c1*c2) and span size (s1*s2), uses the maximum pooling convolution function, utilizes the filter to perform convolution operation on the three-dimensional point space cloud, and takes the maximum value in the convolution window as the pooling result, thereby downsampling the features of the three-dimensional point cloud space to a two-dimensional matrix. Specifically, c1 and c2 refer to the length and width of the convolution kernel, c1*c2 can be 2*2 or 3*3; correspondingly, s1 and s2 are the length and width of the span, s1*s2 can be 3*3 or 5*5. In this embodiment, the size and span of the convolution kernel are set to be different, thereby smoothly removing the data cavity caused by the missing sampling points, which is more conducive to high-resolution sampling. However, the present invention does not impose any limitation on this. After obtaining the two-dimensional matrix, step S3 is executed to construct a Mask library based on the two-dimensional matrix to form the input data of the mass spectrometry inference model. Specifically, a mask is used to cover part of the pixel points of the two-dimensional matrix to generate a plurality of mask images with different mask ratios and sizes consistent with the size of the two-dimensional matrix. The plurality of mask images constitute a mask library, and the mask ratio mask_ratio refers to the ratio of 0 values to 1 values in the mask image. Specifically, based on the set different mask ratios mask_ratio, the two-dimensional matrix is masked one by one to automatically and randomly generate a mask image. However, the present invention does not impose any restrictions on this. In other implementations, mask plates with different mask ratios mask_ratio can also be pre-made, and the two-dimensional matrix corresponding to the target sample can be covered based on a plurality of pre-generated mask plates, thereby quickly generating a plurality of mask images of the two-dimensional matrix to form a mask library. After obtaining the mask library, a plurality of target mask images are extracted from the mask library using a boost extraction method, and each target mask image is multiplied and fused with the two-dimensional matrix to form the input data of the mass spectrometry inference model. A mask library is randomly pre-generated according to the preset mask ratio, and then the mask image is extracted through boosting for feature contribution calculation of different samples, thereby reducing the generation calculation time of the mask image and effectively improving the speed of global feature spectrum inference.
[0099] Then, step S4 is executed to infer the metabolic profile based on the predicted perturbation probability corresponding to each target mask image obtained by the mass spectrometry inference model. This embodiment provides a possible implementation of step S4, as shown in Figure 11. The specific steps are as follows:
[0100] In step S41, each target mask image is multiplied by a two-dimensional matrix and then input into a trained mass spectrometry inference model based on a convolutional neural network to obtain the perturbed predicted probability pred(x(t,r)*mask) corresponding to each target mask image.
[0101] In this step, the construction of the trained convolutional neural network-based mass spectrometry inference model includes:
[0102] (1) Construct a training set, validation set, and test set for the mass spectrometry inference model and include data from different sources as an external test set, using sample attributes as classification labels. (2) Construct a convolutional neural network (CNN) model and train it using the training set. (3) Evaluate the model's performance on the validation and test sets. If performance is poor, adjust the model structure and hyperparameters and retrain. (4) Save the model with the highest accuracy and robustness to form a trained CNN-based mass spectrometry inference model.
[0103] In step S42, based on the predicted probability after perturbation of multiple target mask images, a contribution calculation function is used to calculate the contribution score of each molecular feature in the target sample corresponding to the two-dimensional matrix to form a contribution heat map. Furthermore, the expression of the contribution calculation function provided in this embodiment is as follows: S = sum(predi(x(t,r)*maski)·maski) / (n*maski_ratio)
[0104] Wherein, maski is the i-th target mask image, predi(x(t,r)*maski) is the predicted probability after perturbation of the i-th target mask image, x(t,r) is the signal strength corresponding to the position (t, r); n is the total number of target mask images, maski_ratio is the mask ratio of the i-th target mask image, and the mask ratio refers to the ratio of 0 values to 1 values in the mask image. Although this embodiment is described using the above contribution calculation function as an example, the present invention is not limited to this.
[0105] Step S43 , obtaining mapping functions t=map1(x), r=map2(y) according to the network structure of the mass spectrum inference model; wherein t is the retention time, and r is the mass-to-charge ratio.
[0106] Step S44, mapping the two-dimensional coordinates (x, y) of the feature contribution heat map to retention time and mass-to-charge ratio;
[0107] Step S45 , filtering out molecular features (t, r) whose contribution scores are less than the score threshold and whose ion intensities are less than the intensity threshold to obtain retained features [(t1, r1), (t2, r2), (t3, r3), ..., (tn, rn)];
[0108] In step S46 , key metabolites are screened based on the retained features, and correlation calculation is performed to infer the metabolic markers and metabolic network patterns of the target sample corresponding to the two-dimensional matrix, thereby generating a global metabolic profile of the metabolites in the target sample.
[0109] The high-resolution global metabolic profile inference provided in this embodiment is performed using the maximum pooling convolution method, which uses the convolution kernel as a filter to perform a convolution operation on the extracted data point cloud, thereby downsampling the features of the three-dimensional point cloud space to the two-dimensional space; the convolution kernel is a small matrix or vector, and the pooling result is the maximum value in the convolution window. The maximum pooling convolution method performs sampling by minimizing the sampling interval to avoid false sparsity noise caused by excessive sampling intervals or missing sampling points, thereby obtaining higher-resolution downsampled input data to retain more original data information. Furthermore, in the sampling process, the convolution kernel size and span are not synchronized, thereby smoothly removing the data cavity caused by missing sampling points, which is more conducive to high-resolution sampling. The feature contribution method based on mask image perturbation not only ensures that the input data size and the two-dimensional matrix size are consistent to obtain high resolution, but also the perturbation prediction can provide a more detailed feature importance explanation for the prediction results on individual samples.
[0110] Corresponding to the above-mentioned high-resolution global metabolic profile inference method, as shown in FIG12 , this embodiment also provides a high-resolution global metabolic profile inference device, which includes an acquisition module 1, a sampling module 2, an input data generation module 3, and a profile inference module 4. The acquisition module 1 executes step S1 in this embodiment to acquire the chromatographic-mass spectrometric raw data of the target sample. The sampling module 2 executes steps S21 to S24 in step S2 in this embodiment, using a filter to perform a convolution operation on the data converted from the chromatographic-mass spectrometric raw data in the three-dimensional point cloud space to downsample the data into a two-dimensional matrix. The input data generation module 3 executes step S3, constructing a mask image with multiple mask ratios and a size consistent with the size of the two-dimensional matrix based on the two-dimensional matrix to form a mask image library; multiple target mask images are selected from the mask image library, and each target mask image is multiplied by the two-dimensional matrix and input into the trained mass spectrometric inference model based on the convolutional neural network for perturbation prediction. The characteristic spectrum inference module 4 executes steps S41 to S46 to obtain the retention information of the target sample corresponding to the currently input two-dimensional matrix based on the disturbance prediction, so as to infer the global metabolic characteristic spectrum of the metabolites in the target sample.
[0111] Each module in the high-resolution global metabolic profile inference device provided in this embodiment sequentially performs steps S1 to S4 to infer the metabolite profile of the target sample. Steps S1 to S4 are as described above in this embodiment and will not be repeated here.
[0112] Correspondingly, this embodiment also provides a computer device, which includes a memory, a processor, and computer instructions stored in the memory and executable on the processor. The processor executes the instructions to implement steps S1 to S4 in this embodiment, or the refined steps S21 to S24 within step S2 based on steps S1 to S4, or the refined steps S41 to S46 within step S4 based on steps S1 to S4.
[0113] Furthermore, this embodiment also provides a storage medium storing computer instructions, which, when executed by a processor, implement steps S1 to S4 in this embodiment, or the refined steps S21 to S24 within step S2 based on steps S1 to S4, or the refined steps S41 to S46 within step S4 based on steps S1 to S4.
[0114] The specific limitations of the high-resolution global metabolic profile inference training device can be found in the limitations of the high-resolution global metabolic profile inference method described above and will not be repeated here. Each module in the high-resolution global metabolic profile inference training device can be implemented in whole or in part through software, hardware, or a combination thereof. Each of the modules can be embedded in or independent of a processor in a computer device in hardware form, or can be stored in a memory in the computer device in software form, so that the processor can call and execute the corresponding operations of each module.
[0115] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The program can be stored in a non-volatile computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).
[0116] In summary, in the method for constructing a lung cancer prediction model based on a global metabolic profile provided by the present invention, metabolites in serum samples are detected using a high-retention, high-resolution liquid chromatography-mass spectrometry technique to obtain three-dimensional chromatographic-mass spectrometry raw data. In the inference of the global profile based on the chromatographic-mass spectrometry raw data, a filter is used to perform a convolution operation on the data converted based on the chromatographic-mass spectrometry raw data in the three-dimensional point cloud space to downsample the data into a two-dimensional matrix; the filter formed by the convolution kernel is a small matrix or vector, which minimizes the sampling spacing to obtain a higher-resolution downsampled two-dimensional matrix, thereby effectively avoiding the false sparsity noise caused by the existing sliding window sampling due to excessive intervals or missing sampling points. Furthermore, after obtaining the two-dimensional matrix, a mask image is set to cover some pixels and the output changes of the mass spectrometry inference model after being perturbed by each target mask image are calculated, thereby analyzing the contribution of each molecular feature and obtaining the global metabolic profiles of different samples based on the retained features. The perturbation-based feature contribution method can generate a global profile at single-pixel resolution, providing a more nuanced interpretation of feature importance for individual sample predictions. The minimum sampling spacing and the single-pixel resolution of the global profile ensure that the inferred global metabolic profile retains high information and has extremely high resolution, providing the necessary conditions for lung cancer prediction.
[0117] Although the present invention has been disclosed above by means of preferred embodiments, this is not intended to limit the present invention. Anyone skilled in the art may make slight changes and modifications without departing from the spirit and scope of the present invention. Therefore, the scope of protection of the present invention shall be determined by the scope of protection required by the claims.
Claims
1. A global metabolic profile inference method, characterized in that, Including: Obtaining the original chromatogram - mass spectrometry data of the target sample; Performing a convolution operation on the data converted from the original chromatogram - mass spectrometry data in the three - dimensional point cloud space using a filter to downsample the data into a two - dimensional matrix; Constructing Mask graphs with multiple masking ratios and the same size as the two - dimensional matrix based on the two - dimensional matrix to form a Mask graph library; Selecting multiple target Mask graphs in the Mask graph library, multiplying each target Mask graph by the two - dimensional matrix and then inputting it into a trained mass spectrometry inference model based on a convolutional neural network for perturbation prediction; Based on the perturbation prediction, obtaining the retention information of the sample corresponding to the currently input two - dimensional matrix to infer the global metabolic characteristic spectrum of metabolites in the target sample.
2. The global metabolic profile inference method according to claim 1, wherein Performing a convolution operation on the original chromatogram - mass spectrometry data of the target sample in the three - dimensional point cloud space using a filter, and downsampling the three - dimensional original chromatogram - mass spectrometry data into a two - dimensional matrix includes: Converting the original data into the.mzml format; Setting the starting retention time T0, the ending retention time Te, the starting mass - to - charge ratio R0, and the ending mass - to - charge ratio Re; Sampling three - dimensional points within the retention time range and the mass - to - charge ratio range from the.mzml format, and each three - dimensional point includes the retention time t, the mass - to - charge ratio r, and the ion intensity i of the original chromatogram - mass spectrometry data; Based on the convolution size and stride size of the preset convolution kernel, using the max - pooling convolution function, performing a convolution operation on the three - dimensional point cloud space using a filter, and taking the maximum value in the convolution window as the pooling result, thereby downsampling the features of the three - dimensional point cloud space to a two - dimensional matrix.
3. The global metabolic profile inference method according to claim 2, wherein The convolution size and stride size of the convolution kernel are different to smoothly remove the data cavities caused by missing sampling points.
4. The global metabolic profile inference method according to claim 1, wherein The steps of inferring the global metabolic characteristic spectrum of metabolites in the target sample based on the perturbation prediction of the mass spectrometry inference model based on a convolutional neural network include: Multiplying each target Mask graph by the two - dimensional matrix and then inputting it into a trained mass spectrometry inference model based on a convolutional neural network to obtain the predicted probability after perturbation corresponding to each target Mask graph; Calculating the contribution score of each molecular feature in the target sample corresponding to the two - dimensional matrix based on the predicted probabilities after perturbation of multiple target Mask graphs to form a contribution heat map; According to the network structure of the mass spectrometry inference model, obtaining the mapping functions t = map1(x), r = map2(y); where t is the retention time and r is the mass - to - charge ratio; Mapping the two - dimensional coordinates (x, y) of the feature contribution heat map to the retention time and the mass - to - charge ratio; Filtering out molecular features with a contribution score less than the score threshold and an ion intensity less than the intensity threshold to obtain the retained features; Screening key metabolites according to the retained features, performing correlation calculations to infer the metabolic markers and metabolic network patterns of the target sample corresponding to the two - dimensional matrix, and then generating the global metabolic characteristic spectrum of metabolites in the target sample.
5. The global metabolic profile inference method according to claim 4, wherein After the predicted probability after perturbation corresponding to each target Mask graph, the following contribution calculation function is used to obtain the contribution score of each molecular feature: S = sum(pred i (x(t,r) * mask i )) · mask i )) / (n * mask i _ratio) Among them, mask i is the i-th target Mask image, pred i (x(t,r)*mask i ) is the predicted probability after perturbation of the i-th target Mask image, and x(t,r) is the signal strength at the corresponding (t, r) position; n is the total number of target Mask images, mask i _ratio is the masking ratio of the i-th target Mask image, and the masking ratio refers to the ratio of 0 values to 1 values in the Mask image.
6. The global metabolic profile inference method according to claim 1, characterized in that Generating multiple Mask graphs with different masking ratios and the same size as the two - dimensional matrix by masking some pixel points of the two - dimensional matrix, and multiple Mask graphs form a Mask graph library. The masking ratio mask_ratio refers to the ratio of 0 values to 1 values in the Mask graph.
7. The global metabolic profile inference method according to claim 1, characterized in that After pre - generating a Mask library, extract the target Mask graph based on boost for calculating the contribution scores of molecular features within different samples.
8. A global metabolic profile inference device, characterized in that, Including: An acquisition module to obtain the original chromatogram - mass spectrometry data of the target sample; A sampling module to perform a convolution operation on the data converted from the original chromatogram - mass spectrometry data in the three - dimensional point cloud space using a filter to downsample the data into a two - dimensional matrix; An input data generation module to construct Mask graphs with multiple mask ratios and the same size as the two - dimensional matrix based on the two - dimensional matrix to form a Mask library; Select multiple target Mask graphs from the Mask library, multiply each target Mask graph by the two - dimensional matrix and then input it into a trained mass spectrometry inference model based on a convolutional neural network for perturbation prediction; A feature spectrum inference module to obtain the retention information of the sample corresponding to the currently input two - dimensional matrix based on the perturbation prediction, so as to infer the global metabolic feature spectrum of metabolites in the target sample.
9. A computer device, which includes a memory, a processor, and computer instructions stored on the memory and executable on the processor. The processor executes the instructions to implement the high - resolution global metabolic spectrum inference method according to any one of claims 1 to 7.
10. A storage medium, which stores computer instructions. When the computer instructions are executed by a processor, the high - resolution global metabolic spectrum inference method according to any one of claims 1 to 7 is implemented.
11. A method for constructing a lung cancer prediction model based on a global metabolic feature spectrum, characterized in that, Including: Obtain human serum samples of different genders and age groups and divide the obtained serum samples into three groups based on three states: lung cancer, benign lung nodules, and health; Extract metabolites in each serum sample and detect the metabolites using liquid chromatography - high - resolution mass spectrometry to obtain the original chromatogram - mass spectrometry data; Infer the global metabolic feature spectrum of metabolites in each serum sample based on the original chromatogram - mass spectrometry data, which includes performing a convolution operation on the data converted from the original chromatogram - mass spectrometry data in the three - dimensional point cloud space using a filter to downsample the data into a two - dimensional matrix; constructing Mask graphs with multiple mask ratios and the same size as the two - dimensional matrix based on the two - dimensional matrix to form a Mask library; Select multiple target Mask graphs from the Mask library, multiply each target Mask graph by the two - dimensional matrix and then input it into a trained mass spectrometry inference model based on a convolutional neural network for perturbation prediction, and obtain the retention information of the sample corresponding to the currently input two - dimensional matrix based on the perturbation prediction, so as to infer the global metabolic feature spectrum of metabolites in each serum sample; Train the selected classifier based on the inferred global metabolic feature spectrum of metabolites in each serum sample to obtain a lung cancer prediction model.
12. The method for constructing a lung cancer prediction model based on a global metabolic feature spectrum according to claim 11, wherein The extraction of metabolites in the serum sample includes: extracting and obtaining metabolite dry samples using a liquid - liquid extraction method; re - dissolving the metabolite dry samples and centrifuging them to prepare a sample to be measured.
13. The method for constructing a lung cancer prediction model based on a global metabolic feature spectrum according to claim 12, wherein Metabolite extraction steps for serum samples: After completely thawing the serum samples on ice, take 50 μL and transfer it into a 1.5 mL EP tube. Add 225 μL of frozen methanol and vortex for 30 seconds; then add 750 μL of frozen methyl tert-butyl ether, vortex for 30 seconds, and then shake at 400 rpm on ice for 1 hour; then add 188 μL of pure water and vortex for 1 minute; centrifuge at 15000×g at 4°C for 10 minutes; after centrifugation, take 125 μL of the lower supernatant from two tubes and transfer it into another EP tube, and dry it with a vacuum freeze dryer. All dried serum metabolite samples are stored in a -80°C refrigerator before testing; reconstitute the dried serum metabolite extract, and take the supernatant after centrifugation to prepare a sample to be tested for detection by liquid chromatography-high resolution mass spectrometry.
14. The method for constructing a lung cancer prediction model based on a global metabolic feature spectrum according to claim 11, wherein The steps of performing a convolution operation on the original chromatogram-mass spectrometry data of each serum sample in the three-dimensional point cloud space using a filter, and downsampling the three-dimensional original chromatogram-mass spectrometry data into a two-dimensional matrix include: Convert the original data into the.mzml format; Set the starting retention time T0, the ending retention time Te, the starting mass-to-charge ratio R0, and the ending mass-to-charge ratio Re; In the retention time range and the mass-to-charge ratio range, sample three-dimensional points from the.mzml format, and each three-dimensional point contains the retention time t, the mass-to-charge ratio r, and the ion intensity i of the original chromatogram-mass spectrometry data; Based on the convolution size and the stride size of the preset convolution kernel, use the max pooling convolution function, and perform a convolution operation on the three-dimensional point cloud space using the filter, and use the maximum value in the convolution window as the pooling result, so as to downsample the features of the three-dimensional point cloud space into a two-dimensional matrix.
15. The method for constructing a lung cancer prediction model based on a global metabolic feature spectrum according to claim 11, characterized in that, The steps of predicting perturbations of the mass spectrometry inference model based on a convolutional neural network to infer the global metabolic feature spectrum of metabolites in each serum sample include: Multiply each target Mask map by the two-dimensional matrix and then input it into the trained mass spectrometry inference model based on a convolutional neural network to obtain the perturbed prediction probability corresponding to each target Mask map; Calculate the contribution score of each molecular feature in the serum sample corresponding to the two-dimensional matrix based on the perturbed prediction probabilities of multiple target Mask maps to form a contribution heat map; According to the network structure of the mass spectrometry inference model, obtain the mapping functions t = map1(x), r = map2(y); where t is the retention time and r is the mass-to-charge ratio; Map the two-dimensional coordinates (x, y) of the feature contribution heat map to the retention time and the mass-to-charge ratio; Filter out molecular features with a contribution score less than the score threshold and an ion intensity less than the intensity threshold to obtain the retained features; Screen key metabolites according to the retained features, and perform correlation calculations to infer the metabolic markers and metabolic network patterns of the serum samples corresponding to each two-dimensional matrix, and then generate the global metabolic feature spectrum of metabolites in each serum sample.
16. The method for constructing a lung cancer prediction model based on a global metabolic feature spectrum according to claim 15, wherein After obtaining the predicted probability of the perturbation corresponding to each target Mask map, the following contribution calculation function is used to obtain the contribution score of each molecular feature: S = sum(pred i (x(t,r) * mask i ) · mask i ) / (n * mask i _ratio) Among them, mask i is the i-th target Mask image, and pred i (x(t,r)*mask i ) is the predicted probability after perturbation of the i-th target Mask image, and x(t,r) is the signal strength at the corresponding (t, r) position; n is the total number of target Mask images quantity, mask i _ratio is the mask ratio of the i-th target Mask image. The mask ratio refers to the ratio of 0 values to 1 values in the Mask image.
17. The method for constructing a lung cancer prediction model based on a global metabolic feature spectrum according to claim 11, wherein Pre-generate a Mask map library and extract target Mask maps based on boost for calculating the contribution scores of molecular features in different samples.
18. An apparatus for constructing a lung cancer prediction model based on a global metabolic feature spectrum, characterized in that, Include: A sample acquisition module, which acquires serum samples of people of different genders and age groups, and divides the obtained serum samples into three groups based on three states: lung cancer, benign lung nodules, and health; A metabolite extraction module extracts metabolites in each serum sample and uses liquid chromatography-high resolution mass spectrometry to detect the metabolites to obtain original chromatogram-mass spectrometry data; A characteristic spectrum inference module infers the global metabolic characteristic spectrum of metabolites in each serum sample based on the original chromatogram-mass spectrometry data. It includes performing a convolution operation on the data converted based on the original chromatogram-mass spectrometry data in a three-dimensional point cloud space using a filter, thereby downsampling the data into a two-dimensional matrix; constructing Mask graphs with multiple masking ratios and the same size as the two-dimensional matrix based on the two-dimensional matrix to form a Mask graph library; Selecting multiple target Mask graphs in the Mask graph library, multiplying each target Mask graph by the two-dimensional matrix and then inputting it into a trained mass spectrometry inference model based on a convolutional neural network for perturbation prediction, and obtaining the retention information of the sample corresponding to the currently input two-dimensional matrix based on the perturbation prediction to infer the global metabolic characteristic spectrum of metabolites in each serum sample; A model construction module trains a selected classifier based on the inferred global metabolic characteristic spectrum of metabolites in each serum sample to obtain a lung cancer prediction model.
19. A lung cancer prediction model, characterized in that, It is constructed by using the method for constructing a lung cancer prediction model based on the global metabolic characteristic spectrum according to any one of claims 11 to 17.
20. A lung cancer prediction device, characterized in that, It includes the lung cancer prediction model according to claim 19.