A digital image encoding method based on metabolomics mass spectrometry data
By grouping and filtering LC-MS data, multi-channel images are generated for training deep learning models, solving the problem of low resolution caused by overlapping mass spectrometry signals, and realizing effective parsing of metabolomics information and classification of biological samples.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SHANGHAI INST OF ORGANIC CHEM CHINESE ACAD OF SCI
- Filing Date
- 2022-12-15
- Publication Date
- 2026-04-28
AI Technical Summary
In existing technologies, when LC-MS data is directly encoded into conventional data images, similar mass spectrometry signals will overlap, resulting in low image resolution, which destroys the original structure of the mass spectrometry data and fails to reflect the original state of metabolomics in LC-MS.
By grouping LC-MS data, generating whole metabolome contour images, and then cutting and stacking them, target patches that meet preset conditions are selected to generate multi-channel images for training deep learning models for biological sample classification.
The mass spectrometry structure of metabolomics mass spectrometry data is preserved, enabling the analysis of metabolite types and levels, thus improving the accuracy of biological sample classification.
Smart Images

Figure CN116183796B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data processing technology, and in particular to a digital image encoding method based on metabolomics mass spectrometry data. Background Technology
[0002] Liquid chromatography-tandem mass spectrometry (LC-MS) data can be a two-dimensional matrix containing mass-to-charge ratio (m / z), chromatographic retention time (RT), and ion signal intensity values. Digital image encoding of mass spectrometry data refers to converting LC-MS data information into an image, which can be used to build deep learning models for disease diagnosis.
[0003] However, if LC-MS data is directly encoded to the size of a regular data image, similar mass spectrometry signals will overlap, resulting in an encoded image that is a mixture of multiple mass spectrometry signals. This leads to an image resolution that is too low, destroys the original mass spectrometry structure of the mass spectrometry data, and fails to reflect the original state of metabolomics in LC-MS. Summary of the Invention
[0004] The purpose of this application is to provide a digital image encoding method, apparatus, and computer device based on metabolomics mass spectrometry data, which can solve the problem in related technologies that the encoded image information is the result of mixing multiple mass spectrometry signals, resulting in low image resolution, destruction of the original mass spectrometry structure of the mass spectrometry data, and inability to reflect the original state of metabolomics in LC-MS.
[0005] To solve the above-mentioned technical problems, this application is implemented as follows:
[0006] In a first aspect, embodiments of this application provide a digital image encoding method based on metabolomics mass spectrometry data, which may include:
[0007] Acquire first-order liquid chromatography-tandem mass spectrometry data;
[0008] According to the preset division conditions, the mass-to-charge ratio in the preset mass range of the first liquid chromatography-tandem mass spectrometry data is grouped to obtain P groups, where P is a positive integer.
[0009] Based on the first mass-to-charge ratio and the first scan index in each of the P groups, a whole metabolome contour image is generated. The first scan index is the sequential identifier of the mass spectrum where the first mass-to-charge ratio is acquired.
[0010] The whole metabolome contour image was cut and stacked to obtain the first multi-channel image;
[0011] Based on the pooling signal intensity and image entropy corresponding to the first multi-channel image, a first target image patch is selected from the first multi-channel image. The target pooling signal intensity of the first target image patch satisfies a first preset condition and the target image entropy of the first target image patch satisfies a second preset condition. The first target image patch is used to train a deep learning model for biological sample classification.
[0012] In one possible embodiment, the step of "generating a whole metabolome contour image based on the first mass-to-charge ratio and the first scan index in each of the P groups" may specifically include:
[0013] According to the arrangement order of the first scan index, the first mass-to-charge ratio in each group is aligned and arranged to obtain the target two-dimensional matrix;
[0014] The image represented by the target two-dimensional matrix is determined as the whole metabolome contour image.
[0015] In another possible embodiment, the step of "slicing and stacking the whole metabolome contour image to obtain a first multi-channel image" mentioned above may specifically include:
[0016] According to the preset segmentation order, the whole metabolome contour image is segmented through the preset segmentation window to obtain N first image patches, where N is a positive integer;
[0017] According to the preset segmentation order, N first image patches are stacked to obtain the first multi-channel image.
[0018] In another possible embodiment, the step of "segmenting the whole metabolome contour image according to a preset segmentation order and through a preset segmentation window to obtain N first image patches" mentioned above may specifically include:
[0019] When the first patch includes a first type of patch and a second type of patch, the whole metabolome contour image is segmented according to a preset segmentation order and through a preset segmentation window to obtain the first type of patch and the edge region. The size of the first type of patch is equal to the size of the preset segmentation window.
[0020] By using a preset fill function, the edge area is filled to obtain a second type of tile, and the size of the second type of tile meets the preset segmentation window size.
[0021] Based on this, the aforementioned step of "selecting the first target patch from the first multi-channel image based on the pooling signal intensity and image entropy corresponding to the first multi-channel image" may specifically include:
[0022] Using a preset pooling signal strength algorithm, the pooling signal strength of each second patch is calculated based on the signal strength of each second patch at its corresponding position in the first multi-channel image; and using a preset image entropy algorithm, the image entropy of each second patch is calculated based on the signal strength distribution probability, wherein the signal strength distribution probability is calculated from the signal strength of each second patch.
[0023] Select a first target pooling signal intensity that meets the first preset condition from multiple pooling signal intensities corresponding to multiple second image patches, and select a first target image entropy that meets the second preset condition from multiple image entropies corresponding to multiple second image patches.
[0024] The patch corresponding to the pooling signal intensity of the first target and the image entropy of the first target is determined as the first target patch.
[0025] In another possible embodiment, before the steps of "selecting a first target pooling signal intensity that satisfies a first preset condition from multiple pooling signal intensities corresponding to multiple second patches, and selecting a first target image entropy that satisfies a second preset condition from multiple image entropies corresponding to multiple second patches", the method may further include:
[0026] When the first liquid chromatography-tandem mass spectrometry (LC-MS / MS) data consists of multiple first LC-MS / MS data, and each of the multiple first LC-MS / MS data corresponds to a first multi-channel image, the average value of the pooling signal intensity of the i-th second patch in the multiple first multi-channel images is determined as the pooling signal intensity of the i-th second patch; and the average value of the image entropy of the i-th second image in the multiple second multi-channel images is determined as the image entropy of the i-th second image, where i is a positive integer.
[0027] Secondly, embodiments of this application provide a metabolite analysis method based on the first aspect, which may include:
[0028] Acquire target liquid chromatography-tandem mass spectrometry data of the biological sample to be tested;
[0029] According to the preset division conditions, the mass-to-charge ratio in the preset mass range of the target liquid chromatography-tandem mass spectrometry data is grouped to obtain V groups, where V is a positive integer.
[0030] Based on the second mass-to-charge ratio and second scan index in each of the V groups, a target whole metabolome contour image is generated. The second scan index is the sequential identifier of the mass spectrum where the second mass-to-charge ratio is acquired.
[0031] The target whole metabolome contour image is segmented and stacked to obtain a third multi-channel image;
[0032] Based on the first target scan index corresponding to the first target patch, a second target patch corresponding to the first target scan index is selected from the third multi-channel image;
[0033] The fourth multi-channel image and the zeroed image, which are stacked from the second target image patches, are respectively input into the target deep learning model to obtain the first classification prediction probability value of the fourth multi-channel image and the second classification prediction probability value of the zeroed image; wherein, the target deep learning model is trained on the second multi-channel image constructed from the first target image patches obtained in the first aspect, and the zeroed image is obtained by zeroing the second target image patches;
[0034] By comparing the first classification prediction probability value and the second classification prediction probability value, the target probability value of the second target patch is obtained. The target probability value is used to characterize the importance of the second target patch in participating in the classification of biological samples.
[0035] In one possible embodiment, after the step of "comparing the first classification prediction probability value and the second classification prediction probability value to obtain the target probability value of the second target patch" mentioned above, the method may further include:
[0036] Based on the second target scan index corresponding to the second target patch, obtain the target chromatographic retention time corresponding to the second target scan index, and extract the mass-to-charge ratio of the target metabolite peak based on the second target group corresponding to the second target patch;
[0037] Based on the target chromatographic retention time and the mass-to-charge ratio of the target metabolite peak, the target secondary mass spectrometry spectrum was extracted from the second liquid chromatography-tandem mass spectrometry data;
[0038] The target metabolite was identified in the standard library using the target chromatographic retention time, the mass-to-charge ratio of the target metabolite peak, and the target secondary mass spectrometry information.
[0039] In another possible embodiment, after the step of "selecting the second target patch corresponding to the first target scan index from the third multi-channel image" mentioned above, the method may further include:
[0040] The signal strength value of each second target patch in the fourth multi-channel image is zeroed out to obtain multiple zeroed images.
[0041] Based on this, the aforementioned step of "comparing the first classification prediction probability value and the second classification prediction probability value to obtain the target probability value of the second target patch" may specifically include:
[0042] The first classification prediction probability value is compared with the second classification prediction probability value of each of the plurality of zeroed images to obtain the target probability value corresponding to each second target image patch.
[0043] In yet another possible embodiment, after the step of "selecting the second target patch from the third multi-channel image" described above, the method may further include:
[0044] Obtain model training samples, which include a second multi-channel image constructed from the first target image patch and the preset classification labels corresponding to the second multi-channel image;
[0045] Input the training samples into the initial deep learning model, train the initial deep learning model until the preset training conditions are met, and obtain the target deep learning model.
[0046] Thirdly, embodiments of this application provide a digital image encoding device based on metabolomics mass spectrometry data, the device including:
[0047] The acquisition module is used to acquire first liquid chromatography-tandem mass spectrometry data;
[0048] The partitioning module is used to group the mass-to-charge ratio within a preset mass range in the first liquid chromatography-tandem mass spectrometry data according to preset partitioning conditions, resulting in P groups, where P is a positive integer;
[0049] The generation module is used to generate a whole metabolome contour image based on the first mass-to-charge ratio and the first scan index in each of the P groups. The first scan index is the sequential identifier of the mass spectrum where the first mass-to-charge ratio is acquired.
[0050] The processing module is used to cut and stack the whole metabolome contour image to obtain the first multi-channel image;
[0051] The filtering module is used to filter a first target image patch from the first multi-channel image based on the pooling signal intensity and image entropy corresponding to the first multi-channel image. The target pooling signal intensity of the first target image patch meets a first preset condition and the target image entropy of the first target image patch meets a second preset condition. The first target image patch is used to train a deep learning model for biological sample classification.
[0052] In one possible embodiment, the digital image encoding device in this application may further include an arrangement module and a first determining module; wherein,
[0053] The arrangement module is used to align and arrange the first mass-to-charge ratio in each group according to the arrangement order of the first scan index to obtain the target two-dimensional matrix;
[0054] The first determining module is used to determine the image represented by the target two-dimensional matrix as a whole metabolome contour image.
[0055] In another possible embodiment, the "processing module" mentioned above can specifically be used for:
[0056] According to the preset segmentation order, the whole metabolome contour image is segmented through the preset segmentation window to obtain N first image patches, where N is a positive integer;
[0057] According to the preset segmentation order, N first image patches are stacked to obtain the first multi-channel image.
[0058] In yet another possible embodiment, the digital image encoding device in this application may further include a padding module; wherein,
[0059] The aforementioned "processing module" can also be used to, in the case that the first image patch includes a first type of image patch and a second type of image patch, to cut the whole metabolome contour image according to a preset segmentation order and through a preset segmentation window to obtain the first type of image patch and the edge region, wherein the size of the first type of image patch is equal to the size of the preset segmentation window;
[0060] The fill module is used to fill the edge area using a preset fill function to obtain a second type of tile. The size of the second type of tile meets the size of the preset segmentation window.
[0061] Based on this, the digital image encoding device in the embodiments of this application may further include a calculation module and a second determination module; wherein,
[0062] The calculation module is used to calculate the pooling signal intensity of each second patch in the first multi-channel image based on the signal intensity of each second patch at the corresponding position in the first multi-channel image using a preset pooling signal intensity algorithm; and to calculate the image entropy of each second patch based on the signal intensity distribution probability using a preset image entropy algorithm, wherein the signal intensity distribution probability is calculated from the signal intensity of each second patch.
[0063] The aforementioned "filtering module" can also be used to filter the first target pooling signal intensity that meets the first preset condition from the multiple pooling signal intensities corresponding to multiple second blocks, and to filter the first target image entropy that meets the second preset condition from the multiple image entropies corresponding to multiple second blocks.
[0064] The second determining module is used to determine the patch corresponding to the first target pooling signal intensity and the first target image entropy as the first target patch.
[0065] In another possible embodiment, the digital image encoding device in this application may further include a third determining module; wherein,
[0066] The third determining module is used to determine the average value of the pooling signal intensity of the i-th second patch in the multiple first liquid chromatography-tandem mass spectrometry data as the pooling signal intensity of the i-th second patch when the first liquid chromatography-tandem mass spectrometry data consists of multiple first liquid chromatography-tandem mass spectrometry data, and each of the multiple first liquid chromatography-tandem mass spectrometry data corresponds to a first multi-channel image; and to determine the average value of the image entropy of the i-th second image in the multiple second multi-channel images as the image entropy of the i-th second image, where i is a positive integer.
[0067] Fourthly, embodiments of this application provide a metabolite analysis device based on the first aspect, the device comprising:
[0068] The acquisition module is used to acquire target liquid chromatography-tandem mass spectrometry data of the biological sample to be tested;
[0069] The partitioning module is used to group the mass-to-charge ratio within a preset mass range in the target liquid chromatography-tandem mass spectrometry data according to preset partitioning conditions, resulting in V groups, where V is a positive integer.
[0070] The generation module is used to generate a target whole metabolome contour image based on the second mass-to-charge ratio and the second scan index in each of the V groups. The second scan index is the sequential identifier of the mass spectrum where the second mass-to-charge ratio is acquired.
[0071] The processing module is used to cut and stack the target whole metabolome contour image to obtain a third multi-channel image;
[0072] The filtering module is used to filter a second target patch from the third multi-channel image that corresponds to the first target scan index, based on the first target scan index corresponding to the first target patch.
[0073] The model module is used to input the fourth multi-channel image and the zeroed image, which are stacked from the second target image patches, into the target deep learning model to obtain the first classification prediction probability value of the fourth multi-channel image and the second classification prediction probability value of the zeroed image; wherein, the target deep learning model is trained on the second multi-channel image constructed from the first target image patches obtained in the first aspect, and the zeroed image is obtained by zeroing the second target image patches;
[0074] The comparison module is used to compare the first classification prediction probability value and the second classification prediction probability value to obtain the target probability value of the second target patch. The target probability value is used to characterize the importance of the second target patch in participating in the classification of biological samples.
[0075] In one possible embodiment, the metabolite analysis device in this application may further include an extraction module and a determination module; wherein,
[0076] The aforementioned "acquisition module" can also be used to acquire the target chromatographic retention time corresponding to the second target scan index based on the second target scan index corresponding to the second target patch, and to extract the mass-to-charge ratio of the target metabolite peak based on the second target group corresponding to the second target patch.
[0077] The extraction module is used to extract the target secondary mass spectrum from the second liquid chromatography-tandem mass spectrometry data according to the target chromatographic retention time and the mass-to-charge ratio of the target metabolite peak.
[0078] The determination module is used to identify target metabolites in a standard library using target chromatographic retention time, target metabolite peak mass-to-charge ratio, and target secondary mass spectrometry information.
[0079] In another possible embodiment, the aforementioned "processing module" can also be used to perform zeroing processing on the signal strength value of each second target block in the plurality of second target blocks in the fourth multi-channel image, respectively, to obtain a plurality of zeroed images;
[0080] Based on this, the aforementioned "comparison module" can be specifically used to compare the first classification prediction probability value with the second classification prediction probability value of each of the plurality of zeroed images to obtain the target probability value corresponding to each second target image patch.
[0081] In yet another possible embodiment, the digital image encoding device in the application embodiments may further include a training module; wherein...
[0082] The aforementioned "acquisition module" can also be used to acquire model training samples, which include a second multi-channel image constructed from the first target image patch and a preset classification label corresponding to the second multi-channel image.
[0083] The training module is used to input training samples into the initial deep learning model, train the initial deep learning model until the preset training conditions are met, and obtain the target deep learning model.
[0084] Fifthly, embodiments of this application provide a computer device including a processor, a memory, and a program or instructions stored in the memory and executable on the processor. When the program or instructions are executed by the processor, they implement the steps of the digital image encoding method based on metabolomics mass spectrometry data as shown in the first aspect, or implement the steps of metabolite analysis based on the first aspect as shown in the second aspect.
[0085] In a sixth aspect, embodiments of this application provide a computer-readable storage medium storing a program or instructions that, when executed by a processor, implement the steps of the digital image encoding method based on metabolomics mass spectrometry data as shown in the first aspect, or implement the steps of metabolite analysis based on the first aspect as shown in the second aspect.
[0086] In a seventh aspect, embodiments of this application provide a chip, which includes a processor and a communication interface coupled to the processor. The processor is used to run programs or instructions to implement the steps of the digital image encoding method based on metabolomics mass spectrometry data as shown in the first aspect, or to implement the steps of metabolite analysis based on the first aspect as shown in the second aspect.
[0087] In summary, the digital image encoding method based on metabolomics mass spectrometry data provided in this application embodiment can group the mass-to-charge ratio within a preset mass range in the acquired first liquid chromatography-tandem mass spectrometry data according to preset division conditions to obtain P groups. Then, based on the first mass-to-charge ratio and the first scan index in each of the P groups, a full metabolomics contour image is generated. Furthermore, the full metabolomics contour image is cut and stacked to obtain a first multi-channel image. Then, based on the pooling signal intensity and image entropy corresponding to the first multi-channel image, a first target patch is selected from the first multi-channel image. The second multi-channel image stacked with the first target patch is used to train a deep learning model for biological sample classification. Therefore, through digital image encoding, metabolomics mass spectrometry data can be converted into a second multi-channel image. This second multi-channel image can retain the mass spectrometry structure of the mass spectrometry data that resolves the types and levels of metabolites in the metabolomics. In addition, the metabolomics information in the second multi-channel image can be resolved. Thus, a deep learning model for biological sample classification can be trained using the second multi-channel image stacked with the first target map tiles. This analysis can help distinguish different biological samples by identifying which metabolites or mass spectrometry signals in the biological samples. Attached Figure Description
[0088] Figure 1 A flowchart illustrating a digital image encoding method based on metabolomics mass spectrometry data provided in this application embodiment;
[0089] Figure 2 A schematic diagram illustrating the segmentation of a whole metabolome contour image using a digital image encoding method based on metabolomics mass spectrometry data, provided in an embodiment of this application.
[0090] Figure 3 An embodiment of this application provides a method based on, as shown in the example Figure 1 Metabolite analysis method based on the image encoding method shown;
[0091] Figure 4This application provides a schematic flowchart for determining the target probability value of a second target patch in an embodiment of the present application.
[0092] Figure 5 A schematic diagram of a digital image encoding device based on metabolomics mass spectrometry data provided in this application embodiment;
[0093] Figure 6 An embodiment of this application provides a method based on, as shown in the example Figure 1 A schematic diagram of the structure of a metabolite analysis device using the image encoding method shown;
[0094] Figure 7 This is a schematic diagram of the structure of a computer device provided in an embodiment of this application. Detailed Implementation
[0095] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0096] The terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such use of data can be interchanged where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class and the number of objects is not limited; for example, a first object can be one or more. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.
[0097] Non-targeted metabolomics based on liquid chromatography-mass spectrometry (LC-MS) can provide comprehensive and unbiased qualitative and quantitative characterization of metabolites in living systems. A single measurement of thousands of metabolites in non-targeted metabolomics contains a wealth of information, reflecting the pathological state of an organism and revealing biomarkers that can be used to diagnose diseases, thus helping to reflect metabolic changes in disease.
[0098] In related technologies, the discovery of metabolite biomarkers requires first screening a series of differentially expressed metabolites, and then constructing a disease diagnostic model based on LC-MS data of the screened metabolites for differential diagnosis. However, for complex diseases with high heterogeneity, such as cancer and cardiovascular diseases, directly compressing LC-MS data can lead to overlap of similar mass spectrometry signals. This results in the input image used to construct the disease diagnostic model being a mixture of multiple signals. This resolution reduction strategy destroys the original structure of the mass spectrometer and fails to reflect the original state of omics information in LC-MS. Consequently, the combination of several metabolites cannot reflect the specific pathological state of the disease, leading to poor reproducibility of metabolite biomarkers across different cohorts.
[0099] Furthermore, metabolomics data processing involves complex processes such as peak detection, peak alignment, and metabolite identification. Various factors, including instrument bias, batch effects, and missing values, can interfere with the analytical results. Therefore, the complexity of metabolomics data processing limits the application of LC-MS-based non-targeted metabolomics techniques in biomarker discovery. Consequently, in scenarios involving complex diseases with significant changes in numerous metabolite levels, and where the pathological progression of the disease needs to be reflected, simply selecting a few metabolites cannot capture the overall metabolic changes.
[0100] To address the aforementioned issues, this application discloses a method for converting LC-MS non-targeted metabolomics data into multi-channel digital images. The encoded digital images can be used as samples for training deep learning models for biological sample classification. This method not only preserves the mass spectrometry data format that can resolve the types and levels of metabolites in the metabolomics data, but also trains deep learning models for biological sample classification using a second multi-channel image stacked with first target tiles. It analyzes which metabolites or mass spectrometry signals in biological samples can help distinguish different biological samples.
[0101] Based on this, the following is in conjunction with the appendix Figures 1-3 The digital image encoding method and metabolite analysis method based on metabolomics mass spectrometry data provided in this application will be described in detail through specific embodiments.
[0102] First, combined Figure 1 This paper provides a detailed explanation of digital image coding methods based on metabolomics mass spectrometry data.
[0103] Figure 1 A flowchart illustrating a digital image encoding method based on metabolomics mass spectrometry data, provided for an embodiment of this application.
[0104] like Figure 1 As shown, the digital image encoding method based on metabolomics mass spectrometry data includes steps 110-150.
[0105] First, in step 110, first liquid chromatography-tandem mass spectrometry (LC-MS / MS) data is acquired. Second, in step 120, the mass-to-charge ratio (M / C ratio) within a preset mass range in the first LC-MS / MS data is grouped according to preset division conditions to obtain P groups, where P is a positive integer. Next, in step 130, a full metabolome contour image is generated based on the first M / C ratio and the first scan index in each of the P groups, where the first scan index is the sequential identifier of the mass spectrum where the first M / C ratio was acquired. Furthermore, in step 140, the full metabolome contour image is segmented and stacked to obtain a first multi-channel image. Then, in step 150, based on the pooling signal intensity and image entropy corresponding to the first multi-channel image, a first target patch is selected from the first multi-channel image. The target pooling signal intensity of the first target patch satisfies a first preset condition, and the target image entropy of the first target patch satisfies a second preset condition. The first target patch is used to train a deep learning model for biological sample classification.
[0106] Therefore, through digital image encoding, metabolomics mass spectrometry data can be converted into a second multi-channel image. This second multi-channel image can retain the mass spectrometry structure of the mass spectrometry data that resolves the types and levels of metabolites in the metabolomics. In addition, the metabolomics information in the second multi-channel image can be resolved. Thus, a deep learning model for biological sample classification can be trained using the second multi-channel image stacked with the first target map tiles. This allows analysis of which metabolites or mass spectrometry signals in the biological samples can help distinguish different biological samples.
[0107] The above steps are explained in detail below:
[0108] First, regarding step 110, in one example, the number of first liquid chromatography-tandem mass spectrometry (LC-MS / MS) data points involved in this embodiment can be one. In this case, the first LC-MS / MS data point can correspond to one biological sample. Similarly, in another example, if the number of first LC-MS / MS data points is multiple, then each of the multiple first LC-MS / MS data points can correspond to one biological sample, that is, there is a one-to-one correspondence between biological samples and first LC-MS / MS data points.
[0109] Furthermore, the LC-MS data (i.e., the first liquid chromatography-tandem mass spectrometry data, the second liquid chromatography-tandem mass spectrometry data, and the target phase chromatography-tandem mass spectrometry data) in the embodiments of this application may have at least three attributes, wherein the at least three attributes may include mass-to-charge ratio, scan index, and ion signal intensity.
[0110] In this context, the ion signal intensity corresponds one-to-one with the mass-to-charge ratio; the scan index is the order in which the mass-to-charge ratio is collected in the mass spectrum. Since the mass spectrum is collected one by one in the actual acquisition process, the order in which the mass spectrum is collected can be called the scan index. Also, since the mass spectrum is used to represent the mass-to-charge ratio, the scan index can correspond to the mass-to-charge ratio. Furthermore, the scan index can be converted to and from the RT according to a preset algorithm, meaning that the scan index also corresponds to the RT.
[0111] Secondly, regarding step 120, the preset division conditions can be set based on a uniform atomic mass.
[0112] For example, for m / z in the first LC-MS data, the width can be 0.01 Daltons (Da) and the data can be grouped within a preset quality range such as 60-1200 Da.
[0113] For example, by grouping according to 0.01Da within the 60-1200Da range, the first group of P groups can include m / z values of 60.00-60.01Da, the second group includes those of 60.01-60.02Da, the third group includes those of 60.02-60.03Da, and so on, until the m / z values of the P groups cover the 60-1200Da quality range. In this way, all m / z values within the preset quality range will fall into different groups.
[0114] Next, regarding step 130, in one or more possible embodiments, step 130 may specifically include:
[0115] Step 1301: Align and arrange the first mass-to-charge ratio in each group according to the arrangement order of the first scan index to obtain the target two-dimensional matrix;
[0116] Step 1302: Determine the image represented by the target two-dimensional matrix as the whole metabolome contour image.
[0117] For example, the scan indices corresponding to the first mass-to-charge ratio in each group are arranged into a two-dimensional matrix according to the arrangement order of the first scan indices. This matrix is called the whole metabolome contour image, which is similar in format to a digital image.
[0118] It should be noted that since digital images can be represented as matrices, two-dimensional arrays are typically used to store image data in computer digital image processing programs. Based on this, for a given LC-MS dataset, a target two-dimensional matrix is generated. This target two-dimensional matrix has a format similar to the image; that is, the first dimension of the target two-dimensional matrix is the scan index, which refers to the order in which the mass-to-charge ratio (m / z) mass spectra are arranged according to the scanning sequence. The second dimension is the grouped m / z, and the target two-dimensional matrix can include all the m / z groups in the mass spectrum. Here, it can also be understood that the above two dimensions—the scan index and the grouped m / z—are equivalent to the length and width of the image.
[0119] Therefore, based on the similarity between LC-MS data and digital image data, LC-MS data can be encoded into digital images through digital image encoding methods, so as to construct the deep learning model used to train biological sample classification as described below.
[0120] Furthermore, regarding step 140, in one or more possible embodiments, step 140 may specifically include:
[0121] Step 1401: According to the preset segmentation order, the whole metabolome contour image is segmented through the preset segmentation window to obtain N first image patches, where N is a positive integer;
[0122] Step 1402: Stack N first image patches according to a preset segmentation order to obtain the first multi-channel image.
[0123] For example, refer to Figure 2 According to the preset segmentation order (e.g.) Figure 2 The middle arrow indicates the vertical segmentation from the top left to the bottom right. Using a preset segmentation window, such as a 224×224 pixel window, the obtained whole metabolome contour image 20 is segmented into N smaller first tiles 21. Then, according to the preset segmentation order, the N first tiles are stacked to form a first multi-channel image with a dimension of 224×224×N, where N represents the number of channels, i.e. the number of tiles in the segmented first tiles.
[0124] It should be noted that the multi-channel images (such as the first multi-channel image, the second multi-channel image, the third multi-channel image, and the fourth multi-channel image) in the embodiments of this application include images with a channel number greater than or equal to 3. Here, in digital image processing, the number of channels represents different colors; for example, RGB color encoding is three channels (red, green, and blue), and CMYK is four channels (cyan, magenta, yellow, and black).
[0125] Here, in the embodiments of this application, the whole metabolome contour image 20 may be divided according to a preset segmentation window, that is, after the whole metabolome contour image is segmented, there is no other remaining part except for the first patch with the size of the preset segmentation window; or the whole metabolome contour image 20 may be divided according to the preset segmentation window, and in addition to the first patch with the size of the preset segmentation window, there is also an edge part that is not large enough to form the first patch with the size of the preset segmentation window.
[0126] Based on this, for cases where there are other remaining possibilities, in one or more possible embodiments, the first block includes a first type of block and a second type of block. Therefore, step 1401 mentioned above may specifically include:
[0127] According to the preset segmentation order, the whole metabolome contour image is segmented through the preset segmentation window to obtain the first type of image patch and edge region. The size of the first type of image patch is equal to the size of the preset segmentation window.
[0128] By using a preset fill function, the edge area is filled to obtain a second type of tile, and the size of the second type of tile meets the preset segmentation window size.
[0129] For example, still refer to Figure 2 The obtained full metabolome contour image 20 is segmented using a preset segmentation window, such as a 224×224 pixel window, into smaller first-class patches and edge portions 22 that are insufficient to form first-class patches with the preset segmentation window size. These edge portions 22 can be filled using a preset fill function to obtain second-class patches with the preset segmentation window size. Alternatively, zero values can be used to fill the edges of the matrix of the full metabolome contour image to obtain second-class patches. Then, to retain more information-rich patches and improve the processing performance of the trained deep learning model for biological sample classification, the first-class patches obtained can be filtered according to this application.
[0130] Based on this, step 150, in one or more possible embodiments, may specifically include:
[0131] Step 1501: Using a preset pooling signal strength algorithm, calculate the pooling signal strength of each second patch based on the signal strength of the corresponding position of each second patch in the first multi-channel image; and using a preset image entropy algorithm, calculate the image entropy of each second patch based on the signal strength distribution probability, wherein the signal strength distribution probability is calculated from the signal strength of each second patch.
[0132] Step 1502: Select a first target pooling signal intensity that meets the first preset condition from multiple pooling signal intensities corresponding to multiple second image patches, and select a first target image entropy that meets the second preset condition from multiple image entropies corresponding to multiple second image patches.
[0133] Step 1503: Determine the patch corresponding to the first target pooling signal intensity and the first target image entropy as the first target patch.
[0134] For example, the preset pooling signal intensity algorithm can be executed using the following formula (1), where PI is the pooled signal intensity (PI) of the second patch, and I mn That is, the signal intensity of each second patch at the corresponding position in the first multi-channel image, which represents the signal intensity at (m,n) in the second patch, where m is the row number of the corresponding position of the second patch and n is the column number of the corresponding position of the second patch.
[0135]
[0136] Furthermore, the preset image entropy algorithm can be executed using the following formulas (2) and (3). The calculation of the image entropy of the second patch can be performed according to the following process:
[0137] First, the signal strength distribution probability pi is calculated using formula (2), where i represents the normalized signal strength value of the maximum signal strength of each data point in the second patch, i is an integer between 0 and 255, and x is the number of times each signal strength value i appears in the entire 224×224 second patch:
[0138]
[0139] Then, the image entropy (H) of the second patch is calculated using formula (3):
[0140]
[0141] In this way, the pooling signal intensity and image entropy of each second patch can be obtained.
[0142] Based on this, if the number of first liquid chromatography-tandem mass spectrometry data is one, the first target patch can be determined from multiple second patches in a first multi-channel image. For example, multiple pooling signal intensities corresponding to multiple second patches in a first multi-channel image can be arranged in descending order, and the top 1000 pooling signal intensities can be determined as the first target pooling signal intensities. Similarly, multiple image entropies corresponding to multiple second images in a first multi-channel image can be arranged in descending order, and the top 1000 image entropies can be determined as the first target image entropies. Then, the patch corresponding to both the first target pooling signal intensity and the first target image entropy is determined as the first target patch, which is an information-rich patch. The first target tile can be understood as a tile in an image formed by encoding a 224×224×N first multi-channel image, so as to train a deep learning model for biological sample classification based on a second multi-channel image after stacking the first target tiles. The dimension of the second multi-channel image can be 224×224×N', where N' represents the number of channels in the encoded second multi-channel image, i.e., the number of tiles in the first target tile.
[0143] In another or more possible embodiments, the first liquid chromatography-tandem mass spectrometry (LC-MS / MS) data comprises multiple first LC-MS / MS data sets, each of which corresponds to a first multichannel image. Based on this, prior to step 1502, the digital image encoding method based on metabolomics mass spectrometry data may further include:
[0144] The average value of the pooling signal intensity of the i-th second patch in multiple first multi-channel images is determined as the pooling signal intensity of the i-th second patch; and the average value of the image entropy of the i-th second image in multiple second multi-channel images is determined as the image entropy of the i-th second image, where i is a positive integer.
[0145] For example, if there are multiple first liquid chromatography-tandem mass spectrometry (LC-MS / MS) data sets, the aforementioned steps are first performed on each of the first LC-MS / MS data sets to obtain a first multi-channel image corresponding to each first LC-MS / MS data set. Then, the pooling signal intensity and image entropy of each second patch in each first multi-channel image are calculated. Next, the average value of the pooling signal intensity and image entropy of the i-th second patch in all first multi-channel images is calculated and determined as the pooling signal intensity and image entropy of the i-th second patch. Then, multiple pooling signal intensities can be arranged in descending order, and the top 1000 pooling signal intensities are determined as the first target pooling signal intensity. Similarly, multiple image entropies can be arranged in descending order, and the top 1000 image entropies are determined as the first target image entropy. Finally, the patch corresponding to both the first target pooling signal intensity and the first target image entropy is determined as the first target patch.
[0146] It should be noted that, in order to shorten the computation time, the steps of calculating the pooling signal strength and the image entropy in this embodiment can be performed in parallel.
[0147] In summary, the digital image encoding method based on metabolomics mass spectrometry data provided in this application converts metabolomics mass spectrometry data into an input format that can be recognized by deep learning models through digital image encoding, thus preserving the original characteristics of the mass spectrometry data. Therefore, it can accurately obtain the first target block of the mass spectrometry structure that preserves the types and levels of metabolites in the mass spectrometry data.
[0148] Based on the aforementioned digital image encoding method based on metabolomics mass spectrometry data, in order to be used for biological sample classification, a deep learning model for biological sample classification can be trained using a second multi-channel image of a first target tile stack. This model can analyze which metabolites or mass spectrometry signals in the biological sample can help distinguish different biological samples. Based on this, this application embodiment also provides a metabolite analysis method, as shown below.
[0149] This application's embodiments are combined with Figure 3 It provides a basis such as Figure 1 The image encoding method shown is a metabolite analysis method.
[0150] like Figure 3 As shown, this metabolite analysis method includes steps 310-370. Specifically, as follows:
[0151] Step 310: Obtain the target liquid chromatography-tandem mass spectrometry data of the biological sample to be tested;
[0152] Step 320: According to the preset division conditions, the mass-to-charge ratio within the preset mass range in the target liquid chromatography-tandem mass spectrometry data is grouped to obtain V groups, where V is a positive integer;
[0153] Step 330: Generate a target whole metabolome contour image based on the second mass-to-charge ratio and the second scan index in each of the V groups. The second scan index is the sequential identifier of the mass spectrum where the second mass-to-charge ratio is acquired.
[0154] Step 340: The target whole metabolome contour image is cut and stacked to obtain a third multi-channel image;
[0155] Step 350: Based on the first target scan index corresponding to the first target patch, select the second target patch corresponding to the first target scan index from the third multi-channel image;
[0156] Step 360: Input the fourth multi-channel image and the zeroed-out image, which are stacked from the second target image patches, into the target deep learning model to obtain the first classification prediction probability value of the fourth multi-channel image and the second classification prediction probability value of the zeroed-out image; wherein, the target deep learning model consists of the components described above. Figure 1 The second multi-channel image is obtained by training the first target patch constructed from the data image encoding method shown, and the zeroing image is obtained by zeroing the second target patch;
[0157] Step 370: Compare the first classification prediction probability value and the second classification prediction probability value to obtain the target probability value of the second target patch. The target probability value is used to characterize the importance of the second target patch in participating in the classification of biological samples.
[0158] This enables a deep learning model for biological sample classification to be trained using a second multi-channel image stacked with first target image tiles, and to analyze which metabolites or mass spectrometry signals in the biological samples can help distinguish different biological samples.
[0159] The above steps are explained in detail below:
[0160] It should be noted that the cutting and stacking processes involved in the embodiments of this application, regardless of whether they are in the following stages: Figure 2 Whether it's the stage of generating the first target patch or the stage of training the deep learning model, the same cutting and stacking steps can be used to process the whole metabolome contour image.
[0161] First, the execution principle of steps 320 to 340 is similar to that of steps 110 to 140 above. For details, please refer to steps 110 to 140, which will not be repeated here.
[0162] Next, in one or more possible embodiments, this application employs a channel occlusion-based method to calculate the importance scores of different channels in a multi-channel image and analyze the metabolite composition of different channels in the target deep learning model, thereby providing a corresponding interpretation of the target deep learning model from a metabolomics perspective. Based on this, after step 350 and before step 360, the metabolite analysis method may further include:
[0163] The signal intensity value of each second target patch in the fourth multi-channel image is zeroed out to obtain multiple zeroed images.
[0164] Based on this, since there are multiple zero-reset images, and each zero-reset image corresponds to a second-class classification prediction probability value, there will be multiple second-class classification prediction probability values at this time. Therefore, step 370 may specifically include:
[0165] The first classification prediction probability value is compared with the second classification prediction probability value of each zeroed image in the multiple zeroed images to obtain the target probability value corresponding to each second target patch.
[0166] For example, such as Figure 4 As shown, the fourth multi-channel image can contain L second target patches (L is a positive integer). The signal strength value of each second target patch is zeroed out to obtain multiple zeroed images, such as zeroed image 1, zeroed image 2, ..., zeroed image L. It should be noted that each zeroing process is applied to one second target patch in the fourth multi-channel image. At this time, the other second target patches in the fourth multi-channel image still retain their original signal strength values.
[0167] Thus, by comparing the first classification prediction probability value of the fourth multi-channel image with the second classification prediction probability value of the zeroed-out image 1, we can obtain the following: Figure 4 The target probability value corresponding to the first second target image is 1. The same applies to other second target image patches, which will not be elaborated here. In this way, L target probability values can be obtained.
[0168] And, in one or more possible embodiments, after step 350 and before step 360, the metabolite analysis method may further include:
[0169] Obtain model training samples, which include a second multi-channel image constructed from the first target image patch and the preset classification labels corresponding to the second multi-channel image;
[0170] Input the training samples into the initial deep learning model, train the initial deep learning model until the preset training conditions are met, and obtain the target deep learning model.
[0171] For example, a second multi-channel image with (224×224×N') dimensions after encoding can be input into an initial deep learning model to train the initial deep learning model until the preset training conditions are met, thereby obtaining a target deep learning model. This target deep learning model is used to output a classification prediction probability value for whether a biological sample to be diagnosed is a disease. The target deep learning model can be used in various classification tasks such as clinical diagnosis.
[0172] It should be noted that the initial deep learning model in the embodiments of this application can be a deep belief network learning model, a stacked autoencoder, a convolutional neural network, a recurrent neural network, a residual network (ResNet), or other deep learning models.
[0173] Then, relating to step 370, in one or more possible embodiments, after step 370, the following may also be included:
[0174] Based on the second target scan index corresponding to the second target patch, obtain the target chromatographic retention time corresponding to the second target scan index, and extract the mass-to-charge ratio of the target metabolite peak based on the second target group corresponding to the second target patch;
[0175] Based on the target chromatographic retention time and the mass-to-charge ratio of the target metabolite peak, the target secondary mass spectrometry spectrum (MS / MS) was extracted from the second liquid chromatography-tandem mass spectrometry data;
[0176] The target metabolite was identified in the standard library using the target chromatographic retention time, the mass-to-charge ratio of the target metabolite peak, and the target secondary mass spectrometry information.
[0177] It should be noted that the mass spectrometry data range of the second liquid chromatography-tandem mass spectrometry data is less than or equal to that of the first liquid chromatography-tandem mass spectrometry data.
[0178] Based on this, in order to better understand the above steps 310 to 370, the present application embodiment will illustrate the above steps 310 to 370 with the following example, the specific content of which is as follows.
[0179] First, for the biological sample to be tested, target liquid chromatography-tandem mass spectrometry data of the biological sample to be tested are collected, and a second target patch is constructed according to steps 320 to 350 above. The fourth multi-channel image of the second target patch is then input into the target deep learning model to obtain the first classification prediction probability value f(x) of the second target patch. The first classification prediction probability value is used to characterize the probability of whether the channel corresponding to the second target patch is a disease factor.
[0180] Next, the occlusion method is used to calculate the classification prediction probability value of each channel corresponding to the second target patch in the fourth multi-channel image, such as... Figure 4 As shown, the signal strength values of all data points in the second target patch corresponding to each channel are zeroed out, resulting in zeroed-out images with the same number of data points as the second target patch. This can be called "occluding" the channel. The multiple zeroed-out images after occlusion are then input into the target deep learning model to obtain the second classification prediction probability value f(x) for each zeroed-out patch. -i ).
[0181] Furthermore, the target probability value of the second target patch, i.e., the importance score R of channel i, is calculated using the following formula (4). i , where R i The predicted probability value f(x) for the first classification and the predicted probability value f(x) for the second classification -i The difference between ) is:
[0182] R i =f(x)-f(x) -i (4)
[0183] Among them, R i The larger the absolute value of R, the greater the contribution of that channel (the second target tile) to the target deep learning model. Furthermore, as in disease diagnosis, a value greater than zero for R... i This suggests that the corresponding pathway is a protective factor in the disease, while an R value less than zero... i This indicates that the corresponding channel is a risk factor.
[0184] Then, you can follow R i The absolute value of the channel is used to filter the importance of the channel. The specific filtering threshold can be adjusted according to the model training results. For the important channels that are selected, the biological significance of the channel is explained according to the following process.
[0185] First, determine the number of metabolite peaks contained in the channel, and according to the target scan index and group of the second target block corresponding to the channel, obtain the target chromatographic retention time corresponding to the target scan index and the mass-to-charge ratio of the target metabolite peak corresponding to the group. According to the target chromatographic retention time and the mass-to-charge ratio of the target metabolite peak, extract the target secondary mass spectrometry spectrum (MS / MS) from the second liquid chromatography-tandem mass spectrometry data. Use the target chromatographic retention time, the mass-to-charge ratio of the target metabolite peak, and the target secondary mass spectrometry information to identify the target metabolite in the standard library.
[0186] Thus, the metabolite analysis method provided in this application embodiment can be applied to the scenario of disease diagnosis based on the metabolic state of the metabolome, enabling the target deep learning model to identify metabolites in these biological samples to achieve disease diagnosis or discrimination, characterizing that these metabolites may play an important role in the disease process, thereby explaining the discrimination principle of the target deep learning model and making the target deep learning model interpretable.
[0187] It should be noted that this deep learning model for training biological sample classification can be specifically applied to disease diagnosis, such as the classification of whether a sample is diseased or not, and can also be applied to metabolite classification, such as the classification of whether a sample is a class A metabolite or not.
[0188] Based on the same inventive concept, this application provides a digital image encoding device based on metabolomics mass spectrometry data, specifically combined with Figure 5 Please provide a detailed explanation.
[0189] Figure 5 This is a schematic diagram of the structure of a digital image encoding device based on metabolomics mass spectrometry data, provided in an embodiment of this application.
[0190] like Figure 5 As shown, the digital image encoding device 500 based on metabolomics mass spectrometry data is applied to computer equipment and may specifically include:
[0191] Acquisition module 501 is used to acquire first liquid chromatography-tandem mass spectrometry data;
[0192] The partitioning module 502 is used to group the mass-to-charge ratio within a preset mass range in the first liquid chromatography-tandem mass spectrometry data according to preset partitioning conditions, resulting in P groups, where P is a positive integer.
[0193] The generation module 503 is used to generate a whole metabolome contour image based on the first mass-to-charge ratio and the first scan index in each of the P groups. The first scan index is the sequential identifier of the mass spectrum where the first mass-to-charge ratio is acquired.
[0194] Processing module 504 is used to cut and stack the whole metabolome contour image to obtain a first multi-channel image;
[0195] The filtering module 505 is used to filter a first target image patch from the first multi-channel image based on the pooling signal intensity and image entropy corresponding to the first multi-channel image. The target pooling signal intensity of the first target image patch satisfies a first preset condition and the target image entropy of the first target image patch satisfies a second preset condition. The first target image patch is used to train a deep learning model for biological sample classification.
[0196] The digital image encoding device 500 based on metabolomics mass spectrometry data provided in the embodiments of this application will be described in detail below.
[0197] In one possible embodiment, the digital image encoding device 500 in this application embodiment may further include an arrangement module and a first determining module; wherein,
[0198] The arrangement module is used to align and arrange the first mass-to-charge ratio in each group according to the arrangement order of the first scan index to obtain the target two-dimensional matrix;
[0199] The first determining module is used to determine the image represented by the target two-dimensional matrix as a whole metabolome contour image.
[0200] In another possible embodiment, the "processing module 504" mentioned above can specifically be used for:
[0201] According to the preset segmentation order, the whole metabolome contour image is segmented through the preset segmentation window to obtain N first image patches, where N is a positive integer;
[0202] According to the preset segmentation order, N first image patches are stacked to obtain the first multi-channel image.
[0203] In yet another possible embodiment, the digital image encoding device 500 in this application embodiment may further include a padding module; wherein,
[0204] The aforementioned "processing module 504" can also be used to, in the case that the first image block includes a first type of image block and a second type of image block, cut the whole metabolome contour image according to a preset segmentation order and through a preset segmentation window to obtain the first type of image block and the edge region, wherein the size of the first type of image block is equal to the size of the preset segmentation window;
[0205] The fill module is used to fill the edge area using a preset fill function to obtain a second type of tile. The size of the second type of tile meets the size of the preset segmentation window.
[0206] Based on this, the digital image encoding device 500 in this embodiment may further include a calculation module and a second determination module; wherein,
[0207] The calculation module is used to calculate the pooling signal intensity of each second patch in the first multi-channel image based on the signal intensity of each second patch at the corresponding position in the first multi-channel image using a preset pooling signal intensity algorithm; and to calculate the image entropy of each second patch based on the signal intensity distribution probability using a preset image entropy algorithm, wherein the signal intensity distribution probability is calculated from the signal intensity of each second patch.
[0208] The aforementioned "filtering module 505" can also be used to filter the first target pooling signal intensity that meets the first preset condition from the multiple pooling signal intensities corresponding to the multiple second blocks, and to filter the first target image entropy that meets the second preset condition from the multiple image entropies corresponding to the multiple second blocks.
[0209] The second determining module is used to determine the patch corresponding to the first target pooling signal intensity and the first target image entropy as the first target patch.
[0210] In another possible embodiment, the digital image encoding device 500 in this application embodiment may further include a third determining module; wherein,
[0211] The third determining module is used to determine the average value of the pooling signal intensity of the i-th second patch in the multiple first liquid chromatography-tandem mass spectrometry data as the pooling signal intensity of the i-th second patch when the first liquid chromatography-tandem mass spectrometry data consists of multiple first liquid chromatography-tandem mass spectrometry data, and each of the multiple first liquid chromatography-tandem mass spectrometry data corresponds to a first multi-channel image; and to determine the average value of the image entropy of the i-th second image in the multiple second multi-channel images as the image entropy of the i-th second image, where i is a positive integer.
[0212] Therefore, the digital image encoding device based on metabolomics mass spectrometry data provided in this application embodiment can group the mass-to-charge ratio within a preset mass range in the acquired first liquid chromatography-tandem mass spectrometry data according to preset division conditions to obtain P groups. Then, based on the first mass-to-charge ratio and the first scan index in each of the P groups, a full metabolomics contour image is generated. Furthermore, the full metabolomics contour image is cut and stacked to obtain a first multi-channel image. Then, based on the pooling signal intensity and image entropy corresponding to the first multi-channel image, a first target patch is selected from the first multi-channel image. The second multi-channel image stacked with the first target patch is used to construct a deep learning model for disease diagnosis. Thus, through digital image encoding, metabolomics mass spectrometry data can be converted into a second multi-channel image. This second multi-channel image can retain the mass spectrometry structure of the mass spectrometry data that resolves the types and levels of metabolites in metabolomics. In this way, the second multi-channel image stacked with the first target patch can be used to train a deep learning model for biological sample classification, such as in disease diagnosis scenarios, to analyze which metabolites or mass spectrometry signals in the biological sample can help distinguish different biological samples.
[0213] Based on the same inventive concept, this application provides a metabolite analysis method and apparatus based on a data image encoding method, specifically combined with... Figure 6 Please provide a detailed explanation.
[0214] Figure 6An embodiment of this application provides a method based on, as shown in the example Figure 1 The diagram shows the structure of a metabolite analysis device based on the image encoding method.
[0215] like Figure 6 As shown, the metabolite analysis device 600 is applied to computer equipment and may specifically include:
[0216] The acquisition module 601 is used to acquire target liquid chromatography-tandem mass spectrometry data of the biological sample to be tested;
[0217] The partitioning module 602 is used to group the mass-to-charge ratio within a preset mass range in the target liquid chromatography-tandem mass spectrometry data according to preset partitioning conditions, resulting in V groups, where V is a positive integer.
[0218] The generation module 603 is used to generate a target whole metabolome contour image based on the second mass-to-charge ratio and the second scan index in each of the V groups. The second scan index is the sequential identifier of the mass spectrum where the second mass-to-charge ratio is acquired.
[0219] Processing module 604 is used to cut and stack the target whole metabolome contour image to obtain a third multi-channel image;
[0220] The filtering module 606 is used to filter the second target patch corresponding to the first target scan index from the third multi-channel image according to the first target scan index corresponding to the first target patch;
[0221] Model module 606 is used to input the fourth multi-channel image and the zeroed image of the second target image block stack into the target deep learning model respectively to obtain the first classification prediction probability value of the fourth multi-channel image and the second classification prediction probability value of the zeroed image; wherein, the target deep learning model is trained by the second multi-channel image constructed from the first target image block obtained by any one of the data image encoding methods in claims 1-6, and the zeroed image is obtained by zeroing the second target image block;
[0222] The comparison module 607 is used to compare the first classification prediction probability value and the second classification prediction probability value to obtain the target probability value of the second target patch. The target probability value is used to characterize the importance of the second target patch in participating in the classification of biological samples.
[0223] The metabolite analysis device 600 provided in the embodiments of this application will be described in detail below.
[0224] In one possible embodiment, the metabolite analysis device 600 in this application embodiment may further include an extraction module and a determination module; wherein,
[0225] The aforementioned "acquisition module 601" can also be used to acquire the target chromatographic retention time corresponding to the second target scan index based on the second target scan index corresponding to the second target patch, and to extract the mass-to-charge ratio of the target metabolite peak based on the second target group corresponding to the second target patch.
[0226] The extraction module is used to extract the target secondary mass spectrum from the second liquid chromatography-tandem mass spectrometry data according to the target chromatographic retention time and the mass-to-charge ratio of the target metabolite peak.
[0227] The determination module is used to identify target metabolites in a standard library using target chromatographic retention time, target metabolite peak mass-to-charge ratio, and target secondary mass spectrometry information.
[0228] In another possible embodiment, the "processing module 604" mentioned above can also be used to perform zeroing processing on the signal strength value of each second target block in the multiple second target blocks in the fourth multi-channel image, so as to obtain multiple zeroed images.
[0229] Based on this, the aforementioned "comparison module" can be used to compare the first classification prediction probability value with the second classification prediction probability value of each zeroed image in the multiple zeroed images to obtain the target probability value corresponding to each second target image patch.
[0230] In yet another possible embodiment, the digital image encoding device 600 in the application embodiment may further include a training module; wherein...
[0231] The aforementioned "acquisition module 601" can also be used to acquire model training samples, which include a second multi-channel image constructed from the first target image patch and a preset classification label corresponding to the second multi-channel image.
[0232] The training module is used to input training samples into the initial deep learning model, train the initial deep learning model until the preset training conditions are met, and obtain the target deep learning model.
[0233] In this way, a deep learning model for classifying biological samples can be trained by stacking the first target image tiles into a second multi-channel image, and the analysis can be performed to determine which metabolites or mass spectrometry signals in the biological samples can help distinguish different biological samples.
[0234] The digital image encoding device and metabolite analysis device based on metabolomics mass spectrometry data in the embodiments of this application can be a device, or a component, integrated circuit, or chip in a computer device. The device can be a mobile computer device or a non-mobile computer device.
[0235] For example, mobile computer devices can be mobile phones, tablets, laptops, handheld computers, in-vehicle computer devices, wearable devices, ultra-mobile personal computers (UMPCs), netbooks, or personal digital assistants (PDAs), etc., while non-mobile computer devices can be servers, network attached storage (NAS), personal computers (PCs), televisions (TVs), ATMs, or self-service machines, etc. The embodiments of this application do not impose specific limitations.
[0236] The embodiments of this application include a digital image encoding device and a metabolite analysis device based on metabolomics mass spectrometry data. The operating system can be Android, iOS, or other possible operating systems; this application does not specifically limit the specific operating system used.
[0237] The digital image encoding device and metabolite analysis device based on metabolomics mass spectrometry data provided in this application embodiment can achieve… Figures 1 to 4 The various processes implemented in the method implementation examples will not be described again here to avoid repetition.
[0238] In summary, the digital image encoding method based on metabolomics mass spectrometry data provided in this application embodiment can group the mass-to-charge ratio within a preset mass range in the acquired first liquid chromatography-tandem mass spectrometry data according to preset division conditions to obtain P groups. Then, based on the first mass-to-charge ratio and the first scan index in each of the P groups, a full metabolomics contour image is generated. Furthermore, the full metabolomics contour image is cut and stacked to obtain a first multi-channel image. Then, based on the pooling signal intensity and image entropy corresponding to the first multi-channel image, a first target patch is selected from the first multi-channel image. The second multi-channel image stacked with the first target patch is used to train a deep learning model for biological sample classification. Therefore, through digital image encoding, metabolomics mass spectrometry data can be converted into a second multi-channel image. This second multi-channel image can retain the mass spectrometry structure of the mass spectrometry data that resolves the types and levels of metabolites in the metabolomics. Furthermore, the metabolomics information in this second multi-channel image can be resolved. This allows a deep learning model for biological sample classification to be trained using the second multi-channel image stacked with the first target image tiles. The model can then analyze which metabolites or mass spectrometry signals in the biological samples can help distinguish between different biological samples.
[0239] Optional, such as Figure 7As shown, this application embodiment also provides a computer device 700, including a processor 701, a memory 702, and a program or instructions stored in the memory 702 and executable on the processor 701. When the program or instructions are executed by the processor 701, they implement the various processes of the above-described embodiments of the digital image encoding method and metabolite analysis method based on metabolomics mass spectrometry data, and can achieve the same technical effect. To avoid repetition, they will not be described again here.
[0240] This application also provides a readable storage medium storing a program or instructions. When the program or instructions are executed by a processor, they implement the various processes of the above-described embodiments of the digital image encoding method and metabolite analysis method based on metabolomics mass spectrometry data, and achieve the same technical effect. To avoid repetition, these will not be described again here.
[0241] The processor is the processor in the computer device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.
[0242] In addition, this application provides another chip, which includes a processor and a communication interface. The communication interface and the processor are coupled. The processor is used to run programs or instructions to implement the various processes of the above-described embodiments of the digital image encoding method and metabolite analysis method based on metabolomics mass spectrometry data, and can achieve the same technical effect. To avoid repetition, it will not be described again here.
[0243] It should be understood that the chip mentioned in the embodiments of this application may also be referred to as a system-on-a-chip, system chip, chip system, or system-on-a-chip, etc.
[0244] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.
[0245] Furthermore, it should be noted that the scope of the methods and apparatus in the embodiments of this application is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. In addition, features described with reference to certain examples may be combined in other examples.
[0246] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods of the various embodiments of this application.
[0247] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.
Claims
1. A digital image encoding method based on metabolomics mass spectrometry data, characterized in that, include: Acquire first-order liquid chromatography-tandem mass spectrometry data; According to preset division conditions, the mass-to-charge ratio within a preset mass range in the first liquid chromatography-tandem mass spectrometry data is grouped to obtain P groups, where P is a positive integer. Based on the first mass-to-charge ratio and the first scan index in each of the P groups, a whole metabolome contour image is generated, wherein the first scan index is the sequential identifier of the mass spectrum where the first mass-to-charge ratio is acquired. The whole metabolome contour image is segmented and stacked to obtain a first multi-channel image; Based on the pooling signal intensity and image entropy corresponding to the first multi-channel image, a first target patch is selected from the first multi-channel image. The target pooling signal intensity of the first target patch satisfies a first preset condition, and the target image entropy of the first target patch satisfies a second preset condition. The first target patch is used to train a deep learning model for biological sample classification. The step of filtering the first target image patch from the first multi-channel image based on the pooling signal intensity and image entropy corresponding to the first multi-channel image includes: Using a preset pooling signal intensity algorithm, the pooling signal intensity of each second patch is calculated based on the signal intensity of each second patch at its corresponding position in the first multi-channel image; and using a preset image entropy algorithm, the image entropy of each second patch is calculated based on the signal intensity distribution probability, wherein the signal intensity distribution probability is calculated from the signal intensity of each second patch; the first liquid chromatography-tandem mass spectrometry data comprises multiple first liquid chromatography-tandem mass spectrometry data, and each of the multiple first liquid chromatography-tandem mass spectrometry data corresponds to a first multi-channel image; The average value of the pooling signal intensity of the i-th second patch in a plurality of first multi-channel images is determined as the pooling signal intensity of the i-th second patch; and the average value of the image entropy of the i-th second image in a plurality of second multi-channel images is determined as the image entropy of the i-th second image, where i is a positive integer; Select a first target pooling signal intensity that meets a first preset condition from multiple pooling signal intensities corresponding to multiple second blocks, and select a first target image entropy that meets a second preset condition from multiple image entropies corresponding to multiple second blocks; determine the block corresponding to the first target pooling signal intensity and the first target image entropy as the first target block.
2. The method according to claim 1, characterized in that, The step of generating a full metabolome contour image based on the first mass-to-charge ratio and the first scan index in each of the P groups includes: According to the arrangement order of the first scan index, the first mass-to-charge ratio in each group is aligned and arranged to obtain the target two-dimensional matrix; The image represented by the target two-dimensional matrix is determined as the whole metabolome contour image.
3. The method according to claim 1, characterized in that, The step of cutting and stacking the whole metabolome contour image to obtain a first multi-channel image includes: According to a preset segmentation order, the whole metabolome contour image is segmented through a preset segmentation window to obtain N first image patches, where N is a positive integer; The N first image blocks are stacked according to the preset segmentation order to obtain the first multi-channel image.
4. The method according to claim 3, characterized in that, The first map block includes a first type of map block and a second type of map block; The whole metabolome contour image is segmented according to a preset segmentation order and through a preset segmentation window to obtain N first image patches, including: According to a preset segmentation order, the whole metabolome contour image is segmented through a preset segmentation window to obtain the first type of image patch and edge region. The size of the first type of image patch is equal to the size of the preset segmentation window. The edge region is filled using a preset fill function to obtain the second type of tile, and the size of the second type of tile satisfies the size of the preset segmentation window.
5. A method for metabolite analysis based on the data image encoding method of any one of claims 1-4, characterized in that, include: Acquire target liquid chromatography-tandem mass spectrometry data of the biological sample to be tested; According to preset division conditions, the mass-to-charge ratio within a preset mass range in the target liquid chromatography-tandem mass spectrometry data is grouped to obtain V groups, where V is a positive integer. Based on the second mass-to-charge ratio and the second scan index in each of the V groups, a target whole metabolome contour image is generated, wherein the second scan index is the sequential identifier of the mass spectrum where the second mass-to-charge ratio is acquired; The target whole metabolome contour image is segmented and stacked to obtain a third multi-channel image; Based on the first target scan index corresponding to the first target patch, a second target patch corresponding to the first target scan index is selected from the third multi-channel image; The fourth multi-channel image and the zeroed image, which are stacked from the second target image patches, are respectively input into the target deep learning model to obtain the first classification prediction probability value of the fourth multi-channel image and the second classification prediction probability value of the zeroed image; wherein, the target deep learning model is trained by the second multi-channel image constructed from the first target image patches obtained by any one of the data image encoding methods in claims 1-4, and the zeroed image is obtained by zeroing the second target image patches; By comparing the first classification prediction probability value and the second classification prediction probability value, the target probability value of the second target patch is obtained. The target probability value is used to characterize the importance of the second target patch in participating in biological sample classification.
6. The method according to claim 5, characterized in that, The method further includes: Based on the second target scan index corresponding to the second target patch, obtain the target chromatographic retention time corresponding to the second target scan index, and extract the mass-to-charge ratio of the target metabolite peak based on the second target group corresponding to the second target patch; According to the target chromatographic retention time and the mass-to-charge ratio of the target metabolite peak, the target secondary mass spectrum is extracted from the second liquid chromatography-tandem mass spectrometry data; The target metabolite is identified in a standard library using the target chromatographic retention time, the mass-to-charge ratio of the target metabolite peak, and the target secondary mass spectrometry information.
7. The method according to claim 5, characterized in that, After selecting the second target patch corresponding to the first target scan index from the third multi-channel image, the method further includes: The signal strength value of each second target patch in the fourth multi-channel image is zeroed out to obtain multiple zeroed images. The step of comparing the first classification prediction probability value and the second classification prediction probability value to obtain the target probability value of the second target patch includes: The first classification prediction probability value is compared with the second classification prediction probability value of each of the plurality of zeroed images to obtain the target probability value corresponding to each second target image patch.
8. The method according to claim 5, characterized in that, The method further includes: Obtain model training samples, which include a second multi-channel image constructed from the first target image patch and a preset classification label corresponding to the second multi-channel image; The training samples of the model are input into the initial deep learning model, and the initial deep learning model is trained until the preset training conditions are met to obtain the target deep learning model.
Citation Information
Patent Citations
Four-dimensional metabonomics data processing method
CN116298036A
Systems and methods for analyzing omics data
US20250157570A1