Multi-modal model-based comprehensive two-dimensional mass spectrum data alkane qualitative method, equipment and medium
By training a multimodal model using simulated and real data, the system automatically identifies the boundaries and center positions of chromatographic peaks, solving the problem of poor accuracy and consistency in the qualitative analysis of saturated alkanes in existing technologies, and realizing automated analysis of full two-dimensional mass spectrometry data.
Patent Information
- Application Number
- CN202511381378.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-25
- Publication Date
- 2025-12-30
AI Technical Summary
Existing two-dimensional mass spectrometry data analysis methods have limitations when processing saturated alkanes, leading to false positive or false negative results. Furthermore, they lack fully automated analysis and have poor consistency and reproducibility of results.
A multimodal model was trained using a combination of simulated and real data. Through joint analysis of image data and fragment ion vectors, the chromatographic peak boundaries and center positions were automatically identified, and qualitative analysis was performed based on fragment ion abundance.
It improves the accuracy and consistency of qualitative analysis of saturated alkanes, reduces the uncertainty caused by manual comparison, and realizes automated analysis of full two-dimensional mass spectrometry data.
Smart Images

Figure CN121237268A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of mass spectrometry, in particular to a full two-dimensional mass spectrometry data alkane qualitative method, device and medium based on a multi-modal model. BACKGROUND
[0002] Full two-dimensional gas chromatography-time of flight mass spectrometry (GCxGC-TOF-MS) is a high-resolution separation and detection technology. It can significantly improve peak capacity and detection sensitivity by using orthogonal separation mechanism, and is widely used in qualitative and quantitative analysis of organic compounds in complex matrix. In the fields of petroleum chemical industry, environmental monitoring and food detection, GCxGC-TOF-MS has become an important tool for studying the composition of complex organic mixtures. In the oil and gas field, saturated alkane compounds are the key analysis objects, and their accurate identification and qualitative analysis are of great significance for sample characteristic analysis.
[0003] However, the existing two-dimensional mass spectrometry data analysis method mainly relies on template matching strategy, which has great limitations in processing saturated alkanes. Because the typical fragment ions (m / z=43, 57, 71) of saturated alkanes are highly overlapped in the spectrum and have similar relative abundances, qualitative analysis relying solely on fragment ions is prone to false positive or false negative results. To avoid false determination, analysts usually need to manually compare multiple factors such as retention time and relative peak position. This process not only consumes time and effort, but also varies between different operators, affecting the consistency and repeatability of the results.
[0004] In addition, there are too many manual intervention steps in the existing data processing process. Baseline calibration, region selection and offset selection need to be manually operated, making it difficult to achieve truly automatic analysis. Due to the introduction of human factors, the consistency and objectivity of the results are often difficult to guarantee, and it is also difficult to standardize the quality control of data between different batches. In the analysis of complex samples, researchers often need to frequently re-establish templates to maintain the normal operation of automatic peak searching and qualitative analysis, which further reduces the overall efficiency. SUMMARY
[0005] Based on the above shortcomings of the prior art, the present application provides a full two-dimensional mass spectrometry data alkane qualitative method, device and medium based on a multi-modal model, which trains a multi-modal model by simulating data and real data to improve the efficiency and accuracy of saturated alkane qualitative analysis of the multi-modal model.
[0006] To solve the above technical problems, the present application discloses a full two-dimensional mass spectrometry data alkane qualitative method based on a multi-modal model, which comprises: obtaining a full two-dimensional flight mass spectrometry analysis result, preprocessing the mass spectrometry analysis result to obtain image data and fragment ion vectors; The image data and corresponding fragment ion vectors are jointly analyzed by a multi-modal model to automatically identify the boundary and center position of the chromatographic peak; the saturated alkane is qualitatively analyzed based on the relative abundance of the fragment ions, the boundary and center position of the chromatographic peak, and a qualitative result is output; The multi-modal model is trained by a training set combining simulation data and real data.
[0007] In some embodiments, the multi-modal model is trained by a training set combining simulation data and real data, comprising: A peak fitting result is established by a tailing Gaussian model, and simulation data is generated step by step; the simulation data includes target peaks, target peaks and blank interference, target peaks and chemical interference, and approximately saturated alkane distribution; The multi-modal model is pre-trained using the simulation data to obtain an initial model; The real data is automatically labeled by the initial model, and a semi-automatic labeled data set is formed by combining correction and clustering algorithms; The initial model is iteratively trained by combining the simulation data and the semi-automatic labeled data to obtain the multi-modal model.
[0008] In some embodiments, the full two-dimensional flight mass spectrometry analysis result is two-dimensional matrix data, the row direction of the two-dimensional matrix data is one-dimensional chromatographic retention time, the column direction is two-dimensional chromatographic retention time, and the matrix element is the signal intensity of all fragment ions detected at the retention time combination.
[0009] In some embodiments, the peak fitting result is established by a tailing Gaussian model, and simulation data is generated step by step, comprising: A target peak is fitted based on the one-dimensional chromatographic retention time and the two-dimensional chromatographic retention time using a tailing Gaussian function to generate a bottom-noise-free matrix; A blank baseline and drift, a chromatographic solvent peak, a chromatographic column loss signal, and instrument noise are superimposed in the bottom-noise-free matrix to obtain a synthetic matrix; The ion channel intensity of the target peak in the synthetic matrix is assigned according to the fragment ion set and its relative abundance and retention index of the saturated alkane target in the spectral library; The synthetic matrix is converted into a two-dimensional chromatography-mass spectrometry simulation spectrum as simulation data.
[0010] In some embodiments, the multi-modal model is trained by a training set combining simulation data and real data, further comprising: When the complexity of the real data is high, a preset constructed thought chain sample is used for supervised fine-tuning, and a generalized reward strategy optimization method is used for inference training.
[0011] In some embodiments, the qualitative analysis of the saturated alkanes based on the relative abundance of the fragment ions, the boundary and center position of the chromatographic peaks further comprises: In combination with the reference point assisted qualitative mechanism, the input target peak retention time and retention index value are received as the reference point, or the non-normal alkane characteristic is identified through automatic reference point search and the retention time and RI value thereof are used as the reference point; The key ion pairs are detected and quantitatively analyzed, and the qualitative analysis is performed in combination with the characteristic ion ratio, one-dimensional chromatographic retention time calibration and two-dimensional relative retention position; The qualitative results are cross-compared according to the peak order and relative position of the compounds.
[0012] In some embodiments, the qualitative results include the compound name, retention time, RI value, fragment ion ratio and confidence.
[0013] In some embodiments, the preprocessing includes non-negative operation, normalization operation and quantization operation.
[0014] The second aspect discloses a computer device, characterized by comprising a processor and a memory; wherein the memory stores a computer program, the computer program is suitable for being loaded and executed by the processor, and the steps of the alkane qualitative method based on the multi-modal model of the full two-dimensional mass spectrum data in any one of the above aspects.
[0015] The third aspect discloses a computer storage medium, which stores a computer program, and the computer program is executed by a processor to implement the alkane qualitative method based on the multi-modal model of the full two-dimensional mass spectrum data in any one of the above aspects.
[0016] Compared with the prior art, the alkane qualitative method based on the multi-modal model of the full two-dimensional mass spectrum data has the following beneficial effects: The alkane qualitative method based on the multi-modal model of the full two-dimensional mass spectrum data provided by the present application can realize the automatic identification of the chromatographic peak boundary and center position, and the qualitative judgment of the saturated alkanes in combination with the relative abundance of the fragment ions, thereby reducing the uncertainty caused by manual comparison. Since the training of the multi-modal model uses a training set combining simulation data and real data, the multi-modal model has large-scale sample learning ability and real data constraints at the same time, thereby improving the accuracy and consistency of the qualitative results. BRIEF DESCRIPTION OF DRAWINGS
[0017] Figure 1 The flowchart of the alkane qualitative method based on the multi-modal model of the full two-dimensional mass spectrum data provided by the present application is shown in the figure; Figure 2This is a schematic diagram of the training process for the alkane qualitative method based on a multimodal model using full two-dimensional mass spectrometry data provided by the present invention. Figure 3 This is another flowchart illustrating the training process in the multimodal model-based method for qualitative analysis of alkane data using full two-dimensional mass spectrometry provided by this invention. Figure 4 This is a schematic diagram of step S2 of the method for qualitative identification of alkane data based on a multimodal model using full two-dimensional mass spectrometry provided by the present invention. Detailed Implementation
[0018] To better understand and implement this invention, the technical solutions in the embodiments of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this invention, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0019] The terms “comprising” and “having” and any variations thereof in this invention are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device that includes a series of steps or modules is not necessarily limited to those steps or modules that are explicitly listed, but may include other steps or modules that are not explicitly listed or that are inherent to such processes, methods, products or devices.
[0020] The present invention discloses a method for qualitative analysis of alkane from full two-dimensional mass spectrometry data based on a multimodal model. By training the multimodal model with simulated data and real data, the efficiency and accuracy of qualitative analysis of saturated alkanes are improved.
[0021] like Figure 1 As shown, this method includes: Step S1: Obtain the full two-dimensional flight mass spectrometry analysis results, and preprocess the mass spectrometry analysis results to obtain image data and fragment ion vectors.
[0022] The results of full two-dimensional flight mass spectrometry analysis include two-dimensional matrix data. In the two-dimensional matrix data, the rows correspond to one-dimensional chromatographic retention times, and the columns correspond to two-dimensional chromatographic retention times. The elements in the matrix represent the fragment ion signal intensity of the total ion current (TIC) or extracted ion current (EIC) detected at the corresponding retention time combination. In practical applications, the two-dimensional matrix data can be converted into heatmaps or contour plots as needed to visually display the peak distribution.
[0023] The preprocessing of the mass spectrometry analysis results includes nonnegation, normalization, and quantization. Nonnegation is performed on the two-dimensional matrix data. Specifically, assuming the input matrix is M, if there are noise signals less than zero in matrix M, its maximum negative value dc is calculated, and dc is subtracted from all elements of the matrix, ensuring that all elements in the processed matrix are nonnegative. This process eliminates negative interference caused by baseline drift or noise. Normalization is then performed on the nonnegative matrix. Since all elements of matrix M are nonnegative, the maximum value of the matrix is used as the denominator, and all elements are divided by this maximum value, scaling the matrix's value range to the [0,1] interval. Normalization ensures the comparability of signal intensities between different samples and provides a uniform scale for the data subsequently input into the model.
[0024] The normalized matrix M is then quantized. Each matrix element is multiplied by (2^bits-1) and converted to an integer value corresponding to the number of bits used in the quantization. For example, in this application, when 16-bit grayscale quantization is selected, the matrix elements are mapped to a preset range, resulting in data files in image formats such as JPG, PNG, or TIF. This processed matrix retains the peak characteristics of the original spectrum and can also serve as input data for deep learning visual models.
[0025] After completing the above preprocessing steps, the fragment ion vector information corresponding to the matrix is retained, i.e., the m / z ion distribution and relative abundance of each peak. By simultaneously obtaining image data and fragment ion vectors, multimodal fusion input of two-dimensional chromatographic peak shape information and mass spectrometry ion characteristics is achieved. The original two-dimensional mass spectrometry data is converted into standardized image data and corresponding fragment ion vectors. The image data facilitates peak recognition and boundary detection by the multimodal visual model, while the fragment ion vectors provide the ion ratios and feature information required for qualitative analysis, enabling the subsequent multimodal model to more accurately distinguish highly overlapping alkane fragment ion signals, reduce false positives or false negatives, and improve the accuracy and consistency of qualitative analysis of saturated alkanes.
[0026] Step S2: Perform joint analysis of the image data and corresponding fragment ion vectors using a multimodal model to automatically identify the boundaries and center positions of chromatographic peaks; based on the relative abundance of fragment ions, one-dimensional chromatographic retention time, and two-dimensional retention position, combined with a benchmark calibration mechanism, qualitative analysis of saturated alkanes is performed, and qualitative results are output.
[0027] In the field of analytical chemistry, the level of data integration and the quality of data labeling are often insufficient to meet the requirements for training multimodal models. Therefore, in this application, the multimodal model is trained using a training set that combines simulated and real data.
[0028] Specifically, such as Figure 2 As shown, the training process includes the following steps: Step S201: Establish peak fitting results using a tailed Gaussian model, and gradually generate simulated data; the simulated data includes the target peak, the target peak and blank interference, the target peak and chemical interference, and the approximately saturated alkane distribution. For example... Figure 3 As shown, the simulation data is generated through the following steps: Step S2011: Based on the one-dimensional chromatographic retention time and the corresponding two-dimensional chromatographic retention time, a tailed Gaussian function is used to fit the target score to generate a noise-free matrix; A tailing Gaussian model is used to fit the target peak in the (RT1, RT2) plane to obtain a noise-free two-dimensional spectral matrix. The tailing Gaussian model is used to characterize the broadening and asymmetric tailing characteristics of the chromatographic peak along the two-dimensional time axis. Parameters include the two-dimensional mean, variance or half-maximum width, and tailing factor. This noise-free matrix represents the idealized peak shape and serves as the basis for subsequent stacking and rendering.
[0029] Step S2012: Superimpose the blank baseline, drift, chromatographic solvent peak, column bleed signal, and instrument noise into the noise-free matrix to obtain the synthesis matrix.
[0030] Interference and noise components, including blank baseline and drift, common chromatographic solvent peaks, column bleed signals, and conventional instrument noise, are progressively superimposed on a noise-free matrix. Conventional instrument noise includes statistical perturbations such as mixable Gaussian noise and Poisson noise, forming a composite matrix that more closely resembles the actual instrument output. By randomly sampling the amplitude, width, occurrence window, and phase relationship of each component, spectral differences under various experimental conditions are covered.
[0031] The models are ordered according to their increasing complexity: those containing only the target peak; those containing the target peak and blank interference; those containing the target peak, blank interference, and chemical interference; and those with overlapping scenarios involving approximately saturated alkane distributions. This combination of simple to complex models systematically improves their robustness in handling peak-to-peak overlaps, background fluctuations, and chemical interference.
[0032] Step S2013: Assign values to the ion channel intensities of the target peaks in the synthesis matrix based on the fragment ion sets of the saturated alkane target analytes and their relative abundance and retention index according to the spectral library.
[0033] In this application, a fragment ion set is constructed for saturated alkanes based on the NIST spectral library. Given the elution time, RI value, response ratio of each characteristic ion, and species name for each target analyte, the ion vectors are superimposed and mapped to the corresponding peak positions in the synthesis matrix to obtain a multi-ion channel intensity distribution consistent with the peak shape. Simultaneously, TIC and multiple EICs are derived. This ensures that the simulated data possesses both realistic peak shapes and reasonable ion ratio characteristics and retention index properties.
[0034] Step S2014: Convert the synthesis matrix into a two-dimensional chromatographic-mass spectrometry simulation spectrum as simulation data.
[0035] After the spectral construction is completed, the synthesized matrix is normalized and quantized, then rendered into an image input readable by the model. The image can be an 8-bit or 16-bit grayscale heatmap. In the simulated data, each target peak is a set of saturated alkane fragment ion peaks generated based on NIST spectral library data, including the bounding box of each peak, peak center coordinates (RT1, RT2), elution time, corresponding RI value, response values of each fragment ion, key ion ratio, and species specialized name. To facilitate multimodal / visual-language training, prompt templates (prompt / answer pairs) matching the task can be generated, with the question and answer fields covering retention time, ion response and target name, RI value, and other tag information.
[0036] The simulated data generated through the above steps covers typical peak shapes and interference scenarios, enabling effective pre-training and supervised fine-tuning of multimodal models without extensive manual peak-by-peak annotation. Combined with automatic annotation and manual correction of real data, this reduces data preparation costs and improves the model's resolvability in complex overlapping peak scenarios. The simulated data can be divided into training, validation, and test sets in an 8:1:1 ratio, with the size adjustable according to task difficulty.
[0037] Step S202: Use the simulated data to pre-train the multimodal model to obtain an initial model; The simulated data is used to pre-train the initial multimodal model, which is the first round of fine-tuning. By pre-tuning the multimodal model with large-scale simulated data, the model learns the mapping relationship between two-dimensional peak structures and corresponding ion features, resulting in the initial model. This initial model can automatically label real mass spectrometry data. It acquires basic analytical capabilities without extensive manual annotation, enabling automatic boundary identification and ion feature assignment of chromatographic peaks in real spectra. It can serve as an automatic label loader, reducing the cost of constructing subsequent real data training sets.
[0038] Step S203: Automatically label the real data using the initial model, and combine it with correction and clustering algorithms to form a semi-automatically labeled dataset.
[0039] Real data is input into the initial model, and the initial model outputs corresponding chromatographic peak label information, including but not limited to: peak boundary position, i.e. peak two-dimensional chromatographic retention time range; peak ion TIC assignment result, peak area, peak height; peak molecular structure information; peak qualitative analysis results, such as compound category or candidate name.
[0040] Traditional unsupervised machine learning algorithms are introduced to perform cluster analysis on the model output results. Specifically, algorithms such as k-means clustering or ART2a neural networks can be used to cluster the model output label results based on ion feature vectors, peak shape parameters, or retention time characteristics. Clusters may group results belonging to the same compound or similar interference together, facilitating subsequent verification and correction.
[0041] Real-world data typically uses actual alkane data. Alkanes exhibit strong regularity in their distribution on two-dimensional chromatograms, demonstrating typical contextual relationships. The retention times of alkanes show continuity and sequence, forming ordered adjacent distributions in two-dimensional space. Furthermore, alkanes belong to the same class of compounds, with intra-group inheritance relationships existing between different molecular weights. This combination of spatial location and group relationships allows the model to learn the ability to identify single peaks during training and grasp the contextual semantic connections between a group of compounds. By introducing real-world alkane data, the model can understand intra-group patterns and relative distributions, learn semantic information about the chemical background, strengthen generalization ability, and improve the overall accuracy of identifying similar compounds in qualitative tasks. Training with a training set composed of both simulated and real data ensures that the simulated and real data complement each other, maintaining the scale of the training data and the chemical rationality and practical value of the training.
[0042] Real-world data is more complex than simulated data, and labeling is difficult and costly. By using an initial model to generate automatic labels first, large-scale preliminary labeling can be completed quickly. Then, clustering algorithms are used for structured organization, reducing the complexity of manual verification. Correction ensures the accuracy and scientific validity of the final labels. The resulting semi-automatic labeling dataset combines scale and quality, suitable for further model fine-tuning. By combining automatic labeling, clustering, and manual correction, a training dataset of a certain size is efficiently constructed while ensuring label accuracy. The initial model can be continuously optimized through iterative training, gradually improving its analytical capabilities under complex peak shapes and overlapping signal conditions.
[0043] Step S204: Combine the simulation data with the semi-automatic calibration data, and iteratively train the initial model to obtain a multimodal model.
[0044] The simulation dataset and the semi-automatic calibration dataset are merged according to a preset ratio to form a new training dataset. This dataset includes both synthetic samples generated through peak fitting and interference simulation, and real samples automatically labeled by the initial model and confirmed by calibration. This combined training method maintains a large sample size while ensuring the authenticity and diversity of the data labels.
[0045] After data preparation, the initial model is iteratively trained using a new training set. During training, the model focuses on learning the regular peak shapes and standard ion characteristics of the simulation data in the early stages, and gradually increases the proportion of real data in the later stages to enhance the model's adaptability to complex peak shapes, overlapping peaks, and noise interference. After multiple rounds of iterative training, a series of progressively optimized model versions can be obtained until the performance indicators meet the expected requirements, resulting in a multimodal model.
[0046] During training, for real-world data with high complexity or poor initial model performance, this embodiment introduces a Chain of Thought (CoT) training mechanism. Several CoT samples containing inference processes are pre-constructed and used as additional fine-tuning data for supervised fine-tuning (SFT) of a specific model version. In this process, the model not only learns the correspondence between inputs and outputs but also the logical links of intermediate inference steps, improving its parsing ability for complex tasks. After CoT-supervised fine-tuning, Generalized Reinforcement Policy Optimization (GRPO) is further employed for inference training. By introducing a reward function to constrain and optimize the model's inference path, it maintains strong generalization ability even when facing unseen complex data. Ultimately, a high-precision multimodal model is obtained that can run stably in real-world application scenarios.
[0047] Based on the relative abundance of fragment ions, one-dimensional chromatographic retention time, and two-dimensional retention position, combined with a benchmark calibration mechanism, saturated alkanes are qualitatively identified, and the qualitative results are output, such as... Figure 4 As shown, it includes: Step S21: Combine the benchmark-assisted qualitative mechanism, receive the input target peak retention time and retention index value as the benchmark, or identify non-n-alkane characteristics through automatic benchmark search and use their retention time and RI value as the benchmark. Generally, users can input the one-dimensional chromatographic retention time, two-dimensional chromatographic retention time, and retention index (RI) value of a known target peak as a reference point to assist in qualitative analysis; or, the multimodal model can identify the fragment ion ratio characteristics of common non-n-alkanes such as benzene, toluene, ethylbenzene, and chlorobenzene through an automatic reference point search mechanism, and use the retention time and RI value as a reference point to correct for retention time shifts under different batches or instrument conditions.
[0048] Step S22: Detect and quantitatively analyze the key ion pairs, and perform qualitative analysis by combining characteristic ion ratios, one-dimensional chromatographic retention time calibration, and two-dimensional relative retention positions; A multimodal model was used to detect and quantify key ion pairs with charge-to-mass ratios (m / z) of 43, 57, and 71. By calculating the proportional relationships between characteristic ions and combining one-dimensional chromatographic retention time calibration with two-dimensional relative retention position verification, preliminary qualitative analysis of saturated alkanes was achieved, avoiding the ambiguity caused by single fragment ions and improving the accuracy of qualitative judgment.
[0049] Step S3: Cross-compare the qualitative results based on the elution order and relative position of the compounds.
[0050] The qualitative results are repeatedly verified based on the elution order and relative position of each compound. If the elution order of a candidate compound does not match its theoretical order, or if there are contradictions in the characteristics of its fragment ions, adjustments and corrections will be made automatically to ensure that the final output qualitative results are consistent in terms of elution order and ion characteristics.
[0051] After completing the qualitative analysis of saturated alkanes using the multimodal model, the output qualitative results include, but are not limited to: compound name, one-dimensional and two-dimensional chromatographic retention times, retention index (RI value), fragment ion ratio, and model prediction confidence level. After generating the results, the above information is structured and output as a standardized analysis report. The report can be output in HTML, XML, or JSON format to ensure compatibility and compatibility with various chromatographic-mass spectrometry (GC-MS) analysis software.
[0052] This application, by simultaneously introducing two-dimensional chromatographic image data and fragment ion vectors into a multimodal model for joint analysis, enables automatic identification of chromatographic peak boundaries and center positions, and qualitative judgment of saturated alkanes based on the relative abundance of fragment ions, thereby reducing the uncertainty caused by manual comparison. Since the multimodal model is trained using a training set combining simulated and real data, it possesses the ability to learn from large-scale samples while also being constrained by real data, thus improving the accuracy and consistency of qualitative results.
[0053] Based on the same inventive concept, the present invention also provides a computer device, comprising: a processor and a memory; wherein the memory stores a computer program adapted to be loaded by the processor and executed the steps of the above-described method for qualitative analysis of alkane based on multimodal model full two-dimensional mass spectrometry data.
[0054] The processing methods for computer devices can be referred to the description of the methods above, and will not be repeated here.
[0055] This application also provides a non-transitory machine-readable storage medium storing an executable program, which, when run by a microprocessor, causes the processor to execute the method provided in the above embodiments.
[0056] This invention discloses a computer-readable storage medium storing a computer program for electronic data interchange, wherein the computer program causes a computer to perform the described methods.
[0057] This invention discloses a computer program product including a non-transitory computer-readable storage medium storing a computer program operable to cause a computer to perform the described method.
[0058] The embodiments described above are merely illustrative. The modules described as separate components may or may not be physically separate. The components shown as modules may or may not be physical modules; that is, they may be located in one place or distributed across multiple network modules. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0059] Through the detailed description of the above embodiments, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, including read-only memory (ROM), random access memory (RAM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), one-time programmable read-only memory (OTPROM), electrically-Erasable Programmable Read-Only Memory (EEPROM), compact disc read-only memory (CD-ROM) or other optical disc storage, disk storage, magnetic tape storage, or any other computer-readable medium that can be used to carry or store data.
[0060] Finally, it should be noted that the embodiments disclosed in this invention are merely preferred embodiments of this invention and are only used to illustrate the technical solutions of this invention, not to limit it. Although this invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this invention.
Claims
1. A method for alkane qualification of full two-dimensional mass spectrometry data based on a multi-modal model, characterized in that, The method comprises: acquiring a full two-dimensional flight mass spectrometry result, preprocessing the mass spectrometry result to obtain image data and fragment ion vectors; performing joint analysis on the image data and corresponding fragment ion vectors by a multi-modal model to automatically identify the boundaries and center positions of chromatographic peaks; performing qualitative analysis on saturated alkanes based on the relative abundance of fragment ions, the boundaries and center positions of chromatographic peaks, and outputting a qualitative result; wherein the multi-modal model is trained by a training set combining simulation data and real data.
2. The full two-dimensional mass spectrometry data alkane qualitative method based on a multi-modal model according to claim 1, characterized in that, The multi-modal model is trained by a training set combining simulation data and real data, comprising: a peak fitting result is established by a tailing Gaussian model, and simulation data is gradually generated; the simulation data includes target peaks, target peaks and blank interference, target peaks and chemical interference, and approximately saturated alkane distribution; the multi-modal model is pre-trained using the simulation data to obtain an initial model; real data is automatically labeled by the initial model, and a semi-automatic calibration data set is formed by combining correction and clustering algorithms; the simulation data and the semi-automatic calibration data are combined to iteratively train the initial model to obtain the multi-modal model.
3. The full two-dimensional mass spectrometry data alkane qualitative method based on a multi-modal model according to claim 2, characterized in that, The full two-dimensional flight mass spectrometry result is two-dimensional matrix data, the row direction of the two-dimensional matrix data is one-dimensional chromatographic retention time, the column direction is two-dimensional chromatographic retention time, and the matrix elements are all fragment ion signal intensities detected at the retention time combination.
4. The method of claim 3, wherein the method is based on a multi-modal model. The peak fitting result is established by a tailing Gaussian model, and simulation data is gradually generated, comprising: a target peak is fitted by a tailing Gaussian function based on the one-dimensional chromatographic retention time and the two-dimensional chromatographic retention time to generate a bottom noise-free matrix; a blank baseline and drift, a chromatographic solvent peak, a chromatographic column loss signal, and instrument noise are superimposed in the bottom noise-free matrix to obtain a synthetic matrix; the ion channel intensity of the target peak in the synthetic matrix is assigned according to the fragment ion set, relative abundance, and retention index of the saturated alkane target in the spectral library; the synthetic matrix is converted into a two-dimensional chromatography-mass spectrometry simulation spectrum as simulation data.
5. The full two-dimensional mass spectrometry data alkane qualitative method based on a multi-modal model according to claim 2, characterized in that, The multi-modal model is trained by a training set combining simulation data and real data, and further comprises: when the complexity of the real data is high, a preset constructed thought chain sample is used for supervised fine tuning, and inference training is performed based on a generalized reward strategy optimization method.
6. The full two-dimensional mass spectrometry data alkane qualitative method based on a multi-modal model according to claim 3, characterized in that, The qualitative analysis of saturated alkanes based on the relative abundance of fragment ions, the boundaries and center positions of chromatographic peaks further comprises: a reference point assisted qualitative mechanism is combined, the input target peak retention time and retention index value are received as reference points, or the retention time and RI value of a non-normal alkane characteristic object identified by automatic reference point searching are used as reference points; key ion pairs are detected and quantitatively analyzed, and qualitative analysis is performed in combination with characteristic ion ratios, one-dimensional chromatographic retention time calibration, and two-dimensional relative retention position; the qualitative result is cross-compared according to the peak order and relative position of the compound.
7. The full two-dimensional mass spectrometry data alkane qualitative method based on a multi-modal model according to claim 6, characterized in that, The qualitative result includes compound name, retention time, RI value, fragment ion ratio, and confidence.
8. The full two-dimensional mass spectrometry data alkane qualitative method based on a multi-modal model according to claim 7, characterized in that, The preprocessing includes non-negative operation, normalization operation, and quantization operation.
9. A computer device, comprising: Comprising: a processor and a memory; wherein the memory stores a computer program which is adapted to be loaded and executed by the processor to perform the steps of the method for alkane qualification of full two-dimensional mass spectrometry data based on multi-modal model according to any one of claims 1-8.
10. A computer storage medium, characterized in that, a computer program stored thereon, which, when executed by a processor, implements the steps of the method for alkane qualification of full two-dimensional mass spectrometry data based on multi-modal model according to any one of claims 1-8.
Citation Information
Patent Citations
Data processing and model construction method based on multi-modal fusion, and related equipment
CN118708969A
Method for analyzing hydrocarbon composition in paraffin sample based on comprehensive two-dimensional gas chromatography-time-of-flight mass spectrometry
CN119064511A
Unmanned aerial vehicle knowledge graph construction method based on multi-modal large model recognition
CN119623593A
Chromatography and mass spectrometry automatic integration method, system, equipment and medium for multiple sample items
CN119780322A
Method for determining hydrocarbon components in mixed aromatic hydrocarbon by using comprehensive two-dimensional gas chromatography-mass spectrometry
CN120214150A