Mass spectrometry structure analysis method, equipment, storage medium and program product
By combining database retrieval and large language models, and using mass spectral similarity to screen reference mass spectra to generate prompt information, the problems of mass spectral structure analysis in existing technologies that cannot infer new molecular structures and the high complexity of deep learning are solved, and efficient mass spectral structure analysis and new data integration are achieved.
Patent Information
- Application Number
- CN202411625285.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-14
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2044-11-14
AI Technical Summary
Existing mass spectrometry structure analysis technology cannot effectively infer the structure of new molecules, and deep learning methods are highly complex when integrating new data.
Combining database retrieval and a large language model, reference mass spectra are screened by calculating mass spectrum similarity, and prompt information is generated and input into the large language model to determine the molecular structure, avoiding frequent model training.
It achieves efficient integration of new data and mass spectrometry structure analysis, improves the generalization ability and flexibility of the model, and reduces the complexity of new data expansion.
Smart Images

Figure CN119517179B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of bioinformatics, and in particular to a mass spectrometry structure analysis method, device, storage medium, and program product. Background Art
[0002] Mass spectrometry is a crucial tool in chemistry and biology, enabling detailed identification of small molecules in complex samples. This capability is essential for a variety of scientific fields, including natural product discovery, human metabolomics, synthetic biology, food science, and environmental research. In mass spectrometry, compounds are ionized to produce charged molecules or molecular fragments. These ions are quantified based on their mass-to-charge ratio ( m / z ) for separation and detection.
[0003] Currently, there are two main technologies related to mass spectrometry structure analysis on the market. One is a database retrieval-based method, represented by MS2DeepScore and MS2Query, and the other is a deep learning-based method, represented by Spec2Mol and MSNovelist. On the one hand, in the database retrieval-based method, after obtaining the mass spectrum of the target compound, it is compared with the mass spectrum data in the reference database and the mass spectrum similarity is calculated. The molecular structure corresponding to the mass spectrum with the highest similarity score is identified as the target structure. On the other hand, in the deep learning-based method, by using the reference database as training data, an end-to-end neural network model is trained to convert the mass spectrum into a molecular structure. Once the model training is completed, it can complete the structural analysis of the unknown mass spectrum.
[0004] Database-based retrieval methods can achieve high accuracy in retrieving mass spectra corresponding to molecular structures already in the reference database. However, they cannot infer the structures of new molecules, and existing mass spectrometry databases only cover a small fraction of molecules discovered in nature. Furthermore, deep learning-based methods require a training phase in which the model learns mass spectrometry structure analysis capabilities from the training data. Therefore, as new data is continuously added to the reference database, the model can only learn this new data through retraining or fine-tuning, which greatly complicates the integration of new data.
[0005] To address the above issues, the industry has not yet proposed a better solution. Summary of the Invention
[0006] The present application provides a mass spectrometry structure analysis method, device, storage medium and program product to at least solve the problems in current related technologies that database retrieval cannot infer the structure of new molecules and that deep learning is very complex in integrating new data.
[0007] In a first aspect, an embodiment of the present application provides a mass spectrometry structure analysis method, comprising: calculating the similarity between each reference mass spectrum in a reference database and a target mass spectrum to be analyzed, and screening at least one similar reference mass spectrum from the each reference mass spectrum based on the calculated similarity; the reference database contains multiple reference mass spectra and corresponding reference molecular structures; generating prompt information based on a preset mass spectrometry structure analysis task instruction, the target mass spectrum and each of the similar reference mass spectra and the corresponding reference molecular structures; and inputting the prompt information into a large language model to determine the target molecular structure corresponding to the target mass spectrum.
[0008] In a second aspect, an embodiment of the present application provides a large language model fine-tuning method for a mass spectrometry structure analysis task, comprising: obtaining a training sample set, each training sample in the training sample set comprising a training mass spectrum and a corresponding training molecular structure; calculating the similarity between a first training mass spectrum and each second training mass spectrum in the training sample set, and screening at least one similar second training mass spectrum from each second training mass spectrum based on the calculated similarity; the first training mass spectrum and the second training mass spectrum respectively belong to different training samples; generating prompt information based on a preset mass spectrometry structure analysis task instruction, the first training mass spectrum and each similar second training mass spectra and the corresponding training molecular structure; using the prompt information as model input information, and using the training molecular structure corresponding to the first training mass spectrum as corresponding label information to fine-tune the large language model.
[0009] In a third aspect, an embodiment of the present application provides an electronic device comprising: at least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the steps of the mass spectrometry structure analysis method of any embodiment of the present application.
[0010] In a fourth aspect, an embodiment of the present application provides a storage medium on which a computer program is stored, characterized in that when the program is executed by a processor, the steps of the mass spectrometry structure analysis method of any embodiment of the present application are implemented.
[0011] In a fifth aspect, an embodiment of the present application provides a computer program product, comprising a computer program / instruction, which, when executed by a processor, implements the steps of the mass spectrometry structure analysis method of any embodiment of the present application.
[0012] The beneficial effects of the embodiments of the present application are:
[0013] By using database retrieval to screen out the reference mass spectra and their corresponding molecular structures that are most similar to the target mass spectra, and then synthesizing prompt information, combined with the powerful understanding ability of the large language model, it is no longer limited to the existing database data and can infer new molecular structures based on similar mass spectra. In addition, the reasoning task of deep learning is transformed into the generation task of the large language model. There is no need to frequently train the model. New data can be directly used by the language model through prompt information once it is included in the database, achieving efficient integration of new data and better mass spectrometry structure analysis effect. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, a brief introduction will be given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0015] Figure 1 A schematic diagram showing the effect of an example of a mass spectrometry structure analysis method according to current related technology is shown;
[0016] Figure 2 A flow chart showing an example of a mass spectrometry structure analysis method according to an embodiment of the present application is shown;
[0017] Figure 3 A flowchart of an example of a method for fine-tuning a large language model for mass spectrometry structure analysis tasks according to an embodiment of the present application is shown;
[0018] Figure 4 An operational flow chart of an example of constructing a training sample set according to an embodiment of the present application is shown;
[0019] Figure 5 An operation diagram of an example of a supervised fine-tuning method based on spectral similarity enhancement is shown;
[0020] Figure 6 Schematic diagrams of the architecture of the two baseline models are shown;
[0021] Figure 7 A schematic diagram showing the effect of an example of a target structure and its corresponding few-sample structure;
[0022] Figure 8 This is a schematic structural diagram of an embodiment of an electronic device of the present application. DETAILED DESCRIPTION
[0023] To make the purpose, technical solutions, and advantages of the embodiments of this application more clear, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the drawings in the embodiments of this application. Obviously, the described embodiments are part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0024] Mass spectrometry is an important analytical technique in various scientific fields, capable of identifying small molecules in complex samples. However, determining the complete molecular structure based on mass spectrometry data, known as mass spectrometry structure elucidation, remains a challenging task.
[0025] In mass spectrometry structure elucidation, a molecular mass spectrum is obtained by recording the mass-to-charge ratio and relative intensity of each ionized molecule or molecular fragment. In the mass spectrum, each ionized molecule or molecular fragment is represented as a peak, with the horizontal axis representing the mass-to-charge ratio and the vertical axis representing the relative intensity. Determining the complete molecular structure based on mass spectrometry data is called mass spectrometry structure elucidation. Mass spectrometry structure elucidation is a many-to-many process. Due to variations in experimental conditions, a single molecule can produce multiple different mass spectra. Furthermore, because mass spectrometry can only provide information on molecular substructure, it is possible for a single mass spectrum to correspond to multiple molecules.
[0026] Figure 1 A schematic diagram showing the effect of an example of a mass spectrometry structure analysis method according to current related technology is shown.
[0027] like Figure 1 As shown, on the one hand, database matching is an important method in the current related technologies of mass spectrometry structure elucidation. After the mass spectrum of the target compound is obtained, it is compared with the mass spectrum in the reference database to calculate the similarity. This process can identify a set of the most matching mass spectra and their corresponding molecular structures, which are considered potential candidates. Commonly used mass spectrometry databases include the Global Natural Products Social Molecular Network (GNPS), the North American Mass Spectral Library (MoNA), the Human Metabolite Database (HMDB) and METLIN. For mass spectra with molecular structures that exist in the reference database, database matching can achieve high-precision structural elucidation. However, it cannot infer the structure of new compounds. In addition, the existing mass spectrometry databases only cover a very small part of the molecules in nature.
[0028] like Figure 1As shown, on the other hand, in the current related technologies, some deep learning-based methods have been proposed for generating "de novo structures" from mass spectra. These deep learning methods use reference databases as training data to develop end-to-end models to convert mass spectra into molecular structures. Once the training is completed, the model can be used for mass spectrometry structure analysis tasks. Compared with database matching methods, deep learning-based "de novo structure" generation not only retains the advantage of accurately retrieving training data, but also demonstrates extensive generalization capabilities and is able to generate new molecular structures. However, deep learning-based methods require a training phase, and the model's structural reasoning ability depends on the results learned from the training data. This means that when new data is continuously added to the database, the model can only learn from the new data through retraining or fine-tuning, making the integration process of new data more complicated.
[0029] To achieve even better results, practitioners in this industry typically consider the following: 1) Expanding the database, specifically collecting more mass spectrometry data and molecular structure information for known compounds to broaden its coverage. 2) Improving deep learning models, such as convolutional neural networks or attention mechanisms, to enhance generalization capabilities. 3) Introducing additional information, primarily spectral data such as nuclear magnetic resonance (NMR) or infrared (IR) spectra, to aid in structural elucidation.
[0030] It should be noted that large language models (LLMs) are one of the most successful applications of deep learning. They are pre-trained on a large amount of text and have extensive world knowledge and some professional chemical knowledge. In addition, large language models have strong reasoning and few-shot learning capabilities. They can extract task-related information from the provided examples during the reasoning phase and use it for reasoning. This feature provides the possibility of combining database matching with deep learning methods. However, directly applying large language models to mass spectrometry structure elucidation remains challenging because they lack specific optimization for this task. In addition, their few-shot learning capabilities may make it difficult to capture the detailed information required for accurate structure elucidation.
[0031] In view of this, Figure 2 A flowchart illustrating an example of a mass spectrometry structure analysis method according to an embodiment of the present application is shown. In this embodiment of the present application, by applying a large language model to mass spectrometry structure analysis and cleverly utilizing the small sample learning capabilities of the large language model to combine the advantages of database retrieval and deep learning, the model has better generalization ability and flexibility.
[0032] like Figure 2As shown, in step S210, the similarity between each reference mass spectrum in the reference database and the target mass spectrum to be analyzed is calculated, and at least one similar reference mass spectrum is screened from each reference mass spectrum according to the calculated similarity.
[0033] Here, the reference database contains multiple reference mass spectra and corresponding reference molecular structures. For example, the reference database stores mass spectrometric data and molecular structure information for multiple known compounds. Each record should include peak positions (m / z values), fragment ion information, possible modifications, and error ranges. Similarity calculation methods can be diverse, such as cosine similarity, Euclidean distance, or Pearson correlation coefficient, and are not intended to be limiting.
[0034] During the screening process for similar reference mass spectra, a threshold screening mechanism or a top-K screening rule can be used. For example, based on a calculated similarity score (e.g., greater than a preset threshold of 0.7), at least one or more reference mass spectra are screened as candidate similar spectra. Alternatively, the top K reference mass spectra in terms of similarity are selected as similar reference mass spectra.
[0035] In step S220, prompt information is generated according to the preset mass spectrum structure analysis task instruction, the target mass spectrum, each similar reference mass spectrum and the corresponding reference molecular structure.
[0036] Here, mass spectrometry structure elucidation task instructions can be used to indicate specific requirements, which can be determined through user interaction or system settings. For example, whether the target molecule requires specific functional groups (such as hydroxyl or carboxyl groups), whether it needs to meet a specific molecular weight range or chemical formula, etc.
[0037] In addition, the screened similar reference mass spectra can be correlated with their corresponding reference molecular structures to extract key information, such as common fragmentation patterns and fragmentation paths, etc. Through this information, the system can construct the relationship between possible fragmentation patterns or characteristic peaks and candidate molecular structures.
[0038] Finally, based on the parsing task, target mass spectrum and similar reference mass spectra, prompt information (Prompt) is automatically generated to help the large language model understand. For example, the prompt information may include clues such as possible functional groups, molecular skeletons, fragmentation paths, etc. in the mass spectrum, and generate a prompt information "Please predict the molecular structure of the target mass spectrum based on this information". The prompt can also be optimized according to needs, for example, "Please predict the molecular structure of the target mass spectrum based on this information, and indicate whether there is a possibility of benzene rings and alkyl chains", etc., and all fall within the scope of implementation of the embodiments of this application.
[0039] In step S230 , the prompt information is input into the large language model to determine the target molecular structure corresponding to the target mass spectrum.
[0040] The large language model can be of various types, such as GPT or Anthropic Claude. Prompt information can be provided to the large language model in the form of JSON or text blocks, allowing it to infer the likely target molecular structure. Furthermore, the large language model can be pre-tuned for enhanced performance in chemical structure prediction and mass spectrometry structure elucidation.
[0041] It should be noted that the reference database can be any database, and can be directly expanded as new data appears, making this solution flexible when adding new data. In this way, the inference task of deep learning is transformed into the generation task of a large language model. Without the need for frequent model training, new data can be directly used by the language model through prompt information as soon as it is included in the database, achieving efficient integration of new data.
[0042] Regarding the implementation details of step S220, in some embodiments, the vertical coordinate of the first mass spectrum peak corresponding to the target mass spectrum is extracted, and the vertical coordinates of the corresponding second mass spectrum peaks are respectively extracted from each similar reference mass spectrum. Furthermore, a prompt message is generated based on the mass spectrum structure analysis task instruction, the vertical coordinates of the first mass spectrum peak, the vertical coordinates of each second mass spectrum peak, and the corresponding reference molecular structure.
[0043] Through the embodiment of the present application, input text assembly is performed after similar spectrum search, and the ordinates of all mass spectrum peaks are first removed. Because molecular structure information is mainly contained in the abscissa of the mass spectrum peak, its chemical meaning is the mass-to-charge ratio of charged molecules or molecular fragments, while the ordinate represents relative intensity, which contains less structural information. In addition, since operations such as noise peak removal have been performed during preprocessing to maximize the use of relative intensity information, and considering that the introduction of relative intensity will increase the fine-tuning and inference cost of the model, the ordinate of the mass spectrum peak is ultimately removed.
[0044] Specifically, after similarity matching, although the ordinate information of the target mass spectrum and the reference mass spectrum has been extracted, all ordinate information is removed during the final prompt information assembly process, and only the abscissa (m / z value) of the mass spectrum peak is retained. The m / z value of the target mass spectrum reflects the mass-to-charge ratio of the fragment, which can directly correspond to a specific fragment or functional group in the molecular structure, while the ordinate only represents the intensity of the peak, and the core of the chemical information is expressed through the abscissa. In addition, the relative intensity of the signal represented by the ordinate may fluctuate significantly under different experimental conditions, which may reduce the stability and accuracy of the analysis. Therefore, in order to improve system performance, you can choose to retain only the m / z value indicated by the abscissa in the prompt information.
[0045] Regarding the description of molecular structure, molecular structure can be represented in various ways, such as SMILES (Simplified Molecular Input Line Entry System), InChI (International Chemical Identifier) molecular diagram, or molecular formula, etc.
[0046] In some embodiments, the molecular structure is represented using a SMILES sequence. SMILES is a string representation of molecular structures that is more concise than graphical representations. For example, the SMILES representation of benzene is "c1ccccc1," which intuitively reflects the bond and ring structures of the compound. SMILES expressions are easily generated by algorithms and can be directly used for virtual screening and automated structure prediction of compounds.
[0047] Regarding the details of step S210 above, in some embodiments, the similarity is calculated using the following formula:
[0048] , formula (1)
[0049] Where, Indicates the similarity between the reference mass spectrum and the target mass spectrum; and Represent the reference mass spectrum and the target mass spectrum, respectively. i The intensity of the peak, n is the number of matching peaks.
[0050] By using the cosine similarity method to calculate the vector angle between the reference mass spectrum and the target mass spectrum, the cosine similarity provides a quantitative index (ranging from 0 to 1). The closer the value is to 1, the more similar the two mass spectra are. This method can assess the similarity between each reference mass spectrum and the target mass spectrum, improving the reliability of screening. In addition, occasional noise peaks or non-characteristic peaks in mass spectrometry data may affect the results. Only matching important peaks has a greater impact on the similarity, thus improving the system's noise resistance.
[0051] Figure 3 A flowchart of an example of a large language model fine-tuning method for mass spectrometry structure analysis tasks according to an embodiment of the present application is shown.
[0052] like Figure 3 As shown, in step S310, a training sample set is obtained, and each training sample in the training sample set includes a training mass spectrum and a corresponding training molecular structure.
[0053] Here, the training sample set can be obtained from a public mass spectrometry database (such as MassBank, NIST or HMDB), or constructed through the laboratory's own mass spectrometry data, and both fall within the scope of implementation of the embodiments of the present application.
[0054] In step S320, the similarity between the first training mass spectrum and each second training mass spectrum in the training sample set is calculated, and at least one similar second training mass spectrum is screened from each second training mass spectrum according to the calculated similarity.
[0055] Here, the first training mass spectrum and the second training mass spectrum belong to different training samples.
[0056] In step S330 , prompt information is generated according to the preset mass spectrum structure analysis task instruction, the first training mass spectrum, each similar second training mass spectrum and the corresponding training molecular structure.
[0057] In step S340 , the prompt information is used as model input information, and the training molecular structure corresponding to the first training mass spectrum is used as corresponding label information to fine-tune the large language model.
[0058] Here, the molecular structure corresponding to the first training mass spectrum is used as the target label to guide the model in learning how to predict the correct molecular structure from the prompt information. Thus, through fine-tuning, the large language model can more accurately understand the relationship between mass spectral features and molecular structure. It can also generalize the knowledge learned during fine-tuning to other similar mass spectrometry data, improving its ability to analyze unknown samples.
[0059] In some embodiments, supervised fine-tuning based on spectral similarity enhancement is implemented to exploit the few-shot learning capability of large language models in mass spectrometry structure analysis. The method includes three main stages: similarity spectrum retrieval, input text assembly, and model fine-tuning. In the similarity spectrum retrieval stage, for each mass spectrum in the training set, a specific mass spectrum similarity algorithm is first used to find its most similar mass spectrum in the training set, excluding itself. k After the similarity search, the input text is assembled, combining the task instructions, the retrieved similar mass spectra, their corresponding SMILES sequences, and the current mass spectrum into the input text. During the model fine-tuning phase, the assembled input text is used as the model input and the target molecule SMILES is used as the model output for fine-tuning.
[0060] Through the embodiments of this application, the accuracy and generalization ability of mainstream models in mass spectrometry structure analysis tasks can be improved, especially when dealing with new molecular structures. The chain reaction it produces can reduce the time and computing power costs of training from scratch and continuing training in related research, thereby improving research efficiency. On a deeper level, the promotion and application of this solution can promote the development of the field of mass spectrometry structure analysis and provide more powerful tools for the discovery of new compounds and drug development.
[0061] Figure 4 An operational flowchart of an example of constructing a training sample set according to an embodiment of the present application is shown.
[0062] like Figure 4 As shown, in step S410, for each training mass spectrum, the molecular fingerprint and molecular formula corresponding to the training mass spectrum are extracted, and the extracted molecular fingerprint and molecular formula are input into the SMILES model to determine the corresponding SMILES representation sequence.
[0063] Here, various molecular fingerprinting algorithms (such as ECFP or MACCS keys) can be used to convert the molecular structure into a binary vector. Each vector bit represents the presence or absence of a structural feature. For example, the ECFP4 method captures patterns in the neighboring atomic graph with a radius of 4, making it suitable for characterizing complex structures. Furthermore, the elemental composition and corresponding quantities of the molecule (e.g., C6H6 represents benzene) can be directly extracted to obtain the molecular formula, which can serve as supplementary information for model input and help generate accurate SMILES sequences.
[0064] In some embodiments, the SMILES model comprises a Transformer encoder and a temporal convolutional decoder. The Transformer encoder converts the input molecular fingerprint and molecular formula into a high-dimensional vector representation, capturing global dependencies within the input data and generating feature embeddings. A multi-head self-attention mechanism improves the model's ability to process relationships between different molecular features, ensuring that information about complex molecular structures is preserved. The temporal convolutional decoder progressively generates a sequence of SMILES representations from the encoder output, using causal convolution to ensure that the order of the SMILES sequence is maintained during the generation process, preventing future information leakage.
[0065] Therefore, by leveraging the Transformer encoder's self-attention mechanism, the model can effectively capture complex molecular structural information and improve the accuracy of SMILES representations. The temporal convolutional decoder maintains sequentiality, ensuring that the generated SMILES representation conforms to chemical rules and avoids grammatically incorrect chemical structures. This allows the model to handle not only simple molecules but also complex cyclic structures and multifunctional molecules, enabling a wider range of applications.
[0066] In step S420, a training sample set is constructed based on each training mass spectrum and the corresponding SMILES representation sequence.
[0067] In some embodiments, training mass spectrometry data (including m / z values) are mapped one-to-one to corresponding SMILES representation sequences. Preferably, noise can be introduced or different mass spectrometer outputs can be simulated to generate more mass spectrometry samples and increase the diversity of the training data. Thus, pairing mass spectrometry data with SMILES representation sequences ensures that the model can learn how to infer molecular structure from mass spectrometry data.
[0068] Through the examples of this application, by combining molecular fingerprints, molecular formulas, and mass spectrometry data, the model can quickly and accurately generate corresponding SMILES representation sequences. Furthermore, by combining the Transformer with a temporal convolutional decoder, the model can handle not only simple molecules but also complex cyclic structures and multifunctional molecules, enabling efficient dataset construction and training, providing a reliable data source foundation for mass spectrometry analysis tasks.
[0069] Regarding the details of fine-tuning for a large language model, in some embodiments, the fine-tuning process can adopt the following loss function:
[0070] , Formula (2)
[0071] Where, represents the loss function, is the predicted SMILES representation sequence length in characters, is the first SMILES representation sequence predicted by the large language model j characters, is the first SMILES representation sequence corresponding to the tag information j characters.
[0072] Here, we use the cross-entropy loss function to fine-tune a large language model to optimize the SMILES representation sequence prediction task. This function automatically calculates the loss and backpropagates the error at each training iteration to optimize the model parameters. The cross-entropy loss function is computationally efficient, suitable for training on large datasets, and facilitates rapid model fine-tuning.
[0073] In the embodiments of this application, the loss function calculates the difference between the predicted SMILES sequence and the target SMILES sequence character by character, evaluating the prediction at each position and guiding the model to make more fine-grained adjustments. By gradually narrowing the gap between the model output and the true label, the model is ensured to accurately generate SMILES expressions that conform to chemical rules. Furthermore, due to the order sensitivity of SMILES sequences, loss calculation and optimization can effectively prevent the model from generating incorrect sequences that do not conform to chemical syntax, such as unclosed loops or incorrect bond connections.
[0074] During the practice of this application, the inventors also explored some potential technical implementation methods. In one potential implementation example, an end-to-end learning system was built, continuing the idea of deep learning methods, using a large language model to learn the mapping relationship between mass spectra and molecular structures. The advantage of this version is that the process is direct and simple to implement, which facilitates the rapid construction of models and the completion of preliminary experimental verification. However, since it does not break away from the paradigm of deep learning methods, it also has the disadvantages of insufficient generalization ability and complex new data expansion.
[0075] In another potential implementation, a database search method is introduced during the inference phase of the large language model, using retrieved similar samples as prompts to the model. This approach has the advantage of using a reference database during the model inference phase, simplifying the addition of new data. However, during fine-tuning, the model does not learn the mapping of samples with similar mass spectra to the target structure. Therefore, the model cannot effectively utilize this prompt information to complete structural analysis during inference, limiting its generalization capabilities.
[0076] In another potential implementation, a small deep learning model is used as an interface for mass spectrometry to introduce a large language model. Imagine using a pre-trained small model to perform preliminary processing on the mass spectrometry data, extract key features, and then provide these features as input to the large language model. This small model can be trained quickly with limited computing resources and is easy to optimize. However, if the small model fails to fully capture all relevant information from the mass spectrometry data, the large language model may receive missing information, affecting the accuracy of the final result.
[0077] Through the embodiments of the present application, by retrieving mass spectra similar to the target mass spectra and their corresponding molecular structures, and introducing them as a small sample of information into the model fine-tuning and inference process, the model can better capture the intrinsic connection between similar mass spectra and their corresponding molecular structures, thereby improving the model's generalization ability in mass spectrometry structure analysis tasks. After the model fine-tuning is completed, the expansion of new data can be completed by updating the reference database in the inference process without retraining, which has greater flexibility.
[0078] In Spectral Similarity-Enhanced Supervised Fine-Tuning (S3FT), this method combines the advantages of database retrieval and deep learning paradigms in mass spectrometry structure analysis. By utilizing the few-sample learning ability of the large language model, the method provided in this application effectively captures the intrinsic connection between similar mass spectra and their corresponding molecular structures, thereby improving the generalization ability and flexibility of the model in the reasoning stage. Experimental results on simulated data show that S3FT outperforms small models and conventional large language models, reaching the most advanced performance level. In addition, the transferability of the method provided in the embodiment of this application was verified by experimental data, and training-free migration from simulated data to experimental data was achieved.
[0079] Figure 5 An operational diagram of an example of a supervised fine-tuning method based on spectral similarity enhancement is shown.
[0080] like Figure 5 As shown, the S3FT method aims to leverage the few-shot learning potential of large language models for mass spectrometry structure elucidation. Specifically, it combines the advantages of database matching and deep learning by retrieving similar mass spectrum-SMILES pairs and using them as input along with the target mass spectrum, thereby enhancing the model's generalization and flexibility. Furthermore, S3FT was evaluated on simulated data, demonstrating that the method achieved state-of-the-art performance (SOTA), outperforming both small models and conventional large language models. Furthermore, an ablation study explored the impact of the number and order of few-shot examples on S3FT's performance. Finally, the transferability of S3FT was verified on experimental data. Models trained using S3FT on simulated data demonstrated strong generalization capabilities on experimental data, achieving transfer without training. The following sections will also explore the impact of different similarity measures on model performance.
[0081] This paper reviews some of the current related technologies. In mass spectrometry structure analysis based on database matching, database matching involves comparing an unknown mass spectrum with known mass spectra in the database to find the best match. However, this approach faces some challenges due to possible discrepancies between mass spectrometry information and molecular structure similarity. To address this issue, researchers are committed to building a reasonable mapping relationship between the two. For example, the DeepMASS model uses a multi-layer perceptron (MLP) to extract features from mass spectra and achieves reliable predictions with a mean absolute error (MAE) of only 0.082. The Spec2Vec method converts mass spectrometry peaks into text to evaluate the similarity of mass spectra. MS2DeepScore predicts molecular structure similarity directly from mass spectra.
[0082] In deep learning-based mass spectrometry structure analysis, MS2Mol is a de novo structure prediction model that uses a Transformer model to generate sequence-to-sequence data. Mass2SMILES uses a Transformer encoder and a temporal convolutional decoder to infer the SMILES string and the number of specific functional groups present in the molecule. It can also predict the type of appendages and estimate the molecular formula. Spec2Mol draws on the idea of speech-to-text (Speech2Text). It first pre-trains a pair of SMILES encoders and decoders on a large-scale SMILES corpus, and then trains the mass spectrometry encoder to align the mass spectrometry encoding with the corresponding SMILES encoding. MSNovelist uses the existing SIRIUS and CSI:FingerID methods to obtain molecular fingerprints and molecular formulas from mass spectra. These data are then used as input to an encoder-decoder RNN model to predict SMILES.
[0083] In the examples of this application, it is proposed for the first time to apply large language models (LLMs) to mass spectrometry structure analysis, which has advantages over traditional deep learning methods.
[0084] In the mass spectrometry structure elucidation task, the goal is to infer the molecular structure from a given mass spectrum. Molecular structures are usually represented by SMILES. Therefore, the mass spectrometry structure elucidation task can be modeled as follows:
[0085] , Formula (3)
[0086] in, Representative i The mass-to-charge ratio (m / z) of each peak, Represents the corresponding peak intensity, the goal is to predict the SMILES string describing the target molecular structure .
[0087] In practical applications, researchers don't rely solely on a single mass spectrum to infer molecular structure. Instead, they seek out reference mass spectrum-SMILES pairs similar to the target mass spectrum. These reference pairs provide a correspondence between mass spectrum peaks and molecular substructures, providing valuable insights into the target structure.
[0088] Database matching searches the entire database during inference, leveraging the tendency for similar mass spectra to share similar structures. However, because database matching methods do not involve learning the mapping from mass spectra to SMILES, they lack the ability to infer novel molecules. In contrast, deep learning methods learn the mapping from mass spectra to SMILES during training, but their end-to-end learning model makes it difficult to capture the inherent connections between similar mass spectra. Furthermore, they are unable to provide similar mass spectra and their corresponding SMILES for the target mass spectrum during inference. Relying solely on mass spectral knowledge acquired during training makes these models challenging to infer novel molecules, which is one reason why existing de novo structure generation methods struggle to achieve outstanding performance.
[0089] For example, Spec2Mol's end-to-end mass spectrometry structure elucidation performance on its test data was only 0.8%, meaning it only generated one correct structure out of 125 mass spectra. MSNovelist's accuracy is higher because its molecular fingerprint prediction system is carefully designed and trained on a large amount of data. It uses the SIRIUS tool to accurately determine the molecular formula of the target mass spectra. Therefore, MSNovelist is more like solving a structure elucidation task based on mass spectra and molecular formulas, which is much simpler than pure mass spectrometry structure elucidation.
[0090] The S3FT method proposed in the present embodiment aims to utilize the small sample learning potential of large language models (LLMs) for mass spectrometry structure analysis. Figure 5 As shown in the figure, it is divided into three main stages: similar mass spectrum search, input text assembly and model fine-tuning.
[0091] Similar mass spectrum search: For each mass spectrum in the training dataset, the most similar mass spectrum is first found in the reference database (i.e., the training dataset) using the mass spectrum similarity algorithm. k The calculation of mass spectrum similarity is much more complicated than that of text similarity. The mass spectrum similarity algorithm includes cosine similarity, and its specific calculation formula can refer to the description of formula (1) above.
[0092] It should be noted that in order to prevent the model from being lazy by directly outputting the corresponding SMILES during training, the mass spectra that are identical to the target mass spectrum are first removed during retrieval, thereby ensuring that the model can extract more fine-grained information.
[0093] Input text assembly: After retrieving similar mass spectra, the ordinates of all mass spectra are removed. Because molecular structural information is mainly embedded in the abscissa of the mass spectrum peak, the abscissa represents the mass-to-charge ratio of the ion fragments and has chemical significance. The ordinate represents the peak intensity and contributes less to structural information. Since preprocessing steps such as noise peak filtering have optimized the use of peak intensities, retaining these intensity values will only introduce a lot of noise and increase the cost of fine-tuning and inference, so the peak intensity is removed. Ultimately, the task instructions, similar mass spectrum-SMILES pairs, and the current mass spectrum are concatenated as input text, represented as:
[0094] , Formula (4)
[0095] in, It's a mission instruction. It is j similar mass spectrum-SMILES pairs, is the target mass spectrum.
[0096] Model fine-tuning: Fine-tune the LLM using the assembled input text and target SMILES string. The fine-tuning process involves updating the model parameters to minimize the loss function, which can be a cross-entropy loss function. For more details, please refer to the description above in conjunction with Equation (2).
[0097] Model evaluation: During evaluation, the input text is assembled in the same way as during fine-tuning, and the LLM is used to predict the SMILES string of the target mass spectrum.
[0098] The S3FT method provided in the embodiments of this application combines the advantages of database matching and deep learning methods, and has the following key advantages:
[0099] Improved generalization: S3FT utilizes both the target mass spectrum and similar mass spectra and their corresponding molecular structures during training. This enables the model to more effectively capture the fine-grained correspondence between mass spectral peaks and molecular substructures. As a result, S3FT exhibits superior generalization compared to traditional deep learning methods, making it more likely to infer novel molecular structures.
[0100] Greater Flexibility: The S3FT method is designed to leverage the potential of large language models (LLMs) for few-shot learning in mass spectrometry structure elucidation. Once this capability is activated through fine-tuning, S3FT demonstrates excellent data scalability. To address the challenge of continuously increasing data, which deep learning methods struggle to cope with, the S3FT method only requires updating the reference database during inference, without retraining the model.
[0101] In order to further demonstrate the advanced nature of the method of the embodiment of the present application, the details of the experimental part of the S3FT method will be expanded below.
[0102] In the embodiments of the present application, the experiment includes three parts: performance evaluation of S3FT on simulated data, ablation study on the impact of the number and order of few samples on the performance of S3FT, and verification of the migration ability of S3FT on experimental data. First, the S3FT experiment is performed on simulated data based on Llama3-8B-Instruct. The reason for choosing simulated data is that the experimental data is small in scale and noisy, while simulated data is easy to obtain and has almost no noise. Performing S3FT on simulated data allows the model to better learn the mapping relationship between mass spectral peaks and molecular substructures without being disturbed by noise. After S3FT, the performance of the model is evaluated on simulated data. Next, the impact of the number and order of few samples in the input text on the performance of S3FT will be described. Finally, without further training, the performance of the model on experimental data is evaluated to confirm its strong generalization ability in practical applications.
[0103] Due to the lack of a unified mass spectrometry structure elucidation benchmark, two types of mass spectrometry data were manually collected: simulated data and experimental data. The simulated data were calculated using CFMID 4.0 (Wang et al., 2021) and randomly partitioned to produce 1 million training data and 4,000 test data. The experimental data were divided into two parts: one containing 20,000 training data points collected from the GNPS and MoNA databases, and the other containing 127 test data points from the CASMI 2016 competition. All mass spectrometry data were collected in positive ion mode, and all molecular structures in both test sets were removed from both training sets.
[0104] Table 1: Composition of the dataset.
[0105]
[0106] In terms of evaluation metrics, because mass spectrometry structure elucidation can sometimes correspond to multiple SMILES for a single mass spectrum, all models are configured to generate a single SMILES list. To assess the match between the predicted SMILES list and the label SMILES, the evaluation metrics defined by Spec2Mol were used. Key metrics include smiles acc and formula acc. Smiles acc measures the accuracy of correctly identifying the label molecule in the SMILES list, while formula acc assesses the accuracy of correctly identifying the label formula. Detailed metrics are shown in Table 2.
[0107] Table 2: Evaluation metrics for inference performance.
[0108]
[0109] Based on the training set of simulated data, S3FT experiments were conducted on the Llama3-8B-Instruct model. The number of given few-shot examples was set to three, and the order was arranged in sequence, that is, the most similar few-shot examples were placed at the bottom of the few-shot part.
[0110] Baseline model:
[0111] - Token model: Use common token encoding to process mass spectrometry data (Elser et al., 2023), treat the horizontal axis of the mass spectrum as a string of tokens, and then generate SMILES through the Transformer encoder and decoder.
[0112] - Bin model: Using another commonly used bin encoding method (Litsa et al., 2023), the horizontal coordinate of the mass spectrum is discretized into a one-dimensional vector of fixed length, and then SMILES is generated through the CNN encoder and Transformer decoder.
[0113] - Original Llama3-8B-Instruct model.
[0114] - Llama3-8B-Instruct model fine-tuning in the zero-shot setting.
[0115] - Llama3-8B-Instruct Model fine-tuned on three random few-shot examples.
[0116] Figure 6 The following diagram shows the architecture of two baseline models. In the Token model, each peak (m / z value) in the mass spectrum is treated as a discrete token, such as 41, 43, 45, and so on, representing different mass-to-charge ratios. These m / z tokens are then fed as input sequences into the Transformer encoder, which converts them into high-dimensional feature representations. This step captures the order and relationships between different tokens to understand the spectral information. The encoded features are then passed to the Transformer decoder to generate the corresponding SMILES molecular structure.
[0117] In the bin model, the m / z values in a mass spectrum are divided into a fixed number of bins. Each bin uses a binary code to indicate the presence or absence of a mass value. For example, positions of m / z 43 and 99 are marked as 1, indicating the presence of these mass values in the spectrum. A convolutional neural network (CNN) is then used to encode these binary inputs. CNNs capture local features through a sliding window and excel at processing high-dimensional, sparse data. The CNN-encoded features are then passed to a Transformer decoder to generate the corresponding SMILES molecular structure.
[0118] In all LLM evaluations, the effects of zero-shot, random few-shot, and similar few-shot on performance are explored. For the similar few-shot evaluation, the number of given few-shot examples is three and the order is arranged sequentially.
[0119] Table 3: Performance of S3FT on simulated data.
[0120]
[0121] It should be noted that Token and Bin represent two baseline small models, respectively. RAW represents the original Llama3-8B-Instruct model, ZSFT represents the fine-tuning of the Llama3-8B-Instruct model in the zero-shot setting, RFSFT represents the fine-tuning of the Llama3-8B-Instruct model in the randomized few-shot setting, and S3FT represents the Llama3-8B-Instruct model using the S3FT method. In the evaluation setting, ZS represents zero-shot evaluation, RandFS represents randomized few-shot evaluation, and SimFS represents similar few-shot evaluation.
[0122] As shown in the experimental results in Table 3, the method of the embodiment of the present application outperforms the small model and other LLMs on the simulated dataset, achieving state-of-the-art performance (SOTA). In addition, the following conclusions are drawn based on the experimental results:
[0123] First, the misaligned formats of the training and evaluation phases weaken the ability of LLM. The two few-shot evaluations of ZSFT (7, 8) perform worse than their zero-shot evaluation (6), while the zero-shot evaluations of RFSFT (9) and S3FT (12) perform worse than their few-shot evaluations (10, 11 and 13, 14).
[0124] Second, models fine-tuned using random few-shot examples are unable to leverage few-shot information to assist the task. Random and similar few-shot evaluations of RFSFT (10, 11) yield essentially identical results, and this result is very close to the zero-shot evaluation of ZSFT (6). RFSFT struggles to extract task-relevant information from random few-shot examples during training, and thus gradually abandons the weight of the few-shot portion in the mass spectrometry structure resolution task, degenerating into learning only the mapping from the target mass spectrum to the target structure.
[0125] Finally, the S3FT paradigm has the potential to be a comprehensive approach for LLMs to outperform small models. The ability to leverage similar few-shot examples to aid task completion is likely applicable to a wide range of other tasks.
[0126] To further explore the impact of the number and order of few-shot examples on S3FT performance, this paper describes the details of an ablation study involving six different settings. These settings vary in the number of few-shot examples (3, 5, or 7) and their order in the input text (sequential or reverse order).
[0127] Table 4: Performance of S3FT on experimental data.
[0128]
[0129] As shown in Table 4, the S3FT model is fine-tuned on the simulated data and directly evaluated on the experimental data by modifying only the reference database. “Generate num” represents the length of the SMILES list generated in the evaluation phase, which is set to 10, 30, and 50 as reference.
[0130] The performance of S3FT in each setting is evaluated on a simulated dataset. The reported evaluation metrics include smiles acc, formula acc, and max fngp cos (maximum fingerprint cosine similarity).
[0131] Table 5: Results of ablation experiments.
[0132]
[0133] As shown in Table 5, Sequential and Reversed respectively place the most similar few-shot (FS) examples at the bottom or top of the few-shot part. The results of the ablation study are consistent with the expected hypothesis and provide valuable insights into the behavior of S3FT:
[0134] - Sequential order is better than reverse order: When a few examples with high similarity are closer to the target mass spectrum, the model is more likely to pay more attention to their internal information. Since similar mass spectra are more likely to have similar structures, increasing the weight of this information will undoubtedly help in inferring the target structure.
[0135] When the order of the minority samples is reversed, increasing their number actually weakens the model's performance: as the number of minority samples increases, the more similar the samples are, the further away they are from the target mass spectrum. The model tends to focus on the minority samples closest to the target mass spectrum, but this information is less helpful in inferring the target structure. As a result, in subsequent training stages, the model gradually ignores the minority sample information, causing it to degenerate into a model similar to that used in the ZSFT setting.
[0136] When the order of the few-shot examples is sequential, increasing their number improves model performance: as the number of few-shot examples increases, the distance between the most similar examples and the target mass spectrum remains constant. However, the additional examples provide more information to the model, ultimately leading to better performance.
[0137] It should be noted that the best performing model in the ablation study, the S3FT model using 7 few-shot examples arranged sequentially, was used to directly evaluate the experimental data to verify the strong generalization ability of S3FT. There are three baseline models:
[0138] - Spec2Mol, based on an encoder-decoder architecture, where the encoder learns mass spectra embeddings and the decoder is pre-trained on a large dataset of chemical structures.
[0139] - Token model, which uses token encoding to process mass spectra and is trained on simulated data.
[0140] - Token model, which uses token encoding to process mass spectrometry. After training on simulated data, it is subsequently trained on experimental data.
[0141] In terms of evaluation, for the S3FT model, the embodiments of this application explore the impact of different similarity metrics on model performance. The main similarity metrics include Cosine, MS2DeepScore, TokenEmbedding, and BinEmbedding. Cosine calculates the cosine similarity between mass spectra. MS2DeepScore is developed based on a machine learning algorithm. TokenEmbedding and BinEmbedding use intermediate vectors derived from encoders of token and bin models trained on simulated data as input for cosine similarity calculation. During the evaluation process, the reference database is set to the training set of experimental data.
[0142] As shown in Table 4, the model trained on simulated data using S3FT demonstrated strong generalization capabilities on experimental data. During the evaluation phase, the model was able to achieve untrained transfer from simulated to experimental data simply by changing the reference database. Different similarity metrics influenced model performance during the evaluation process. Using the BinEmbedding setting, the model achieved the highest values for smiles acc and formula acc, significantly outperforming other metrics. This demonstrates that BinEmbedding more effectively captures the inherent similarities between mass spectra, thereby improving the model's performance in structure elucidation.
[0143] Figure 7 A schematic diagram showing the effect of an example of a target structure and its corresponding few-sample structure.
[0144] To demonstrate how few-shot examples can provide valuable reference information for target structures, this embodiment of the application conducts a case study in the evaluation of experimental data. Figure 7 The target structure and its corresponding minority-sample structure are shown, with the shared substructures between the target and minority-sample structures highlighted. The figure clearly shows that the minority-sample structures contain different substructures of the target structure. For example, minority-sample structures 1, 2, and 4 contain substructures such as hydroxyl groups, carboxyl groups, and benzene rings, while minority-sample structures 3 and 6 contain the rest of the target structure. This diversity enables the model to infer the target structure by piecing together information from the minority-sample examples.
[0145] The S3FT method provided in the examples of this application combines the advantages of database matching and deep learning for mass spectrometry structure elucidation tasks. Experimental results on simulated data demonstrate that S3FT outperforms small models and conventional large language models, achieving state-of-the-art performance (SOTA). Furthermore, the examples of this application validate the method's ability to transfer to experimental data by directly evaluating the S3FT model trained on simulated data. The model demonstrates strong generalization capabilities on experimental data, achieving training-free transfer from simulated to experimental data. This demonstrates the potential of the methods of the examples of this application for real-world applications.
[0146] As a further optimization implementation, data from other spectroscopic techniques such as nuclear magnetic resonance (NMR) or infrared spectroscopy (IR) can also be integrated to further improve the accuracy and reliability of the structure elucidation process.
[0147] It should be noted that, for the aforementioned method embodiments, for the sake of simplicity of description, they are all expressed as a series of combined actions, but those skilled in the art should be aware that this application is not limited by the order of the actions described, because according to this application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by this application. In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, please refer to the relevant description of other embodiments.
[0148] In some embodiments, an embodiment of the present application provides a non-volatile computer-readable storage medium, which stores one or more programs including execution instructions, and the execution instructions can be read and executed by an electronic device (including but not limited to a computer, server, or network device, etc.) to execute any of the above-mentioned mass spectrometry structure analysis methods of the present application.
[0149] In some embodiments, the embodiments of the present application also provide a computer program product, which includes a computer program stored on a non-volatile computer-readable storage medium, and the computer program includes program instructions. When the program instructions are executed by a computer, the computer executes any of the above-mentioned mass spectrometry structure analysis methods.
[0150] In some embodiments, an embodiment of the present application also provides an electronic device, comprising: at least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform a mass spectrometry structure analysis method.
[0151] Figure 8 FIG. 1 is a schematic diagram of the hardware structure of an electronic device for performing a mass spectrometry structure analysis method according to another embodiment of the present application. Figure 8 As shown, the device includes:
[0152] One or more processors 810 and memory 820, Figure 8 A processor 810 is taken as an example.
[0153] The device for performing the mass spectrometry structure analysis method may further include: an input device 830 and an output device 840 .
[0154] The processor 810, the memory 820, the input device 830 and the output device 840 may be connected via a bus or other means. Figure 8 The bus connection is taken as an example.
[0155] Memory 820, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules, such as the program instructions / modules corresponding to the mass spectrometry structure analysis method in the embodiments of this application. Processor 810 executes the non-volatile software programs, instructions, and modules stored in memory 820 to execute various server functional applications and data processing, thereby implementing the mass spectrometry structure analysis method in the above-described method embodiment.
[0156] The memory 820 may include a program storage area and a data storage area, wherein the program storage area may store an operating system and applications required for at least one function; the data storage area may store data created based on the use of the electronic device, etc. In addition, the memory 820 may include a high-speed random access memory and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other non-volatile solid-state storage device. In some embodiments, the memory 820 may optionally include a memory remotely located relative to the processor 810, and these remote memories may be connected to the electronic device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0157] The input device 830 can receive input digital or character information and generate signals related to user settings and function control of the electronic device. The output device 840 can include a display device such as a display screen.
[0158] The one or more modules are stored in the memory 820 and, when executed by the one or more processors 810, perform the mass spectrometry structure analysis method in any of the above method embodiments.
[0159] The above-mentioned product can execute the method provided in the embodiment of this application, and has the functional modules and beneficial effects corresponding to the execution method. For technical details not fully described in this embodiment, please refer to the method provided in the embodiment of this application.
[0160] The electronic devices of the embodiments of the present application exist in various forms, including but not limited to:
[0161] (1) Mobile communication devices: These devices are characterized by their mobile communication capabilities and their primary purpose is to provide voice and data communications. These terminals include smartphones, multimedia phones, feature phones, and low-end phones.
[0162] (2) Ultra-mobile personal computer devices: These devices fall under the category of personal computers and have computing and processing capabilities, and generally also have mobile Internet access. These terminals include PDAs, MIDs, and UMPCs.
[0163] (3) Portable entertainment devices: These devices can display and play multimedia content. They include audio and video players, handheld game consoles, e-books, smart toys, and portable car navigation devices.
[0164] (4) Other onboard electronic devices with data interaction functions, such as onboard computer devices installed in vehicles.
[0165] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of this embodiment.
[0166] Through the description of the above embodiments, those skilled in the art will clearly understand that each embodiment can be implemented using software plus a general hardware platform, or of course, hardware. Based on this understanding, the essence of the above technical solution, or the portion that contributes to the relevant technology, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, or an optical disk, and includes a number of instructions for causing a computer device (such as a personal computer, server, or network device) to execute the methods described in each embodiment or certain portions of the embodiments.
[0167] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A mass spectrometry structure analysis method comprising: Calculating the similarity between each reference mass spectrum in the reference database and the target mass spectrum to be analyzed, and selecting at least one similar reference mass spectrum from the reference mass spectra according to the calculated similarity; The reference database comprises a plurality of reference mass spectra and corresponding reference molecular structures; Generate prompt information according to the preset mass spectrum structure analysis task instruction, the target mass spectrum and each of the similar reference mass spectra and the corresponding reference molecular structures; Inputting the prompt information into a large language model to determine the target molecular structure corresponding to the target mass spectrum; The step of generating prompt information according to the preset mass spectrum structure analysis task instruction, the target mass spectrum, and each of the similar reference mass spectra and the corresponding reference molecular structures includes: Extracting the first mass spectrum peak ordinate corresponding to the target mass spectrum, and extracting the corresponding second mass spectrum peak ordinate from each of the similar reference mass spectra; Prompt information is generated according to the mass spectrum structure analysis task instruction, the first mass spectrum peak ordinate, each of the second mass spectrum peak ordinates, and the corresponding reference molecular structure.
2. The method according to claim 1, wherein The molecular structure is represented by SMILES sequence.
3. The method according to claim 1, wherein The similarity is calculated by the following formula: Where Cosine socre represents the similarity between the reference mass spectrum and the target mass spectrum; I r,i and I q,i represent the intensities of the i-th peak in the reference mass spectrum and the target mass spectrum, respectively, and n is the number of matching peaks.
4. A method for fine-tuning a large language model for mass spectrometry structure analysis tasks, comprising: Acquire a training sample set, wherein each training sample in the training sample set comprises a training mass spectrum and a corresponding training molecular structure; Calculating the similarity between the first training mass spectrum and each second training mass spectrum in the training sample set, and selecting at least one similar second training mass spectrum from each second training mass spectrum according to the calculated similarity; The first training mass spectrum and the second training mass spectrum belong to different training samples respectively; Generate prompt information according to the preset mass spectrum structure analysis task instruction, the first training mass spectrum, each of the similar second training mass spectra and the corresponding training molecular structures; Using the prompt information as model input information and the training molecular structure corresponding to the first training mass spectrum as corresponding label information to fine-tune the large language model; The step of generating prompt information according to the preset mass spectrometry structure analysis task instruction, the first training mass spectrum, each of the similar second training mass spectra, and the corresponding training molecular structures includes: Extracting the first mass spectrum peak ordinate corresponding to the first training mass spectrum, and extracting the corresponding second mass spectrum peak ordinate from each of the similar second training mass spectra; Prompt information is generated according to the mass spectrum structure analysis task instruction, the first mass spectrum peak ordinate, each of the second mass spectrum peak ordinates, and the corresponding training molecular structure.
5. The method according to claim 4, wherein The obtaining of the training sample set includes: Acquire multiple training mass spectra; For each of the training mass spectra, extracting a molecular fingerprint and a molecular formula corresponding to the training mass spectrum, and inputting the extracted molecular fingerprint and molecular formula into a SMILES model to determine a corresponding SMILES representation sequence; the SMILES model includes a Transformer encoder and a time convolution decoder; A training sample set is constructed based on each of the training mass spectra and the corresponding SMILES representation sequence.
6. The method according to claim 5, wherein: The fine-tuning loss function is defined as: Where L represents the loss function, is the predicted SMILES representation sequence length in characters, is the j-th character of the SMILES representation sequence predicted by the large language model, and S[j] is the j-th character of the SMILES representation sequence corresponding to the label information.
7. A storage medium having a computer program stored thereon, wherein: When the program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.
8. An electronic device comprising: At least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the steps of the method according to any one of claims 1 to 6.
9. A computer program product comprising a computer program / instruction, which implements the steps of the method according to any one of claims 1 to 6 when the computer program / instruction is executed by a processor.
Citation Information
Patent Citations
Test case set generation method, device and equipment and computer readable storage medium
CN111708703A
Extended corpus generation method and device in target field and electronic equipment
CN112541076A