Construction method and device of non-target public spectrogram database and electronic equipment

By integrating multiple open source spectral databases, supplementing and filtering data, and constructing a non-target public spectral database, the problem of inconsistent information in different databases was solved, the number and quality of spectra were improved, and the completeness and accuracy of metabolite annotation information were enhanced.

CN120673860APending Publication Date: 2025-09-19BEIJING NOVOGENE TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510871597.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-26
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

Existing non-targeted metabolomics spectrum databases cannot be used directly due to different research objects, missing, redundant and conflicting information, which affects the precision and accuracy of metabolite detection.

Method used

By acquiring multiple open source spectral databases, extracting metabolite information, supplementing and filtering the data, and integrating them into a non-target public spectral database, the quantity and quality of spectra can be improved.

Benefits of technology

The increase in the number of spectra and the improvement in their quality have enhanced the completeness and accuracy of metabolite annotation information, covering different biological species such as humans, animals, plants, and microorganisms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120673860A_ABST
    Figure CN120673860A_ABST
Patent Text Reader

Abstract

The invention provides a construction method and device of a non-target public spectrogram database and electronic equipment. The construction method comprises the steps of obtaining a plurality of open source spectrogram databases; extracting metabolite information of each metabolite from a plurality of open source spectrogram databases; wherein the metabolite information at least comprises a spectrogram and annotation information; performing data supplementation and data filtering on the extracted metabolite information to obtain target metabolite information; and integrating the target metabolite information to obtain a non-target public spectrogram database. According to the method, the number and quality of the spectrograms in the spectrogram database are improved, and meanwhile, the integrity and accuracy of metabolite annotation information are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of non-targeted metabolomics, and in particular to a method, device and electronic equipment for constructing a non-target public spectrum database. Background Art

[0002] Untargeted metabolomics requires detecting as many metabolites as possible in biological samples. The number of metabolites and the quality of the spectra in the spectral database directly affect the precision and accuracy of metabolite detection. Currently, there are many open-source spectral databases for metabolomics available online, but different databases target different research subjects (e.g., humans, animals, plants, microorganisms, etc.), and contain missing, redundant, and conflicting information (e.g., the same metabolites may exist in humans and animals, but different databases may describe them differently), making them difficult to use directly. Summary of the Invention

[0003] In view of this, the purpose of the present invention is to provide a method, device and electronic device for constructing a non-target public spectrum database, so as to improve the quantity and quality of spectra in the spectrum database, and at the same time improve the completeness and accuracy of metabolite annotation information.

[0004] In order to achieve the above object, the technical solution adopted by the present invention is as follows: In a first aspect, the present invention provides a method for constructing a non-target public spectrogram database, comprising: obtaining multiple open source spectrogram databases; extracting metabolite information of each metabolite from the multiple open source spectrogram databases; wherein the metabolite information includes at least: spectra and annotation information; performing data supplementation and data filtering on the extracted metabolite information to obtain target metabolite information; and integrating the target metabolite information to obtain a non-target public spectrogram database.

[0005] Optionally, extracting metabolite information of each metabolite from multiple open source spectrogram databases includes: searching for corresponding data from multiple open source spectrogram databases according to preset keywords to obtain the metabolite information of each metabolite.

[0006] Optionally, data supplementation and data filtering are performed on the extracted metabolite information to obtain target metabolite information, including: supplementing the metabolite information based on preset data supplementation rules; filtering the metabolite information based on preset data filtering rules.

[0007] Optionally, the metabolite information is supplemented based on preset data supplementation rules, including: supplementing missing information in the metabolite information based on the ID index association between different open source spectrum databases; if there are differences between the metabolite information of the same metabolite extracted from different open source spectrum databases, retaining the metabolite information according to the first database priority order of the preset open source spectrum database; among them, the first database priority of the PubChem database is the highest, followed by the HMDB database; if the metabolite information is missing ClassI~III classification data, the ClassI~III classification data is searched from the ClassyFire database based on InChIKey; if the metabolite information is missing Chinese name, the Chinese name in the preset database shall prevail, or the Chinese name shall be obtained using the preset AI model.

[0008] Optionally, the metabolite information is filtered based on preset data filtering rules, including: if the English name of the metabolite is the same as the secondary spectrum, then the metabolite from the open source spectrum database with the highest priority is retained according to the second database priority order of the preset open source spectrum database; the metabolites whose mass-to-charge ratio of the secondary spectrum is not within the first preset range are filtered; the metabolites whose mass-to-charge ratio has less than a preset number of digits after the decimal point in the secondary spectrum are filtered; the metabolites without daughter ions in the secondary spectrum are filtered; the metabolites whose ion number exceeds the ion threshold are filtered; the metabolites whose secondary spectrum intensity is zero are filtered. Filter metabolites; for metabolites whose secondary spectrum intensity is within the second preset range, filter metabolites whose intensity does not include the first intensity; for metabolites whose secondary spectrum intensity is greater than the first intensity, filter metabolites whose intensity is less than or equal to the second intensity; filter metabolites with the same mass-to-charge ratio in the secondary spectrum; modify metabolites whose summed ion form does not meet the preset requirements; filter metabolites whose English names exceed the preset length and whose English names are garbled; if the InChIKey of the metabolites is the same, metabolites from the HMDB database are retained first.

[0009] Optionally, integrating the target metabolite information to obtain a non-target public spectrum database also includes: integrating the target metabolite information into an XCMS recognition format, and integrating the annotation information in the target metabolite information into an information analysis process recognition format to obtain a non-target public spectrum database.

[0010] In a second aspect, the present invention provides a device for constructing a non-target public spectrogram database, comprising: a data acquisition module for acquiring multiple open source spectrogram databases; an information extraction module for extracting metabolite information of each metabolite from multiple open source spectrogram databases; wherein the metabolite information includes at least: spectra and annotation information; a data processing module for supplementing and filtering the extracted metabolite information to obtain target metabolite information; and a database construction module for integrating the target metabolite information to obtain a non-target public spectrogram database.

[0011] Optionally, the information extraction module is specifically used to: according to preset keywords, search corresponding data from multiple open source spectrum databases to obtain metabolite information of each metabolite.

[0012] In a third aspect, the present invention provides an electronic device comprising a processor and a memory, wherein the memory stores computer-executable instructions that can be executed by the processor, and the processor executes the computer-executable instructions to implement the steps of any one of the methods provided in the first aspect above.

[0013] In a fourth aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the computer program executes the steps of any one of the methods provided in the first aspect.

[0014] The present invention brings the following beneficial effects: The method, device, and electronic device for constructing the non-target public spectrogram database provided by the present invention first obtain multiple open source spectrogram databases; then extract metabolite information for each metabolite from the multiple open source spectrogram databases; wherein the metabolite information includes at least: spectra and annotation information; then the extracted metabolite information is supplemented and filtered to obtain target metabolite information; finally, the target metabolite information is integrated to obtain a non-target public spectrogram database. In the above method, data from multiple Kaiyuan spectrogram databases can be integrated, covering different biological species such as humans, animals, plants, and microorganisms, thereby increasing the number of spectra; at the same time, through data filtering and data supplementation, low-quality spectra are filtered out and missing information is supplemented, thereby improving the quality of spectra and the integrity and accuracy of metabolite annotation information.

[0015] Other features and advantages of the present invention will be described in the following description, and in part will become apparent from the description, or understood by practicing the present invention. The purposes and other advantages of the present invention are realized and obtained by the structures particularly pointed out in the description, claims and drawings.

[0016] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the preferred embodiments are specifically listed below and described in detail with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the specific embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0018] Figure 1 A flowchart of a method for constructing a non-target public spectrum database provided in an embodiment of the present invention; Figure 2 A schematic structural diagram of a device for constructing a non-target public spectrum database provided by an embodiment of the present invention; Figure 3 A schematic structural diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0019] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of them. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0020] Currently, non-targeted metabolite identification in metabolites mainly relies on two types of databases: one is similar to the commercial software Compound Discoverer, which is equipped with the mzCloud spectral database and the internal Thermo Scientific mzVault spectral database; the other is the company's own spectral database. However, the information accuracy of the spectral database equipped with Compound Discoverer is not high, and the number of substances that can be accurately identified in the sample is relatively small. In addition, due to different chromatographic conditions, parameters such as RT cannot be reused, which may have an adverse effect on the results. The company's own spectral database is a non-public database and cannot be used directly. The open source spectral databases on the Internet are oriented to different research objects (for example, only for humans, animals, plants, microorganisms, etc.), and the information contained is missing, redundant, and conflicting (for example, the same metabolites exist between humans and animals, but the descriptions in different databases are different), and cannot be used directly.

[0021] Based on this, the embodiments of the present invention provide a method, device, and electronic device for constructing a non-target public spectrogram database, which can improve the quantity and quality of spectra in the spectrogram database, while improving the completeness and accuracy of metabolite annotation information.

[0022] To facilitate understanding of this embodiment, a method for constructing a non-target public spectrum database disclosed in an embodiment of the present invention is first introduced in detail. This method can be executed by electronic devices such as smart phones, computers, tablet computers, etc. Figure 1 The flowchart of a method for constructing a non-target public spectrum database is shown, which illustrates that the method mainly includes the following steps S101 to S103: Step S101: Acquire multiple open source spectrogram databases.

[0023] In one embodiment, the following open source spectrogram databases may be obtained: The Reference Metabolome Database for PlantsefMetaPlant (RefMetaPlant, https: / / www.biosino.org / RefMetaDB / ); The Human Metabolome Database (HMDB, https: / / hmdb.ca / ); The Microbial Metabolites Database (MiMeDB, https: / / mimedb.org / ); RIKEN MSn spectral database for phytochemicals (ReSpect, http: / / spectra.psc.riken.jp / ); MassBank of North America (MoNA, https: / / mona.fiehnlab.ucdavis.edu / ).

[0024] Step S102: extracting metabolite information of each metabolite from multiple open source spectrum databases.

[0025] In one embodiment, metabolite information for each metabolite can be obtained by searching for corresponding data from multiple open source spectrogram databases according to preset keywords. The metabolite information includes at least: spectra and annotation information, and missing data is replaced by -. Specifically, the preset keywords include at least: Name: English name of metabolite; ChineseName: Chinese name of metabolite; IonMode: acquisition mode, P means positive mode acquisition, N means negative mode acquisition; MSMS: secondary spectrum; m / z: mass-to-charge ratio; CE: Collision Energy; Adduct: adduct ionic form; Formula: molecular formula of metabolite; CAS: metabolite CAS number; Molecular Weight: molecular weight of metabolites; Other_Name: other English names of metabolites; ClassI: first-level classification of metabolites; ClassII: secondary classification of metabolites; ClassIII: third-level classification of metabolites; KEGG_ID: metabolite KEGG database number; KEGG_MapID: KEGG database pathway number where the metabolite is located; HMDB_ID: metabolite HMDB database number; PubChemID: metabolite PubChem database number; SMILES: originated from the PubChem database and uses a single line of text to express the structure of a compound; InChIKey: Derived from the PubChem database, it represents a molecular representation with a fixed length of 25 characters; Lipidmaps_ID: metabolite Lipidmaps database ID.

[0026] Step S103: performing data supplementation and data filtering on the extracted metabolite information to obtain target metabolite information.

[0027] In one embodiment, in order to improve the accuracy of the data, the extracted metabolite information can be supplemented and filtered according to preset rules to remove low-quality data and improve the accuracy and completeness of the data.

[0028] Step S104: Integrate the target metabolite information to obtain a non-target public spectrum database.

[0029] In one embodiment, the target metabolite information can be integrated into an XCMS recognition format, and the annotation information in the target metabolite information can be integrated into an information analysis process recognition format to obtain a non-target public spectrum database.

[0030] The method for constructing the above-mentioned non-target public spectrum database provided by the embodiment of the present invention can integrate data from multiple Kaiyuan spectrum databases, covering different biological species such as humans, animals, plants, and microorganisms, thereby increasing the number of spectra; at the same time, through data filtering and data supplementation, low-quality spectra are filtered out, and missing information is supplemented, thereby improving the spectrum quality and the integrity of metabolite annotation information.

[0031] In one embodiment, in the aforementioned step S103, i.e., when the extracted metabolite information is supplemented and filtered to obtain target metabolite information, the following methods may be used, including but not limited to: First, the metabolite information is supplemented based on the preset data supplementation rules.

[0032] In specific implementation, the metabolite information is supplemented based on the preset data supplementation rules, which at least includes the following processes: (1) Based on the ID index association between different open source spectral databases, the missing information in the metabolite information is supplemented. Specifically, through the ID index association of different open source spectral databases, such as KEGG_ID, HMDB_ID, PubChemID, Lipidmaps_ID, etc., the missing information in the metabolite information is supplemented. For example, if metabolite A extracted from the HMDB database is missing information a, and metabolite A in the PubChem database is retrieved based on HMDB_ID and includes information a, then the information extracted from the HMDB database is supplemented based on the metabolite A in the PubChem database including information a.

[0033] (2) If there are differences in the metabolite information of the same metabolite extracted from different open source spectral databases, the metabolite information will be retained according to the preset first database priority order of the open source spectral databases; among them, the first database priority of the PubChem database is the highest, and the HMDB database is second. Specifically, if there is a conflict in information (i.e., the metabolite information extracted from different open source spectral databases is different), the metabolite information will be determined according to the first database priority order of the open source spectral databases. Specifically, the first database priority order is: the information in the PubChem database is first, and the information in the HMDB database is second.

[0034] (3) If the metabolite information is missing Class I-III classification data, the Class I-III classification data will be searched from the ClassyFire database based on the InChIKey. Specifically, for missing Class I-III classification data, the ClassyFire (http: / / classyfire.wishartlab.com / ) classification will be used as the basis, and the InChIKey can be used to enter https: / / cfb.fiehnlab.ucdavis.edu / for search.

[0035] (4) If the Chinese name of the metabolite is missing, the Chinese name in the preset database shall prevail, or the Chinese name shall be obtained using the preset AI model. Specifically, for data with missing Chinese names, the Chinese name on Chemsrc.com (https: / / www.chemsrc.com / ) shall prevail, followed by the Chinese name obtained by translation using the AI ​​model.

[0036] Then, the metabolite information is filtered based on the preset data filtering rules.

[0037] In specific implementation, filtering metabolite information based on preset data filtering rules includes at least the following processes: (1) If the English name and secondary spectrum of the metabolite are the same, the metabolite from the open source spectrum database with the highest priority will be retained according to the preset second database priority order of the open source spectrum database. Specifically, if the name and MSMS of the metabolite are exactly the same, the metabolite from the open source spectrum database with the highest priority will be retained according to the second database priority order of the open source spectrum database. The second database priority order is: HMDB database, MiMeDB database, MoNA database, ReSpect database, and RefMetaPlant database.

[0038] (2) Filter out metabolites whose mass-to-charge ratios in the secondary spectra are not within the first preset range. Specifically, the mass spectrometer scan range is 50-1500 (i.e., the first preset range), and filter out metabolites whose mass-to-charge ratios are outside this range.

[0039] (3) Filter out the metabolites whose mass-to-charge ratio in the secondary spectrum has less than a preset number of digits after the decimal point. Specifically, filter out the metabolites whose mass-to-charge ratio in the secondary spectrum has less than four decimal places after the decimal point.

[0040] (4) Filter out metabolites that have no daughter ions in the secondary spectrum. Specifically, filter out metabolites that have only parent ions but no daughter ions in the secondary spectrum.

[0041] (5) Filter out metabolites whose ion counts exceed the ion threshold. Specifically, filter out metabolites whose ion counts are greater than 20,000 groups of ions.

[0042] (6) Filter out metabolites whose secondary spectrum intensity is zero. Specifically, filter out metabolites whose secondary spectrum intensity is zero.

[0043] (7) For metabolites whose intensities in the secondary spectra are within the second preset range, the metabolites whose intensities do not include the first intensity are filtered out. Specifically, for metabolites whose intensities are in the range of 1 to 100 (i.e., the second range), the metabolites whose intensities do not include 100 (i.e., the first intensity) are filtered out.

[0044] (8) For metabolites with intensities greater than the first intensity in the secondary spectrum, metabolites with intensities less than or equal to the second intensity are filtered out. Specifically, for metabolites with intensities greater than 100, metabolites with intensities less than or equal to 5000 (i.e., the second intensity) are filtered out.

[0045] (9) Filter out metabolites with the same mass-to-charge ratio in the secondary spectra. Specifically, if the mass-to-charge ratio in the secondary spectra is repeated, the metabolite is filtered out.

[0046] (10) Modify metabolites whose summed ion form does not meet the preset requirements. Specifically, the summed ion form needs to be written in the form of square brackets [] and judged as + or - according to ionMode (acquisition mode). If it does not meet the requirements, the summed ion form of the metabolite is modified. If the summed ion conflicts with the acquisition mode, it is filtered out.

[0047] (11) Filter out metabolites whose English names exceed the preset length and whose English names are garbled. Specifically, filter out metabolites whose English names are longer than 70 characters and whose English names contain special characters and are garbled.

[0048] (12) If the InChIKeys of metabolites are the same, the metabolites from the HMDB database will be retained first. Specifically, if the InChIKeys are exactly the same, the metabolites from the HMDB database will be retained first.

[0049] The above-mentioned method provided by the embodiment of the present invention can ensure the balance between the number of metabolites and the quality of spectra through data supplementation and data filtering, and maximize the completeness of metabolite annotation information. Compared with existing databases, the non-target public spectra database constructed by the embodiment of the present invention has the following advantages: (1) a large number of metabolites and spectra, integrating five open source spectra databases covering different biological species such as humans, animals, plants, and microorganisms, and achieving more than 374,000 spectra and more than 147,000 substances; (2) through data filtering rules, low-quality spectra are filtered out, and the spectra quality is high; (3) through data supplementation rules, missing information is supplemented to the greatest extent, and the metabolite annotation information is more complete.

[0050] In an embodiment of the present invention, since the spectrum database itself contains a wide range of species, if only a certain type of species is studied, the database can be split according to the source, such as into an animal spectrum database or a plant spectrum database, and metabolite detection can be performed on a certain type of species.

[0051] Regarding the method for constructing a non-target public spectrum database provided in the above embodiment, the present invention also provides a device for constructing a non-target public spectrum database, see Figure 2 The schematic diagram of the structure of a non-target public spectrum database construction device is shown, indicating that the device mainly includes the following parts: A data acquisition module 201 is used to acquire multiple open source spectrogram databases; An information extraction module 202 is configured to extract metabolite information of each metabolite from a plurality of open source spectrogram databases; wherein the metabolite information includes at least spectra and annotation information; The data processing module 203 is used to supplement and filter the extracted metabolite information to obtain target metabolite information; The database construction module 204 is used to integrate the target metabolite information to obtain a non-target public spectrum database.

[0052] The above-mentioned non-target public spectrum database construction device provided by the embodiment of the present invention can integrate data from multiple Kaiyuan spectrum databases, covering different biological species such as humans, animals, plants, and microorganisms, thereby increasing the number of spectra; at the same time, through data filtering and data supplementation, low-quality spectra are filtered out, and missing information is supplemented, thereby improving the spectrum quality and the integrity of metabolite annotation information.

[0053] In one embodiment, the information extraction module 202 is specifically configured to: search for corresponding data from a plurality of open source spectrogram databases according to preset keywords to obtain metabolite information of each metabolite.

[0054] In one embodiment, the data processing module 203 is specifically configured to: supplement the metabolite information based on a preset data supplementation rule; and filter the metabolite information based on a preset data filtering rule.

[0055] In one embodiment, the data processing module 203 is specifically used to: supplement the missing information in the metabolite information based on the ID index association between different open source spectrogram databases; if there are differences between the metabolite information of the same metabolite extracted from different open source spectrogram databases, retain the metabolite information according to the first database priority order of the preset open source spectrogram database; among them, the first database priority of the PubChem database is the highest, followed by the HMDB database; if the ClassI~III classification data is missing in the metabolite information, the ClassI~III classification data is searched from the ClassyFire database based on InChIKey; if the Chinese name is missing in the metabolite information, the Chinese name in the preset database shall prevail, or the Chinese name shall be obtained using the preset AI model.

[0056] In one embodiment, the data processing module 203 is specifically used to: if the English name of the metabolite is the same as the secondary spectrum, retain the metabolite from the open source spectrum database with the highest priority according to the second database priority order of the preset open source spectrum database; filter the metabolites whose mass-to-charge ratio of the secondary spectrum is not within the first preset range; filter the metabolites whose mass-to-charge ratio has less than the preset number of digits after the decimal point in the secondary spectrum; filter the metabolites without daughter ions in the secondary spectrum; filter the metabolites whose ion number exceeds the ion threshold; filter the metabolites whose secondary spectrum intensity is zero Filtering; for metabolites whose secondary spectrum intensity is within the second preset range, filter out metabolites whose intensity does not include the first intensity; for metabolites whose secondary spectrum intensity is greater than the first intensity, filter out metabolites whose intensity is less than or equal to the second intensity; filter out metabolites with the same mass-to-charge ratio in the secondary spectrum; modify metabolites whose summed ion form does not meet the preset requirements; filter out metabolites whose English names exceed the preset length and whose English names are garbled; if the InChIKey of the metabolites is the same, give priority to retaining metabolites from the HMDB database.

[0057] In one embodiment, the database construction module 204 is specifically used to integrate the target metabolite information into an XCMS recognition format, and integrate the annotation information in the target metabolite information into an information analysis process recognition format to obtain a non-target public spectrum database.

[0058] It should be noted that the implementation principles and technical effects of the apparatus provided in the embodiments of the present invention are the same as those of the aforementioned method embodiments. For the sake of brevity, any details not mentioned in the apparatus embodiments are referred to the corresponding contents of the aforementioned method embodiments. The specific numerical values ​​provided in the implementation of the present invention are merely exemplary and are not intended to be limiting.

[0059] An embodiment of the present invention further provides an electronic device. Specifically, the electronic device includes a processor and a storage device. The storage device stores a computer program, and when the computer program is executed by the processor, it executes the method described in any one of the above embodiments.

[0060] Figure 3 This is a structural diagram of an electronic device provided in an embodiment of the present invention. The electronic device 100 includes: a processor 30, a memory 31, a bus 32 and a communication interface 33, wherein the processor 30, the communication interface 33 and the memory 31 are connected via the bus 32; the processor 30 is used to execute an executable module stored in the memory 31, such as a computer program.

[0061] Memory 31 may include high-speed random access memory (RAM) and may also include non-volatile memory, such as at least one disk storage device. Communication between the system network element and at least one other network element is achieved through at least one communication interface 33 (which may be wired or wireless), and may utilize the Internet, a wide area network, a local area network, a metropolitan area network, or the like.

[0062] The bus 32 may be an ISA bus, a PCI bus, or an EISA bus. The bus may be divided into an address bus, a data bus, a control bus, and the like. For ease of representation, Figure 3 Only one bidirectional arrow is used in the diagram, but this does not mean that there is only one bus or one type of bus.

[0063] Among them, the memory 31 is used to store programs, and the processor 30 executes the program after receiving the execution instruction. The method executed by the device for flow process definition disclosed in any embodiment of the above-mentioned embodiment of the present invention can be applied to the processor 30 or implemented by the processor 30.

[0064] The processor 30 may be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method may be completed by hardware integrated logic circuits or software instructions in the processor 30. The processor 30 may be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it may also be a digital signal processing unit (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It may implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present invention. The general-purpose processor may be a microprocessor or any conventional processor. The steps of the method disclosed in conjunction with the embodiments of the present invention may be directly implemented and executed by a hardware decoding processor, or by a combination of hardware and software modules in the decoding processor. The software module may be located in a storage medium well-known in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or the like. The storage medium is located in the memory 31 , and the processor 30 reads the information in the memory 31 and completes the steps of the above method in combination with its hardware.

[0065] The computer program product of the readable storage medium provided in the embodiment of the present invention includes a computer-readable storage medium storing program code. The instructions included in the program code can be used to execute the method described in the previous method embodiment. The specific implementation can be referred to the previous method embodiment and will not be repeated here.

[0066] If the functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage media include various media capable of storing program code, such as USB flash drives, mobile hard drives, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical disks.

[0067] Finally, it should be noted that the above-described embodiments are only specific implementation methods of the present invention, which are used to illustrate the technical solutions of the present invention, rather than to limit them. The scope of protection of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the above-described embodiments, those skilled in the art should understand that any person skilled in the art can modify or easily conceive of changes to the technical solutions described in the above-described embodiments within the technical scope disclosed by the present invention, or replace some of the technical features therein with equivalents. Such modifications, changes, or replacements do not deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should be included in the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be based on the scope of protection of the claims.

Claims

1. A method for constructing a non-target public spectrum database, characterized in that: include: Access to multiple open source spectrogram databases; Extracting metabolite information of each metabolite from the plurality of open source spectrum databases; wherein the metabolite information includes at least: spectrum and annotation information; Performing data supplementation and data filtering on the extracted metabolite information to obtain target metabolite information; The target metabolite information is integrated to obtain a non-target public spectrum database.

2. The method according to claim 1, characterized in that Metabolite information for each metabolite was extracted from multiple open source spectrogram databases, including: According to the preset keywords, corresponding data are searched from the plurality of open source spectrogram databases to obtain the metabolite information of each metabolite.

3. The method according to claim 1, characterized in that The extracted metabolite information is supplemented and filtered to obtain target metabolite information, including: Supplementing the metabolite information based on preset data supplementation rules; The metabolite information is filtered based on preset data filtering rules.

4. The method according to claim 3, characterized in that The metabolite information is supplemented based on preset data supplementation rules, including: Based on the ID index association between different open source spectrum databases, the missing information in the metabolite information is supplemented; If there are differences between the metabolite information of the same metabolite extracted from different open source spectrogram databases, the metabolite information will be retained according to the preset first database priority order of the open source spectrogram database; among which, the first database priority of the PubChem database is the highest, followed by the HMDB database; If the Class I to III classification data is missing in the metabolite information, the Class I to III classification data is searched from the ClassyFire database based on InChIKey; If the Chinese name is missing in the metabolite information, the Chinese name in the preset database shall prevail, or the Chinese name shall be obtained using the preset AI model.

5. The method according to claim 3, characterized in that Filtering the metabolite information based on preset data filtering rules includes: If the English name and secondary spectrum of the metabolite are the same, the metabolite from the open source spectrum database with the highest priority is retained according to the second database priority order of the preset open source spectrum database; Filtering metabolites whose mass-to-charge ratios in the secondary spectrum are not within a first preset range; Filtering the metabolites whose retained digits after the decimal point of the mass-to-charge ratio of the secondary spectrum are less than a preset digit; Filtering metabolites without daughter ions in the secondary spectrum; Metabolites with ion counts exceeding the ion threshold were filtered out; Filtering metabolites with zero intensity in the secondary spectrum; For metabolites whose intensities in the secondary spectrum are within a second preset range, filtering metabolites whose intensities do not include the first intensity; For metabolites whose intensity in the secondary spectrum is greater than the first intensity, filtering metabolites whose intensity is less than or equal to the second intensity; Filtering metabolites with the same mass-to-charge ratio in the secondary spectrum; Modify metabolites whose adduct ion forms do not meet the preset requirements; Filter out metabolites whose English names exceed a preset length and whose English names display garbled characters; If the InChIKey of metabolites is the same, the metabolite from the HMDB database will be retained first.

6. The method according to claim 1, characterized in that The target metabolite information is integrated to obtain a non-target public spectrum database, which also includes: The target metabolite information is integrated into an XCMS recognition format, and the annotation information in the target metabolite information is integrated into an information analysis process recognition format to obtain a non-target public spectrum database.

7. A device for constructing a non-target public spectrum database, characterized in that: include: Data acquisition module, used to obtain multiple open source spectrogram databases; An information extraction module, configured to extract metabolite information of each metabolite from the plurality of open source spectrogram databases; wherein the metabolite information includes at least: spectra and annotation information; A data processing module is used to supplement and filter the extracted metabolite information to obtain target metabolite information; The database construction module is used to integrate the target metabolite information to obtain a non-target public spectrum database.

8. The device according to claim 7, characterized in that The information extraction module is specifically used for: According to the preset keywords, corresponding data are searched from the plurality of open source spectrogram databases to obtain the metabolite information of each metabolite.

9. An electronic device, characterized in that: The method comprises a processor and a memory, wherein the memory stores computer-executable instructions that can be executed by the processor, and the processor executes the computer-executable instructions to implement the steps of the method according to any one of claims 1 to 6.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are performed.

Citation Information

Patent Citations

  • Construction method and device of plant metabolite database, medium and terminal

    CN113643768A

  • Construction method and application of intestinal microorganism related metabolite spectrum database

    CN115083528A

  • Metabonomics MS / MS database based on liquid chromatography-mass spectrometry and construction method

    CN120108577A