Substance identification apparatus, method, computer device, program product, and medium

By employing a two-stage comparison method, utilizing protein fingerprinting and differential information, candidate reference objects are screened and accurately compared, thus solving the problem of inaccurate microbial identification results and achieving higher identification accuracy.

CN122224314APending Publication Date: 2026-06-16ZYBIO INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-13
Publication Date
2026-06-16

AI Technical Summary

Technical Problem

In the prior art, when the mass spectrometry detection results of the microorganism to be tested have a high similarity to the mass spectrometry detection results of multiple known categories of microorganisms, or when the sample contains multiple microorganisms, the microorganism identification results are inaccurate.

Method used

A two-stage comparison method was adopted. First, the protein fingerprint of the reference object was compared with the protein fingerprint of the object to be tested to screen out candidate reference objects. Then, a second similarity comparison was performed based on the protein difference information of the candidate reference objects to determine the identification result of the object to be tested.

Benefits of technology

It improves the accuracy of microbial identification, clarifies the similarity between the test object and each candidate reference object, and solves the problem of uncertainty in identification results caused by interference from multiple microorganisms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122224314A_ABST
    Figure CN122224314A_ABST
Patent Text Reader

Abstract

The application discloses a substance identification device, method, computer equipment, program product and medium, and first similarity comparison is performed on a protein fingerprint of a to-be-tested object and standard protein fingerprints of reference objects to obtain a plurality of candidate reference objects similar to the to-be-tested object; second similarity comparison is performed on the to-be-tested object and the candidate reference objects based on protein difference information of the candidate reference objects, and an identification result of the to-be-tested object is determined based on comparison results of the candidate reference objects. The protein difference information corresponding to each candidate reference object determined based on the first comparison is used for secondary similarity comparison on the to-be-tested object, and the identification result of the to-be-tested object is determined based on the secondary comparison result, so that the application of the protein difference information solves the problem that the identification result of the to-be-tested object cannot be determined due to the high similarity of protein fingerprints of a plurality of known types of microorganisms or the presence of a plurality of microorganisms in the to-be-tested object, and the accuracy of microorganism identification is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure generally relates to the field of microbial detection technology, and specifically to a substance identification device, method, computer equipment, program product, and medium. Background Technology

[0002] In existing technologies, the identification results of the microorganisms to be tested can be determined by applying mass spectrometry ionization technology. Specifically, matrix-assisted laser desorption / ionization technology can be used to determine the mass spectrometry detection results of the microorganisms to be tested, and the species of the microorganisms to be tested can be determined by comparing the mass spectrometry detection results with the standard mass spectrometry detection results of known categories of microorganisms.

[0003] However, when the mass spectrometry results of the microorganism to be tested are highly similar to the standard mass spectrometry results of multiple known categories of microorganisms, or when the microorganism sample to be tested contains multiple microorganisms, it is easy to be unable to determine the species of the microorganism to be tested, thus leading to a certain error in the identification results of the microorganism to be tested.

[0004] Therefore, the poor accuracy of microbial identification results in existing technologies is a problem that urgently needs to be solved. Summary of the Invention

[0005] In view of the above-mentioned defects or deficiencies in the prior art, it is desirable to provide a substance identification device, method, computer equipment, program product and medium that compares the test object with reference objects based on protein difference information, so as to more accurately determine the degree of similarity between the test object and each reference object, thereby obtaining more accurate identification results of the test object and improving the accuracy of current microbial identification.

[0006] In a first aspect, the present invention provides a substance identification device, comprising a detection module for detecting a test object and obtaining its protein fingerprint, a storage module for storing microbial identification-related data, and a data processing module for retrieving data from the storage module for data processing. The data processing module is configured as follows: The protein fingerprint of the test object is compared with the standard protein fingerprint of each reference object to obtain multiple candidate reference objects that are similar to the test object. The candidate reference objects include at least two types, and the two types include different species or different subspecies. Determine the protein difference information for each candidate reference object; the protein difference information is the protein feature expression of the candidate reference object, and the protein difference information is used to distinguish the candidate reference object from one or more other candidate reference objects; Based on the protein difference information of the candidate reference objects, a second similarity comparison is performed between the test object and the candidate reference objects, and the identification result of the test object is determined based on the comparison results of each candidate reference object.

[0007] In a second aspect, a method for identifying substances is provided, comprising: performing a first similarity comparison between the protein fingerprint spectrum of the test object and the standard protein fingerprint spectrum of each reference object to obtain multiple candidate reference objects similar to the test object, wherein the candidate reference objects include at least two types, and the two types include different species or different subspecies; Determine the protein difference information for each candidate reference object; the protein difference information is the protein feature expression of the candidate reference object, and the protein difference information is used to distinguish the candidate reference object from one or more other candidate reference objects; Based on the protein difference information of the candidate reference objects, a second similarity comparison is performed between the test object and the candidate reference objects, and the identification result of the test object is determined based on the comparison results of each candidate reference object.

[0008] Thirdly, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, it performs the steps performed by the substance identification device as described in the first aspect.

[0009] Fourthly, a computer-readable storage medium is provided having a computer program stored thereon, characterized in that, when executed by a processor, the program performs the steps performed by the substance identification device as described in the first aspect.

[0010] Fifthly, a computer program product is provided, which includes instructions that, when executed, cause the steps performed by the substance identification device in the first aspect to be performed.

[0011] Compared to existing technologies that directly compare the target substance with known categories of microorganisms to determine its identification result, which introduces a certain degree of error, the substance identification device, method, computer equipment, program product, and medium provided in this application can perform a two-stage comparison of the target substance, greatly improving identification accuracy. Firstly, a first similarity comparison can be performed between the standard protein fingerprint spectrum of a reference object and the protein fingerprint spectrum of the target substance, eliminating some reference objects and obtaining candidate reference objects that are relatively similar to the target substance. Secondly, based on the protein difference information corresponding to each candidate reference object determined in the first comparison, a second similarity comparison is performed on the target substance to obtain its identification result. Since protein difference information can distinguish different candidate reference objects, comparing the test object with the candidate reference objects again based on protein difference information can further clarify the degree of similarity between the test object and each candidate reference object. This allows for a more accurate identification result of the test object based on the degree of similarity with each candidate reference object, thus solving the problem that the identification result of the test object cannot be determined due to the high similarity of protein fingerprints of multiple known categories of microorganisms or the presence of multiple microorganisms in the test object, thereby improving the accuracy of microbial identification. Attached Figure Description

[0012] Other features, objects, and advantages of this application will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings: Figure 1 A schematic diagram of the substance identification device provided in the embodiments of this application; Figure 2 A schematic flowchart of a substance identification method provided in an embodiment of this application; Figure 3 A schematic flowchart of another substance identification method provided in an embodiment of this application; Figure 4 A schematic diagram of the identification results of the test object provided in the embodiments of this application; Figure 5 A schematic diagram of the structure of a computer device provided in an embodiment of this application. Detailed Implementation

[0013] The present application will now be described in further detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and not intended to limit it. Furthermore, it should be noted that, for ease of description, only the parts relevant to the invention are shown in the accompanying drawings.

[0014] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other. The present application will now be described in detail with reference to the accompanying drawings and embodiments. Furthermore, the term "and / or" in this document is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. The terms "first" and "second," etc., in the specification and claims of the embodiments of this application are used to distinguish different objects, not to describe a specific order of objects.

[0015] First, the terminology used in this application will be explained as follows: (1) Mass Spectrometry: This is an analytical technique that identifies and quantifies the substances present in a analyte by measuring the mass-to-charge ratio of each ion in the analyte. The mass-to-charge ratio is a symbol describing the ratio between the mass and charge of an ion. For example, it can be M / Z, where M is the standard proton mass, and the unit can be u or Da; Z is the charge of the ion. Its working principle is as follows: First, the analyte is ionized by electron bombardment, chemical ionization, or matrix-assisted laser desorption / ionization. Then, the charged ions are accelerated by an electric field to gain kinetic energy. The ion beam is then passed through a magnetic field to deflect the ions according to their mass-to-charge ratio. Finally, the number of ions is measured using a detector.

[0016] (2) Protein fingerprinting: Also known as protein fingerprint sequence mapping, it is a spectrum drawn based on the mass-to-charge ratio of each ion after separating and identifying proteins in a biological sample. It can display information such as the molecular weight and content of various proteins in the sample. Protein fingerprinting can characterize the relationship between the mass-to-charge ratio m / z of ions and the ion abundance (i.e., intensity). The peaks with higher ion abundance (i.e., higher intensity peaks) in the protein fingerprint spectrum can be called characteristic peaks in the protein fingerprint spectrum.

[0017] Figure 1 This is a schematic diagram of the substance identification device provided in an embodiment of this application. Figure 1 As shown, the substance identification device includes: a detection module 101, a storage module 102, and a data processing module 103.

[0018] For example, the detection module 101 can be used to perform mass spectrometry analysis on the analyte and obtain the protein fingerprint spectrum of the analyte. The analyte can be a microbial sample, and the detection module 101 can include a mass spectrometry identification instrument using matrix-assisted laser desorption / ionization technology (i.e., MALDI-TOF-MS). Based on this, the substance placed into the mass spectrometry identification instrument can be a substance produced after processing an in vitro sample. The storage module 102 can be used to store data related to the identification of substances in this application. Among them, the data related to microbial identification can be protein fingerprints obtained by the detection module 101, protein fingerprints and / or genetic information of various known categories of reference objects (e.g., known categories of microorganisms), and other relevant information. The data processing module 103 can be used to call relevant data in the storage module 102 and use the relevant data to perform data processing to execute the substance identification method described in the embodiments of this application. The data processing module 103 may be, for example, a computer device.

[0019] In a specific implementation, the detection module 101 can perform mass spectrometry detection on the test object to obtain the protein fingerprint spectrum of the test object. Further, the data processing module 103 compares the protein fingerprint spectrum of the test object with the protein fingerprint spectrum of other known categories of microorganisms stored in the storage module 102 to determine the identification result of the test object.

[0020] However, when the protein fingerprint of the microorganism to be tested is highly similar to the standard mass spectrometry detection results of multiple known categories of microorganisms, or when the sample of the microorganism to be tested contains multiple microorganisms, it is easy to make it impossible to determine the species of the microorganism to be tested. Secondly, when the microorganism to be tested contains multiple microorganisms, the mutual interference between the microorganisms can also easily lead to certain errors in the identification results of the microorganism to be tested.

[0021] Based on this, embodiments of this application provide a method for substance identification, which can be performed by... Figure 1 The method or apparatus is executed by the substance identification device shown, specifically by the data processing module 103. On one hand, the method or apparatus can perform a similarity comparison between the standard protein fingerprint spectrum of a reference object and the protein fingerprint spectrum of the object to be tested, eliminating some reference objects and obtaining candidate reference objects that are relatively similar to the object to be tested. On the other hand, based on the protein difference information corresponding to each candidate reference object determined in the first comparison, a second similarity comparison is performed on the object to be tested to obtain the identification result of the object to be tested. Since protein difference information can distinguish different candidate reference objects, comparing the object to be tested with the candidate reference objects again based on the protein difference information can further clarify the degree of similarity between the object to be tested and each candidate reference object, thereby obtaining a more accurate identification result of the object to be tested based on the degree of similarity with each candidate reference object. This solves the problem that the identification result of the object to be tested cannot be determined due to the high similarity of protein fingerprint spectra of multiple known categories of microorganisms or the presence of multiple microorganisms in the object to be tested, thus improving the accuracy of microbial identification.

[0022] For example, Figure 2This is a schematic flowchart of the substance identification method provided in the embodiments of this application. Figure 2 As shown, the data processing module 103 is configured to perform the following steps: Step S201: Perform a first similarity comparison between the protein fingerprint of the object to be tested and the standard protein fingerprint of each reference object to obtain multiple candidate reference objects that are similar to the object to be tested. Optionally, the candidate reference objects may include at least two types, which may include different species or different subspecies.

[0023] In this embodiment, the protein fingerprint of the test object is first compared with the standard protein fingerprints of various known categories of reference objects to screen out candidate reference objects similar to the test object. This first similarity comparison eliminates some reference objects, significantly narrowing the comparison range for the subsequent second similarity comparison, thereby improving identification efficiency to some extent.

[0024] In one possible implementation, the analyte may include at least one microorganism. Microorganisms of the same category exhibit essentially consistent mass-to-charge ratios and ion abundance relationships among the ions produced by the breakdown of various proteins during mass spectrometry analysis; that is, the protein fingerprint spectra of microorganisms of the same category are also essentially consistent. Therefore, by obtaining the protein fingerprint spectra of the analyte microorganism through mass spectrometry analysis, and by comparing the protein fingerprint image of the analyte microorganism with the protein fingerprint spectra of various known categories of microorganisms, the analyte microorganism can be identified. Therefore, in this embodiment, the protein fingerprint spectra of the analyte are first compared with the standard protein fingerprint spectra of each reference object to identify microorganisms similar to the analyte microorganism, i.e., the candidate reference objects described in this embodiment.

[0025] In one possible implementation, standard protein fingerprints of each reference object can be obtained from the standard fingerprint library stored in the storage module 102. Each reference object can be a known category of microorganism.

[0026] For example, a large number of protein fingerprints of known microorganisms can be pre-acquired and stored in storage module 102 to construct a standard fingerprint library composed of standard protein fingerprints of various microorganisms.

[0027] In one possible implementation, N reference objects with the highest similarity to the test object can be determined based on the similarity results between the test object and each reference object. Here, N is an integer greater than 1. Furthermore, candidate reference objects for the test object can be determined based on these N reference objects.

[0028] For example, the similarity results between the test object and each reference object can characterize the degree of similarity between the test object and each reference object. Specifically, the similarity results can include similarity scores between the test object and each reference object. The magnitude of the similarity score is positively correlated with the degree of similarity; that is, the higher the similarity score between the test object and a certain reference object, the higher the similarity between the test object and that reference object.

[0029] For example, the similarity results between the test object and the reference object can be obtained by comparing the similarity between the protein fingerprint of the test object and the standard protein fingerprint of each reference object.

[0030] Specifically, the similarity score between the test object and each reference object can be determined by comparing the protein fingerprint of the test object with the standard protein fingerprint of each reference object. If the similarity score between the test object and each reference object is positively correlated with the degree of similarity between the test object and each reference object, the similarity scores can be sorted in descending order, and the N reference objects corresponding to the top N similarity scores in the descending order can be identified as candidate reference objects.

[0031] One possible implementation is to use a scoring method to compare the similarity between the protein fingerprint of the test object and the standard protein fingerprint of each reference object.

[0032] Specifically, the similarity information of the bar-shaped peaks corresponding to the protein fingerprint spectrum of the test object can be compared with the similarity information of the standard protein fingerprint spectrum of each reference object. When the protein fingerprint spectrum of the test object contains the same characteristic peak as the protein fingerprint spectrum of a certain reference object, a score is added to the reference object; otherwise, a score is deducted. Then, the similarity result between the test object and each reference object is obtained based on the specific score.

[0033] In this embodiment of the application, candidate reference objects for the object to be tested can be determined based on the above N reference objects in the following three ways.

[0034] Method 1: N reference objects can be directly identified as candidate reference objects.

[0035] For example, when the reference objects A, B, and C with the highest similarity to the test object (i.e., the highest scores) are determined based on the similarity scores between the test object and each reference object, reference objects A, B, and C can be identified as candidate reference objects.

[0036] Method 2: Reference objects of the same type as the N reference objects can be used as candidate reference objects.

[0037] For example, the reference objects of the same type can be other subspecies belonging to the same bacterial species as the N reference objects. Specifically, when A1 and A2 with the highest similarity to the test object are determined based on the similarity scores between the test object and each reference object, and A1 and A2 belong to the same subspecies under the same bacterial species, subspecies A1 and A2, as well as subspecies A3 under that bacterial species, can be determined as candidate reference objects.

[0038] Method 3: When there is a target reference object belonging to the reference object combination among the N reference objects, candidate reference objects can be determined based on the reference object combination and the N reference objects.

[0039] The reference object combination can include multiple reference objects with similar protein feature expression. Protein feature expression can be the expression of ions from microbial protein degradation in a protein fingerprint, such as the relationship between ion mass-to-charge ratio and ion abundance, i.e., characteristic peaks in the protein fingerprint. Therefore, multiple reference objects with similar protein feature expression can be multiple reference objects with similar protein fingerprints.

[0040] For example, the reference object combination can be a list of similar bacteria, or a combination of individual bacteria in the list of similar bacteria.

[0041] Specifically, when reference objects A, B, and C with the highest similarity to the test object are determined based on the similarity scores between the test object and each reference object, and reference objects A and B belong to the same reference object combination, reference objects A, B, D, E, and C in the reference object combination to which reference objects A and B belong can all be determined as candidate reference objects.

[0042] Taking the reference object combination as a list of similar bacteria as an example, when there is a target reference object A in the N reference objects determined based on the similarity scores between the test object and each reference object, all reference objects included in the list of similar bacteria 1 and the N reference objects other than reference object A can be determined as candidate reference objects.

[0043] Optionally, when reference objects A, B, and C with the highest similarity to the test object are determined based on their similarity scores with each reference object, and reference object A belongs to reference object combination 1 and reference object B belongs to reference object combination 2, then reference objects A, D, and E from reference object combination 1 (to which reference object A belongs), B, F, and G from reference object combination 2 (to which reference object B belongs), and reference object C can all be identified as candidate reference objects. Note that there is no necessary connection between reference object combination 1 and reference object combination 2.

[0044] Taking the reference object combination as a list of similar bacteria as an example, when there is a target reference object A in the list of similar bacteria 1 and a target reference object B in the list of similar bacteria 2 among the N reference objects determined based on the similarity scores between the test object and each reference object, all reference objects included in the lists of similar bacteria 1 and 2, as well as the reference objects other than reference objects A and B in the N reference objects, can be determined as candidate reference objects.

[0045] It should be noted that the information of the reference object combination to which each of the above reference objects belongs, as well as the information of other reference objects in the reference object combination, can be obtained from the storage module 102 (that is, the storage module 102 pre-stores reference object combination information).

[0046] [Creating and updating reference object combinations] For example, reference object combination information is used to characterize multiple reference object combinations and the protein fingerprint map corresponding to each reference object combination, wherein the reference object combination includes multiple reference objects with similar protein feature expression.

[0047] In one possible implementation, based on the first similarity comparison, a new reference object combination can be created or an existing reference object combination can be updated based on the N reference objects with the highest similarity to the object to be tested.

[0048] For example, update information is generated based on the new or updated reference object combination and the protein fingerprint of the reference objects recorded in the combination. The update information can characterize the update status of an existing reference object combination or a newly created reference object combination, or the update information can be added to the reference object combination information.

[0049] For example, based on the N reference objects / candidate reference objects obtained after the first similarity comparison and the corresponding gene information or kinship information of the reference objects / candidate reference objects, the existing reference object combination can be updated or a new reference object combination can be created.

[0050] Specifically, when creating a new reference object combination or updating an existing reference object combination based on the N reference objects with the highest similarity to the test object obtained through the first similarity comparison, it can be first determined whether each of the N reference objects belongs to the existing reference object combination. Then, the genetic information or kinship information of each reference object that does not belong to the existing reference object combination is compared with the reference objects in the existing reference object combination. If the comparison result indicates that the reference object that does not belong to the existing reference object combination is similar to the reference object in the existing reference object combination, then the reference object is added to the reference object combination to which the similar reference object belongs, and the reference object combination is updated.

[0051] Alternatively, if the comparison results indicate that a reference object not belonging to the existing reference object combination is not similar to the reference objects in the existing reference object combination, then the genetic information or kinship information of the reference object is compared with the genetic information or kinship information of the other reference objects not belonging to the existing reference object combination. If the comparison results indicate that the reference object is similar to the other reference objects not belonging to the existing reference object combination, then a new reference object combination is created based on the reference object and the reference objects similar to the reference object.

[0052] Optionally, in the above comparison, the specific methods for determining whether there is similarity between reference objects through genetic information or kinship information include: when the similarity of the genetic information of a reference object with other reference objects reaches a preset threshold or the kinship information is consistent, it can be determined that the reference object is similar to other reference objects. It is understood that the average nucleotide identity (ANI) in the genetic information can be used to determine whether there is similarity between reference objects.

[0053] For example, when reference subjects with an average genomic genetic similarity of 94% or higher essentially belong to the same species, reference subjects with an average genomic genetic similarity of 93%, 92%, 91%, 90%, 89%, 88%, 87%, 86%, 85%, 84%, 83%, 82%, 81%, 80%, etc., can be defined as similar reference subjects based on the accuracy of mass spectrometry identification in actual use. It should be noted that various methods can be used to determine the similarity between reference subjects using genetic information, and this application does not impose any restrictions on this.

[0054] Correspondingly, if a certain mass spectrometer has high identification accuracy and is only prone to confusing bacterial species with an average genomic genetic similarity of 93% or more, then the above-mentioned preset threshold can be set to 93%; while based on research findings, the average genomic genetic similarity between microorganisms listed as the same reference group is 80%, so the above-mentioned preset threshold can be set to 80%.

[0055] In this embodiment of the application, when determining candidate reference objects similar to the test object, the reference object with the highest similarity to the test object can be created as a reference object combination, and the specific reference objects and their protein fingerprints recorded in the reference object combination can be recorded and stored in the storage module 102.

[0056] For example, the N reference objects with the highest similarity to the object to be tested can be created as a new reference object combination, or the existing reference object combination can be updated based on the N reference objects with the highest similarity to the object to be tested, so as to form updated information based on the newly created reference object combination or the updated reference object combination and the protein fingerprint of each reference object in the reference object combination, thereby adding the updated information to the reference object combination information.

[0057] For example, when reference objects A, B, and C with the highest similarity to the object under test are determined based on the similarity results between the object under test and each reference object, reference objects A, B, and C can be created as reference object combination 1.

[0058] Specifically, the protein fingerprints of the aforementioned reference object combination 1 and reference objects A, B, and C can be added to the reference object combination information. Here, reference object combination 1 can be understood as a list of similar objects (e.g., a list of similar bacteria).

[0059] Preferably, if the similarity of the gene information of reference objects A, B, and C reaches a preset threshold or the kinship information of reference objects A, B, and C is consistent, then reference objects A, B, and C are used to create or update reference object combination 1; otherwise, no creation or update is performed.

[0060] For example, when reference object A belongs to the existing reference object combination 1 while reference objects B and C do not belong to reference object combination 1, if the similarity of the genetic information of reference objects A, B, and C reaches a preset threshold or the kinship information of reference objects A, B, and C is consistent, then reference objects B and C are added to reference object combination 1 to update reference object combination 1; otherwise, no update is performed.

[0061] For example, when reference objects A, B, and C do not belong to the existing reference object combination, if the comparison of gene information or kinship information between each reference object by the data processing module 103 can determine that the similarity between the gene information of reference objects A, B, and C and the gene information of one or more reference objects in the existing reference object combination 1 reaches a preset threshold, then reference objects A, B, and C are added to the reference object combination 1 to update the reference object combination 1.

[0062] For example, when reference objects A1, A2, and B1 with the highest similarity to the test object are determined based on the similarity results between the test object and each reference object, and reference objects A1 and A2 belong to the same reference object combination 2, reference object B1 can be added to reference object combination 2 to update reference object combination 2.

[0063] Step S202: Determine the protein difference information of each candidate reference object; the protein difference information is the protein feature expression of the candidate reference object, and the protein difference information is used to distinguish the candidate reference object from one or more other candidate reference objects.

[0064] In this embodiment of the application, after performing the first similarity comparison in step S201, multiple candidate reference objects similar to the test object can be identified. Since the candidate reference objects are all highly similar to the test object and there is a certain degree of similarity between them, in order to perform a more accurate second similarity comparison, the protein difference information of each candidate reference object can be obtained in step S202 so that the test object and each candidate reference object can be compared based on the protein difference information to obtain a more accurate identification result.

[0065] For example, protein difference information of a candidate reference object can be used to distinguish that candidate reference object from one or more other candidate reference objects. For instance, protein difference information can be the differential characteristic peaks in the protein fingerprint of a candidate reference object, the genome of a candidate reference object that is different from one or more other candidate reference objects, or the protein gene of a candidate reference object that is different from one or more other candidate reference objects.

[0066] It should be noted that, among the various methods listed in this application for determining the protein difference information of each candidate reference object and performing a second similarity comparison based on the protein difference information, defining the protein difference information as any of the above methods will not affect the understanding and implementation of the technical solution of this application.

[0067] The following text focuses on the example of protein difference characteristics, specifically the differential peaks.

[0068] For example, the differential characteristic peaks of candidate reference objects, or combinations of differential characteristic peaks composed of differential characteristic peaks of candidate reference objects, can be used to uniquely characterize candidate reference objects.

[0069] Specifically, taking candidate reference objects A, B, and C as an example, when the protein fingerprint spectrum of candidate reference object A contains a characteristic peak 1, while the protein fingerprint spectra of candidate reference objects B and C do not contain a characteristic peak 1, then characteristic peak 1 can be determined as the unique differential characteristic peak representing candidate reference object A. That is, characteristic peak 1 is used to distinguish candidate reference object A from the other two candidate reference objects, and characteristic peak 1 is the protein differential information of candidate reference object A.

[0070] Optionally, if, based on gene alignment, protein alignment, or protein fingerprint comparison, it is determined that the protein fingerprint of candidate reference object A includes characteristic peaks 1, 2, and 3, the protein fingerprint of candidate reference object B includes characteristic peaks 2 and 3, and the protein fingerprint of candidate reference object C includes characteristic peaks 1 and 3, then it can be concluded that the difference characteristic peak between candidate reference object A and candidate reference object B is characteristic peak 1 in the protein fingerprint, and the difference characteristic peak between candidate reference object A and candidate reference object C is characteristic peak 2 in the protein fingerprint. In this case, characteristic peak 1 or characteristic peak 2 alone cannot distinguish candidate reference object A from the other two candidate reference objects. Instead, a combination of characteristic peaks 1, 2, and 3 or a combination of characteristic peaks 1 and 2 (i.e., the difference characteristic peak combination) is needed to uniquely characterize candidate reference object A. In other words, the difference characteristic peak combination is used to distinguish candidate reference object A from the other two candidate reference objects.

[0071] It should be noted that, in the above example, although a single characteristic peak cannot uniquely characterize a candidate reference object, if a combination of difference characteristic peaks can uniquely characterize a candidate reference object, then characteristic peak 1, characteristic peak 2, or characteristic peak 3, as characteristic peaks constituting the combination of difference characteristic peaks, can all be identified as the difference characteristic peaks described in this application.

[0072] Secondly, although characteristic peak 1 or the combination of characteristic peaks 1 and 3 can only distinguish between candidate reference object 1 and candidate reference object 2, but not between candidate reference object 1 and candidate reference object 3 (i.e. characteristic peak 1 can only distinguish between candidate reference object and the other candidate reference object), characteristic peak 1 or the combination of characteristic peaks 1 and 3 also belongs to the protein difference information between candidate reference object and other candidate reference objects described in this application.

[0073] Optionally, when the protein fingerprint spectra of all reference objects in reference object combination 1 contain characteristic peak 1, while the protein fingerprint spectra of all reference objects in reference object combination 2 do not contain characteristic peak 1, then characteristic peak 1 can be identified as the difference characteristic peak between reference object combination 1 and reference object combination 2. That is, characteristic peak 1 can be used as a candidate reference object that can be used to uniquely characterize different categories. For example, if candidate reference object A belongs to reference object combination 1 and candidate reference object B belongs to reference object combination 2, then characteristic peak 1 can be the difference characteristic peak between candidate reference objects A and B.

[0074] Optionally, the differential characteristic peaks among candidate reference objects can also be characteristic peaks corresponding to mutually exclusive gene fragments among the candidate reference objects. It should be noted that the same gene fragment exists only once in a single organism; therefore, the protein characteristic expression of this gene fragment in a single organism has only one molecular weight. Based on this, the protein characteristic expression (i.e., characteristic peaks) corresponding to mutually exclusive gene fragments can be identified as differential characteristic peaks.

[0075] For example, taking "single eyelid" and "double eyelid" as examples, the gene expression fragments of "single eyelid" and "double eyelid" are mutually exclusive. When the translation result of gene fragment 1 of candidate reference object A is double eyelid and the translation result of gene fragment 2 of candidate reference object B is single eyelid, the characteristic peak 1 of gene fragment 1 in the protein fingerprint can be used as the differential characteristic peak of candidate reference object A, and the characteristic peak 2 of gene fragment 2 in the protein fingerprint can be used as the differential characteristic peak of candidate reference object B.

[0076] In one possible implementation, for each candidate reference object, protein expression differential analysis can be performed on each candidate reference object to obtain protein differential information of the candidate reference objects.

[0077] Specifically, after identifying candidate reference subjects, the data processing module 103 can use forward compilation to perform protein expression differential analysis on the candidate reference subjects to determine protein difference information in real time. Forward compilation can be based on characteristic proteins obtained from the genome encoding of the candidate reference subjects, or it can be based on protein analysis results from a protein database.

[0078] Taking the protein difference information as the differential characteristic peaks in the protein fingerprint spectrum corresponding to a certain candidate reference as an example, the determination of differential characteristic peaks is explained: First, differential characteristic peaks of candidate reference subjects are determined based on their genomes: For example, characteristic proteins of candidate reference objects can be determined first based on the genome of the candidate reference objects, and differential proteins of each candidate reference object can be determined based on the characteristic proteins of each candidate reference object. In this way, the characteristic peaks corresponding to the differential proteins in the corresponding protein fingerprints can be determined as the differential characteristic peaks of each candidate reference object.

[0079] Optionally, differentially expressed genomes among the candidate reference objects can be identified first based on their genomes, and differentially expressed proteins of each candidate reference object can be identified based on their differentially expressed genomes. The characteristic peaks corresponding to the differentially expressed proteins in the corresponding protein fingerprints can then be identified as the differential characteristic peaks of each candidate reference object.

[0080] Optionally, the characteristic proteins of the candidate reference objects can be determined first based on their genomes, and then the characteristic peaks of each candidate reference object in the corresponding protein fingerprint can be determined based on the characteristic proteins of each candidate reference object, thereby identifying the characteristic peaks of each candidate reference object in the corresponding protein fingerprint as the differential characteristic peaks of each candidate reference object.

[0081] Secondly, differential characteristic peaks of candidate reference materials are determined based on protein databases: For example, characteristic proteins of candidate reference objects can be obtained from a protein database, and the characteristic peaks corresponding to the characteristic proteins of each candidate reference object in the corresponding protein fingerprint spectrum can be identified as the differential characteristic peaks of each candidate reference object.

[0082] Optionally, characteristic proteins of candidate reference objects can be obtained from a protein database to determine the differential proteins of each candidate reference object based on the characteristic proteins of each candidate reference object, thereby determining the characteristic peaks corresponding to the differential proteins of each candidate reference object in the corresponding protein fingerprint as the differential characteristic peaks of each candidate reference object.

[0083] Third, the differential characteristic peaks of candidate reference objects are determined based on their protein fingerprint profiles: For example, the protein fingerprints of each candidate reference object can be compared to obtain the differential characteristic peaks of each candidate reference object.

[0084] In one possible implementation, before determining the protein difference information of each candidate reference object, interfering features in the protein fingerprint profiles of each candidate reference object can be removed to eliminate interference when determining the protein difference information and improve the accuracy of the determination. The interfering features can be common features in the protein fingerprint profiles of each candidate reference object.

[0085] Specifically, the expression of the same protein features in the protein fingerprints of each candidate reference object can be ignored, and the differential characteristic peaks of each candidate reference object can be determined based on the protein fingerprints after ignoring the expression of the same protein features.

[0086] In one possible implementation, during the process of determining the characteristic proteins of candidate reference objects based on their genomes, the expression deficiencies of theoretically differentially expressed peaks can be corrected. These theoretically differentially expressed peaks can be the differentially expressed peaks corresponding to the genomes of the candidate reference objects.

[0087] Specifically, it is possible to determine the protein expression loss at the differential characteristic peaks in the genome of candidate reference objects, and to correct the differential characteristic peaks in the protein fingerprint based on the protein expression loss.

[0088] Step S203: Based on the protein difference information of the candidate reference objects, a second similarity comparison is performed between the test object and the candidate reference objects, and the identification result of the test object is determined based on the comparison results of each candidate reference object.

[0089] In this embodiment of the application, since the protein feature expression characterized by the protein difference information can distinguish different candidate reference objects, when the test object and the candidate reference objects are compared again based on the protein difference information of the candidate reference objects determined in step S202, the degree of similarity between the test object and each candidate reference object can be further clarified, so as to obtain a more accurate identification result of the test object based on the degree of similarity with each candidate reference object.

[0090] In one possible implementation, a second similarity comparison is performed between the test object and each candidate reference object based on the protein difference information of the candidate reference objects to obtain the identification result of the test object. The second similarity comparison can be implemented using a similarity scoring method.

[0091] For example, the mass-to-charge ratio (m / z value), intensity, and weight information of the difference characteristic peaks corresponding to protein difference information in the protein fingerprint spectrum can be adjusted to increase the importance of the difference characteristic peaks in determining the similarity comparison results between the test object and the candidate reference object, so as to obtain a more accurate comparison result between the test object and the candidate reference object based on the difference characteristic peaks of the candidate reference object.

[0092] Specifically, for each candidate reference object, the scoring weight of the differential characteristic peaks of the candidate reference object can be increased. Based on the weight and amplitude of each characteristic peak in the protein fingerprint spectrum of the candidate reference object, the similarity score between the candidate reference object and the test object is scored to obtain the similarity score between the candidate reference object and the test object. Based on the similarity score of the candidate reference object, the matching result between the test object and the candidate reference object is output.

[0093] For example, when comparing the bar-shaped peak information corresponding to the protein fingerprint spectrum of the test object with the protein fingerprint spectrum of the candidate reference object, a scoring operation can be performed based on the weight and amplitude of the characteristic peak when the two have the same characteristic peak, so as to obtain the similarity score between the candidate reference object and the test object; when the protein fingerprint spectrum of the test object shows the difference characteristic peak of the candidate reference object, a scoring operation can be performed based on the increased scoring weight (e.g., product coefficient).

[0094] Optionally, for each candidate reference object, the candidate reference object and the test object can be compared based on the differential characteristic peaks of the candidate reference object, so as to output the matching result between the test object and the candidate reference object based on whether the protein fingerprint spectrum of the test object contains the same characteristic peaks as the differential characteristic peaks.

[0095] For example, when comparing the bar-shaped peak information corresponding to the protein fingerprint spectrum of the test object with the protein fingerprint spectrum of the candidate reference object, if the protein fingerprint spectrum of the test object contains the same characteristic peak as the difference characteristic peak, it can be determined that the test object matches the candidate reference object; otherwise, it is determined that the test object does not match the candidate reference object.

[0096] For example, the identification result of the test object can be determined based on the matching result between the test object and the candidate reference object.

[0097] Specifically, when the similarity score between the test object and the candidate reference object is higher than the preset threshold, it can be determined that the test object and the candidate reference object are a match. At this time, the candidate reference object that matches the test object can be determined as the identification result of the test object.

[0098] Optionally, when the similarity score between the test object and the candidate reference object is the highest, it can be determined that the test object and the candidate reference object are the same, and at this time the candidate reference object can be identified as the test object.

[0099] In one possible implementation, different types of candidate reference objects can correspond to different types of identification results for the objects to be tested.

[0100] For example, when multiple candidate reference objects include at least several subspecies belonging to the same bacterial species, the identification result of the test object corresponds to the subspecies.

[0101] Optionally, when multiple candidate reference objects include at least multiple strains belonging to the same genus, the identification result of the test object corresponds to the bacterial species.

[0102] Optionally, when multiple candidate reference objects include at least multiple species belonging to the same complex microbial community, the identification result to be tested corresponds to the microbial species.

[0103] Optionally, when multiple candidate reference objects include at least multiple bacteria belonging to the same list of similar bacteria, the identification result of the test object corresponds to the bacteria.

[0104] Compared to existing technologies that directly determine the identification result of the test object by comparing it with known categories of microorganisms, the substance identification device provided in this application can perform a second similarity comparison of the test object based on the protein difference information corresponding to each candidate reference object determined in the first comparison, after using the standard protein fingerprint spectrum of the reference object to perform a first similarity comparison of the protein fingerprint spectrum of the test object. Since the protein feature expression characterized by the protein difference information can distinguish different candidate reference objects, the application of the protein difference information can further clarify the degree of similarity between the test object and each candidate reference object, so as to obtain a more accurate identification result of the test object based on the degree of similarity with each candidate reference object. This solves the problem that the identification result of the test object cannot be determined due to the high similarity of the protein fingerprint spectra of multiple known categories of microorganisms, and improves the accuracy of microbial identification.

[0105] [Mixed Reference Results] In another embodiment of this application, the identification result of the test object may also include a mixed reference object, such as a mixed bacteria.

[0106] For example, the substance identification device provided in this application has a detection module for detecting the test object and obtaining its protein fingerprint spectrum, a storage module for storing substance identification-related data, and a data processing module for calling data from the storage module for data processing, wherein the data processing module is configured as follows: The protein fingerprint of the test object is compared with the standard protein fingerprint of each reference object to obtain multiple candidate reference objects that are similar to the test object. The candidate reference objects include at least two types, and the two types include different species or different subspecies.

[0107] For example, the protein fingerprint of the test object is compared with each reference object to determine the N reference objects with the highest similarity as candidate reference objects, where N is an integer greater than 1. Based on the N reference objects, mixed information is judged, and the ordinary identification result or mixed reference result of the test object is output based on the mixed information.

[0108] In one possible implementation, when the output is a mixed reference result, the test object can be identified as a mixed bacteria. It is understood that the method for outputting ordinary identification results is basically the same as the methods in other embodiments described above. For example, the identification result is obtained by performing a second similarity comparison between the test object and N reference objects (i.e., candidate reference objects) based on protein difference information, as shown in step S203. The following mainly explains the implementation method for judging mixed information and the output of mixed reference results.

[0109] Optionally, the method for determining mixed information based on N reference objects and outputting a normal identification result or a mixed reference result is as follows: determine whether there are multiple reference objects among the N reference objects that do not belong to the same combination of reference objects. If so, output a mixed reference result; otherwise, output a normal identification result. It should be noted that during this process, the determination result of mixed information (i.e., "yes" or "no") can be output, or the determination result of mixed information can be omitted, and only the mixed reference result or normal identification result needs to be output according to this logic.

[0110] Optionally, the method for determining mixed information based on N reference objects and outputting a normal identification result or a mixed reference result is as follows: determine whether there are multiple reference objects among the N reference objects that do not belong to the same group of similar reference objects. If so, output a mixed reference result; otherwise, output a normal identification result. It should be noted that during this process, the determination result of mixed information (i.e., "yes" or "no") can be output, or the determination of mixed information can be omitted, and only the mixed reference result or normal identification result needs to be output according to this logic.

[0111] Among them, the same group of similar reference objects refers to the reference objects that meet the similarity judgment conditions among N reference objects; meeting the similarity judgment conditions refers to one or more of the following conditions: (a) multiple reference objects belong to the same reference object combination, (b) or the genetic information similarity of multiple reference objects reaches a certain threshold, or (c) multiple reference objects belong to different subspecies within the same species.

[0112] For example, when reference objects with an average genomic genetic similarity of 94% or more are basically of the same species, reference objects with an average genomic genetic similarity of 93%, 92%, 91%, 90%, 89%, 88%, 87%, 86%, 85%, 84%, 83%, 82%, 81%, 80%, etc. can be defined as similar reference objects based on the accuracy of mass spectrometry identification.

[0113] Correspondingly, when a mass spectrometer is set to be judged as similar to the reference object if it meets any one of the three similarity judgment conditions (a), (b), and (c) above, and the combination of reference objects in the mass spectrometer is relatively complete, the above threshold can be set to a value close to 94%, such as 93%, 92%, 91%, etc.; while when a mass spectrometer basically does not include the combination of reference objects, the above threshold is determined according to the identification accuracy of the mass spectrometer.

[0114] Optionally, based on N reference objects, the mixed information is judged and the mixed reference result is output. In the mixed reference result, at least one result is obtained by performing a second similarity comparison based on the protein difference information of multiple candidate reference objects.

[0115] Furthermore, if all candidate reference objects are mixed reference objects, the first similarity comparison result is output.

[0116] Furthermore, if not all candidate reference objects are mixed reference objects, a mixed reference result is output, which includes the result of the second similarity comparison.

[0117] For example, the method for judging mixed information based on N reference objects and outputting a normal identification result or a mixed reference result is: If Q out of the N reference objects meet the similarity determination criteria, then the Q reference objects are a group of similar reference objects, where 1≤Q≤N; If N = Q, then the test object is compared with the candidate reference objects of N reference objects to obtain the identification result of the test object; since the N reference objects belong to the same group of similar reference objects, it is only necessary to distinguish the similar reference objects, without considering the situation of multiple substances or multiple microorganisms, that is, the ordinary identification result is output. If N≠Q, then output a hybrid reference result, which contains at least two categories of reference objects, and at least two categories do not belong to the same group of similar reference objects.

[0118] It is understandable that the Q reference objects do not all need to meet one of the three conditions (a), (b), and (c) above. For example, if three reference objects A, B, and C belong to the same reference object combination, B and C belong to different subspecies of the same species, or the genetic information similarity between B and C reaches a certain threshold, then A, B, and C obviously belong to the same group of similar reference objects.

[0119] Optionally, in the steps of creating or updating the combined information of the reference objects mentioned above, the update is performed based on the judgment result of the mixed information.

[0120] Specifically, the remaining steps for creating or updating the reference object combination information are the same, namely, adding update information to the reference object combination information. The update information is used to characterize the reference object combination created or updated based on N reference objects, as well as the protein fingerprint of the N reference objects.

[0121] The update based on the judgment result of mixed information refers to determining whether there are multiple reference objects that do not belong to the same group of similar reference objects among the N reference objects. If so, select any group of similar reference objects to update the reference object combination. If there are reference objects that do not belong to the same reference object combination in the same group of similar reference objects, then update each reference object in this group of similar reference objects to the same reference object combination.

[0122] In this embodiment of the application, by judging the mixed information of N reference objects, the identification result of the test object can be clearly determined when the test object is mixed with multiple microorganisms, so as to avoid the omission of the identification result of the test object.

[0123] [Examples of various output mixed reference results] As mentioned above, the criteria for determining mixed information are whether multiple reference objects belong to the same set of reference objects or whether they belong to the same set of similar reference objects. The criteria for belonging to the same set of similar reference objects include belonging to the same set of reference objects, the genetic information similarity of multiple reference objects reaching a certain threshold, or multiple reference objects belonging to different subspecies within the same species. Of course, multiple conditions can also be selectively used for judgment. For example, two candidate reference objects may be determined not to belong to the same set of similar reference objects and their genetic information similarity may be below a certain threshold, resulting in a mixed reference result. This mixed reference result includes two candidate reference objects. It is understandable that other combinations of judgment conditions can also be used to determine whether multiple reference objects belong to the same set of similar reference objects.

[0124] The following examples illustrate the method for determining mixed information. However, the embodiments of this application do not limit the method for determining mixed information.

[0125] First, the specific implementation of "judging mixed information based on whether candidate reference objects belong to the same combination of reference objects" is as follows: For example, if the first similarity comparison result between the test object and each reference object includes candidate reference objects A and B, and neither A nor B belongs to any combination of reference objects, then the second similarity comparison can be skipped, and the identification result can be directly output. The identification result must contain at least A and B.

[0126] When the first similarity comparison result between the test object and each reference object includes candidate reference objects A1, A2, and B, and A1 and A2 belong to a reference object combination, while B does not belong to any reference object combination, a second similarity comparison can be performed between the test object and the candidate reference objects A1, A2, or the reference object combination corresponding to A1 and A2. The second similarity comparison result and B constitute the identification result of the test object, and the identification result must contain at least B.

[0127] Optionally, when the first similarity comparison result between the test object and each reference object includes candidate reference objects A1, A2, and B, and A1 and A2 belong to one reference object combination a and B belong to another reference object combination b, a second similarity comparison can be performed between the test object and the candidate reference objects in reference object combination a, and a second similarity comparison can be performed between the test object and the candidate reference objects in reference object combination b. Based on the second similarity comparison, a mixed reference result is output, which includes at least one reference object in reference object combination a and at least one reference object in reference object combination b.

[0128] Optionally, when the first similarity comparison results between the test object and each reference object include candidate reference objects A1, A2, and A3, and A1, A2, and A3 all belong to the same group of reference objects, the identification result of the test object can be determined by comparing the second similarity between the test object and the candidate reference objects A1, A2, and A3 respectively.

[0129] Optionally, when the first similarity comparison results between the test object and each reference object include candidate reference objects A1, A2, B1, and B2, where A1 and A2 belong to the same reference object combination and B1 and B2 belong to another reference object combination, the identification result can be obtained by comparing the test object with the candidate reference objects A1, A2, B1, and B2 respectively. The identification result contains at least one of A1 and A2 and at least one of B1 and B2.

[0130] Optionally, when the first similarity comparison results between the test object and each reference object include candidate reference objects A1, A2, B1, B2, and C, where A1 and A2 belong to the same reference object combination, B1 and B2 belong to another reference object combination, and C does not belong to any reference object combination, the identification result can be obtained by comparing the test object with the candidate reference objects A1, A2, B1, and B2 respectively. The identification result must contain at least one of A1 and A2, at least one of B1 and B2, and at least one of C.

[0131] Optionally, when the first similarity comparison results between the test object and each reference object include candidate reference objects A, B, and C, and A, B, and C do not belong to the same reference object combination, the candidate reference objects A, B, and C can be directly used as the identification results of the test object.

[0132] If the first similarity comparison result between the test object and each reference object only contains A, and A does not belong to any combination of reference objects, then a second similarity comparison is not required, and the identification result is directly output.

[0133] Second, the specific implementation of "judging mixed information based on the similarity of gene information of candidate reference objects" is as follows: For example, when the first similarity comparison results between the test object and each reference object include candidate reference objects A and B, the gene information of A and B is obtained and the gene similarity is compared. If the gene similarity is equal to or lower than a certain threshold, a mixed reference result is output, which must contain at least A and B.

[0134] When the first similarity comparison results between the test object and each reference object include candidate reference objects A1, A2, and B, if the gene similarity between A1 and A2 is higher than a certain threshold, and the gene similarity between B and A1, and between B and A2 is equal to or lower than a certain threshold, then the output mixed reference result will contain at least B, and also contain the second similarity comparison results between the test object and A1 and A2.

[0135] Other possible scenarios are not listed here.

[0136] Third, the specific implementation logic of "judging whether candidate reference objects belong to different subspecies of the same species by mixing information" is consistent with the judgment logic of the first case, and will not be elaborated here.

[0137] Fourth, the specific implementation of "using multiple conditions to judge mixed information" includes: Optionally, when the first similarity comparison results between the test object and each reference object include candidate reference objects A and B, if the similarity between the genes of A and B is equal to or lower than a certain threshold, and A and B do not belong to the same reference object combination or different subspecies of the same species, then a mixed reference result is output, which must contain at least A and B.

[0138] The above only illustrates some implementation methods for judging mixed information. Other methods can also be used to judge the mixed information of multiple reference objects. In other judgment methods, only one principle needs to be met, that is, if there are objects that do not belong to the same combination of reference objects or do not belong to the same group of similar reference objects, then the mixed reference result is output. Each combination of reference objects or each group of similar reference objects needs to have a corresponding result in the mixed reference result.

[0139] [Validation of Hybrid Reference Results] In another embodiment of this application, the mixed reference results are verified based on the gene information, protein information, or protein fingerprint of the candidate reference object to eliminate the influence of mixing multiple substances on mass spectrometry detection. It should be noted that the verification of the mixed reference results generally occurs before outputting the mixed reference results.

[0140] The verification method can be one or more of the following methods: (1) Determine whether the similarity of gene information of each reference object in the mixed reference results is lower than a certain threshold; (2) Compare the gene of a certain reference object in the mixed reference results with the information of the other reference objects in the mixed reference results. Specifically, based on the gene inference of the mutually interfering protein feature expression, remove the mutually interfering protein feature expression in the standard protein fingerprint of the reference object (or set the weight to zero), and then compare the protein fingerprint after removing the interfering protein feature expression with the protein fingerprint of the test protein.

[0141] For example, if the mixed reference results contain two bacteria, A and B, the gene information of bacteria A and B is compared to identify and remove interfering protein features (or the genes corresponding to these protein expressions). Based on this, a set of distinguishing feature peaks is obtained after removing the interfering features of bacteria B from bacteria A. The set of distinguishing feature peaks is then used to perform a third similarity comparison on bacteria A to verify bacteria A. Similarly, bacteria B can also be verified. (3) Compare the protein information of a certain reference object in the mixed reference results (the set of characteristic proteins obtained from the protein data) with the protein information of the other reference objects in the mixed reference results. Specifically, remove the protein feature expressions that interfere with each other (or make the weight zero), and then compare the protein fingerprint after removing the interference protein feature expressions with the protein fingerprint of the test protein for the third similarity comparison. (4) Compare the protein fingerprint spectrum of a certain reference object in the mixed reference results with the protein fingerprint spectrum of the other reference objects in the mixed reference results, remove the protein feature expression that interferes with each other (or set the weight to zero), and then compare the protein fingerprint spectrum after removing the interference protein feature expression with the protein fingerprint spectrum to be tested for the third similarity comparison.

[0142] Optionally, the identification result is determined based on the result of the first similarity comparison and the verification result. Optionally, in verification method (1), if the similarity of the gene information of the corresponding species of each reference object in the mixed reference result is lower than a certain threshold, then the mixed reference result is correct. Optionally, in verification method (2) or (3) or (4), the result of the first similarity comparison and the result of the third similarity comparison are directly compared. If the two are consistent or the result of the third similarity comparison is not lower than the preset threshold, then the mixed reference result is correct. It should be noted that the consistency of the two here does not mean that the scores or similarity are the same, but only that the categories of the results are the same. For example, if the result of the first similarity comparison is A and B, and the result of the third similarity comparison is still A and B, then the result is correct.

[0143] Alternatively, this can be understood as removing protein feature expressions from the protein fingerprint map corresponding to a certain category in the mixed reference result that are identical to those of other categories in the mixed reference result, thus obtaining a mixed verification fingerprint map; then, a third similarity comparison is performed between the test object and the mixed verification fingerprint map, and the identification result is determined based on the results of the first similarity comparison and the third similarity comparison. This verification of the upcoming mixed reference result is particularly necessary when judging mixed information without considering the similarity of candidate reference object gene information.

[0144] Optionally, the method for removing protein feature expressions that are identical to other categories in the protein fingerprint of a certain category in the mixed reference result to obtain the mixed verification fingerprint is as follows: Based on the comparison of gene information of the species corresponding to the mixed reference results with the gene information of the other reference objects for the same species, and based on gene deduction, the protein feature expressions that interfere with each other between the mixed reference object and the other reference objects are removed from the standard protein fingerprint of the mixed reference object; or, Based on the protein data information in the protein database corresponding to the mixed reference results, the protein data information of each mixed reference object is compared, and interfering protein feature expressions are removed; or, The standard protein fingerprints corresponding to each category in the mixed reference results are compared, and interfering characteristic peaks are removed. For example, protein feature expression identical to other references in the protein fingerprint of the mixed reference object can be removed. That is, mutual interference can be understood as identical peaks. Of course, mutual interference can also be understood as peaks that result in the same or similar detected characteristic peaks.

[0145] Display of hybrid reference results In another embodiment of this application, the substance identification device provided by this application has a detection module for detecting the test object and obtaining its protein fingerprint spectrum, a storage module for storing substance identification-related data, a data processing module for calling data in the storage module for data processing, and a display module for outputting identification results. The data processing module is configured to: perform a first similarity comparison between the protein fingerprint of the object to be tested and the standard protein fingerprint of each reference object to obtain the first similarity comparison result; The display module is configured to display either a standard identification result or a mixed reference result based on the first similarity comparison result from the data processing module.

[0146] Optionally, the display module may display ordinary identification results or mixed reference results based on the first similarity comparison results of the data processing module. This may further include: the data processing module obtaining multiple candidate reference objects similar to the object to be tested based on the first similarity comparison results. The candidate reference objects include at least two types, and the two types may include different species or different subspecies. The system determines mixed information based on multiple candidate reference objects, and outputs either a standard identification result or a mixed reference result for the test object based on this mixed information. The specific method for determining mixed information is the same as described above and will not be repeated here.

[0147] In another embodiment of this application, the output mixed reference results can also be marked and displayed.

[0148] For example, when outputting a mixed reference result, the data processing module 103 can generate a prompt message to indicate that the output result is a mixed reference result.

[0149] For example, the data processing module 103 can mark and display the mixed reference results; wherein, it can mark and display each category of reference objects in the mixed reference results, or mark and display similar reference objects in the mixed reference results, or mark reference objects in the mixed reference results that do not belong to the same combination of reference objects or the same group of similar reference objects.

[0150] For example, when the mixed reference result includes candidate reference objects A1, A2, and B, and candidate reference objects A1 and A2 are similar reference objects, while candidate reference object B is not similar to either candidate reference objects A1 or A2, then candidate reference objects A1, A2, and B can all be marked and displayed, or candidate reference objects A1 and B can be marked, or candidate reference objects A2 and B can be marked; or candidate reference objects A1 and A2 can be marked. Specifically, candidate reference objects A1 and A2 can be indicated to be similar reference objects through at least one of the marking methods of text marking, image marking, or color marking.

[0151] Preferably, in the two methods of labeling candidate reference object A1 and candidate reference object B, and labeling candidate reference object A2 and candidate reference object B, if the similarity between candidate reference object A1 and candidate reference object B is higher than the similarity between candidate reference object A2 and candidate reference object B, then candidate reference object A1 and candidate reference object B are selected to be labeled.

[0152] In this embodiment of the application, by marking the mixed reference results, it is not only clear that the output result is a mixed reference result, but also clear which combination the mixed reference result belongs to, so as to facilitate clinical judgment.

[0153] In another embodiment of this application, another method for substance identification is also provided. For example, Figure 3 This is a schematic flowchart of another substance identification method provided in the embodiments of this application, as shown below. Figure 3 As shown, the method includes the following steps: Step S301: Obtain the protein fingerprint of the object to be tested.

[0154] Step S302: Perform a first similarity comparison between the protein fingerprint of the object to be tested and the standard protein fingerprint of the reference object in the protein fingerprint library to obtain a primary target set similar to the object to be tested.

[0155] For example, the primary target set may include multiple candidate reference objects similar to the object to be tested, wherein each candidate reference object may be a secondary target in the primary target set.

[0156] Specifically, each candidate reference object includes at least two types, wherein the two types include different species or different subspecies.

[0157] Step S303: Based on the protein difference information of each secondary target in the primary target set, perform a second similarity comparison between the test object and the secondary targets.

[0158] For example, protein expression differential analysis can be performed on each secondary target to obtain protein difference information for each secondary target. This protein difference information can be characteristic peaks in the standard protein fingerprint of a reference object.

[0159] For example, Vibrio cholerae and Vibrio mimicus have similar protein expression. When the sample to be tested is Vibrio cholerae or Vibrio mimicus, the protein fingerprint of the sample to be tested, after similarity comparison using existing methods, has a very close similarity score to Vibrio cholerae and Vibrio mimicus, making accurate differentiation impossible. Table 1 is a protein difference information table provided in the embodiments of this application. As shown in Table 1, the primary target set can include Vibrio cholerae and Vibrio mimicus, that is, the primary target set can be a combination of similar bacteria of Vibrio cholerae and Vibrio mimicus.

[0160] Table 1. Protein Differential Information Table

[0161] Specifically, when the secondary target is the L31 protein, the differential characteristic peaks of Vibrio cholerae and Vibrio mimicry can be the characteristic peaks corresponding to a mass-to-charge ratio of 7956 Da and 7970 Da. The protein fingerprints of Vibrio cholerae and Vibrio mimicry can be obtained using the ESX2600 mass spectrometer from Zhongyuan Huiji.

[0162] For example, the mass-to-charge ratio (m / z value), intensity, and weight information of the differential characteristic peaks in the protein fingerprint spectrum corresponding to the protein difference information can be adjusted to increase the proportion of the differential characteristic peaks in determining the similarity comparison results between the test object and the candidate reference object through the relevant data of the adjusted differential characteristic peaks.

[0163] Specifically, for each candidate reference object, the scoring weight of the differential characteristic peaks of the candidate reference object can be increased. Based on the weight and amplitude of each characteristic peak in the protein fingerprint spectrum of the candidate reference object, the similarity score between the candidate reference object and the test object is scored to obtain the similarity score between the candidate reference object and the test object. Based on the similarity score of the candidate reference object, the matching result between the test object and the candidate reference object is output.

[0164] For example, when comparing the bar-shaped peak information corresponding to the protein fingerprint of the test object with the protein fingerprint of the secondary target, a scoring operation can be performed based on the weight and amplitude of the characteristic peak when the two have the same characteristic peak, so as to obtain the similarity score between the secondary target and the test object; when the protein fingerprint of the test object shows the difference characteristic peak of the secondary target, a scoring operation can be performed based on the increased scoring weight (e.g., product coefficient).

[0165] Optionally, for each secondary target, the secondary target and the test object can be compared based on the differential characteristic peaks of the secondary target. The matching result between the test object and the secondary target can be output based on whether the protein fingerprint spectrum of the test object contains the same characteristic peaks as the differential characteristic peaks.

[0166] Step S304: Based on the second similarity comparison results between the object to be tested and the secondary target, determine the identification result of the object to be tested.

[0167] For example, Figure 4 This is a schematic diagram of the identification result of the test object provided in an embodiment of this application, such as... Figure 4 As shown, when the primary target set is a combination of similar bacteria such as Vibrio cholerae and Vibrio mimicus, the similarity score between the candidate reference object and the object to be tested can be obtained by increasing the scoring weight of each difference feature peak. Figure 4 The similarity score shown indicates that the identification result of the test object can be Vibrio mimicus.

[0168] It should be noted that, as shown in Table 1 above, the protein differences between Vibrio cholerae and Vibrio mimicry are quite similar (i.e., the differences in the protein fingerprints of Vibrio cholerae and Vibrio mimicry are small). This makes it difficult for existing identification methods to distinguish between Vibrio cholerae and Vibrio mimicry. However, the method described in this application, which determines the similarity scores between candidate reference objects and test objects based on the differential protein information of each test object, can make the differences between test objects and candidate reference objects more obvious.

[0169] The following is for reference. Figure 5 , Figure 5 A schematic diagram of a computer device suitable for implementing embodiments of this application is shown, such as... Figure 5 As shown, the computer device 500 includes a central processing unit (CPU) 501, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 502 or a program loaded from a storage section 505 into a random access memory (RAM) 503. The RAM 503 also stores various programs and data required for the system's operating instructions. The CPU 501, ROM 502, and RAM 503 are interconnected via a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.

[0170] The following components are connected to the input / output (I / O) interface 505: an input section 506 including a keyboard, mouse, etc.; an output section 507 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and speakers, etc.; a storage section 508 including a hard disk, etc.; and a communication section 509 including a network interface card such as a LAN card, modem, etc. The communication section 509 performs communication processing via a network such as the Internet. A drive 510 is also connected to the input / output (I / O) interface 505 as needed. A removable medium 511, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed on the drive 510 as needed so that computer programs read from it can be installed into the storage section 508 as needed.

[0171] Specifically, according to embodiments of this application, the flowchart above refers to... Figures 2-3Any of the described processes can be implemented as a computer software program. For example, embodiments of this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowchart. In such an embodiment, the computer program contains program code for performing the methods shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via communication section 509, and / or installed from removable medium 511. When the computer program is executed by central processing unit (CPU) 501, it performs the functions defined in the system of this application.

[0172] It should be noted that the computer-readable medium shown in this application can be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this application, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this application, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. Computer-readable signal media can also be any computer-readable medium compatible with computer-readable storage media, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wireless, wire, optical fiber, RF, etc., or any suitable combination thereof.

[0173] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operational instructions of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two connected blocks may actually be executed substantially in parallel, or they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified functions or operational instructions, or using a combination of dedicated hardware and computer instructions.

[0174] The units or modules described in the embodiments of this application can be implemented in software or hardware. The described units or modules can also be housed in a processor; for example, a processor may be described as including a semantic extraction unit, a weight allocation unit, and a determination unit. The names of these units or modules do not necessarily constitute a limitation on the unit or module itself.

[0175] On the other hand, this application also provides a computer-readable storage medium, which may be included in the computer device described in the above embodiments, or may exist independently and not assembled into the computer device. The aforementioned computer-readable storage medium stores one or more programs that, when used by one or more processors, execute the methods described in this application. For example, it may execute... Figures 2-3 Each step of any of the methods shown.

[0176] This application provides a computer program product including instructions that, when executed, cause the method described in this application to be performed. For example, it can execute... Figures 2-3 Each step of any of the methods shown.

[0177] The above description is merely a preferred embodiment of this application and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of disclosure in this application is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the foregoing disclosed concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features with similar functions disclosed in this application.

Claims

1. A substance identification device, characterized in that, The device includes a detection module for detecting the test object and obtaining its protein fingerprint, a storage module for storing data related to substance identification, and a data processing module for calling data from the storage module for data processing. The data processing module is configured as follows: The protein fingerprint of the test object is compared with the standard protein fingerprint of each reference object to obtain multiple candidate reference objects that are similar to the test object. The candidate reference objects include at least two types, and the two types include different species or different subspecies. Determine the protein difference information of each of the candidate reference objects; the protein difference information is the protein feature expression of the candidate reference object, and the protein difference information is used to distinguish the candidate reference object from one or more other candidate reference objects; Based on the protein difference information of the candidate reference objects, a second similarity comparison is performed between the test object and the candidate reference objects, and the identification result of the test object is determined based on the comparison results of each candidate reference object.

2. The apparatus according to claim 1, characterized in that, The step of performing a first similarity comparison between the protein fingerprint of the test object and the standard protein fingerprint of each reference object to obtain multiple candidate reference objects similar to the test object includes: The protein fingerprint of the object to be tested is compared with the standard protein fingerprint of each reference object to obtain the similarity result between each reference object and the object to be tested; the similarity result corresponds to the similarity level. Based on the similarity results, N reference objects with the highest similarity are obtained, and the candidate reference objects are determined based on the N reference objects; where N is an integer greater than 1.

3. The apparatus according to claim 2, characterized in that, The step of determining the candidate reference object based on the N reference objects includes: The N reference objects are selected as candidate reference objects; or... Select reference objects of the same type from the N reference objects as candidate reference objects; or... If there is a target reference object belonging to the reference object combination among the N reference objects, then the candidate reference object is determined based on the reference object combination and the N reference objects; wherein, the reference object combination includes multiple reference objects with similar protein feature expression.

4. The apparatus according to claim 2, characterized in that, The data processing module is also configured to: Add update information to the reference object combination information, wherein the update information is used to characterize the reference object combination created or updated based on the N reference objects, and the protein fingerprint of the N reference objects; The reference object combination information is used to characterize multiple reference object combinations and the protein fingerprint map corresponding to each reference object combination. The reference object combination includes multiple reference objects with similar protein feature expression.

5. The apparatus according to claim 1, characterized in that, The determination of protein difference information for each of the candidate reference objects includes: Protein expression differential analysis was performed on each of the candidate reference objects to obtain protein difference information for each candidate reference object.

6. The apparatus according to any one of claims 1-5, characterized in that, The step of determining the protein difference information of each of the candidate reference objects, wherein the protein difference information is the protein feature expression of the candidate reference object, and the protein difference information is used to distinguish the candidate reference object from one or more other candidate reference objects; Based on the protein difference information of the candidate reference objects, a second similarity comparison is performed between the test object and the candidate reference objects, and the identification result of the test object is determined based on the comparison results of each candidate reference object, including: The genomes of each candidate reference object are obtained; characteristic proteins of the candidate reference objects are determined based on their genomes; differentially expressed proteins of each candidate reference object are determined based on their characteristic proteins; differentially expressed characteristic peaks of each candidate reference object are determined based on the characteristic peaks corresponding to the differentially expressed proteins in the corresponding protein fingerprint profiles; a second similarity comparison is performed between the test object and the candidate reference objects based on the differentially expressed characteristic peaks of the candidate reference objects; and the identification result of the test object is determined based on the comparison results of each candidate reference object; or... The genomes of each candidate reference object are obtained; differentially expressed genomes are identified among the candidate reference objects based on their genomes; differentially expressed proteins are identified among the candidate reference objects based on their differentially expressed genomes; differentially expressed characteristic peaks are identified among the candidate reference objects based on the characteristic peaks corresponding to the differentially expressed proteins in their respective protein fingerprints; a second similarity comparison is performed between the test object and the candidate reference objects based on the differentially expressed characteristic peaks; and the identification result of the test object is determined based on the comparison results of each candidate reference object. Alternatively, The genomes of each candidate reference object are obtained; characteristic proteins of the candidate reference objects are determined based on their genomes; characteristic peaks of each candidate reference object in corresponding protein fingerprints are determined based on their characteristic proteins; differential characteristic peaks of each candidate reference object are determined based on their characteristic peaks in corresponding protein fingerprints; a second similarity comparison is performed between the test object and the candidate reference objects based on the differential characteristic peaks of the candidate reference objects; and the identification result of the test object is determined based on the comparison results of each candidate reference object; or... The method involves: obtaining characteristic proteins of the candidate reference objects from a protein database; determining differentially expressed proteins of each candidate reference object based on their characteristic proteins; determining differential characteristic peaks of each candidate reference object based on the characteristic peaks corresponding to the differentially expressed proteins in the corresponding protein fingerprints; performing a second similarity comparison between the test object and the candidate reference objects based on the differential characteristic peaks of the candidate reference objects; and determining the identification result of the test object based on the comparison results of each candidate reference object. Alternatively... The method involves: obtaining characteristic proteins of the candidate reference objects from a protein database; determining the characteristic peaks of each candidate reference object in the corresponding protein fingerprint based on these characteristic proteins; determining the differential characteristic peaks of each candidate reference object based on these differential characteristic peaks; performing a second similarity comparison between the test object and the candidate reference objects based on these differential characteristic peaks; and determining the identification result of the test object based on the comparison results of each candidate reference object. Alternatively... The protein fingerprints of each candidate reference object are compared to obtain the differential characteristic peaks of each candidate reference object. Based on the differential characteristic peaks of the candidate reference objects, a second similarity comparison is performed between the test object and the candidate reference objects. Based on the comparison results of each candidate reference object, the identification result of the test object is determined. The differential feature peaks of the candidate reference object, or combinations of differential feature peaks formed by the differential feature peaks of the candidate reference object, are used to uniquely characterize the candidate reference object.

7. The apparatus according to claim 6, characterized in that, The differential characteristic peaks of the candidate reference object are the characteristic peaks corresponding to the gene segments of the candidate reference object that are mutually exclusive with one or more other candidate reference objects.

8. The apparatus according to claim 6, characterized in that, Determining the protein difference information for each of the candidate reference objects includes: Ignore the same protein feature expression in the protein fingerprint of each candidate reference object, and determine the differential feature peaks of each candidate reference object based on the protein fingerprint after ignoring the same protein feature expression.

9. The apparatus according to claim 6, characterized in that, The step of obtaining the genomes of each of the candidate reference objects, determining the characteristic proteins of the candidate reference objects based on the genomes of the candidate reference objects, determining the differentially expressed proteins of each of the candidate reference objects based on the characteristic proteins of each of the candidate reference objects, and determining the differential characteristic peaks of each of the candidate reference objects based on the characteristic peaks corresponding to the differentially expressed proteins in the corresponding protein fingerprints, further includes: The protein expression loss at the differential characteristic peak in the genome is determined, and the differential characteristic peak in the protein fingerprint is corrected based on the protein expression loss.

10. The apparatus according to any one of claims 1-9, characterized in that, The step of performing a second similarity comparison between the test object and the candidate reference objects based on the protein difference information of the candidate reference objects, and determining the identification result of the test object based on the comparison results of each candidate reference object, includes: For each candidate reference object, the scoring weight of the difference characteristic peak of the candidate reference object is increased; Based on the weights and amplitudes of each characteristic peak in the protein fingerprint of the candidate reference object, a similarity score is assigned to the candidate reference object and the object to be tested to obtain the similarity result between the object to be tested and the candidate reference object.

11. The apparatus according to claim 9, characterized in that, The step of performing a second similarity comparison between the test object and the candidate reference objects based on the protein difference information of the candidate reference objects, and determining the identification result of the test object based on the comparison results of each candidate reference object, includes: For each candidate reference object, the candidate reference object and the object to be tested are compared based on the difference feature peaks of the candidate reference object; Based on whether the protein fingerprint of the test object contains the same characteristic peak as the differential characteristic peak, the matching result between the test object and the candidate reference object is output.

12. The apparatus according to any one of claims 1-11, characterized in that, The plurality of candidate reference objects includes at least multiple subspecies belonging to the same bacterial species, and the identification result of the test object corresponds to the subspecies; or, The plurality of candidate reference objects includes at least multiple strains belonging to the same genus, and the identification result of the test object corresponds to the bacterial species; or, The plurality of candidate reference objects includes at least multiple species belonging to the same complex microbial community, and the identification result to be tested corresponds to the microbial species; or, The multiple candidate reference objects include at least multiple bacteria belonging to the same list of similar bacteria, and the identification result of the object to be tested corresponds to the bacteria.

13. A method for identifying a substance, characterized in that, include: The protein fingerprint of the test object is compared with the standard protein fingerprint of each reference object to obtain multiple candidate reference objects that are similar to the test object. The candidate reference objects include at least two types, and the two types include different species or different subspecies. Determine the protein difference information of each of the candidate reference objects; the protein difference information is the protein feature expression of the candidate reference object, and the protein difference information is used to distinguish the candidate reference object from one or more other candidate reference objects; Based on the protein difference information of the candidate reference objects, a second similarity comparison is performed between the test object and the candidate reference objects, and the identification result of the test object is determined based on the comparison results of each candidate reference object.

14. A computer device comprising a memory and a processor, the memory storing instructions, the processor executing the instructions to perform the steps performed by the substance identification apparatus of any one of claims 1-12 and the method of claim 13.

15. A computer program product, characterized in that, The computer program product includes instructions that, when executed, cause the steps performed by the substance identification apparatus as claimed in any one of claims 1-12 to be implemented, and the method as claimed in claim 13 to be implemented.

16. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps performed by the substance identification device in any one of claims 1-12 and the method described in claim 13.