Lipid identification method, apparatus, and system, computer-readable storage medium
By using data fitting and spectral matching methods, the problems of long time consumption and high false positive rate in existing lipid identification methods have been solved, achieving efficient and accurate lipid identification.
Patent Information
- Application Number
- CN202411603940.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-11
- Publication Date
- 2025-11-07
- Estimated Expiration
- 2044-11-11
AI Technical Summary
Existing lipid identification methods are time-consuming and prone to misjudgment, often requiring experienced experts to spend several weeks on them, and identification methods based on retention time are prone to high false positive rates.
By obtaining the identification results of lipid samples, their unsaturation information and subclass are determined. Data fitting is performed using retention time and mass-to-charge ratio as variables to construct a lipid group intelligent model. The "fragment tree" hierarchical spectral library and hash table are used to optimize spectral matching, and identification is performed by combining various derivatization mass spectrometry methods.
It improves the accuracy and efficiency of lipid identification, reduces false positive results, shortens identification time, and reduces reliance on expert experience.
Smart Images

Figure CN119541680B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the technical field of computer, and particularly relates to a lipid identification method, device and system, and a computer readable storage medium. BACKGROUND
[0002] Lipids are a class of organic compounds composed of carbon, hydrogen and oxygen elements, usually containing long-chain hydrocarbons (fatty acids) and connected with polar groups (hydroxyl, phosphate groups, etc.). The standard representation of lipids is composed of "lipid subclass, total carbon number: double bond number, chain carbon number: double bond number", for example, PC 36:3 (PC 18:1_18:2). Lipids are high in synthesis cost and numerous, even the same subclass with different side chain combinations can reach tens of millions of kinds, so there is a lack of comprehensive standards.
[0003] Mass spectrometry is an analysis method widely used in the fields of environment, biology, medicine, etc. First-order and second-order information of different substances can be obtained through mass spectrometry instruments. The first-order information includes mass-to-charge ratio, retention time, corresponding precursor, etc. of the substance, and the second-order information includes mass-to-charge ratio and response intensity of a series of fragment ions. SUMMARY
[0004] According to a first aspect of the present disclosure, a lipid identification method is provided, comprising: obtaining an identification result of each lipid sample of a plurality of lipid samples; determining the unsaturation information and the belonging subclass of each lipid sample according to the identification result of each lipid sample; performing data fitting on the plurality of lipid samples based on the unsaturation information and the belonging subclass, with the retention time and the mass-to-charge ratio of each lipid sample as variables; and determining the accuracy of the identification result of each lipid sample according to the result of the data fitting.
[0005] In some embodiments, the data fitting on the plurality of lipid samples based on the unsaturation information and the belonging subclass, with the retention time and the mass-to-charge ratio of each lipid sample as variables, comprises: mapping each lipid sample into a rectangular coordinate system with the retention time and the mass-to-charge ratio as coordinates, wherein each lipid sample is represented by a sample point in the coordinate system; and performing data fitting on the plurality of lipid samples according to the relationship among the unsaturation information, the belonging subclass, the retention time and the mass-to-charge ratio.
[0006] In some embodiments, the data fitting of the plurality of lipid samples according to the relationship among the unsaturation information, the belonging subclass, the retention time and the mass-to-charge ratio comprises: selecting a starting point satisfying a condition from the sample points according to the relationship among the unsaturation information, the belonging subclass, the retention time and the mass-to-charge ratio to add the starting point to a first set to obtain a preliminary fitting curve; constructing a sub-coordinate system with the starting point as the origin, and searching for a target sample point satisfying the condition in the first and third quadrants of the sub-coordinate system corresponding to the starting point to add the target sample point to the first set; and whenever the number of the target sample points added to the first set reaches a threshold, re-determining the preliminary fitting curve according to the first set.
[0007] In some embodiments, the selecting a starting point satisfying a condition from the sample points according to the relationship among the unsaturation information, the belonging subclass, the retention time and the mass-to-charge ratio to add the starting point to a first set to obtain a preliminary fitting curve comprises: selecting a starting point satisfying a condition from the sample points according to the relationship among the unsaturation information, the belonging subclass, the retention time and the mass-to-charge ratio, and the measured mass-to-charge ratio of the sample point and the mass-to-charge ratio obtained by data fitting to add the starting point to a first set to obtain a preliminary fitting curve.
[0008] In some embodiments, the unsaturation information comprises at least one of a total unsaturation number and a branch-chain unsaturation number, and the condition comprises at least one of: a fitting degree of a curve fitted according to the sample points satisfying the condition exceeding a first threshold; on the curve fitted according to the sample points satisfying the condition, for a first sample point and a second sample point with the same total unsaturation number, in a case where the retention time of the first sample point is greater than the retention time of the second sample point, the mass-to-charge ratio of the first sample point is greater than the mass-to-charge ratio of the second sample point; on the curve fitted according to the sample points satisfying the condition, for a third sample point and a fourth sample point with the same branch-chain unsaturation number, in a case where the retention time of the third sample point is greater than the retention time of the fourth sample point, the mass-to-charge ratio of the third sample point is greater than the mass-to-charge ratio of the fourth sample point.
[0009] In some embodiments, the determining the accuracy of the identification result of each lipid sample according to the result of the data fitting comprises: determining the accuracy of the identification result of each lipid sample according to whether the sample point is in the first set and the measured mass-to-charge ratio of the sample point and the mass-to-charge ratio obtained by data fitting.
[0010] In some embodiments, the determining the accuracy of the identification result of each lipid sample according to whether the sample point is in the first set and the measured mass-to-charge ratio of the sample point and the data-fitted mass-to-charge ratio comprises: determining a loss value of each sample point in the first set according to the measured mass-to-charge ratio and the data-fitted mass-to-charge ratio; determining a first loss value as the maximum loss value of the sample points in the first set; determining a second loss value as the minimum loss value of the sample points in the first set; and for each sample point in the first set, determining the accuracy of the identification result of the lipid sample corresponding to the sample point according to the loss value of the sample point, the first loss value and the second loss value.
[0011] In some embodiments, the determining the accuracy of the identification result of each lipid sample according to whether the sample point is in the first set and the measured mass-to-charge ratio of the sample point and the data-fitted mass-to-charge ratio comprises: determining a confidence interval of the target curve; and for the sample points not in the first set, determining the accuracy of the identification result of the corresponding lipid sample according to whether the sample point is in the confidence interval.
[0012] In some embodiments, the determining the accuracy of the identification result of each lipid sample according to whether the sample point is in the first set and the measured mass-to-charge ratio of the sample point and the data-fitted mass-to-charge ratio comprises: determining a confidence interval of the target curve; and for the sample points not in the first set, determining the accuracy of the identification result of the corresponding lipid sample according to whether the sample point is in the confidence interval.
[0013] In some embodiments, the unsaturation information comprises a total unsaturation number, and the data fitting of the plurality of lipid samples based on the unsaturation information and the sub-class comprises: dividing the plurality of lipid samples according to the sub-class and the total unsaturation number to obtain a plurality of second sets, wherein the lipid samples in each second set belong to the same sub-class and have the same total unsaturation number; and for each second set, performing data fitting with the retention time and the mass-to-charge ratio of the lipid samples in the second set as variables to obtain a first fitting curve corresponding to the sub-class and the total unsaturation number of the second set.
[0014] In some embodiments, the data fitting, for each second set, with the retention time and mass-to-charge ratio of the lipid samples therein as variables, to obtain a first fitting curve corresponding to the sub-class and the total unsaturation number of the second set, comprises: in a case that the number of the lipid samples of the first total unsaturation number of the specified sub-class is less than a threshold value, determining the first fitting curve corresponding to the first total unsaturation number of the specified sub-class according to the first fitting curve corresponding to the second total unsaturation number of the specified sub-class of the second set, wherein the first total unsaturation number and the second total unsaturation number are different.
[0015] In some embodiments, the data fitting, for each second set, with the retention time and mass-to-charge ratio of the lipid samples therein as variables, to obtain a first fitting curve corresponding to the sub-class and the total unsaturation number of the second set, comprises: in a case that the number of the lipid samples of the first total unsaturation number of the specified sub-class is less than a threshold value, determining the first fitting curve corresponding to the first total unsaturation number of the specified sub-class according to the first fitting curve corresponding to the second total unsaturation number of the specified sub-class of the second set, wherein the first total unsaturation number and the second total unsaturation number are different.
[0016] In some embodiments, the unsaturation information further comprises an unsaturation distribution, and the data fitting, for each lipid sample, with the retention time and mass-to-charge ratio of the lipid sample as variables, based on the unsaturation information and the sub-class to which the lipid sample belongs, comprises: dividing the plurality of lipid samples according to the sub-class, the total unsaturation number and the unsaturation distribution to obtain a plurality of third sets, wherein the lipid samples in each third set belong to the same sub-class and have the same total unsaturation number and the same unsaturation distribution; and for each third set, data fitting with the retention time and mass-to-charge ratio of the lipid samples therein as variables to obtain a second fitting curve corresponding to the sub-class, the total unsaturation number and the unsaturation distribution of the third set.
[0017] In some embodiments, the determining, according to the result of the data fitting, the accuracy of the identification result of each lipid sample, comprises: determining a first accuracy according to the first fitting curve corresponding to each lipid sample; determining a second accuracy according to the second fitting curve corresponding to each lipid sample; and determining the accuracy of the identification result according to the first accuracy and the second accuracy.
[0018] In some embodiments, the obtaining the identification result of each lipid sample of the plurality of lipid samples comprises: identifying the species of each lipid sample as the identification result of each lipid sample by comparing the spectrum of each lipid sample with a reference spectrum.
[0019] In some embodiments, the identifying the category of each lipid sample by comparing the spectrum of each lipid sample with the reference spectrum includes: calculating a first score of each lipid sample according to a number of matched characteristic peaks of the spectrum of each lipid sample and the reference spectrum; calculating a second score of each lipid sample according to an importance of the matched characteristic peaks of the spectrum of each lipid sample and the reference spectrum; determining the reference spectrum matched with the spectrum of each lipid sample according to the first score and the second score of each lipid sample; and determining the category of each lipid sample as the identification result of each lipid sample according to the matched reference spectrum.
[0020] According to a second aspect of the present disclosure, a lipid identification device is provided, including: an acquisition module configured to acquire an identification result of each lipid sample of a plurality of lipid samples; an information determination module configured to determine unsaturation information and a belonging subcategory of each lipid sample according to the identification result of each lipid sample; a fitting module configured to perform data fitting on the plurality of lipid samples based on the unsaturation information and the belonging subcategory with the retention time and the mass-to-charge ratio of each lipid sample as variables; and an accuracy determination module configured to determine an accuracy of the identification result of each lipid sample according to a result of the data fitting.
[0021] According to a third aspect of the present disclosure, a lipid identification device is provided, including: a memory; and a processor coupled to the memory, the processor being configured to execute a lipid identification method according to some embodiments of the present disclosure based on instructions stored in the memory.
[0022] According to a fourth aspect of the present disclosure, a lipid identification system is provided, including: a lipid identification device according to some embodiments of the present disclosure.
[0023] According to a fifth aspect of the present disclosure, a computer-readable storage medium is provided, having computer program instructions stored thereon, the instructions being executed by a processor to implement a lipid identification method according to some embodiments of the present disclosure.
[0024] According to a sixth aspect of the present disclosure, a computer program product is provided, including computer program instructions, the computer program instructions being executed by a processor to implement a lipid identification method according to some embodiments of the present disclosure. BRIEF DESCRIPTION OF DRAWINGS
[0025] The accompanying drawings, which form a part of the specification, illustrate embodiments of the present disclosure and, together with the description, serve to explain the principles of the present disclosure.
[0026] The present disclosure can be more clearly understood and appreciated from the following detailed description, taken in conjunction with the following drawings of which:
[0027] Figure 1 a flowchart showing a method of lipid identification according to some embodiments of the present disclosure;
[0028] Figure 2 a schematic diagram showing lipid identification according to some embodiments of the present disclosure;
[0029] Figure 3 a block diagram showing a lipid identification apparatus according to some embodiments of the present disclosure;
[0030] Figure 4 a block diagram showing a lipid identification apparatus according to some embodiments of the present disclosure;
[0031] Figure 5 a block diagram showing a computer system for implementing some embodiments of the present disclosure;
[0032] Figures 6a-6c a user interface of a lipid identification system for implementing some embodiments of the present disclosure. DETAILED DESCRIPTION
[0033] Various exemplary embodiments of the present disclosure will now be described in detail with reference to the accompanying drawings. It should be noted that the relative arrangement of the components and steps set forth in the embodiments, numerical expressions, and numerical values are not limiting to the scope of the present disclosure unless specifically stated otherwise.
[0034] It should be understood, of course, that the various embodiments of the disclosure are merely examples of the various application of the principles of the present disclosure. Accordingly, those skilled in the art will recognize that the disclosure has other applications in other environments. Moreover, the use of the terms "preferably," "preferably," "more preferably," "most preferably" and other terms beginning with the word "preferably" indicates that the described feature is not necessarily required. For example, that a particular silyl group is "preferably" a trimethylsilyl group indicates that a trimethylsilyl group is a preferred silyl group, but that other silyl groups can also be used.
[0035] The following description of at least one example embodiment is merely exemplary in nature and is in no way intended to limit the scope of the disclosure, its application, or uses.
[0036] Techniques, methods, and apparatus known to those of ordinary skill in the relevant art can not be discussed in detail herein. However, where appropriate, such techniques, methods, and apparatus can be considered as part of the present disclosure.
[0037] In all of the examples shown and discussed herein, any specific values should be interpreted as merely exemplary and not as a limitation. Thus, other examples of the exemplary embodiments can have different values.
[0038] It should be noted that like reference numerals and letters in the various figures indicate similar items, and thus, once an item is defined in one figure, it is not necessary to discuss it further in subsequent figures.
[0039] It is an important application of mass spectrometry technology to identify substances by comparing mass spectrometry data with standard library and characteristic peak spectrum library information. The lipid identification method in the related technology uses manual inspection of characteristic spectrum and equivalent carbon number model of retention time to comprehensively identify lipids. This method is time-consuming and prone to misjudgment, and often requires 3-5 years of experience of experts to spend several weeks to complete the identification of a mass spectrometry file.
[0040] The identification method based on similarity matching algorithm and absolute retention time ignores the actual meaning of characteristic peaks, causing characteristic redundancy. At the same time, since the retention time is affected by experimental conditions, only relying on absolute retention time for judgment is easy to lead to high false positive rate in the identification result.
[0041] The present disclosure provides a lipid identification method, device and system, and computer readable storage medium, which can improve the accuracy of lipid identification.
[0042] Figure 1 A flowchart of a lipid identification method according to some embodiments of the present disclosure is shown.
[0043] As shown in Figure 1 The lipid identification method includes steps S1-S4. In some embodiments, the lipid identification method is performed by a lipid identification device.
[0044] In step S1, the identification result of each lipid sample of a plurality of lipid samples is obtained. In step S2, according to the identification result of each lipid sample, the unsaturation information and the belonging subclass of each lipid sample are determined. In step S3, taking the retention time and mass-to-charge ratio of each lipid sample as variables, the plurality of lipid samples are data fitted based on the unsaturation information and the belonging subclass. In step S4, according to the result of data fitting, the accuracy of the identification result of each lipid sample is determined.
[0045] The identification result is, for example, the type of the lipid sample, such as oleic acid, linoleic acid, etc. After the type of the lipid sample is identified, its structure can be determined, and thus its saturation information and belonging subclass can be determined. The retention time and mass-to-charge ratio of the lipid sample are obtained through instrument experiments.
[0046] There is a certain rule between the retention time and mass-to-charge ratio of a plurality of lipid samples with the same unsaturation information and / or subclass. If the identification result is incorrect, the lipid sample will deviate from these rules to some extent. Therefore, the accuracy of the identification result of each lipid sample can be determined by data fitting of a plurality of lipid samples according to the relationship between the retention time, mass-to-charge ratio, unsaturation information and belonging subclass.
[0047] For example, if two lipid samples belong to the same subcategory and have the same unsaturation information, there is a certain rule between the retention time and the mass-to-charge ratio of the two, for example, the lipid sample with a longer retention time has a larger mass-to-charge ratio.
[0048] The lipid identification method of the present disclosure takes the retention time and mass-to-charge ratio of each lipid sample as variables, and performs data fitting on a plurality of lipid samples based on the unsaturation information and the subcategory to which each lipid sample belongs, thereby determining the accuracy of the identification result of each lipid sample. The lipid identification method can be used to identify false positive identification results, thereby improving the accuracy of lipid identification.
[0049] The process of obtaining an identification result according to some embodiments of the present disclosure is described below.
[0050] According to the lipid characteristic spectrum, a “fragment tree” hierarchical spectrum library of 166.3 million lipids can be constructed. The hierarchical spectrum library includes five levels. The first level mainly includes the precursor ion peaks of lipids and the peaks caused by the loss of the head group of the substance; the second level mainly includes the mass spectrometry peaks of the side chain of the lipid; the third level includes the mass spectrometry peaks caused by the neutral loss of the precursor ion and the head group; the fourth level includes the mass spectrometry peaks caused by the neutral loss of the side chain of the lipid; and the fifth level includes the fingerprint spectrum generated by the reverse metabolomics algorithm.
[0051] For example, the spectrum is iteratively calculated according to the fragmentation rules in R language, and the spectrum is stored in the form of a list, and each element in the list is the spectrum information of a substance at a certain level. Finally, it is stored in the “rda” file format. Through the construction of the above-mentioned “fragment tree” hierarchical spectrum library, the feature redundancy caused by the similarity algorithm is reduced. In addition, in the construction of the spectrum library, the spectrum of the derivative is added, which is beneficial to the subsequent combination of various derivative mass spectrometry methods for further identification of the structure of the lipid (such as sn position isomerism and carbon-carbon double bond position isomerism).
[0052] By deeply mining the characteristics of the lipid mass spectrometry spectrum, a hierarchical library of 166 million can be constructed. The database is iteratively calculated, such as carbon number from 10-80 and unsaturation from 0-20, so as to be able to cover many potential lipid substances, i.e. non-endogenous or unreported chain composition.
[0053] The reference spectrum stored in the “fragment tree” hierarchical spectrum library can be used for comparison with the spectrum of the lipid sample.
[0054] In some embodiments, by comparing the spectrum of each lipid sample with the reference spectrum, the species of each lipid sample is identified as the identification result of each lipid sample.
[0055] For example, the hash table and binary search can be combined to optimize the data structure and perform the comparison between the spectrum of the lipid sample and the reference spectrum in the library. By combining the hash table and the binary search, the spectrum matching can be realized at a speed of 10-100 billion times per second, thereby improving the matching speed.
[0056] In some embodiments, identifying the species of each lipid sample by comparing the spectrum of each lipid sample with the reference spectrum includes: calculating a first score according to the number of matched characteristic peaks of the spectrum of each lipid sample and the reference spectrum; calculating a second score of each lipid sample according to the importance of the matched characteristic peaks of the spectrum of each lipid sample and the reference spectrum; determining the reference spectrum matched with the spectrum of each lipid sample according to the first score and the second score; and determining the species of each lipid sample according to the matched reference spectrum as the identification result of each lipid sample.
[0057] The first score is denoted as Score matched , and the second score is denoted as Score ratio . The following describes how to calculate the scores.
[0058] In a spectrum matching task, a reference spectrum in the library has n1 characteristic peaks, and the spectrum to be measured has n2 peaks. The formula is as follows.
[0059]
[0060] wherein s i represents the mass-to-charge ratio of the i-th peak of the reference spectrum. s' j represents the mass-to-charge ratio of the j-th peak in the spectrum of the lipid sample (the response of the mass spectrum peak is denoted as intensity j ). I(x) is an indicator function, which is 1 when the condition is met, and 0 otherwise. threshold ppm represents the acceptable mass-to-charge ratio error.
[0061] Score matched evaluates the matching of the characteristic peaks of the spectrum of the lipid sample with the reference spectrum in the library. For example, a spectrum in the library has 8 characteristic peaks, and 6 peaks in the spectrum of the lipid sample match the 8 characteristic peaks, and the Score matched value is 0.75. Within the allowed mass-to-charge ratio error, if multiple peaks match the same characteristic peak, only one can be counted.
[0062] Score ratio evaluates whether the 6 matched characteristic peaks represent the main information of the entire spectrum of the lipid sample. For example, the total response value of the 6 characteristic peaks is 0.7, and the total response value of the spectrum of the lipid sample is 1, and the Score ratio value is 0.7.
[0063] Due to the deviation of mass spectrometry instruments in the actual detection of substances, the matching of the above-mentioned characteristic peaks is all under the comparison of the allowed mass-to-charge ratio error threshold ppm
[0064] For a task of identifying n substances, the time complexity of the matching module after using the hash table is O(mΔ), where m represents the number of spectrum graphs in the library, and Δ represents the allowed error of matching.
[0065] In some embodiments, based on the unsaturation information and the belonging subclass, a plurality of lipid samples are data fitted with the retention time and mass-to-charge ratio of each lipid sample as variables, including: mapping each lipid sample to a rectangular coordinate system with the retention time and mass-to-charge ratio as coordinates, wherein each lipid sample is represented by a sample point in the coordinate system; according to the relationship between the unsaturation information, the belonging subclass, the retention time and the mass-to-charge ratio, the plurality of lipid samples are data fitted.
[0066] For example, in the coordinate system, the x-axis is the retention time and the y-axis is the mass-to-charge ratio, each lipid sample can be mapped to a point in the rectangular coordinate system, and the label of the point is, for example, the specific name of the lipid determined according to the identification result. According to the unsaturation information and the belonging subclass, the lipid samples are divided into sets. For the lipid samples belonging to the same set, the data is fitted according to the retention time and the mass-to-charge ratio.
[0067] The present disclosure statistically obtains three rules of equivalent carbon number model (ECN), equivalent chain composition carbon number model (ESCN) and intra-subclass unsaturation parallel (IUP).
[0068] ECN requires that substances in a lipid subclass with the same total unsaturation increase in retention time as the mass-to-charge ratio increases, and the changes in mass-to-charge ratio and retention time show a monotonically increasing quadratic function trend.
[0069] ESCN requires that substances in a lipid subclass with the same chain unsaturation composition increase in retention time as the mass-to-charge ratio increases, and the changes in mass-to-charge ratio and retention time show a monotonically increasing quadratic function trend.
[0070] IUP requires that the fitted quadratic curves of different total unsaturation in the same lipid subclass do not intersect in the closed interval.
[0071] After more than 100 data sets from different experimental conditions and different species are statistically analyzed, the proportion of samples meeting the above three conditions is about 90%, and the above three conditions lay the foundation for the intelligent cooperation behavior of sub-individuals in the construction of the lipid swarm intelligence model. The lipid swarm intelligence model can be used to remove false positives in identification.
[0072] In some embodiments, the data fitting of the plurality of lipid samples according to the relationship among the unsaturation information, the belonging subclass, the retention time and the mass-to-charge ratio comprises: selecting a starting point meeting the condition from the sample points to join a first set according to the relationship among the unsaturation information, the belonging subclass, the retention time and the mass-to-charge ratio, to obtain a preliminary fitting curve; constructing a sub-coordinate system with the starting point as the origin, and searching for a target sample point meeting the condition in the first and fourth quadrants of the sub-coordinate system corresponding to the starting point to join the first set; whenever the number of the target sample points newly joining the first set reaches a threshold, re-determining the preliminary fitting curve according to the first set; and fitting a target curve according to the first set.
[0073] For example, the lipid group intelligent model first greedily finds the optimal point combination as a starting point, and then performs a global search from the optimal starting point combination. After completing the search, a polynomial is used to fit the curve of all points meeting the search condition, and points falling within the 95% confidence interval of the fitted curve are found back.
[0074] After obtaining the starting point set, a global search is performed in the first third quadrant with the points in the starting point set, and points meeting the loss function are expanded and connected. That is, the starting point and the ending point in the starting point set are the starting point and the ending point of the local optimal solution, and other points meeting the condition in the global are not taken. All points in the first quadrant with the local ending point as the origin or all points in the third quadrant with the starting point as the origin are traversed, and the points in these quadrants are included in the current set one by one to check the loss function value and determine whether the condition is met.
[0075] For example, after determining the starting point set, a local pruning method is adopted. For each starting point, a sub-coordinate system is established with the starting point as the origin, and a global search is performed only on the points in the first and third quadrants of the sub-coordinate system, thereby saving the complexity of the global search. If a sample point falls in the first third quadrant, the line connecting the sample point and the starting point is an increasing function, and therefore it can be further determined whether the sample point meets the condition.
[0076] For sample points not falling in any first third quadrant of a sub-coordinate system, it is difficult to meet the condition, and therefore the above condition specific determination can not be used.
[0077] Meanwhile, whether it is the determination of the starting point or the determination of the global search, the same loss function is used.
[0078] The threshold of the number of target sample points newly joining the first set can be 1, that is, a new fitting is performed every time a point is found.
[0079] After the search is completed, a target curve can be fitted according to the first set.
[0080] In some embodiments, the unsaturation information comprises at least one of a total unsaturation number, a branch-chain unsaturation number, the condition of data fitting comprises at least one of: a goodness of fit of a curve fitted according to the sample points satisfying the condition exceeds a first threshold value; on the curve fitted according to the sample points satisfying the condition, for a first sample point and a second sample point with the same total unsaturation number, in a case where the retention time of the first sample point is greater than the retention time of the second sample point, the mass-to-charge ratio of the first sample point is greater than the mass-to-charge ratio of the second sample point; on the curve fitted according to the sample points satisfying the condition, for a third sample point and a fourth sample point with the same branch-chain unsaturation number, in a case where the retention time of the third sample point is greater than the retention time of the fourth sample point, the mass-to-charge ratio of the third sample point is greater than the mass-to-charge ratio of the fourth sample point. The goodness of fit represents the degree of fitting of the fitted curve to the observation values.
[0081] wherein, "on the curve fitted according to the sample points satisfying the condition, for a first sample point and a second sample point with the same total unsaturation number, in a case where the retention time of the first sample point is greater than the retention time of the second sample point, the mass-to-charge ratio of the first sample point is greater than the mass-to-charge ratio of the second sample point; on the curve fitted according to the sample points satisfying the condition, for a third sample point and a fourth sample point with the same branch-chain unsaturation number, in a case where the retention time of the third sample point is greater than the retention time of the fourth sample point, the mass-to-charge ratio of the third sample point is greater than the mass-to-charge ratio of the fourth sample point" is ECN and ESCN, the identification result of the sample conforming to the two rules is high in accuracy.
[0082] The total unsaturation number can be understood as the unsaturation value of the whole organic molecule, which reflects the total amount of all unsaturated structures (such as double bonds, triple bonds, rings, etc.) in the molecule.
[0083] The "branch-chain unsaturation number" is the unsaturation contributed by the part (such as the branch) other than the main chain (or the main skeleton) in the molecule.
[0084] The corresponding loss function is expressed as follows.
[0085]
[0086] wherein, I() represents an indicator function, S represents a candidate set of points satisfying the condition, n represents the number of points satisfying the condition, represents the goodness of fit of the fitted curve obtained by this fitting, mz i and mz j respectively represent the mass-to-charge ratio of the i th point and the j th point, rt i and rt j respectively represent the retention time of the i th point and the j th point.[rt 1 , rt 2represents a specified closed interval. i and j are positive integers.
[0087] loss'1 requires that the retention time and mass-to-charge ratio of the candidate set of sample points have a certain functional form, loss'2 requires that the fitting curve of the candidate set of sample points has a monotonically increasing change trend on the closed interval [rt 1 , rt 2 ], and loss'3 requires that the fitting curve of the candidate set is continuously present without jump points and discontinuous points on the closed interval [rt 1 , rt 2 ].
[0088] The sample points in the candidate set should simultaneously satisfy loss'1, loss'2, loss'3, and loss' is equal to 1. When loss' is equal to 1, the two rules of ECN or ESCN are satisfied, that is, the substances in the same lipid subclass with the same total unsaturation degree have a monotonically increasing quadratic function trend in the change of mass-to-charge ratio and retention time; the substances in the same lipid subclass with the same chain unsaturation degree have a monotonically increasing quadratic function trend in the change of mass-to-charge ratio and retention time.
[0089] When calculating the loss i of multiple points in batch form in matrix form, the following formula can be used to replace the calculation formula of the above loss i .
[0090]
[0091] In some embodiments, according to the relationship between the unsaturation degree information, the belonging subclass, the retention time and the mass-to-charge ratio, a starting point that satisfies the condition is selected from the sample points to join the first set to obtain a preliminary fitting curve, including: according to the relationship between the unsaturation degree information, the belonging subclass, the retention time and the mass-to-charge ratio, the measured mass-to-charge ratio of the sample point and the data fitting obtained mass-to-charge ratio, a starting point that satisfies the condition is selected from the sample points to join the first set to obtain a preliminary fitting curve.
[0092] Wherein, the measured mass-to-charge ratio of the sample point is used to plot the sample point in the coordinate system and data fitting. The "data fitting obtained mass-to-charge ratio" refers to the estimated value of the mass-to-charge ratio calculated by substituting the retention time rt i of the sample point into the fitting obtained curve, that is, the estimated value obtained after the fitting curve.
[0093] Taking a set of lipid samples of a certain total unsaturation degree number of a certain subclass as an example, traversing the set, according to the following loss function, a starting point that satisfies the condition in the set is found.
[0094]
[0095]
[0096] wherein, represents the estimated value of the mass-to-charge ratio of the i-th point by fitting a function, and the remaining parameters are defined as before.
[0097] The starting points can be selected according to the loss i For example, the points with loss i less than a threshold k are selected as the starting points, or the N points with the smallest loss i are selected as the starting points.
[0098] The starting points are combinations of multiple points obtained by greedy enumeration, for example, 3-5 points, and the fitting curve is for example a polynomial fitting.
[0099] After obtaining the starting points, a three-quadrant global search is performed according to the starting points, as described above. After ending the global search, all points in the first set are fitted, and the loss function loss i is used to calculate the loss value of each point.
[0100] In the process of fitting, the sample set can be divided according to two division methods and fitted respectively, and the two division methods include:
[0101] The samples with the same sub-class and the same total unsaturation number are divided into a set;
[0102] The samples with the same sub-class, the same total unsaturation number, and the same unsaturation distribution are divided into a set.
[0103] For the above two division methods, the scores after fitting can be calculated respectively, that is, a sample can have two scores corresponding to the above two division methods.
[0104] First, the way of dividing the samples with the same sub-class and the same total unsaturation number into a set is introduced.
[0105] In some embodiments, the unsaturation information includes a total unsaturation number, and the data fitting of the plurality of lipid samples based on the unsaturation information and the sub-class to which the lipid sample belongs, with the retention time and the mass-to-charge ratio of each lipid sample as variables, includes: dividing the plurality of lipid samples according to the sub-class and the total unsaturation number to obtain a plurality of second sets, wherein the lipid samples in each second set of the plurality of second sets belong to the same sub-class and have the same total unsaturation number; and for each second set, performing data fitting with the retention time and the mass-to-charge ratio of the lipid samples in the second set as variables to obtain a first fitting curve corresponding to the sub-class and the total unsaturation number of the second set.
[0106] For example, the lipid samples of the same subclass and with the same total unsaturation number belong to one set, and the lipid samples of different sets are fitted respectively. That is, if there are lipid samples of multiple subclasses or multiple total unsaturation numbers, multiple curves can be fitted.
[0107] In some embodiments, for each second set, the data fitting is performed with the retention time and mass-to-charge ratio of the lipid samples therein as variables to obtain a first fitting curve corresponding to the subclass and total unsaturation number of the second set, including: in a case where the number of lipid samples of a specified subclass at a first total unsaturation number is less than a threshold value, determining the first fitting curve corresponding to the first total unsaturation number of the specified subclass according to the first fitting curve of the corresponding second set at a second total unsaturation number of the specified subclass, wherein the first total unsaturation number and the second total unsaturation number are different.
[0108] For example, sometimes there are too few optional sample points of a certain total unsaturation number of a certain subclass, which cannot be fitted, and the analytical equation thereof can be inversely solved from the fitting results of other unsaturation degrees belonging to the same subclass.
[0109] In some embodiments, the determining the first fitting curve corresponding to the first total unsaturation number of the specified subclass according to the first fitting curve of the corresponding second set at the second total unsaturation number of the specified subclass in the case where the number of lipid samples of the specified subclass at the first total unsaturation number is less than the threshold value includes: determining, according to the first fitting curve of the second total unsaturation number of the specified subclass, a curve of the lipid samples at the first total unsaturation number of the specified subclass that does not intersect the first fitting curve of the corresponding second set at the second total unsaturation number of the specified subclass as the first fitting curve corresponding to the first total unsaturation number of the specified subclass.
[0110] For example, the search iteration process is skipped by the IUP rule. The IUP requires that the fitted quadratic curves of different total unsaturation degrees within the same lipid subclass do not intersect in the closed interval. In the case where there is only one sample point (rt i , mz i ) of a certain total unsaturation degree, the analytical formula thereof is inversely solved as follows:
[0111]
[0112] wherein g(x) is the fitting equation of other unsaturation degrees of the subclass, f(x) represents the functional relationship between the mass-to-charge ratio and the retention time, that is, the analytical result to be solved, that is, the change trend of the point with too few optional points leading to the inability to perform fitting. f(x) should pass through or infinitely approach the point and be a quotient of g(x). q(x) can be a constant division term, for example, a constant a. C is a constant term.
[0113] The following describes a way of dividing samples of the same subclass and having the same total unsaturation number and the same unsaturation distribution into one set.
[0114] In some embodiments, the unsaturation information further comprises an unsaturation distribution, and the data fitting based on the unsaturation information and the subclass for each lipid sample with the retention time and the mass-to-charge ratio as variables comprises: dividing the plurality of lipid samples according to the subclass, the total unsaturation number and the unsaturation distribution to obtain a plurality of third sets, wherein the lipid samples in each third set of the plurality of third sets belong to the same subclass and have the same total unsaturation number and the same unsaturation distribution; and for each third set, performing data fitting with the retention time and the mass-to-charge ratio of the lipid samples therein as variables to obtain a second fitting curve corresponding to the subclass, the total unsaturation number and the unsaturation distribution of the third set.
[0115] The method of obtaining the first fitting curve and the second fitting curve is similar, and the main difference lies in the set division manner.
[0116] In some embodiments, the accuracy of the identification result of each lipid sample is determined according to the result of the data fitting, comprising: determining the accuracy of the identification result of each lipid sample according to whether the sample point is in the first set and the measured mass-to-charge ratio of the sample point and the mass-to-charge ratio obtained by data fitting.
[0117] For example, the accuracy of the identification result of the sample point in the first set is greater than that of the sample point not in the first set. The closer the measured mass-to-charge ratio of the sample point is to the mass-to-charge ratio obtained by data fitting, the higher the accuracy is.
[0118] In some embodiments, the accuracy of the identification result of each lipid sample is determined according to whether the sample point is in the first set and the measured mass-to-charge ratio of the sample point and the mass-to-charge ratio obtained by data fitting, comprising: determining the confidence interval of the target curve; and for the sample point not in the first set, determining the accuracy of the identification result of the corresponding lipid sample according to whether the sample point is in the confidence interval. For example, the points falling within the 95% confidence interval of the fitting curve are retrieved. Although the accuracy of the sample point not in the first set is less than that of the sample in the first set, the identification result of the sample falling within the 95% confidence interval of the fitting curve also has certain reference value, and therefore, these points are retrieved and the accuracy thereof is calculated.
[0119] In some embodiments, for sample points not in the first set, determining the accuracy of the identification result of the corresponding lipid sample based on whether the sample point is in the confidence interval includes: determining the accuracy of the identification result of the lipid sample corresponding to the sample point not in the first set but in the confidence interval as a first value; and determining the accuracy of the identification result of the lipid sample corresponding to the sample point not in the confidence interval and not in the first set as a second value, wherein the first value is greater than the second value.
[0120] For example, the accuracy of the identification results for sample points within the confidence interval is less than the accuracy of sample points within the first set, but greater than the accuracy of sample points that are neither in the confidence interval nor in the first set.
[0121] In some embodiments, determining the accuracy of the identification result for each lipid sample based on whether the sample point is in a first set and the mass-to-charge ratio measured by the sample point and the mass-to-charge ratio obtained by data fitting includes: determining the loss value of each sample point in the first set based on the measured mass-to-charge ratio and the mass-to-charge ratio obtained by data fitting; determining the maximum loss value of the multiple sample points in the first set as the first loss value; determining the minimum loss value of the multiple sample points in the first set as the second loss value; and for each sample point in the first set, determining the accuracy of the identification result of the lipid sample corresponding to that sample point based on the loss value, the first loss value, and the second loss value.
[0122] For example, building PHS i The function standardizes the results, limiting the score to between 0 and 1, while ensuring that lower loss function values result in higher scores. (PHS) i The definition is as follows.
[0123]
[0124] in:
[0125] loss max =max(loss p p = 1, 2, 3...
[0126] loss min =min(loss p p = 1, 2, 3...
[0127] p represents the index of the sample in the first set. The lipidomic intelligent model uses lipid subclasses and different unsaturation compositions within subclasses as classification units, and performs the above search and calculate the score PHS in both types of datasets. i PHS i The value of i can be 1 or 2, representing the scores obtained by fitting two different classification methods.
[0128] In some embodiments, according to the result of data fitting, the accuracy of the identification result of each lipid sample is determined, including: determining a first accuracy according to the first fitting curve corresponding to each lipid sample; determining a second accuracy according to the second fitting curve corresponding to each lipid sample; and determining the accuracy of the identification result according to the first accuracy and the second accuracy.
[0129] For example, the first accuracy PHS1 and the second accuracy PHS2 are weightedly averaged to obtain the accuracy PHS of the identification result. z .
[0130] Only one of PHS1 and PHS2 can also be selected as the accuracy PHS of the identification result. z .
[0131] According to PHS z , the accuracy of the identification result can be calculated. For example, PHS z , Score matched , Score ratio are weightedly summed to obtain a final accuracy score, final_score.
[0132] By comprehensively considering Score matched , Score ratio , and PHS z , the coverage and accuracy of the identification are improved.
[0133] At the same time, based on the search model described above, the judgment rule is self-learned from the retention time distribution of the data, and the dependence on the platform is small.
[0134] Figure 2 A schematic diagram of lipid identification according to some embodiments of the present disclosure is shown.
[0135] As shown in Figure 2 , some embodiments of the present disclosure mainly relate to constructing a hierarchical spectrum library, fast matching, intelligent search of lipid groups, etc. Hash table is a common data structure that realizes fast data storage and retrieval by mapping key-value pairs to specific positions in an array, and has the advantages of efficient lookup, saving computing space, supporting out-of-order access, and good dynamic expandability.
[0136] The group intelligence model framework, i.e., the machine learning method based on cooperation of multiple agents introduced above, can learn the intelligent cooperation behavior of sub-agents, and has the advantages of robustness, expandability, and strong interpretability.
[0137] The identification method of some embodiments of the present disclosure can be applied to process data generated by a mass spectrometer, for example, data generated by a mass spectrometer such as Thermo Scientific Orbitrap Exploris 240 MS system, Thermo Scientific Orbitrap Exploris 120 MS system, Agilent 6546 Q-TOF MS system, Xevo G2-XS Q-TOF MS system, SCIEX ZenoTOF 7600 MS system, etc. can be identified using the method mentioned in the present disclosure.
[0138] To verify the effectiveness of the identification method of some embodiments of the present disclosure, the inventors removed the results with an accuracy score final_score less than 2 by judging the dynamic retention time, and obtained the top 500 credible results. And through the review of four lipid identification experts, the false discovery rate of the identification results was checked, and the results are shown in the following table.
[0139] Table 1
[0140]
[0141] Among them, the false discovery rate = the number of false identifications / total number of identifications.
[0142] Taking the value "78" in the second row and the second column as an example, the number of samples with a score threshold ≥ 2.8 is 78, which is consistent with the correct identification result in the next row of the column, and the false discovery rate is 0%. That is, the identification results of samples with a score threshold ≥ 2.8 are all correct, which shows that the accuracy of the identification results calculated according to the identification method of some embodiments of the present disclosure can be used to judge whether the identification is correct or not.
[0143] Figure 3 A block diagram of a lipid identification device according to some embodiments of the present disclosure is shown.
[0144] As shown in Figure 3 , the lipid identification device includes an acquisition module 31, an information determination module 32, a fitting module 33, and an accuracy determination module 34.
[0145] The acquisition module 31 is configured to acquire the identification result of each lipid sample of a plurality of lipid samples, for example, by performing step S1 as shown in Figure 1 .
[0146] The information determination module 32 is configured to determine the unsaturation information and the sub-class of each lipid sample according to the identification result of each lipid sample, for example, by performing step S2 as shown in Figure 1 .
[0147] The fitting module 33 is configured to perform data fitting on the plurality of lipid samples based on the unsaturation information and the belonging subclass with the retention time and the mass-to-charge ratio of each lipid sample as variables, for example, to perform step S3 as shown in Figure 1
[0148] The accuracy determining module 34 is configured to determine the accuracy of the identification result of each lipid sample according to the result of the data fitting, for example, to perform step S4 as shown in Figure 1
[0149] Figure 4 A block diagram of a lipid identification apparatus according to some other embodiments of the present disclosure is shown.
[0150] As shown in Figure 4 The lipid identification apparatus 4 comprises a memory 41 and a processor 42 coupled to the memory 41, the memory 41 being configured to store instructions for performing the lipid identification method. The processor 42 is configured to perform the lipid identification method in any of the embodiments of the present disclosure based on the instructions stored in the memory 41.
[0151] Figure 5 A block diagram of a computer system for implementing some embodiments of the present disclosure is shown.
[0152] As shown in Figure 5 The computer system 50 can be in the form of a general purpose computing device. The computer system 50 comprises a memory 510, a processor 520 and a bus 500 connecting different system components.
[0153] The memory 510 can comprise, for example, system memory, non-volatile storage media and the like. The system memory, for example, stores an operating system, application programs, a Boot Loader and other programs and the like. The system memory can comprise volatile storage media, such as random access memory (RAM) and / or cache memory. The non-volatile storage media, for example, stores instructions for performing the lipid identification method in any of the embodiments of the present disclosure. The non-volatile storage media includes, but is not limited to, magnetic storage media, optical storage media, flash memory and the like.
[0154] The processor 520 can be implemented in the form of a general purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic, and the like discrete hardware components. Accordingly, each module such as the determining module and the determining module can be implemented by a central processing unit (CPU) running instructions stored in the memory for performing the corresponding steps, or by a dedicated circuit for performing the corresponding steps.
[0155] Bus 500 can use any of the various bus architectures. For example, bus architectures include, but are not limited to, Industry Standard Architecture (ISA) bus, Micro Channel Architecture (MCA) bus, and Peripheral Component Interconnect (PCI) bus.
[0156] The computer system 50 may also include an input / output interface 530, a network interface 540, and a storage interface 550. These interfaces 530, 540, and 550, as well as the memory 510 and processor 520, can be connected via a bus 500. The input / output interface 530 provides a connection interface for input / output devices such as a monitor, mouse, and keyboard. The network interface 540 provides a connection interface for various networked devices. The storage interface 550 provides a connection interface for external storage devices such as floppy disks, USB flash drives, and SD cards.
[0157] Various aspects of this disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus, and computer program products according to embodiments of this disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations thereof, can be implemented by computer-readable program instructions.
[0158] These computer-readable program instructions are provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable device to produce a machine, such that execution of the instructions by the processor produces means for implementing the functions specified in one or more boxes of the flowchart and / or block diagram.
[0159] These computer-readable program instructions are also readablely stored in a computer-readable storage medium. These instructions cause a computer to work in a particular manner to produce an article of manufacture, including instructions that implement the functions specified in one or more boxes in a flowchart and / or block diagram.
[0160] This disclosure also provides a lipid identification system, including: a lipid identification device according to any of the embodiments of this disclosure.
[0161] Figures 6a-6c A user interface for implementing a lipid identification system according to some embodiments of the present disclosure is shown.
[0162] The following is combined with Figures 6a-6c This paper describes a process for identifying lipids using an identification system (hereinafter referred to as LipidIN) according to some embodiments of the present disclosure.
[0163] To avoid feature redundancy caused by similarity algorithms, LipidIN first constructs a "fragment tree" hierarchical spectral library of 166.3 million lipids based on lipid feature spectra.
[0164] Users convert the mass spectrometer's output files to mzML format using MSConvert software, then upload the file. LipidIN converts the mzML format and extracts the spectral information. The identification system in some embodiments of this disclosure uses standard mzML format as input and has high versatility.
[0165] Then, primary and secondary spectrum information extraction is performed. Open the website, go to LipidIN, click the "analysis" option, left-click, enter the parameter name "ESI" under "param key," and enter the corresponding mode "p" under "param value." ESI can input three modes (p represents positive ion mode, n1 represents formate mobile phase, and n2 represents acetate mobile phase). In this example, "p" is entered. Continue left-clicking, enter the parameter name "MS2" under "param key," and enter the corresponding parameter value "0.01" under "param value," representing filtering secondary spectrum peaks whose response intensity is less than 1% of the highest peak. Left-click "upload," select "data.zip," and upload the data. Wait for the data upload to complete, left-click "working," wait for the operation to finish, left-click "operation records," find "result.zip," and left-click to download the results. A total of 5601 valid primary and secondary spectrum ion pairs were extracted from the example file.
[0166] Then, perform a fast matching. Based on the extracted spectral information, continue processing the data and name the extracted spectral information result data.zip. Left-click on "upload" to select the upload folder, then left-click again to add a parameter. Enter the parameter name "ESI" under "param key" and the corresponding mode "p" under "param value". Left-click again, enter the parameter name "PPM1" under "param key" and the corresponding parameter value "10" under "param value", representing setting the charge-to-mass ratio of the first-level spectrum to allow an error of ±10 ppm (threshold). ppm Continue left-clicking, enter the parameter name PPM2 under "paramkey", and enter the corresponding parameter value 30 under the corresponding "param value". This represents setting the charge-to-mass ratio of the secondary spectrum to allow an error of ±30 ppm (threshold). ppm After the data upload is complete, left-click on "working" and wait for the process to finish. Left-click on "operation records" to find "result.zip" and left-click to download the results. In this example, 676 peaks were effectively identified, resulting in 2028 candidate identification results, with matching taking 0.04 seconds.
[0167] The results of spectrum information extraction and fast matching are placed under the data folder, compressed into data.zip, then left-click upload, select data.zip to upload, and wait for the upload to complete. Left-click working, wait for the running to end, left-click operation records to find result.zip, and left-click download to get the running results. The lipid group intelligent model generates an accuracy score for the identification results by judging the dynamic retention time.
[0168] The 1.66 billion fraction library of the above identification system can be compatible with multiple mass spectrometry identification software, and the identification framework is also applicable to various mass spectrometry files in mzML format and MSP format spectra.
[0169] The present disclosure optimizes the selection of multiple parameters, and only three parameters (primary mass-to-charge ratio error threshold, secondary mass-to-charge ratio error threshold, ionization mode) can complete the identification process, reducing the influence of multiple parameters on the identification results and the complexity of the operation.
[0170] The present disclosure can adopt a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects.
[0171] Those skilled in the art can understand that in the description of the steps of the above-mentioned embodiments, the order of the actual execution of the steps is not strictly limited, and the order of the steps in the actual operation can be adjusted according to the functional requirements and the inherent logical relationship.
[0172] In the above description of various embodiments, emphasis is placed on highlighting the differences between different embodiments, and the same points or similarities of various embodiments can be understood by mutual reference. In order to keep it simple, repeated descriptions are not made.
[0173] Through the lipid identification method, device and system, and computer readable storage medium in the above embodiments, the accuracy of lipid identification is improved.
[0174] So far, the lipid identification method, device and system, and computer readable storage medium according to the present disclosure have been described in detail. In order to avoid obscuring the concept of the present disclosure, some details known in the art are not described. Those skilled in the art can fully understand how to implement the technical solutions disclosed herein according to the above description.
Claims
1. A method for identifying lipids, comprising: obtaining identification results of each of a plurality of lipid samples; determining, according to the identification result of each lipid sample, unsaturation information and a sub-class of each lipid sample; performing data fitting on the plurality of lipid samples based on the unsaturation information and the sub-class, with retention time and mass-to-charge ratio of each lipid sample as variables, comprising: mapping each lipid sample into a rectangular coordinate system with retention time and mass-to-charge ratio as coordinates, wherein each lipid sample is represented by a sample point in the coordinate system; selecting a starting point from the sample points that meets a condition according to a relationship among the unsaturation information, the sub-class, the retention time and the mass-to-charge ratio, and adding the starting point to a first set to obtain a preliminary fitting curve, wherein the unsaturation information comprises at least one of total unsaturation number and branch-chain unsaturation number; constructing a sub-coordinate system with the starting point as the origin, and searching for a target sample point that meets the condition in the first and third quadrants of the sub-coordinate system corresponding to the starting point, and adding the target sample point to the first set; each time the number of target sample points added to the first set reaches a threshold, re-determining the preliminary fitting curve according to the first set; fitting a target curve according to the first set; determining accuracy of the identification result of each lipid sample according to a result of the data fitting, wherein the condition comprises at least one of the following: a goodness of fit of a curve fitted according to the sample points that meet the condition exceeds a first threshold; on the curve fitted according to the sample points that meet the condition, for a first sample point and a second sample point with the same total unsaturation number, in a case where the retention time of the first sample point is greater than the retention time of the second sample point, the mass-to-charge ratio of the first sample point is greater than the mass-to-charge ratio of the second sample point; on the curve fitted according to the sample points that meet the condition, for a third sample point and a fourth sample point with the same branch-chain unsaturation number, in a case where the retention time of the third sample point is greater than the retention time of the fourth sample point, the mass-to-charge ratio of the third sample point is greater than the mass-to-charge ratio of the fourth sample point.
2. The method of lipid identification of claim 1, wherein, the selecting a starting point from the sample points that meets a condition according to a relationship among the unsaturation information, the sub-class, the retention time and the mass-to-charge ratio, and adding the starting point to a first set to obtain a preliminary fitting curve, comprises: the selecting a starting point from the sample points that meets a condition according to a relationship among the unsaturation information, the sub-class, the retention time and the mass-to-charge ratio, and adding the starting point to a first set to obtain a preliminary fitting curve, comprises:
3. The method of lipid identification according to claim 1, wherein, the determining accuracy of the identification result of each lipid sample according to a result of the data fitting, comprises: the determining accuracy of the identification result of each lipid sample according to a result of the data fitting, comprises:
4. The method of lipid identification according to claim 3, wherein, the determining accuracy of the identification result of each lipid sample according to a result of the data fitting, comprises: determining a loss value of each sample point in the first set according to the measured mass-to-charge ratio and the data-fitted mass-to-charge ratio; determine the maximum loss value of the plurality of sample points in the first set as a first loss value; determine the minimum loss value of the plurality of sample points in the first set as a second loss value; for each sample point in the first set, determine the accuracy of the identification result of the lipid sample corresponding to the sample point according to the loss value of the sample point, the first loss value, and the second loss value.
5. The method of lipid identification according to claim 3, wherein, The accuracy of the identification result of each lipid sample is determined according to whether the sample point is in the first set, and the mass-to-charge ratio measured by the sample point and the mass-to-charge ratio fitted by the data, comprising: determining the confidence interval of the target curve; for the sample points not in the first set, determining the accuracy of the identification result of the corresponding lipid sample according to whether the sample point is in the confidence interval.
6. The method of lipid identification according to claim 5, wherein, The accuracy of the identification result of the corresponding lipid sample is determined according to whether the sample point is in the confidence interval for the sample points not in the first set. The accuracy of the identification result of the lipid sample corresponding to the sample point not in the first set and in the confidence interval is determined as a first value. The accuracy of the identification result of the lipid sample corresponding to the sample point not in the confidence interval and the first set is determined as a second value, wherein the first value is greater than the second value.
7. The method of lipid identification according to claim 1, wherein, The unsaturation information includes the total number of unsaturation degrees, and the plurality of lipid samples are fitted based on the unsaturation information and the subcategory with the retention time and mass-to-charge ratio of each lipid sample as variables, comprising: According to the subcategory and the total number of unsaturation degrees, the plurality of lipid samples are divided to obtain a plurality of second sets, wherein the lipid samples in each second set belong to the same subcategory and have the same total number of unsaturation degrees; for each second set, data fitting is performed with the retention time and mass-to-charge ratio of the lipid samples therein as variables to obtain a first fitting curve corresponding to the subcategory and total number of unsaturation degrees of the second set.
8. The method of lipid identification of claim 7, wherein, The accuracy of the identification result of the corresponding lipid sample is determined according to whether the sample point is in the confidence interval for the sample points not in the first set. In the case that the number of lipid samples of the first total unsaturation degree of the specified subcategory is less than the threshold value, the first fitting curve corresponding to the first total unsaturation degree of the specified subcategory is determined according to the first fitting curve of the second set corresponding to the second total unsaturation degree of the specified subcategory, wherein the first total unsaturation degree and the second total unsaturation degree are different.
9. The method of lipid identification according to claim 8, wherein, In the case that the number of lipid samples of the first total unsaturation degree of the specified subcategory is less than the threshold value, the first fitting curve corresponding to the first total unsaturation degree of the specified subcategory is determined according to the first fitting curve of the second set corresponding to the second total unsaturation degree of the specified subcategory, wherein the first total unsaturation degree and the second total unsaturation degree are different. According to the first fitting curve corresponding to the second total unsaturation degree of the specified subcategory, the curve intersecting the first fitting curve of the second set corresponding to the second total unsaturation degree of the specified subcategory and passing through the lipid samples of the first total unsaturation degree of the specified subcategory is determined as the first fitting curve corresponding to the first total unsaturation degree of the specified subcategory.
10. The method of lipid identification of claim 8, wherein, The unsaturation information further comprises an unsaturation distribution, and the data fitting of the plurality of lipid samples based on the unsaturation information and the subcategory with the retention time and the mass-to-charge ratio of each lipid sample as variables comprises: dividing the plurality of lipid samples according to the subcategory, the total number of unsaturations and the unsaturation distribution, to obtain a plurality of third sets, wherein the lipid samples in each third set belong to the same subcategory and have the same total number of unsaturations and the same unsaturation distribution; for each third set, performing data fitting with the retention time and the mass-to-charge ratio of the lipid samples in the third set as variables to obtain a second fitting curve corresponding to the subcategory, the total number of unsaturations and the unsaturation distribution of the third set.
11. The method of lipid identification of claim 10, wherein, The accuracy of the identification result of each lipid sample is determined according to the result of the data fitting, comprising: determining a first accuracy according to the first fitting curve corresponding to each lipid sample; determining a second accuracy according to the second fitting curve corresponding to each lipid sample; determining the accuracy of the identification result according to the first accuracy and the second accuracy.
12. The method of lipid identification according to claim 1, wherein, The identification result of each lipid sample of the plurality of lipid samples is obtained, comprising: identifying the type of each lipid sample by comparing the spectrum of each lipid sample with a reference spectrum as the identification result of each lipid sample.
13. The method of lipid identification according to claim 12, wherein, The type of each lipid sample is identified by comparing the spectrum of each lipid sample with a reference spectrum as the identification result of each lipid sample, comprising: calculating a first score of each lipid sample according to the number of matching characteristic peaks of the spectrum of each lipid sample and the reference spectrum; calculating a second score of each lipid sample according to the importance of the matching characteristic peaks of the spectrum of each lipid sample and the reference spectrum; determining the reference spectrum matching the spectrum of each lipid sample according to the first score and the second score of each lipid sample; determining the type of each lipid sample according to the matching reference spectrum as the identification result of each lipid sample.
14. A lipid identification device, comprising: an acquisition module configured to obtain an identification result of each lipid sample of a plurality of lipid samples; an information determination module configured to determine unsaturation information and a subcategory of each lipid sample according to the identification result of each lipid sample; a fitting module configured to perform data fitting of the plurality of lipid samples based on the unsaturation information and the subcategory with the retention time and the mass-to-charge ratio of each lipid sample as variables, comprising: mapping each lipid sample to a rectangular coordinate system with the retention time and the mass-to-charge ratio as coordinates, wherein each lipid sample is represented by a sample point in the coordinate system; selecting a starting point meeting the conditions from the sample points to join a first set to obtain a preliminary fitting curve according to the relationship among the unsaturation information, the subcategory, the retention time and the mass-to-charge ratio, wherein the unsaturation information comprises at least one of the total number of unsaturations and the number of chain unsaturations; constructing a sub-coordinate system with the starting point as the origin, and searching for a target sample point meeting the conditions in the first and third quadrants of the sub-coordinate system corresponding to the starting point to join the first set; re-determine the preliminary fitting curve according to the first set, whenever the number of target sample points newly added to the first set reaches a threshold value; fit the target curve according to the first set; an accuracy determination module configured to determine the accuracy of the identification result of each lipid sample according to the result of the data fitting, wherein the condition comprises at least one of the following: a goodness of fit of the curve fitted according to the sample points satisfying the condition exceeds a first threshold value; on the curve fitted according to the sample points satisfying the condition, for a first sample point and a second sample point with the same total unsaturation number, in a case where the retention time of the first sample point is greater than the retention time of the second sample point, the mass-to-charge ratio of the first sample point is greater than the mass-to-charge ratio of the second sample point; on the curve fitted according to the sample points satisfying the condition, for a third sample point and a fourth sample point with the same branch unsaturation number, in a case where the retention time of the third sample point is greater than the retention time of the fourth sample point, the mass-to-charge ratio of the third sample point is greater than the mass-to-charge ratio of the fourth sample point.
15. A lipid identification device, comprising: a memory; and a processor coupled to the memory, the processor being configured to perform the lipid identification method according to any one of claims 1 to 13 based on instructions stored in the memory.
16. A lipid identification system, comprising: the lipid identification device according to claim 14 or 15.
17. A computer readable storage medium having stored thereon computer program instructions, which instructions, when executed by a processor, implement the lipid identification method according to any one of claims 1 to 13.
18. A computer program product comprising computer program instructions, which instructions, when executed by a processor, implement the lipid identification method according to any one of claims 1 to 13.
Citation Information
Patent Citations
Method for identifying lipid component structure in lipidomics based on CID fragmentation
CN107228908A
Methods for identifying fungi
US20160215322A1