Method for identifying fragrance base formula based on lc-ms / ms data of natural fragrance raw materials
By analyzing LC-MS/MS data, we constructed characteristic spectra of fragrance monomers and a data matrix of the least dissimilar monomers in descending order, which solved the problem of accurate identification of complex natural fragrance raw materials and compound fragrance bases, and achieved effective identification of the composition of fragrance monomers.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- CHINA TOBACCO HUNAN IND CORP
- Filing Date
- 2025-06-06
- Publication Date
- 2026-07-10
AI Technical Summary
Existing technologies make it difficult to accurately identify complex natural fragrance raw materials or their compound formulations, leading to difficulties in identifying the formulations after the compounding of fragrance monomers.
Using a method based on LC-MS/MS data, the probability of flavoring monomers in the test sample is calculated by constructing the characteristic spectra of flavoring monomers and the data matrix of the least dissimilar monomers in descending order, and the composition of flavoring monomers is identified.
It enables precise identification of complex natural fragrance raw materials and compound fragrance bases, effectively identifies the composition of fragrance monomers, and simplifies the identification process of fragrance monomer compounding.
Smart Images

Figure CN120636584B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of fragrance raw material formulation analysis technology, and in particular to a method for identifying fragrance base formulations based on LC-MS / MS data of natural fragrance raw materials. Background Technology
[0002] Flavorings and fragrances in cigarette products play a role in modifying or masking defects in tobacco leaves, compensating for and enhancing the physicochemical properties of cigarettes, smoothing the aroma, reducing the irritation of smoke, making the smoke more pleasant, improving the taste, adding a pleasing tobacco aroma, and suppressing pungent irritation and unpleasant odors, thereby satisfying various consumer needs. Flavoring raw materials refer to water-soluble additives applied to tobacco sheets. They are generally composed of sugars, natural fragrances, humectants, combustion aids, and mildew inhibitors. The purpose of adding these materials is to reduce the irritation of smoke, make the smoke more pleasant, improve the taste, enhance the aroma of cigarettes, and simultaneously enhance their physical properties. Flavoring raw materials are mostly compounded from plant-based natural fragrances. Natural fragrances mainly use the flowers, branches, leaves, herbs, roots, bark, stems, seeds, or fruits of aromatic plants as raw materials, and are produced through different processes as essential oils, extracts, tinctures, balsams, and resin extracts and refined products. In the extraction of natural fragrances, water, ethanol, and propylene glycol are commonly used as extraction media. While large molecules such as cellulose, lignin, polypeptides, and proteins are difficult to extract, smaller molecules such as sugars, organic acids, amino acids, alkaloids, polyphenols, sterols, pigments, vitamins, lipids, minerals, alcohols, aldehydes, and ketones are also extracted. This results in highly complex compositions of fragrance monomers and functional flavor bases blended from monomers, forming a typical complex system. This makes the analysis and characterization of natural fragrance raw materials and the identification of monomer formulations in blended functional flavor bases extremely challenging. Developing analytical characterization techniques for natural fragrance raw materials and using feature analysis of characterization databases to achieve accurate aroma identification of blended flavor bases is of significant value for the imitation and substitution of functional flavor bases, intelligent fragrance creation, and digital tobacco flavoring design.
[0003] Currently, the industry primarily relies on traditional separation and classification techniques for the analysis and characterization of natural flavoring raw materials used in tobacco. Volatile and semi-volatile components are extracted with organic solvents and then analyzed by GC-MS. Water-soluble sugars are derivatized using a two-step silanization process involving hydroxylamine hydrochloride, anhydrous pyridine, N,O-bis(trimethylsilyl)trifluoroacetamide, and trimethylchlorosilane, followed by GC-MS analysis. Water-soluble organic acids are derivatized by methyl esterification and then analyzed by GC-MS. Furthermore, phenolic components in flavoring raw materials can be characterized using LC-DAD, taking advantage of their UV absorption properties. However, these methods for analyzing and characterizing sugars, acids, phenols, and volatile / semi-volatile components in natural flavoring raw materials do not provide sufficient information on the sample composition, resulting in unclear characteristics and hindering accurate aroma identification and formulation (flavor base formulation) recognition after blending individual flavoring monomers. Therefore, developing a new method or algorithm for the accurate identification of complex natural flavoring raw materials or their blends is particularly necessary. Summary of the Invention
[0004] In view of this, the technical problem to be solved by the present invention is to provide a method for identifying fragrance base formulations based on LC-MS / MS data of natural fragrance raw materials.
[0005] This invention provides a method for identifying the composition of fragrance monomers in a fragrance monomer complex sample based on LC-MS / MS data, comprising the following steps:
[0006] Step 1: Obtain the LC-MS / MS data matrix X of N flavor monomers in the flavor monomer composite sample; where N is the number of flavor monomers used in the flavor monomer composite sample; the i-th row X(i,:) of X represents the LC-MS / MS data of the flavor monomer in the i-th (i=1,2,…,N) flavor monomer; the j-th element X(i,j) of X(i,:) is the mass spectrum peak signal intensity of the j-th chemical component in the i-th flavor monomer.
[0007] Step 2: Based on the data matrix X, set p to 0.6~1.0, find the differential chemical components of each flavoring monomer relative to other flavoring monomers, and obtain the characteristic spectrum of each flavoring monomer;
[0008] Step 3: Based on the characteristic spectra of each flavoring monomer, construct a data matrix of the least dissimilar monomers of each flavoring monomer in descending order;
[0009] Step 4: Based on the data matrix of the least dissimilar monomers of each flavoring monomer arranged in descending order, calculate the probability of each flavoring monomer existing in the flavoring monomer composite sample to be tested, and obtain the flavoring monomer composition of the flavoring monomer composite sample to be tested.
[0010] The mass spectrum peak signal intensity includes: mass spectrum peak area or mass spectrum peak height;
[0011] The fragrance monomers are extracts or refined products of plants or their organs or tissues.
[0012] In step 2 of the method described in this invention, the value of parameter p should be such that the monomer characteristic spectrum The number of elements is in the range of 3×N to 5×N. For the embodiments described in this application, the parameter is set to p = 0.8. It is worth noting that as long as the value of parameter p makes the monomer characteristic spectrum... The number of elements is in the range of 3×N to 5×N, and the change in its value has no significant impact on the final formula identification result. Experiments showed that when the p-value is between 0.6 and 1.0, there is no significant difference between the formula identification results of the current embodiment data.
[0013] In step 2, the characteristic spectrum of the i-th flavoring monomer (i = 1, 2, ..., N) in each flavoring monomer is obtained. Includes the following steps:
[0014] Step a: Standardize the mass spectrometry peak signal intensity of a certain flavoring monomer i (i = 1, 2, ..., N) from among the N flavoring monomers that may be used to prepare the flavoring monomer composite sample, to obtain X. norm (i,:);
[0015] Step b, select X norm The mass spectrum peak signal intensity in (i,:) is greater than p×median(X). norm The chemical composition of (i,:) is used to form the main component sequence vector MainCompSeqNum, with the corresponding sequence numbers of the chemical components as the main component sequence number vectors. i,1 ;where median(X) norm (i,:) is X norm The median value of (i,:);
[0016] Step c: Calculate the sum of the LC-MS / MS data vectors of all flavor monomers except the i-th flavor monomer among the N flavor monomers, and normalize them to obtain x. sum,norm ;
[0017] Step d: Select x sum,norm The intensity of the medium mass spectrum peak signal is greater than p×median(x) sum,norm The chemical composition of the component is used to construct a main component sequence vector MainCompSeqNum, which is composed of the sequence numbers corresponding to the chemical components. i,2 ;
[0018] Step e: Set MainCompSeqNum i,1and MainCompSeqNum i,2 The chemical components commonly contained in MainCompSeqNum are numbered from MainCompSeqNum i,1 After removing the sub-components, the remaining chemical component indices form the characteristic component indices vector CharacCompSeqNum for the i-th monomer. i The characteristic spectrum of the i-th fragrance monomer is then... The index of X(i,:) is equal to CharacCompSeqNum i The mass spectrometry peak signal composition of the chemical composition of the elements, that is,
[0019] Furthermore, in step 2a, the standardization method is maximum value normalization; the formula for maximum value normalization is: X notm (i,:) = X(i,:) / max(X(i,:));
[0020] In step 2c, x sum,norm The formula for calculating x is: sum,norm =x sum / max(x sum ),in
[0021] In step 3 of the method described in this invention, the data matrix of the least dissimilar monomers relative to other monomers is constructed by calculating Euclidean distance, Jaccard distance, Mahalanobis distance, Minkowski distance, or Chebyshev distance. In a specific embodiment of this invention, Euclidean distance is used.
[0022] Before calculating the Euclidean distance, the process also includes normalizing the characteristic spectral data of each spice monomer.
[0023] The principle for constructing the data matrix of the least dissimilar monomers relative to other monomers in descending order is: the least dissimilar monomers are those with the largest Euclidean distance.
[0024] Furthermore, in step 3 of this invention, a DantiSortingMatrix is constructed, which is a descending-order sorted data matrix of the least dissimilar monomers of the i-th flavoring monomer in each flavoring monomer. i Includes the following steps:
[0025] A) To and X(k,CharacCompSeqNum i (k = 1, 2, ..., N; k ≠ i) are preprocessed by normalization, and the Euclidean distance d between the two after normalization is calculated. i,k ;
[0026] B) Select the Euclidean distance d between the i-th flavor monomer and the i-th flavor monomer. i,k The largest spice monomer, and the European distance d i,k The largest single spice data as shown in the DantiSortingMatrix i The first line;
[0027] C) Calculate the Euclidean distance between each of the remaining fragrance monomers and the i-th fragrance monomer and the selected fragrance monomers;
[0028] D) Based on the principle that the least dissimilar flavoring monomer has the largest minimum Euclidean distance, select the flavoring monomers least dissimilar to the i-th flavoring monomer and the selected flavoring monomers from the remaining flavoring monomers, and add their data to the DaniSortingMatrix in sequence. i In the middle, until DantiSortingMatrix i It contains data for all flavor monomers except for the i-th flavor monomer.
[0029] In step 4 of this invention, the probability of the i-th flavoring monomer in each flavoring monomer being present in the flavoring monomer composite sample to be tested is Poss. i The calculation method for (i = 1, 2, ..., N) includes calculation method 1, which includes the following steps:
[0030] i) Let ProjMatrix k For DantiSortingMatrix i The matrix consisting of the first k rows of data (k = 1, 2, ..., N-1); x test The sample containing the fragrance base to be tested has an index equal to CharacCompSeqNum. i Row vector data composed of mass spectrometry peak signals of the chemical composition of elements in the middle;
[0031] ii) Calculate Poss using the following formula i (i = 1, 2, ..., N)
[0032] Res1 k =x test ×(I-(ProjMatrix k ) + ×ProjMatrix k )
[0033]
[0034] Coeff=[Coeff(1),Coeff(2),…,Coeff(N-1)]
[0035] Poss i =max(Coeff)
[0036] Where I is an m×m identity matrix (m is CharacCompSeqNum) i (the number of elements); + Represents the Moore-Penrose generalized inverse; cov(Res1) k Res2 k ) represents Res1 k With Res1 k The covariance between them; σ(Res1) k ) and σ(Res2 k ) represent Res1 k With Res1 k The standard deviation.
[0037] In step 4 of this invention, the probability of the i-th flavoring monomer in each flavoring monomer being present in the flavoring monomer composite sample to be tested is Poss. i The calculation method for (i = 1, 2, ..., N) includes calculation method 2, which includes the following steps:
[0038] i) Let ProjMatrix1 k For DantiSortingMatrix i The matrix consisting of the first k rows of data (k = 1, 2, ..., N-1), x test The sample containing the fragrance base to be tested has an index equal to CharacCompSeqNum. i Row vector data composed of mass spectrometry peak signals of the chemical composition of elements in the middle;
[0039] ii) Calculate Poss using the following formula i (i = 1, 2, ..., N)
[0040] Res1 k =x test ×(I-(ProjMatrix1 k ) + ×ProjMatrix1 k )
[0041] Res2 k =x tess ×(I-(ProjMatrix2 k ) + ×ProjMatrix2 k )
[0042]
[0043] DiffVal=[DiffVal(1),DiffVal(2),…,DiffVal(N-1)]
[0044] Poss i =max(DiffVal)
[0045] Where I is an m×m identity matrix, and m is the CharacCompSeqNum i The number of elements; + Represents the Moore-Penrose generalized inverse; Res1′ k Res1 k The transpose operation; Res2′ k Res2 k The transpose operation.
[0046] In step 4 of this invention, the probability of the i-th flavoring monomer in each flavoring monomer being present in the flavoring monomer composite sample to be tested is Poss. i The calculation method for (i = 1, 2, ..., N) includes calculation method 3, which includes the following steps:
[0047] i) Let ProjMatrix1 k For DantiSortingMatrix i The matrix consisting of the first k rows of data (k = 1, 2, ..., N-1), x test The sample containing the fragrance base to be tested has an index equal to CharacCompSeqNum. i Row vector data composed of mass spectrometry peak signals of the chemical composition of the elements, ProjMatrix2 k =[ProjMatrix1 k ;x test ];
[0048] ii) Calculate Poss using the following formula i (i = 1, 2, ..., N)
[0049]
[0050] DiffVal=[DiffVal(1),DiffVal(2),…,DiffVal(N-1)]
[0051] Poss i =max(DiffVal)
[0052] Where I is an m×m identity matrix (m is CharaCompSeqNum)i (Number of elements); the superscript '+' represents the Moore-Penrose generalized inverse; Res1′ k Res1 k The transpose operation; Res2′ k Res2 k The transpose operation.
[0053] In the method described in this invention, the extract or purified product of the plant or its organ or tissue includes one or more of the following:
[0054] Plum extract, raisin extract, fig extract, apricot extract, jujube tincture, jujube essential oil, apple extract, wolfberry extract, carob extract, monk fruit tincture, tamarind extract, hawthorn tincture, sea buckthorn extract, dandelion fluid extract, fenugreek tincture, Hangzhou white chrysanthemum tincture, tomato extract, tree flower tincture, iris tincture, valerian root tincture, valerian root extract, chicory extract, cocoa extract, alfalfa extract, Roman chamomile extract, vanilla extract, hops tincture, maple extract, tobacco Maillard reaction product, angelica extract, angelica pubescens extract, licorice fluid extract, malt extract, turkey extract, cardamom extract, tamarind extract, rue extract, grape extract, jujube extract, compound tobacco extract, malt extract, hawthorn extract, Virginia flue-cured tobacco extract, Zimbabwean tobacco extract, Yunnan tobacco extract, Brazilian tobacco extract and / or dried plum extract.
[0055] The spice monomer composite sample includes at least two or more combinations of extracts or refinements of the plants or their organs or tissues described in this invention.
[0056] Furthermore, the flavor monomer complex sample also includes solvents and sugars.
[0057] In a specific embodiment of the present invention, the fragrance monomer composite sample can be any one of fragrance bases 1 to 12 shown in Table 2, more specifically, it can be fragrance base-1, fragrance base-2, fragrance base-3, fragrance base-4 or fragrance base-12, etc., and the fragrance monomer in the step refers to the fragrance monomer shown in Table 1.
[0058] The solvents include, but are not limited to, propylene glycol (PG); the sugars include, but are not limited to, high fructose.
[0059] In this invention, because the LC-MS / MS data of each natural fragrance monomer (fragrance monomer) contains information on thousands of chemical components, and these chemical components are not unique to a single natural fragrance monomer but are common to multiple natural fragrance monomers, it is impossible to determine whether a functional fragrance base sample's formulation uses a certain natural fragrance monomer simply by observing the presence of certain chemical components in its LC-MS / MS data. This makes identifying the monomer formulation of a compound functional fragrance base based on the LC-MS / MS data of natural fragrance monomers and compound functional fragrance base samples extremely challenging. Currently, no corresponding monomer formulation identification method has been developed domestically or internationally. This invention provides for the first time a method or algorithm for identifying natural fragrance monomers (fragrance monomers) in functional fragrance base samples (fragrance monomer compound samples) based on LC-MS / MS data.
[0060] Based on the characteristics that natural fragrance raw materials are mainly plant-based and mainly composed of water-soluble components, this invention develops an algorithm for identifying compound fragrance base monomer formulations based on LC-MS / MS data of natural fragrance raw materials using LC-MS / MS analysis and characterization methods. The method described in this invention has simple model parameter settings and fast calculation speed, and can effectively identify and characterize the monomer formulations in fragrance bases. Attached Figure Description
[0061] Figure 1 The flowchart of the method of the present invention is shown. Detailed Implementation
[0062] This invention provides a method for identifying fragrance base formulations based on LC-MS / MS data of natural fragrance raw materials. Those skilled in the art can refer to the content of this document and appropriately modify the process parameters to achieve the desired result. It should be particularly noted that all similar substitutions and modifications are obvious to those skilled in the art and are considered to be included in this invention. The methods and applications of this invention have been described through preferred embodiments. Those skilled in the art can obviously modify or appropriately change and combine the methods and applications described herein without departing from the content, spirit, and scope of this invention to realize and apply the technology of this invention.
[0063] Liquid chromatography-tandem mass spectrometry (LC-MS / MS) is a commonly used tool for qualitative and quantitative analysis of substances. LC-MS / MS-based data analysis is a highly efficient technique that allows for in-depth exploration of the chemical composition of biological samples. This technique can comprehensively identify and accurately quantify various components in a sample.
[0064] This invention is the first to propose an effective algorithm for determining the monomer formulation in an unknown fragrance base sample when acquiring liquid chromatography-tandem mass spectrometry data of natural monomeric fragrances and unknown fragrance base samples.
[0065] The value of the model parameter p described in this invention should be such that the single-unit characteristic spectrum... The number of elements is in the range of 3×N to 5×N.
[0066] As described above, the present invention is a method for determining the monomer formulation in an unknown fragrance base sample based on LC-MS / MS data of natural monomeric fragrance raw materials and unknown fragrance base samples.
[0067] The advantages of the method described in this invention can be summarized as follows:
[0068] 1. This invention has only one model parameter p, the value of which only needs to make the single-unit characteristic spectrum... The number of elements is in the range of 3×N to 5×N, thus it has the advantage of simple model parameter setting;
[0069] 2. This invention significantly reduces the computational load by constructing characteristic component spectra of each monomer relative to other monomers, thus offering the advantage of high computational speed;
[0070] 3. Each step of the calculation in this invention has its corresponding spectroscopic theoretical basis, thus having the advantage of a sound theoretical foundation.
[0071] The 47 natural raw materials and their numbers in this invention are shown in Table 1, and the 12 self-prepared fragrance base samples (natural raw material compound samples) and their numbers are shown in Table 2:
[0072] Table 1.47 Flavor Monomers and Their Numbers
[0073]
[0074]
[0075] Table 2.12 Fragrance Base Samples (Fragrance Monomer Composite Samples) and Their Numbers
[0076] Serial Number Fragrance base name Sample number 1 Xiangji 1 XJ-1 2 Xiangji 2 XJ-2 3 Scented Base 3 XJ-3 4 Scented 4 XJ-4 5 Xiangji 5 XJ-5 6 Xiangji 6 XJ-6 7 Xiangji 7 XJ-7 8 Xiangji 8 XJ-8 9 Xiangji 9 XJ-9 10 Xiangji 10 XJ-10 11 Xiangji 11 XJ-11 12 Xiangji 12 XJ-12
[0077] The formula for Fragrance Base-1 in Table 2 is shown in Table 3. The formulas for other fragrance base samples are similar, all composed of the various natural raw materials shown in Table 1.
[0078] Table 3. Formulation of Fragrance Base-1 (One of the samples of fragrance monomer complex)
[0079] Individual number Fragrance raw materials Proportion 1 Plum extract 10 2 Raisin extract 5 water 4 High fructose 35 PG 46 total 100
[0080] The test materials used in this invention are all common commercially available products. The invention is further illustrated below with reference to embodiments:
[0081] Example 1: Formula Identification of Self-Prepared Fragrance Base Samples
[0082] The data used in this embodiment are LC-MS / MS data of a batch of natural fragrance raw material samples. The data includes mass spectrometry signal intensity data of 1195 chemical components in 47 natural raw materials (Table 1) and 12 formulated fragrance base samples (Table 2).
[0083] I. Algorithm for analyzing Xiangji LC-MS / MS data
[0084] 1. Steps of the Xiangji LC-MS / MS data parsing algorithm
[0085] The LC-MS / MS data of natural fragrance raw material samples were processed using the following algorithm to identify the natural fragrance composition of the formulated fragrance base samples. The parameter in this experiment was set to p = 0.8. The algorithm flowchart is shown below. Figure 1 The specific steps are as follows:
[0086] The formulation identification method based on aroma LC-MS / MS data of the present invention includes the following steps ( Figure 1 ):
[0087] (1) First, obtain the LC-MS / MS data matrix X of N monomers (fragrance monomers) that may be used to prepare the fragrance monomer composite sample; the i-th row X(i,:) represents the LC-MS / MS data of the i-th monomer (i=1,2,…,N); the j-th element X(i,j) of X(i,:) is the mass spectrum peak signal intensity of the j-th chemical component in the i-th monomer.
[0088] (2) Based on the LC-MS / MS data of N monomers, find the chemical components with certain characteristics in content relative to other monomers according to the strategies shown in a) to e) below (Note: the selection of model parameter p is involved here), and then use the mass spectrum peak signal intensity of these components to form the characteristic spectrum of the monomer.
[0089] a) First, the mass spectrum peak signal data (row vector) X(i,:) (i=1,2,…,N) of the i-th monomer is standardized to obtain X norm (i,:)(X norm (i,:) = X(i,:) / max(X(i,:)), where max(X(i,:)) is the maximum element value of X(i,:).
[0090] b) Select X norm The mass spectrum peak signal intensity in (i,:) is greater than p×median(X). norm The chemical composition of (i,:)) (Note: median(X) norm (i,:) is X norm The median value of (i,:) is used to form the main component index vector MainCompSeqNum.i,1 .
[0091] c) Calculate the sum of the LC-MS / MS data vectors of all monomers except the i-th monomer among the N monomers, and then standardize them to obtain...
[0092] d) Select x sum,norm The peak area in the medium mass spectrum is greater than p × median(x) sum,norm The chemical composition of the component is determined, and its corresponding serial numbers are used to form the main component serial number vector MainCompSeqNum. i,2 ;
[0093] e) Set MainCompSeqNum i,1 and MainCompSeqNum i,2 The common chemical component numbers contained in it are from MainCompSeqNum i,1 After removing the sub-components, the remaining chemical component indices form the characteristic component indices vector CharacCompSeqNum for the i-th monomer. i Then the characteristic spectrum of the i-th monomer The index of X(i,:) is equal to CharacCompSeqNum i The mass spectrometry peak signal composition of the chemical composition of the elements, that is,
[0094] (3) Construct the 'least dissimilar monomer descending order data matrix' of the i-th monomer according to the following strategies A) to D) . i .
[0095] A) To and X(k,CharacCompSeqNum i Perform normalization preprocessing on (k = 1, 2, ..., N; k ≠ i) and calculate the Euclidean distance d between the two after normalization preprocessing. i,k ;
[0096] B) Select the Euclidean distance d between the monomer and the 'i-th monomer' from the monomer set. i,k The largest monomer, and its mass spectrometry data as the DantiSortingMatrix i The first line;
[0097] C) Calculate the Euclidean distance between each individual in the remaining individual set and 'the i-th individual and the selected individual';
[0098] D) According to the rule of 'the least dissimilar is the one with the largest minimum Euclidean distance', select the monomers least dissimilar to 'the i-th monomer and the selected monomers' from the remaining monomer set, and add their mass spectrometry data to the DaniSortingMatrix in sequence. i In the middle, until DantiSortingMatrix i It contains LC-MS / MS data for all monomers except the i-th monomer (each row represents a monomer).
[0099] (4) Calculate the 'probability' Poss of the presence of the i-th monomer in the fragrance sample to be tested according to the strategies shown in i) to ii) below. i .
[0100] i) Let ProjMatrix k For DantiSortingMatrix i The matrix consisting of the first k rows of data (k = 1, 2, ..., N-1); x test The sample containing the fragrance base to be tested has an index equal to CharacCompSeqNum. i Row vector data composed of mass spectrometry peak signals of the chemical composition of elements in the middle;
[0101] ii) Calculate Poss using the following formula i (i = 1, 2, ..., N)
[0102] Res1 k =x test ×(I-(ProjMatrix k ) + ×ProjMatrix k )
[0103]
[0104] Coeff=[Coeff(1),Coeff(2),…,Coeff(N-1)]
[0105] Poss i =max(Coeff)
[0106] Where I is an m×m identity matrix (m is CharacCompSeqNum) i (number of elements); the superscript '+' represents the Moore-Penrose generalized inverse; cov(Res1) k Res2 k ) represents Res1 j With Res1 k The covariance between them; σ(Res1) k) and σ(Res2 k ) represent Res1 k With Res1 k The standard deviation.
[0107] Based on the 'probability value' of the monomer Poss i Sort all monomers (i = 1, 2, ..., N) from largest to smallest to obtain the possible formulation analysis results of the fragrance sample to be tested.
[0108] 2. The probability value 'Poss' in the Xiangji LC-MS / MS data analysis algorithm i Replaceable algorithms
[0109] Equivalent method 1:
[0110] i) Let ProjMatrix1 k For DantiSortingMatrix i A matrix consisting of the first k rows of data (k = 1, 2, ..., N-1) x test The sample containing the fragrance base to be tested has an index equal to CharacCompSeqNum. i Row vector data composed of mass spectrometry peak signals of the chemical composition of elements in the middle;
[0111] ii) Calculate Poss using the following formula i (i = 1, 2, ..., N)
[0112] Res1 k =x test ×(I-(ProjMatrix1 k ) + ×ProjMatrix1 k )
[0113] Res2 j =x test ×(I-(ProjMatrix2 k ) + ×ProjMatrix2 k )
[0114]
[0115] DiffVal=[DiffVal(1),DiffVal(2),…,DiffVal(N-1)]
[0116] Poss i =max(DiffVal)
[0117] Where I is an m×m identity matrix (m is CharacCompSeqNum) i (Number of elements); the superscript '+' represents the Moore-Penrose generalized inverse; Res1′ k Res1 k The transpose operation; Res2′ k Res2 k The transpose operation.
[0118] Equivalent method 2:
[0119] i) Let ProjMatrix1 k For DantiSortingMatrix i The matrix consisting of the first k rows of data (k = 1, 2, ..., N-1), x test The sample containing the fragrance base to be tested has an index equal to CharacCompSeqNum. i Row vector data composed of mass spectrometry peak signals of the chemical composition of the elements, ProjMatrix2 k =[ProjMatrix1 k ;x test ];
[0120] ii) Calculate Poss using the following formula i (i = 1, 2, ..., N)
[0121]
[0122] DiffVal=[DiffVal(1),DiffVal(2),…,DiffVal(N-1)]
[0123] Poss i =max(DiffVal)
[0124] Where I is an m×m identity matrix (m is CharacCompSeqNum) i (Number of elements); the superscript '+' represents the Moore-Penrose generalized inverse; Res1′ k Res1 k The transpose operation; Res2′ k Res2 k The transpose operation.
[0125] II. Algorithm Analysis Results of Fragrance Sample Formulation
[0126] The above algorithm was used to identify the fragrance base samples in Table 2. Tables 3 to 25 show the formulation identification results of the above fragrance base samples using the fragrance base sample LC-MS / MS data analysis algorithm proposed in this invention. From the formulation identification results shown in Tables 3 to 25, it can be seen that the fragrance base LC-MS / MS data analysis algorithm proposed in this invention can effectively identify the monomers used in the formulation of the fragrance base samples.
[0127] For example, Fragrance Base-1 is a blend of two monomers: plum extract and raisin extract. The LC-MS / MS data analysis algorithm for Fragrance Base-1 indicates that these two monomers are more likely to be present in Fragrance Base-1 than other monomers, accurately reflecting the formulation of Fragrance Base-1. Fragrance Base-3 is a blend of four monomers: plum extract, raisin extract, fig extract, and apricot extract. The LC-MS / MS data analysis algorithm for Fragrance Base-3 ranks these four monomers among the top five most likely to be present, effectively identifying the monomer formulation of the fragrance base sample. For fragrance base samples with more complex formulations, such as Fragrance Base-11, which is formulated with 10 monomers, the LC-MS / MS data analysis algorithm can rank the 8 monomers actually used in the formulation among the top 12, also showing a relatively ideal monomer identification effect.
[0128] The ability of fragrance LC-MS / MS data analysis algorithms to identify monomers is positively correlated with the monomer content in the fragrance. For example, plum extract and raisin extract contain 10% and 5% respectively in fragrance-1, and the fragrance LC-MS / MS data analysis algorithm ranks the probability of their presence in fragrance-1 as the first and second most likely, respectively. Plum extract and raisin extract contain 5% and 10% respectively in fragrance-2, and the fragrance LC-MS / MS data analysis algorithm detects this change in content and accordingly changes its ranking order. Because the characteristics of different monomer LC-MS / MS data differ relative to other monomer LC-MS / MS data, the sensitivity of fragrance LC-MS / MS data analysis algorithms to the identification of different monomers varies significantly. For example, malt extract contains 10% in fragrance-4, but the fragrance LC-MS / MS data analysis algorithm fails to detect its presence. Dandelion fluid extract accounted for only 2% of Flavor-10, yet the Flavor-10 LC-MS / MS data analysis algorithm correctly identified it, ranking it second. It's worth noting that the Flavor-10 LC-MS / MS data analysis algorithm appears to be quite sensitive to various tobacco extracts (such as compound tobacco extracts, Virginia flue-cured tobacco extracts, Zimbabwean tobacco refinements, Yunnan tobacco refinements, Brazilian tobacco extracts, and tobacco Maillard reactants), correctly identifying all tobacco extract monomers from Flavor-9 to Flavor-12. This may be related to the strong characteristic features of LC-MS / MS data for tobacco extracts.
[0129] In summary, the natural fragrance raw material LC-MS / MS data analysis algorithm proposed in this invention can effectively identify the monomeric fragrance raw materials contained in compound fragrance base samples. Its ability to identify monomers is related not only to the content of the monomers in the fragrance base sample, but also to the strength of the characteristic features of the monomer LC-MS / MS data.
[0130] Table 4. Actual formulation of Fragrance Base-1
[0131] Individual number Fragrance raw materials Proportion 1 Plum extract 10 2 Raisin extract 5 water 4 High fructose 35 PG 46 total 100
[0132] Table 5. Analysis results of Scent-1 (ranked from highest to lowest probability of containing monomers, top 20).
[0133]
[0134]
[0135] Table 6. Actual formulation of Fragrance Base-2
[0136] Individual number Fragrance raw materials Proportion 1 Plum extract 5 2 Raisin extract 10 water 4 High fructose 35 PG (Propylene Glycol) 46 total 100
[0137] Table 7. Analysis results of sangyl-2 (ranked from highest to lowest probability of containing monomers, top 20).
[0138] Ranking 1 2 3 4 5 6 7 8 9 10 11 12 13 Individual serial number 2 1 10 6 45 5 47 44 31 3 27 15 25 Ranking 14 15 16 17 18 19 20 Individual serial number 37 26 46 18 36 19
[0139] Table 8. Actual formulation of Fragrance Base-3
[0140] serial number Fragrance raw materials Proportion 1 Plum extract 10 2 Raisin extract 10 3 Fig extract 10 4 Apricot extract 10 water 2 High fructose 30 PG 28 total 100
[0141] Table 9. Analysis results of Scent-3 (ranked from highest to lowest probability of containing monomers, top 20).
[0142] Ranking 1 2 3 4 5 6 7 8 9 10 11 12 13 Individual serial number 9 4 3 1 2 40 47 31 20 10 16 37 6 Ranking 14 15 16 17 18 19 20 Individual serial number 5 22 39 38 44 19 35
[0143] Table 10. Actual formulation of Fragrance Base-4
[0144]
[0145]
[0146] Table 11 Analysis results of sangyl-4 (ranked from highest to lowest probability of containing monomers, top 20).
[0147] Ranking 1 2 3 4 5 6 7 8 9 10 11 12 13 Individual serial number 14 11 10 22 25 12 1 27 24 35 32 7 30 Ranking 14 15 16 17 18 19 20 Individual serial number 17 2 8 16 40 36 44
[0148] Table 12. Actual formulation of Fragrance Base-5
[0149] serial number Fragrance raw materials Proportion 8 wolfberry extract 10 9 Carob extract 10 13 Sea buckthorn extract 10 28 Maple extract 10 38 grape products 10 41 Malt extract (for research use) 15 water 8 High fructose 12 PG 15 total 100
[0150] Table 13 Analysis results of fentanyl-5 (ranked from highest to lowest probability of containing monomers, top 20).
[0151] Ranking 1 2 3 4 5 6 7 8 9 10 11 12 13 Individual serial number 8 13 28 5 38 15 11 41 33 2 9 46 1 Ranking 14 15 16 17 18 19 20 Individual serial number 31 45 4 14 10 39 17
[0152] Table 14. Actual formulation of Fragrance Base-6
[0153]
[0154]
[0155] Table 15 Analysis results of fennel-6 (ranked from highest to lowest probability of containing monomers, top 20).
[0156] Ranking 1 2 3 4 5 6 7 8 9 10 11 12 13 Individual serial number 5 15 12 26 10 1 2 6 31 39 47 8 27 Ranking 14 15 16 17 18 19 20 Individual serial number 25 42 24 3 28 37 17
[0157] Table 16. Actual Formulation of Fragrance Base-7
[0158] serial number Fragrance raw materials Proportion 9 Carob extract 10 11 Tamarind extract 10 14 Dandelion extract 10 23 Cocoa extract 10 28 Maple extract 10 29 Tobacco Maillard reactants 10 41 Malt extract (for research use) 10 38 grape products 10 water 5 High fructose 7 PG 8 total 100
[0159] Analysis results of sangyl-7 (ranked from highest to lowest probability of containing monomers, top 20).
[0160] Ranking 1 2 3 4 5 6 7 8 9 10 11 12 13 Individual serial number 29 28 14 23 38 11 12 5 9 22 10 24 32 Ranking 14 15 16 17 18 19 20 Individual serial number 15 1 4 17 41 33 30
[0161] Table 18. Actual Formulation of Fragrance Base-8
[0162]
[0163]
[0164] Analysis results of sangyl-8 (ranked from highest to lowest probability of containing monomers, top 25).
[0165] Ranking 1 2 3 4 5 6 7 8 9 10 11 12 13 Individual serial number 21 1 5 13 38 14 4 41 47 33 12 42 6 Ranking 14 15 16 17 18 19 20 21 22 23 24 25 Individual serial number 39 10 2 31 9 26 30 3 46 45 35 22
[0166] Table 20. Actual Formulation of Fragrance Base-9
[0167] serial number Fragrance raw materials Proportion 1 Plum extract 8 11 Tamarind extract 8 3 Fig extract 8 14 Dandelion extract 8 28 Maple extract 8 29 Tobacco Maillard reactants 8 41 Malt extract (for research use) 8 38 grape products 8 40 Compound tobacco extract 8 43 Virginia tobacco extract 8 water 1 High fructose 8 PG 11 total 100
[0168] Analysis results of sangyl-9 (ranked from highest to lowest probability of containing monomers, top 25).
[0169] Ranking 1 2 3 4 5 6 7 8 9 10 11 12 13 Individual serial number 29 14 28 17 40 46 38 43 10 45 11 15 31 Ranking 14 15 16 17 18 19 20 21 22 23 24 25 Individual serial number 12 44 1 24 13 22 3 8 41 5 27 37
[0170] Table 22. Actual formulation of Fragrance Base-10
[0171]
[0172]
[0173] Analysis results of sangyl-10 (ranked from highest to lowest probability of containing monomers, top 25).
[0174] Ranking 1 2 3 4 5 6 7 8 9 10 11 12 13 Individual serial number 29 14 17 28 40 46 38 11 10 43 45 44 13 Ranking 14 15 16 17 18 19 20 21 22 23 24 25 Individual serial number 31 15 41 12 8 2 5 33 24 22 25 6
[0175] Table 24. Actual Formulation of Fragrance Base-11
[0176] serial number Fragrance raw materials Proportion 40 Compound tobacco extract 4 43 Virginia tobacco extract 4 44 Zimbabwean tobacco products 4 45 Yunnan Tobacco Refined Products 4 46 Brazilian tobacco extract 4 29 Tobacco Maillard reactants 4 41 Malt extract (for research use) 4 36 Tamarind extract 4 9 Carob extract 4 38 grape products 4 water 2 High fructose 25 PG 33 total 100
[0177] Table 25. Analysis results of Sangil-11 (ranked from highest to lowest probability of containing monomers, top 25).
[0178] Ranking 1 2 3 4 5 6 7 8 9 10 11 12 13 Individual serial number 29 17 40 45 46 38 44 43 24 10 11 9 31 Ranking 14 15 16 17 18 19 20 21 22 23 24 25 Individual serial number 4 14 27 8 28 13 2 30 41 6 39 5
[0179] Table 26. Actual formulation of Fragrance Base-12
[0180] serial number Fragrance raw materials Proportion 40 Compound tobacco extract 1 43 Virginia tobacco extract 2 44 Zimbabwean tobacco products 3 45 Yunnan Tobacco Refined Products 4 46 Brazilian tobacco extract 5 29 Tobacco Maillard reactants 6 41 Malt extract (for research use) 7 36 Tamarind extract 8 9 Carob extract 9 38 grape products 10 water 5 High fructose 15 PG 25 total 100
[0181] Table 27. Analysis results of sangyl-12 (ranked from highest to lowest probability of containing monomers, top 25).
[0182] Ranking 1 2 3 4 5 6 7 8 9 10 11 12 13 Individual serial number 29 46 40 45 38 44 17 24 11 9 43 14 4 Ranking 14 15 16 17 18 19 20 21 22 23 24 25 Individual serial number 5 31 27 28 36 15 2 39 1 10 26 25
[0183] The above are merely preferred embodiments of the present invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.
Claims
1. A method for identifying the composition of fragrance monomers in a fragrance monomer complex sample based on LC-MS / MS data, characterized in that, Includes the following steps: Step 1: Obtain the LC-MS / MS data matrix X of N flavor monomers that may be used to prepare flavor monomer composite samples; where N is the number of flavor monomers that may be used when preparing flavor monomer composite samples; the i-th row X(i,:) of X represents the LC-MS / MS data vector of the i-th flavor monomer; the j-th element X(i,j) of X(i,:) is the mass spectrum peak signal intensity of the j-th chemical component in the i-th flavor monomer. Step 2: Based on the data matrix X, set p to 0.6~1.0, find the differential chemical components of each flavoring monomer relative to other flavoring monomers, and obtain the characteristic spectrum of each flavoring monomer; Step 3: Based on the characteristic spectra of each flavoring monomer, construct a data matrix of the least dissimilar monomers of each flavoring monomer in descending order; Step 4: Based on the data matrix of the least dissimilar monomers for each flavoring monomer arranged in descending order, calculate the probability of each flavoring monomer existing in the flavoring monomer composite sample to be tested, and obtain the flavoring monomer composition of the flavoring monomer composite sample to be tested; in, The mass spectrum peak signal intensity includes: mass spectrum peak area or mass spectrum peak height; The fragrance monomers are extracts or refined products of plants or their organs or tissues; In step 4, the probability of the i-th flavor monomer in each flavor monomer being present in the flavor monomer composite sample to be tested is Poss. i The calculation method includes calculation method 1, which includes the following steps: i) Let ProjMatrix k For DantiSortingMatrix i A matrix consisting of the first k rows of data, where ;X text The sample containing the fragrance base to be tested has an index equal to CharacCompSeqNum. i Row vector data composed of mass spectrometry peak signals of the chemical composition of elements in the middle; ii) Calculate Poss using the following formula i : Where I is an m×m identity matrix, and m is the CharacCompSeqNum i The number of elements; + Represents the Moore-Penrose generalized inverse; cov( Res1 k Res2 k ) represents Res1 k With Res1 k The covariance between them; σ(Res1) k ) and σ( Res2 k ) represent Res1 k With Res1 k The standard deviation.
2. The method according to claim 1, characterized in that, In step 2, the characteristic spectrum of the i-th flavoring monomer in each flavoring monomer is obtained. Includes the following steps: Step a: Standardize the LC-MS / MS data vector of the i-th flavor monomer to obtain X. norm (i,:); where, Step b, select X norm The mass spectrum peak signal intensity in (i,:) is greater than p × median (X). norm The chemical composition of (i,:) is used to form the main component sequence vector MainCompSeqNum, with the corresponding sequence numbers of the chemical components as the main component sequence number vectors. i,1 ;where median (X norm (i,:) is X norm The median value of (i,:); Step c: Calculate the sum of the LC-MS / MS data vectors of all flavor monomers except the i-th monomer among the N flavor monomers, and normalize them to obtain x. sum,norm ; Step d: Select x sum,norm The intensity of the medium mass spectrum peak signal is greater than p × median (x sum,norm The chemical composition of the component is used to construct a main component sequence vector MainCompSeqNum, which is composed of the sequence numbers corresponding to the chemical components. i,2 ; Step e: Set MainCompSeqNum i,1 and MainCompSeqNum i,2 The chemical components commonly contained in MainCompSeqNum are numbered from MainCompSeqNum i,2 After removing the sub-components, the remaining chemical component indices form the characteristic component indices vector CharacCompSeqNum for the i-th monomer. i The characteristic spectrum of the i-th monomer The index of X(i,:) is equal to CharacCompSeqNum i The mass spectrometry peak signal composition of the chemical composition of the elements, i.e., the... .
3. The method according to claim 2, characterized in that, In step a, the X norm The formula for calculating (i,:) is: .
4. The method according to claim 3, characterized in that, In step c, x sum,norm The calculation formula is: ,in .
5. The method according to claim 4, characterized in that, In step 3: The data matrix of the least dissimilar monomers relative to other flavor monomers was constructed by calculating the Euclidean distance. Before calculating the Euclidean distance, the process also includes normalizing the characteristic spectral data of each spice monomer. The principle for constructing the data matrix of the least dissimilar monomers relative to other monomers in descending order is: the least dissimilar monomers are those with the largest Euclidean distance.
6. The method according to claim 5, characterized in that, Construct a data matrix of the least dissimilar monomers of the i-th flavoring monomer in descending order. Includes the following steps: A) To and X(k, CharacCompSeqNum i Normalization preprocessing, and calculation of the Euclidean distance d between the two after normalization. i,k ,in B) Select the Euclidean distance d between the i-th flavor monomer and the i-th flavor monomer. i,k The largest spice monomer, and the European distance d i,k The largest flavor monomer data as The first line; C) Calculate the Euclidean distance between each of the remaining fragrance monomers and the i-th fragrance monomer and the selected fragrance monomers; D) Following the principle that the least dissimilar flavoring monomer has the largest minimum Euclidean distance, select the flavoring monomers least dissimilar to the i-th flavoring monomer and the already selected flavoring monomers from the remaining flavoring monomers, and add their data sequentially to the list. In the middle, until It contains data for all flavor monomers except for the i-th flavor monomer.
7. The method according to claim 6, characterized in that, In step 4, the probability of the i-th flavor monomer in each flavor monomer being present in the flavor monomer composite sample to be tested is Poss. i The calculation method also includes calculation method 2, which includes the following steps: i) Let ProjMatrix1 k For DantiSortingMatrix i A matrix consisting of the first k rows of data, where , X text The sample containing the fragrance base to be tested has an index equal to CharacCompSeqNum. i Row vector data composed of mass spectral peak signals of the chemical composition of elements; ii) Calculate Poss using the following formula i : Where I is an m×m identity matrix, and m is the CharacCompSeqNum i The number of elements; + Represents the Moore-Penrose generalized inverse; Res1 ′ k Res1 k The transpose operation of Res2; ′ k Res2 k The transpose operation.
8. The method according to claim 6, characterized in that, In step 4, the probability of the i-th flavor monomer in each flavor monomer being present in the flavor monomer composite sample to be tested is Poss. i The calculation method also includes calculation method 3, which includes the following steps: i) Let ProjMatrix1 k For DantiSortingMatrix i A matrix consisting of the first k rows of data, where X text The sample containing the fragrance base to be tested has an index equal to CharacCompSeqNum. i Row vector data composed of mass spectrometry peak signals of the chemical composition of elements in the middle. ; ii) Calculate Poss using the following formula i : Where I is an m×m identity matrix, and m is the CharacCompSeqNum i The number of elements; + Represents the Moore-Penrose generalized inverse; Res1 ′ k Res1 k The transpose operation of Res2; ′ k Res2 k The transpose operation.
9. The method according to any one of claims 1 to 8, characterized in that, The extracts or purified products of the plant or its organs or tissues include any of the following: Plum extract, raisin extract, fig extract, apricot extract, jujube tincture, jujube essential oil, apple extract, wolfberry extract, carob extract, monk fruit tincture, tamarind extract, hawthorn tincture, sea buckthorn extract, dandelion fluid extract, fenugreek tincture, Hangzhou white chrysanthemum tincture, tomato extract, tree flower tincture, iris tincture, valerian root tincture, valerian root extract, chicory extract, cocoa extract, alfalfa extract, Roman chamomile extract, vanilla extract, hops tincture, maple extract, tobacco Maillard reaction product, angelica extract, angelica pubescens extract, licorice fluid extract, malt extract, turkey extract, cardamom extract, tamarind extract, rue extract, grape extract, jujube extract, compound tobacco extract, malt extract, hawthorn extract, Virginia flue-cured tobacco extract, Zimbabwean tobacco extract, Yunnan tobacco extract, Brazilian tobacco extract and / or dried plum extract; The spice monomer complex sample comprises at least two or more combinations of extracts or refinements of the plant or its organs or tissues.
Citation Information
Patent Citations
Method for evaluating and controlling quality of tobacco essence and perfume
CN116908350A
Analytical method of functional essence base formula
CN117877607A