Methods, apparatus, equipment, and storage media for calculating the similarity of viscous tobacco flavorings and fragrances.

By using dilution treatment and LC-MS spectral analysis, the similarity of viscous tobacco flavorings and fragrances was calculated, solving the problem of evaluating the quality stability of flavorings and fragrances and providing a simple and efficient method for similarity calculation.

CN116825225BActive Publication Date: 2025-10-28CHINA TOBACCO HEBEI INDUSTRIAL CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310512361.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-05-08
Publication Date
2025-10-28
Estimated Expiration
2043-05-08

AI Technical Summary

Technical Problem

Existing technologies make it difficult to quickly and easily compare the similarity of viscous tobacco flavorings and fragrances, which makes it difficult to evaluate the quality stability of flavorings and fragrances.

Method used

The dilution pretreatment was performed using ethanol and water or a mixture of ethanol, water and propylene glycol. The spectral similarity of the fragrance and flavor was calculated by converting the spectral data into a two-dimensional plaintext matrix using LC-MS. After noise reduction, time windows were divided, and the abundance, cosine value, abundance ratio and similarity weight of the mass-to-nucleus ratio were calculated.

Benefits of technology

It enables accurate evaluation of flavor and fragrance similarity even when the components are unclear, simplifies the operation process, improves calculation efficiency, reduces reliance on human experience, has a wide range of applications, and is highly accurate.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116825225B_ABST
    Figure CN116825225B_ABST
Patent Text Reader

Abstract

This invention relates to a method, apparatus, device, and storage medium for calculating the similarity of viscous tobacco flavorings and fragrances. The method includes: performing dilution pretreatment on two viscous tobacco flavorings and fragrances respectively to obtain corresponding LC-MS spectra; converting the LC-MS spectra into two-dimensional plaintext data matrices; performing noise reduction processing on the two plaintext data matrices on the retained time series to obtain denoised data matrices; dividing the two denoised data matrices into multiple time windows on the retained time series respectively, calculating the sum of abundance of a single time window and the sum of abundance of all time windows for the i-th time window; calculating the cosine value, abundance ratio, and similarity weight of the i-th time window using the abundance sum of the mass-to-nucleus ratios corresponding to the two plaintext data matrices, and using these to calculate the similarity of the spectra of the two viscous tobacco flavorings and fragrances. This method can quickly and easily compare the similarity of different viscous tobacco flavorings and fragrances.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of tobacco flavoring technology, and in particular to a method, apparatus, equipment, and storage medium for calculating the similarity of viscous tobacco flavorings. Background Technology

[0002] Flavorings and fragrances are essential raw materials in cigarette production. Adding them to cigarettes improves taste and flavor, imparts characteristic aromas, and enhances the smoking experience. The stability of flavoring and fragrance quality significantly impacts the stability of cigarette style. Currently, the quality of flavorings and fragrances is primarily controlled through physical indicators and chemical composition. Flavorings and fragrances come in a wide variety, and their physicochemical indicators and chemical compositions vary considerably. Physical indicators such as boiling point, solubility, color, relative density, refractive index, acid value, and volatile components have relatively low sensitivity and specificity for flavorings and fragrances. Chemical composition analysis mainly employs methods such as gas chromatography, high-performance liquid chromatography, and gas / high-performance liquid chromatography-mass spectrometry.

[0003] Currently, GC-MS is the most common method for analyzing tobacco flavorings and fragrances. This method is suitable for analyzing most volatile and semi-volatile compounds in tobacco flavorings and fragrances, and can analyze the chemical structure information of the main chemical components in the fingerprint spectrum by comparing with a standard mass spectrometry library. However, GC-MS has certain requirements for sample flowability. For viscous flavorings and fragrances with poor flowability, liquid-liquid extraction or other methods are required for sample dilution pretreatment. Commonly used solvents such as dichloromethane, ethanol, ethyl acetate, and n-hexane have poor solubility and low extraction efficiency for viscous flavorings and fragrances, making it difficult to obtain their effective components. Incorrect operation may also damage the injection needle, chromatographic column, and detector. In addition, the content of volatile and semi-volatile compounds in viscous tobacco flavorings and fragrances is low, so the information on effective components obtained by GC-MS is very limited. To obtain more information on effective components, LC-MS analysis can be considered. However, since there is no commercially available standard mass spectrometry library for LC-MS, the qualitative analysis of compounds requires comparison with the retention time, parent ion, and secondary fragment ion fragmentation patterns of standards. This process is cumbersome, making it difficult to achieve qualitative analysis of all components. Therefore, how to quickly and easily compare the similarity of different viscous tobacco flavorings and fragrances to evaluate their stability has become a major concern for those skilled in the art. Summary of the Invention

[0004] To address or partially address the problems existing in related technologies, this invention provides a method, apparatus, device, and storage medium for calculating the similarity of viscous tobacco flavorings and fragrances.

[0005] First, this invention provides a method for calculating the similarity of viscous tobacco flavorings, which includes the following steps:

[0006] S11. Using the same diluting solvent and the same dilution ratio, perform dilution pretreatment on two viscous tobacco flavorings and fragrances respectively; the diluting solvent is: a mixed solvent of ethanol and water, or a mixed solvent of ethanol, water and propylene glycol.

[0007] S12. Under the same conditions, obtain LC-MS spectra of two viscous tobacco flavorings and fragrances that have undergone dilution pretreatment; convert the LC-MS spectra into two-dimensional plaintext data matrices to obtain plaintext data matrices dat1[m,n] and dat2[m,n] corresponding to the two viscous tobacco flavorings and fragrances; the elements of the plaintext data matrices are abundances, n is the number of sequences on the mass-to-nucleus ratio sequence, and m is the number of sequences on the retention time series;

[0008] S13. Perform noise reduction processing on the retained time series for the n mass-nucleus ratios in the plaintext data matrix dat1[m,n] to obtain the noise-reduced data matrix dat1_dn[m,n]; perform noise reduction processing on the plaintext data matrix dat2[m,n] in the same way to obtain the noise-reduced data matrix dat2_dn[m,n];

[0009] S14. Divide the denoised data matrix dat1_dn[m, n] into multiple time windows on the retained time series; let the window width of each time window be 2t, the step size of each time window be half the window width t, and the number of time windows be wn. At this time, the data matrix of the i-th time window is dat1_wd. i [2t, n], 1 ≤ i ≤ wn;

[0010] For the data matrix dat1_wd of the i-th time window i The calculation is performed on [2t, n] to obtain the mass-to-nucleus ratio abundance and dat1_wd for each mass-to-nucleus ratio in the i-th time window. i _ms, and, obtain the single-time-window abundance and dat1_wd of the i-th time window. i _tms;

[0011] Calculate the sum of the abundance of each time window for wn time windows to obtain the total abundance of all time windows corresponding to the plaintext data matrix dat1[m,n] and dat1_tms;

[0012] S15. Referring to step S14, calculate the denoised data matrix dat2_dn[m, n] to obtain the mass-nucleus ratio abundance and dat2_wd of each mass-nucleus ratio in the i-th time window corresponding to the plaintext data matrix. i _ms, the single-time-window abundance of the i-th time window, and dat2_wd i_tms, and abundance for all time windows and dat2_tms;

[0013] S16, through the mass-to-nucleus ratio abundance and dat1_wd i _ms and dat2_wd i _ms calculates the cosine value cos(i) of the i-th time window;

[0014] Through the single time window abundance and dat1_wd i _tms and dat2_wd i _tms calculates the abundance ratio tms(i) for the i-th time window;

[0015] Through the single time window abundance and dat1_wd i _tms and dat2_wd i _tms, and the abundance of all time windows and dat1_tms and dat2_tms are used to calculate the similarity weight w(i) of the i-th time window;

[0016] S17. Calculate the similarity sml of the spectra of two viscous tobacco flavorings and fragrances based on the cosine value cos(i), abundance ratio tms(i), and similarity weight w(i) of the i-th time window, where i ranges from 1 to wn.

[0017] Further, the diluting solvent is a mixed solvent of ethanol and water with a volume ratio of (0.2-5):1, or a mixed solvent of ethanol, water and propylene glycol with a volume ratio of (0.2-5):(0.2-5):1.

[0018] Further, in step S14, the mass-to-nucleus ratio abundance and dat1_wd of each mass-to-nucleus ratio in the i-th time window are... i _ms is calculated as follows:

[0019] For the data matrix dat1_wd of the i-th time window i Summing the data in the [2t, n] time series yields an n-dimensional vector dat1_wd. i _ms[n] represents the abundance of each mass-to-nucleus ratio in the i-th time window, and dat1_wd. i _ms;

[0020] In step S14, the single-time-window abundance of the i-th time window and dat1_wd i _tms is calculated as follows:

[0021] For dat1_wd i Summation of data on mass-to-nucleus ratio sequences in _ms Get dat1_wd i_tms represents the sum of abundance in a single time window for the i-th time window.

[0022] Further, in step S16, according to formula I, the mass-to-nucleus ratio abundance and dat1_wd are used to determine the mass-to-nucleus ratio abundance. i _ms and dat2_wd i _ms calculates the cosine value cos(i) of the i-th time window:

[0023]

[0024] Further, in step S16, according to Equation II, the abundance of the single time window and dat1_wd are used... i _tms and dat2_wd i _tms calculates the abundance ratio tms(i) for the i-th time window;

[0025]

[0026] Further, in step S16, according to Equation III, the abundance of the single time window and dat1_wd are used... i _tms and dat2_wd i _tms, and the abundance of all time windows and dat1_tms and dat2_tms are used to calculate the similarity weight w(i) of the i-th time window;

[0027]

[0028] Further, in step S17, the similarity sml of the spectra of the two viscous tobacco flavorings is calculated according to formula IV based on the cosine value cos(i), abundance ratio tms(i), and similarity weight w(i) of the i-th time window;

[0029]

[0030] Accordingly, the present invention also provides a similarity calculation device for viscous tobacco flavorings, comprising:

[0031] The data acquisition and conversion unit is used to acquire LC-MS spectra of two diluted and pretreated viscous tobacco flavorings under the same conditions; and convert the spectra into two-dimensional plaintext data matrices to obtain plaintext data matrices dat1[m,n] and dat2[m,n] corresponding to the two viscous tobacco flavorings; the elements of the plaintext data matrices are abundances, n is the number of sequences on the mass-to-nucleus ratio sequence, and m is the number of sequences on the retention time series;

[0032] The data processing unit performs noise reduction on the n mass ratios in the plaintext data matrix dat1[m,n] over the retained time series to obtain a denoised data matrix dat1_dn[m,n]. It then performs the same noise reduction on the plaintext data matrix dat2[m,n] to obtain a denoised data matrix dat2_dn[m,n]. The denoised data matrix dat1_dn[m,n] is then divided into multiple time windows over the retained time series. Let the window width of each time window be 2t, the step size of each time window be half the window width t, and the number of time windows be wn. In this case, the data matrix of the i-th time window is dat1_wd. i [2t, n], 1 ≤ i ≤ wn; then for the data matrix dat1_wd of the i-th time window. i The calculation is performed on [2t, n] to obtain the mass-to-nucleus ratio abundance and dat1_wd for each mass-to-nucleus ratio in the i-th time window. i _ms, and, obtain the single-time-window abundance and dat1_wd of the i-th time window. i _tms; calculate the sum of the abundance of each time window for wn time windows to obtain the total abundance of all time windows corresponding to the plaintext data matrix dat1[m,n], and then calculate the denoised data matrix dat2_dn[m,n] in the same way to obtain the mass-to-nucleus ratio abundance of each mass-to-nucleus ratio in the i-th time window corresponding to the plaintext data matrix, dat2_wd. i _ms, the single-time-window abundance of the i-th time window, and dat2_wd i _tms, and abundance for all time windows and dat2_tms;

[0033] The similarity calculation unit is used to calculate the similarity based on the mass-to-nucleus ratio abundance and dat1_wd. i _ms and dat2_wd i The cosine value cos(i) of the i-th time window is calculated using _ms, and the abundance of the single time window and dat1_wd are used as the basis for the calculation. i _tms and dat2_wd i The abundance ratio tms(i) for the i-th time window is calculated using _tms; then, the abundance ratio for the single time window and dat1_wd are used as the basis for the comparison. i _tms and dat2_wd i _tms, and the abundance of all time windows and dat1_tms and dat2_tms are used to calculate the similarity weight w(i) of the i-th time window; finally, the similarity sml of the spectra of the two viscous tobacco flavorings is calculated based on the cosine value cos(i), abundance ratio tms(i) and similarity weight w(i) of the i-th time window, where i takes the value from 1 to wn.

[0034] The present invention also provides a similarity calculation device for viscous tobacco flavorings and fragrances, comprising:

[0035] Memory, used to store computer programs;

[0036] A processor for invoking and executing the computer program to implement the steps of any of the methods described above.

[0037] The present invention also provides a computer-readable storage medium including a software program adapted for execution by a processor of the steps of the method as described in any of the above.

[0038] The similarity calculation method for viscous tobacco flavorings provided by this invention can include the following beneficial effects:

[0039] 1) This method evaluates similarity based on LC-MS spectra and does not require the identification of every component in the spectrum or the precise measurement of each component. Only the corresponding LC-MS spectrum needs to be obtained. This allows for the comparison of similarity between different viscous tobacco flavorings and fragrances even when the full chemical composition is still unknown. This enables the quality control of viscous tobacco flavorings and fragrances. This method does not require complex sample pretreatment and is simple and efficient to operate.

[0040] 2) This method first converts the LC-MS spectra of flavorings and fragrances into a plaintext data matrix containing abundance information of ions with different mass-to-nucleus ratios at different retention times. After noise reduction, the matrix is ​​divided into multiple time windows for calculation to address the retention time offset issue between two spectra at the same retention time. Simultaneously, setting the time window movement step size to half the window width allows as much data with retention time deviations as possible to be contained within a single time window, further improving the inclusiveness of subsequent calculations regarding the aforementioned retention time deviations. Subsequently, a certain similarity weight is assigned based on the substance content within the time window; time windows containing substances with higher content are given higher weights in similarity calculations, while time windows containing substances with lower content are given lower weights. Finally, the similarity between the spectra of two viscous tobacco flavorings and fragrances is calculated using the cosine value, abundance ratio, and similarity weight of each time window. Therefore, this method can accurately and objectively reflect the similarity between two viscous tobacco flavorings and fragrances, and the calculation process can be completely handled by a computer, greatly improving the efficiency of similarity calculation and saving labor costs.

[0041] 3) This method calculates similarity by directly comparing spectral data instead of peak selection, solving the problem of inaccurate qualitative and quantitative analysis of overlapping peaks, embedded peaks, and small peaks. It also addresses the issue of poor reproducibility and repeatability caused by reliance on experience for qualitative and quantitative analysis by testing personnel. Furthermore, this method does not require high peak separation, allowing for a shorter spectral acquisition time. In summary, the similarity calculation method for viscous tobacco flavorings provided by this invention has the advantages of high accuracy, high efficiency, wide applicability, and low degree of human experience involvement.

[0042] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit the invention. Attached Figure Description

[0043] The above and other objects, features and advantages of the present invention will become more apparent from the more detailed description of exemplary embodiments of the invention in conjunction with the accompanying drawings, wherein the same reference numerals generally represent the same parts.

[0044] Figure 1 This is a schematic diagram illustrating the steps of a method for calculating the similarity of viscous tobacco flavorings and fragrances according to an embodiment of the present invention;

[0045] Figure 2 These are the LC-MS spectra of ten viscous fragrances in Example 1 of this invention;

[0046] Figure 3 This is a schematic diagram of the structure of a similarity calculation device for viscous tobacco flavorings and fragrances shown in an embodiment of the present invention;

[0047] Figure 4 This is a hardware structure block diagram of a similarity calculation device for viscous tobacco flavorings and fragrances, as shown in an embodiment of the present invention. Detailed Implementation

[0048] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0049] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of the invention. The singular forms “a,” “the,” and “the” used in this invention and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to and includes any or all possible combinations of one or more of the associated listed items.

[0050] It should be understood that although the terms "first," "second," "third," etc., may be used in this invention to describe various information, this information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, without departing from the scope of this invention, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Thus, features defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this invention, "a plurality of" means two or more, unless otherwise explicitly specified.

[0051] To quickly and easily compare the similarity of different viscous tobacco flavorings and fragrances in order to evaluate their stability, this invention provides a method for calculating the similarity of viscous tobacco flavorings and fragrances. Please refer to [link to relevant documentation]. Figure 1 The method includes the following steps:

[0052] S11. Using the same diluting solvent and the same dilution ratio, perform dilution pretreatment on two viscous tobacco flavorings and fragrances respectively; the diluting solvent is: a mixed solvent of ethanol and water, or a mixed solvent of ethanol, water and propylene glycol.

[0053] S12. Under the same conditions, obtain LC-MS spectra of two viscous tobacco flavorings and fragrances that have undergone dilution pretreatment; convert the LC-MS spectra into two-dimensional plaintext data matrices to obtain plaintext data matrices dat1[m,n] and dat2[m,n] corresponding to the two viscous tobacco flavorings and fragrances; the elements of the plaintext data matrices are abundance, which is the number of sequences on the mass-to-nucleus ratio sequence, and m is the number of sequences on the retention time series;

[0054] S13. Perform noise reduction processing on the retained time series for the n mass-nucleus ratios in the plaintext data matrix dat1[m,n] to obtain the noise-reduced data matrix dat1_dn[m,n]; perform noise reduction processing on the plaintext data matrix dat2[m,n] in the same way to obtain the noise-reduced data matrix dat2_dn[m,n];

[0055] S14. Divide the denoised data matrix dat1_dn[m, n] into multiple time windows on the retained time series; let the window width of each time window be 2t, the step size of each time window be half the window width t, and the number of time windows be wn. At this time, the data matrix of the i-th time window is dat1_wd. i [2t, n], 1 ≤ i ≤ wn;

[0056] For the data matrix dat1_wd of the i-th time window iThe calculation is performed on [2t, n] to obtain the mass-to-nucleus ratio abundance and dat1_wd for each mass-to-nucleus ratio in the i-th time window. i _ms, and, obtain the single-time-window abundance and dat1_wd of the i-th time window. i _tms;

[0057] Calculate the sum of the abundance of each time window for wn time windows to obtain the total abundance of all time windows corresponding to the plaintext data matrix dat1[m,n] and dat1_tms;

[0058] S15. Referring to step S14, calculate the denoised data matrix dat2_dn[m, n] to obtain the mass-nucleus ratio abundance and dat2_wd of each mass-nucleus ratio in the i-th time window corresponding to the plaintext data matrix. i _ms, the single-time-window abundance of the i-th time window, and dat2_wd i _tms, and abundance for all time windows and dat2_tms;

[0059] S16, through the mass-to-nucleus ratio abundance and dat1_wd i _ms and dat2_wd i _ms calculates the cosine value cos(i) of the i-th time window;

[0060] Through the single time window abundance and dat1_wd i _tms and dat2_wd i _tms calculates the abundance ratio tms(i) for the i-th time window;

[0061] Through the single time window abundance and dat1_wd i _tms and dat2_wd i _tms, and the abundance of all time windows and dat1_tms and dat2_tms are used to calculate the similarity weight w(i) of the i-th time window;

[0062] S17. Calculate the similarity sml of the spectra of two viscous tobacco flavorings and fragrances based on the cosine value cos(i), abundance ratio tms(i), and similarity weight w(i) of the i-th time window, where i ranges from 1 to wn.

[0063] Those skilled in the art will understand that the viscous tobacco flavorings described in this application refer to tobacco flavorings that are semi-solid or relatively viscous liquid at room temperature (because they contain a large amount of non-volatile components, mainly sugars). Specifically, they can include: extracts, extracts, or compound flavorings containing extracts and / or extracts, such as licorice extract, osmanthus extract, benzoin extract, St. John's bread extract, plum extract, cocoa extract, etc. Because viscous tobacco flavorings contain a large amount of non-volatile components, the effective information obtained by GC-MS is limited. In the method provided by this invention, the two viscous tobacco flavorings are first diluted and pretreated according to step S11. The purpose of the dilution and pretreatment is to facilitate subsequent LC-MS analysis. Generally, viscous tobacco flavorings have good solubility in water; therefore, some detection methods first dissolve the flavoring sample in water, then extract the components in the aqueous phase using conventional solvents such as n-hexane or dichloromethane, and finally perform GC-MS analysis. Viscous flavorings dissolved in water cannot be directly injected into gas chromatography. This is primarily because water has a high coefficient of expansion, which can easily overload the liner during injection, affecting peak area reproducibility. Furthermore, commonly used columns for flavoring analysis, such as the DB-5 non-polar column, are easily damaged by water at high temperatures, reducing column efficiency and leading to decreased instrument sensitivity and repeatability. However, for conventional C18 liquid chromatography columns, the presence of water does not affect the analytical results. Therefore, a mixed solution of ethanol, water, and propylene glycol in a certain proportion can be used to dissolve viscous flavoring samples before LC-MS analysis to evaluate the similarity of viscous tobacco flavorings. In this step, the solvent used is a mixture of ethanol and water, or a mixture of ethanol, water, and propylene glycol. Preferably, the solvent is a mixture of ethanol and water with a volume ratio of (0.2–5):1, or a mixture of ethanol, water, and propylene glycol with a volume ratio of (0.2–5):(0.2–5):1. Furthermore, the preferred dilution ratio for viscous fragrances is 1:(1-50). After the above dilution pretreatment, the homogeneity of the sample can be improved by methods such as ultrasound, vortexing, and centrifugation.

[0064] After diluting and pretreating the two viscous tobacco flavorings according to step S11, the LC-MS spectra of the two viscous tobacco flavorings can be obtained according to step S12. Specifically, under the same conditions, the two viscous tobacco flavoring samples are analyzed and detected using liquid chromatography-mass spectrometry (LC-MS) to obtain the corresponding LC-MS spectra. Then, the LC-MS spectra are converted into a two-dimensional plaintext data matrix. The elements of the plaintext data matrix are abundances, where n is the number of sequences in the mass-to-nucleus ratio sequence and m is the number of sequences in the retention time sequence. This abundance can reflect the abundance of ions with different mass-to-nucleus ratios at different retention times. This step can be specifically: first, the LC-MS spectra are converted into two-dimensional data, where the horizontal axis is the mass-to-nucleus ratio and the vertical axis is the retention time. The two-dimensional data can be shown in Table 1.

[0065] Table 1 Two-dimensional data

[0066]

[0067]

[0068] Then, abundance information is extracted from the two-dimensional database to obtain a two-dimensional plaintext data matrix dat. u [m,n], u=1 or 2, the two-dimensional plaintext data matrix dat obtained from Table 1 u [m,n] are as follows:

[0069]

[0070] Following the above method, plaintext data matrices dat1[m,n] and dat2[m,n] corresponding to the LC-MS spectra of two viscous tobacco flavorings can be obtained. Then, step S13 is performed to denoise the two plaintext data matrices to reduce noise in the data. Preferably, step S13 uses the moving median method to denoise the n mass-to-nucleus ratios in the plaintext data matrix dat1[m,n] on the retained time series; the same method is used to denoise the plaintext data matrix dat2[m,n] to obtain the denoised data matrix dat2_dn[m,n]. This method can effectively correct baseline drift, reduce the impact of background noise data on the calculation results, and improve data accuracy.

[0071] Step S14 involves first dividing the denoised data matrix processed in step S13 into time windows, and then processing the resulting time window data matrices to obtain the abundance sum of a single time window and the abundance sum of all time windows. Specifically, this step sets the width of the time window to 2t and the step size of the time window to half the window width t, thus dividing the original denoised data matrix into wn time window matrices. In subsequent calculations and comparisons, calculations are performed in units of time windows. The main purpose of dividing the denoised data matrix into time windows in this step is to solve the problem that the actual effluent substances may not be exactly the same at the same retention time, i.e., there is a retention time offset between the two spectra to be compared. For example, if similarity is compared at each time point, and at the same retention time, the first spectrum has already produced a peak, while the second spectrum produces the same peak only after 100 retention time points, then at these 100 points, both the coherent similarity and the abundance ratio will be very small. However, if the same material peaks from two spectra are grouped into the same time window and the similarity is calculated after ion summation, the two spectra meet the condition of material correspondence, thus solving the time offset problem. Furthermore, calculating in units of time windows can reduce computational load. Setting the time window movement step size to half the window width t ensures that adjacent time windows have some data overlap, allowing as much data with retention time deviation as possible to be contained within a single time window, thereby improving the tolerance of subsequent calculations for the aforementioned retention time deviation. Those skilled in the art will understand that there may be cases where the number of retained time series is not divisible by 2t, i.e., the width of the last time window is less than 2t. In such cases, the portion of the window width less than 2t can be padded with 0, or the time window can be discarded, and the preceding time window can be used as the last time window. After this step, the data matrix of the i-th time window of the plaintext data matrix dat1[m,n] can be obtained as dat1_wd. i [2t, n], specifically as follows:

[0072]

[0073] After dividing the time window, the data matrix dat1_wd for the i-th time window... i The calculation is performed on [2t, n] to obtain the mass-to-nucleus ratio abundance and dat1_wd for each mass-to-nucleus ratio in the i-th time window. i _ms, and, obtain the single-time-window abundance and dat1_wd of the i-th time window. i _tms. Preferably, this step specifically involves:

[0074] For dat1_wd i Summing the data in the [2t, n] time series yields an n-dimensional vector dat1_wd.i _ms[n] represents the abundance of each mass-to-nucleus ratio in the i-th time window, and dat1_wd. i _ms; that is, for dat1_wd i Summing the data in each column of [2t, n] yields an n-dimensional vector dat1_wd. i _ms[n], the n-dimensional vector dat1_wd i _ms[n] represents the aforementioned mass-to-nucleus ratio and dat1_wd. i _ms;

[0075] For dat1_wd i Summation of data on mass-to-nucleus ratio sequences in _ms Get dat1_wd i _tms, which is the sum of abundance within a single time window of the i-th time window. That is, for an n-dimensional vector dat1_wd i Summing the data in _ms[n] yields the single-time-window abundance and dat1_wd for the i-th time window. i _tms.

[0076] The aforementioned steps can obtain the abundance sum of each single time window. Then, the abundance sums of the single time windows of each time window are summed, that is, the sum of the abundance sums of the single time windows of wn time windows is calculated, which yields the abundance sum of all time windows dat1_tms corresponding to the plaintext data matrix dat1[m,n].

[0077] Then, following step S14 above, the denoised data matrix dat2_dn[m, n] is calculated in the same way to obtain the mass-to-nucleus ratio abundance and dat2_wd of each mass-to-nucleus ratio in the i-th time window corresponding to the plaintext data matrix (i.e., the plaintext data matrix corresponding to the aforementioned denoised data matrix). i _ms, the single-time-window abundance of the i-th time window, and dat2_wd i _tms, as well as the abundance of all time windows and dat2_tms.

[0078] Step S16 above uses the mass-to-nucleus ratio and dat1_wd obtained from the preceding steps. i _ms and dar2_wd i The step of calculating the cosine value cos(i) of the i-th time window; this cosine value cos(i) is used as one of the parameters for evaluating the similarity sml of the spectra of two viscous tobacco flavorings. This step is preferably performed according to Equation I using the above mass-to-nucleus ratio abundance and dat1_wd. i _ms and dat2_wd i _ms calculates the cosine value cos(i) of the i-th time window:

[0079]

[0080] Step S16 also uses the above-mentioned single-time-window abundance and dat1_wd i _tms and dat2_wd i The abundance ratio tms(i) of the i-th time window is calculated using _tms; this abundance ratio tms(i) of the i-th time window is also used as one of the parameters for evaluating the similarity sml of the spectra of two viscous tobacco flavorings; this step preferably follows Equation II by using the above single time window abundance and dat1_wd i _tms and dat2_wd i _tms calculates the abundance ratio tms(i) for the i-th time window; those skilled in the art will understand that in Equation II, min means taking the minimum value.

[0081]

[0082] Step S16 also uses single-time-window abundance and dat1_wd i _tms and dat2_wd i The similarity weight w(i) of the i-th time window is calculated using the abundance of all time windows and dat1_tms and dat2_tms. The similarity weight w(i) is also used as one of the parameters for evaluating the similarity sml of the spectra of two viscous tobacco flavorings. The purpose of setting this similarity weight w(i) is to assign a higher weight to time windows containing substances with higher content and a lower weight to time windows containing substances with lower content, thereby improving the accuracy of the similarity results. For example, a smaller weight is assigned to a time window containing small peaks (lower substance content) in two spectra, resulting in a smaller weight proportion in the similarity calculation; conversely, a higher weight is assigned to a time window containing large peaks (higher substance content) in two spectra, resulting in a larger weight proportion in the similarity calculation. In this step, it is preferable to use Equation III to calculate the similarity using the abundance of a single time window and dat1_tms. i _tms and dat2_wd i _tms, and the abundance of all time windows and dat1_tms and dat2_tms are used to calculate the similarity weight w(i) of the i-th time window;

[0083]

[0084] The cosine value cos(i), abundance ratio tms(i), and similarity weight w(i) for the i-th time window corresponding to the two viscous tobacco flavorings have been obtained through the aforementioned steps. Then, according to step S17, the similarity sml of the spectra of the two viscous tobacco flavorings is calculated based on the aforementioned data. The cosine value cos(i), abundance ratio tms(i), and similarity weight w(i) of the i-th time window can be used to calculate the similarity contribution at the i-th time window. Then, the similarity contributions of all time windows (i ranges from 1 to wn) are summed to obtain the similarity sml of the spectra of the two viscous tobacco flavorings; that is, the similarity sml of the two spectra should be the sum of cos(i)*tns(i)*w(i) at wn time windows. Specifically, in this step, it is preferable to calculate the similarity sml of the spectra of two viscous tobacco flavorings and fragrances according to Equation IV based on the cosine value cos(i), abundance ratio tms(i), and similarity weight w(i) of the i-th time window.

[0085]

[0086] The similarity between two viscous tobacco flavorings can be calculated using the above method. In this application, the two viscous tobacco flavorings can be a standard viscous tobacco flavoring and the other a sample flavoring to be tested. In this case, the similarity score (sml) calculated according to the method of this invention can be used to evaluate whether the quality of the sample flavoring is acceptable. Alternatively, the two viscous tobacco flavorings can be different batches of the same type of viscous tobacco flavoring. In this case, the similarity score (sml) calculated according to the method of this invention can be used to evaluate whether the quality of the flavoring is consistent between different batches. Furthermore, this method can also be applied to fields such as selecting alternative samples from a large number of samples and determining sample usability.

[0087] Those skilled in the art will understand that the similarity described in this application is actually a pairwise comparison of the abundance of the mass-to-nucleus ratio. The dimensions must remain consistent; therefore, the mass-to-nucleus ratio corresponding to each row of the two plaintext data matrices should also be consistent, i.e., n in dat1[m,n] and dat2[m,n] should be the same. There may be millisecond-level deviations in the retained time series. These millisecond-level deviations refer to the time deviations recorded at the same point (e.g., the 5000th time point) in two spectra under the same instrument conditions. These deviations are negligible in the calculation; therefore, m in dat1[m,n] and dat2[m,n] can be considered essentially the same. If the mass-to-nucleus ratio ranges are not completely consistent during data acquisition, there are two strategies: one is to pad the missing mass-to-nucleus ratio data with 0, and the other is to only take the overlapping mass-to-nucleus ratios.

[0088] As can be seen from the above, the similarity calculation method for viscous tobacco flavorings provided by this invention has the following advantages:

[0089] 1) This method evaluates similarity based on LC-MS spectra and does not require the identification of every component in the spectrum or the precise measurement of each component. Only the corresponding LC-MS spectrum needs to be obtained. This method can achieve similarity comparison of different viscous tobacco flavorings and fragrances even when the full chemical composition is not yet known, thereby enabling quality control of viscous tobacco flavorings and fragrances. This method is simple and efficient to operate.

[0090] 2) This method first converts the LC-MS spectra of flavorings and fragrances into a plaintext data matrix containing abundance information of ions with different mass-to-nucleus ratios at different retention times. After noise reduction, the matrix is ​​divided into multiple time windows for calculation to address the retention time offset issue between two spectra at the same retention time. Simultaneously, setting the time window movement step size to half the window width allows as much data with retention time deviations as possible to be contained within a single time window, further improving the inclusiveness of subsequent calculations regarding the aforementioned retention time deviations. Subsequently, a certain similarity weight is assigned based on the substance content within the time window; time windows containing substances with higher content are given higher weights in similarity calculations, while time windows containing substances with lower content are given lower weights. Finally, the similarity between the spectra of two viscous tobacco flavorings and fragrances is calculated using the cosine value, abundance ratio, and similarity weight of each time window. Therefore, this method can accurately and objectively reflect the similarity between two viscous tobacco flavorings and fragrances, and the calculation process can be completely handled by a computer, greatly improving the efficiency of similarity calculation and saving labor costs.

[0091] 3) This method calculates similarity by directly comparing spectral data instead of selecting peaks, which solves the problem of difficulty in accurately identifying and quantifying overlapping peaks, embedded peaks, and small peaks; it also solves the problem of poor reproducibility and repeatability caused by the reliance of testing personnel on experience for qualitative and quantitative analysis; and this method does not have high requirements for peak separation, which can appropriately shorten the time for spectral acquisition.

[0092] In summary, the similarity calculation method for viscous tobacco flavorings provided by this invention has the advantages of high accuracy, high efficiency, wide applicability, and low degree of human experience involvement.

[0093] The technical solution of the present invention will be further described below with reference to specific embodiments:

[0094] Example 1

[0095] Ten batches of viscous flavorings from a certain cigarette brand were selected as test samples, and their information is listed in Table 2. An ethanol:water solution of 1:1 (volume ratio) was prepared as a diluent. 0.5g of flavoring (accurate to 0.1mg) was weighed and placed in a 15mL centrifuge tube, and 10mL of the diluent was added to dissolve it. All samples were placed in an ultrasonic water bath and sonicated for 30 minutes.

[0096] Table 2 Information on Flavor Samples

[0097]

[0098] Each type of viscous fragrance sample was analyzed by liquid chromatography-mass spectrometry (LC-MS) to obtain the corresponding LC-MS spectra. The analytical conditions were as follows:

[0099] Chromatographic conditions

[0100] The chromatographic column was an Agilent ZORBAX Eclipse Plus C18 column (2.1×50m 1.8-Micron); the mobile phase was 0.1% formic acid water (A)-acetonitrile (B), gradient elution, the elution conditions are shown in Table 3, the flow rate was 0.4 mL / min, the column temperature was 40℃, and the injection volume was 1 μL.

[0101] Table 3 LC-MS Elution Conditions

[0102]

[0103] Mass spectrometry conditions

[0104] Electrospray ionization source, positive ion mode, scanning range (m / z) 100-2000, nozzle voltage 1500V, capillary voltage 4000V, atomizing gas flow rate 12L / min, drying gas temperature 300℃, sheath gas temperature 250℃, sheath gas flow rate 11L / min.

[0105] The acquired LC-MS spectra were processed as follows:

[0106] 1) Convert the ten spectra into two-dimensional data. Each two-dimensional data has 8115 rows of time series and 1900 columns of mass-to-nucleus ratio series. Due to the large amount of data, only a portion of the two-dimensional data of flavor A1 is shown in this embodiment. The partial two-dimensional data of the LC spectrum of flavor A1 is listed in Table 4. Then, the abundance information in the two-dimensional data is extracted (i.e., the header of Table 4 is deleted, and the data in bold in Table 4 is retained) to obtain the corresponding plaintext data matrix.

[0107] Table 4. Partial two-dimensional data of the LC-MS spectrum of fragrance A1.

[0108]

[0109] 2) For each plaintext data matrix, the background noise of the spectrum is calculated and subtracted using the moving median method, with the window width set to 4000, to obtain the corresponding denoised data matrix.

[0110] 3) Set the time window length to 200 (converted according to the shifter sampling rate, the retention time is about 50s), and the step size to 100 (the retention time is about 25s). According to the aforementioned steps S14 to S16, use GNU OCTAVE data analysis software to process and calculate each denoised data matrix to obtain the cosine value cos(i), abundance tms(i), and similarity weight w(i) of each time window for fragrances A1 to A10. According to step S17, compare the above fragrances pairwise to calculate the similarity of different batches of fragrance A. The results are shown in Table 5.

[0111] Table 5. Similarity of different batches of flavoring A

[0112]

[0113]

[0114] As shown in Table 5, the similarity of the flavorings prepared in different batches was higher than 92%, indicating good formulation stability.

[0115] Corresponding to the method embodiment, another aspect of the present invention also provides a similarity calculation device for viscous tobacco flavorings and fragrances. Please refer to [link to relevant documentation]. Figure 3 The device is for use with Figure 1 The apparatus corresponding to the method described in the corresponding embodiment is implemented through a virtual device. Figure 1 In the corresponding embodiments, the virtual modules constituting the similarity calculation device for viscous tobacco flavorings can be executed by electronic devices, such as network devices, terminal devices, or servers. Specifically, the similarity calculation device for viscous tobacco flavorings in this embodiment includes:

[0116] Data acquisition and conversion unit 01 is used to acquire LC-MS spectra of two viscous tobacco flavorings and fragrances that have undergone dilution pretreatment under the same conditions; and convert the spectra into two-dimensional plaintext data matrices to obtain plaintext data matrices dat1[m,n] and dat2[m,n] corresponding to the two viscous tobacco flavorings and fragrances; the elements of the plaintext data matrices are abundance, n is the number of sequences on the mass-to-nucleus ratio sequence, and m is the number of sequences on the retention time series;

[0117] Data processing unit 02 is used to perform noise reduction processing on the n mass ratios in the plaintext data matrix dat1[m,n] on the retained time series to obtain a noise-reduced data matrix dat1_dn[m,n]. The same noise reduction processing is performed on the plaintext data matrix dat2[m,n] to obtain a noise-reduced data matrix dat2_dn[m,n]. Then, the noise-reduced data matrix dat1_dn[m,n] is divided into multiple time windows on the retained time series. Let the window width of each time window be 2t, the step size of each time window be half the window width t, and the number of time windows be wn. At this time, the data matrix of the i-th time window is dat1_wd. i [2t, n], 1 ≤ i ≤ wn; then for the data matrix dat1_wd of the i-th time window. i The calculation is performed on [2t, n] to obtain the mass-to-nucleus ratio abundance and dat1_wd for each mass-to-nucleus ratio in the i-th time window. i _ms, and, obtain the single-time-window abundance and dat1_wd of the i-th time window. i _tms; calculate the sum of the abundance of each time window for wn time windows to obtain the total abundance of all time windows corresponding to the plaintext data matrix dat1[m,n], and then calculate the denoised data matrix dat2_dn[m,n] in the same way to obtain the mass-to-nucleus ratio abundance of each mass-to-nucleus ratio in the i-th time window corresponding to the plaintext data matrix, dat2_wd. i _ms, the single-time-window abundance of the i-th time window, and dat2_wd i _tms, and abundance for all time windows and dat2_tms;

[0118] Similarity calculation unit 03 is used to calculate the similarity based on the mass-to-nucleus ratio and dat1_wd. i _ms and dat2_wd i The cosine value cos(i) of the i-th time window is calculated using _ms, and the abundance of the single time window and dat1_wd are used as the basis for the calculation. i _tms and dat2_wd i The abundance ratio tms(i) for the i-th time window is calculated using _tms; then, the abundance ratio for the single time window and dat1_wd are used as the basis for the comparison. i _tms and dat2_wd i _tms, and the abundance of all time windows and dat1_tms and dat2_tms are used to calculate the similarity weight w(i) of the i-th time window; finally, the similarity sml of the spectra of the two viscous tobacco flavorings is calculated based on the cosine value cos(i), abundance ratio tms(i) and similarity weight w(i) of the i-th time window, where i takes the value from 1 to wn.

[0119] It should be noted that the specific implementation methods and technical effects in the embodiments of the present invention can be referred to Figure 1 The corresponding methods will not be elaborated here.

[0120] Corresponding to the method embodiments, this invention also provides a similarity calculation device for viscous tobacco flavorings and fragrances. This device includes: a memory for storing computer programs; and a processor for calling and executing the computer programs to implement the steps of the method described in the above embodiments. This similarity calculation device for viscous tobacco flavorings and fragrances can be a terminal, a server, etc. The server can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The terminal can be a smartphone, tablet, laptop, desktop computer, etc., but is not limited to these.

[0121] An example diagram of the hardware structure block diagram of the similarity calculation device for viscous tobacco flavorings provided in this application embodiment is shown below. Figure 4 As shown, it may include:

[0122] Processor 1, communication interface 2, memory 3, and communication bus 4;

[0123] The processor 1, communication interface 2, and memory 3 communicate with each other via communication bus 4.

[0124] Optionally, communication interface 2 can be an interface of a communication module, such as the interface of a GSM module;

[0125] Processor 1 may be a central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of this application.

[0126] Memory 3 may include high-speed RAM memory, and may also include non-volatile memory, such as at least one disk storage device.

[0127] Specifically, processor 1 is used to execute the computer program stored in memory 3 to perform the following steps:

[0128] S11. Using the same diluting solvent and the same dilution ratio, perform dilution pretreatment on two viscous tobacco flavorings and fragrances respectively; the diluting solvent is: a mixed solvent of ethanol and water, or a mixed solvent of ethanol, water and propylene glycol.

[0129] S12. Under the same conditions, obtain LC-MS spectra of two viscous tobacco flavorings and fragrances that have undergone dilution pretreatment; convert the LC-MS spectra into two-dimensional plaintext data matrices to obtain plaintext data matrices dat1[m,n] and dat2[m,n] corresponding to the two viscous tobacco flavorings and fragrances; the elements of the plaintext data matrices are abundances, n is the number of sequences on the mass-to-nucleus ratio sequence, and m is the number of sequences on the retention time series;

[0130] S13. Perform noise reduction processing on the retained time series for the n mass-nucleus ratios in the plaintext data matrix dat1[m,n] to obtain the noise-reduced data matrix dat1_dn[m,n]; perform noise reduction processing on the plaintext data matrix dat2[m,n] in the same way to obtain the noise-reduced data matrix dat2_dn[m,n];

[0131] S14. Divide the denoised data matrix dat1_dn[m, n] into multiple time windows on the retained time series; let the window width of each time window be 2t, the step size of each time window be half the window width t, and the number of time windows be wn. At this time, the data matrix of the i-th time window is dat1_wd. i [2t, n], 1 ≤ i ≤ wn;

[0132] For the data matrix dat1_wd of the i-th time window i The calculation is performed on [2t, n] to obtain the mass-to-nucleus ratio abundance and dat1_wd for each mass-to-nucleus ratio in the i-th time window. i _ms, and, obtain the single-time-window abundance and dat1_wd of the i-th time window. i _tms;

[0133] Calculate the sum of the abundance of each time window for wn time windows to obtain the total abundance of all time windows corresponding to the plaintext data matrix dat1[m,n] and dat1_tms;

[0134] S15. Referring to step S14, calculate the denoised data matrix dat2_dn[m, n] to obtain the mass-nucleus ratio abundance and dat2_wd of each mass-nucleus ratio in the i-th time window corresponding to the plaintext data matrix. i _ms, the single-time-window abundance of the i-th time window, and dat2_wd i _tms, and abundance for all time windows and dat2_tms;

[0135] S16, through the mass-to-nucleus ratio abundance and dat1_wd i _ms and dat2_wd i_ms calculates the cosine value cos(i) of the i-th time window;

[0136] Through the single time window abundance and dat1_wd i _tms and dat2_wd i _tms calculates the abundance ratio tms(i) for the i-th time window;

[0137] Through the single time window abundance and dat1_wd i _tms and dat2_wd i _tms, and the abundance of all time windows and dat1_tms and dat2_tms are used to calculate the similarity weight w(i) of the i-th time window;

[0138] S17. Calculate the similarity sml of the spectra of two viscous tobacco flavorings and fragrances based on the cosine value cos(i), abundance ratio tms(i), and similarity weight w(i) of the i-th time window, where i ranges from 1 to wn.

[0139] The above-described product can execute the method provided in the embodiments of the present invention, and has the corresponding functional modules and beneficial effects of executing the method. Technical details not described in detail in this embodiment can be found in the similarity calculation method for viscous tobacco flavorings provided in the embodiments of the present invention.

[0140] In this embodiment of the invention, a computer-readable storage medium is also provided, which can store a program suitable for execution by a processor, the program being used to perform the following steps:

[0141] S11. Using the same diluting solvent and the same dilution ratio, perform dilution pretreatment on two viscous tobacco flavorings and fragrances respectively; the diluting solvent is: a mixed solvent of ethanol and water, or a mixed solvent of ethanol, water and propylene glycol.

[0142] S12. Under the same conditions, obtain LC-MS spectra of two viscous tobacco flavorings and fragrances that have undergone dilution pretreatment; convert the LC-MS spectra into two-dimensional plaintext data matrices to obtain plaintext data matrices dat1[m,n] and dat2[m,n] corresponding to the two viscous tobacco flavorings and fragrances; the elements of the plaintext data matrices are abundances, n is the number of sequences on the mass-to-nucleus ratio sequence, and m is the number of sequences on the retention time series;

[0143] S13. Perform noise reduction processing on the retained time series for the n mass-nucleus ratios in the plaintext data matrix dat1[m,n] to obtain the noise-reduced data matrix dat1_dn[m,n]; perform noise reduction processing on the plaintext data matrix dat2[m,n] in the same way to obtain the noise-reduced data matrix dat2_dn[m,n];

[0144] S14. Divide the denoised data matrix dat1_dn[m, n] into multiple time windows on the retained time series; let the window width of each time window be 2t, the step size of each time window be half the window width t, and the number of time windows be wn. At this time, the data matrix of the i-th time window is dat1_wd. i [2t, n], 1 ≤ i ≤ wn;

[0145] For the data matrix dat1_wd of the i-th time window i The calculation is performed on [2t, n] to obtain the mass-to-nucleus ratio abundance and dat1_wd for each mass-to-nucleus ratio in the i-th time window. i _ms, and, obtain the single-time-window abundance and dat1_wd of the i-th time window. i _tms;

[0146] Calculate the sum of the abundance of each time window for wn time windows to obtain the total abundance of all time windows corresponding to the plaintext data matrix dat1[m,n] and dat1_tms;

[0147] S15. Referring to step S14, calculate the denoised data matrix dat2_dn[m, n] to obtain the mass-nucleus ratio abundance and dat2_wd of each mass-nucleus ratio in the i-th time window corresponding to the plaintext data matrix. i _ms, the single-time-window abundance of the i-th time window, and dat2_wd i _tms, and abundance for all time windows and dat2_tms;

[0148] S16, through the mass-to-nucleus ratio abundance and dat1_wd i _ms and dat2_wd i _ms calculates the cosine value cos(i) of the i-th time window;

[0149] Through the single time window abundance and dat1_wd i _tms and dat2_wd i _tms calculates the abundance ratio tms(i) for the i-th time window;

[0150] Through the single time window abundance and dat1_wd i _tms and dat2_wd i _tms, and the abundance of all time windows and dat1_tms and dat2_tms are used to calculate the similarity weight w(i) of the i-th time window;

[0151] S17. Calculate the similarity sml of the spectra of two viscous tobacco flavorings and fragrances based on the cosine value cos(i), abundance ratio tms(i), and similarity weight w(i) of the i-th time window, where i ranges from 1 to wn.

[0152] The above-described product can execute the methods provided in the embodiments of the present invention, and has the corresponding functional modules and beneficial effects for executing the methods. Technical details not described in detail in this embodiment can be found in the methods provided in other embodiments of the present invention.

[0153] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0154] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. Furthermore, the couplings or direct couplings or communication connections shown or discussed may be indirect couplings or communication connections through interfaces, devices, or units, and may be electrical, mechanical, or other forms.

[0155] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0156] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0157] It should be understood that in the embodiments of this application, the various embodiments and features can be combined with each other to solve the aforementioned technical problems.

[0158] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0159] The above description of the disclosed embodiments enables those skilled in the art to make or use this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for calculating the similarity of viscous tobacco flavorings, characterized in that, Including the following steps: S11. Using the same diluting solvent and the same dilution ratio, perform dilution pretreatment on two viscous tobacco flavorings and fragrances respectively; the diluting solvent is: a mixed solvent of ethanol and water, or a mixed solvent of ethanol, water and propylene glycol. S12. Under the same conditions, obtain LC-MS spectra of two viscous tobacco flavorings and fragrances that have undergone dilution pretreatment; convert the LC-MS spectra into two-dimensional plaintext data matrices to obtain plaintext data matrices dat1[m,n] and dat2[m,n] corresponding to the two viscous tobacco flavorings and fragrances; the elements of the plaintext data matrices are abundances, n is the number of sequences on the mass-to-nucleus ratio sequence, and m is the number of sequences on the retention time series; S13. Perform noise reduction processing on the retained time series for the n mass-nucleus ratios in the plaintext data matrix dat1[m,n] to obtain the noise-reduced data matrix dat1_dn[m,n]; perform noise reduction processing on the plaintext data matrix dat2[m,n] in the same way to obtain the noise-reduced data matrix dat2_dn[m,n]; S14. Divide the denoised data matrix dat1_dn[m, n] into multiple time windows on the retained time series; let the window width of each time window be 2t, the step size of each time window be half the window width t, and the number of time windows be wn. At this time, the data matrix of the i-th time window is dat1_wd. i [2t, n], 1 ≤ i ≤ wn; For the data matrix dat1_wd of the i-th time window i The calculation is performed on [2t, n] to obtain the mass-to-nucleus ratio abundance and dat1_wd for each mass-to-nucleus ratio in the i-th time window. i _ms, and, obtain the single-time-window abundance and dat1_wd of the i-th time window. i _tms; Calculate the sum of the abundance of each time window for wn time windows to obtain the total abundance of all time windows corresponding to the plaintext data matrix dat1[m,n] and dat1_tms; S15. Referring to step S14, calculate the denoised data matrix dat2_dn[m, n] to obtain the mass-nucleus ratio abundance and dat2_wd of each mass-nucleus ratio in the i-th time window corresponding to the plaintext data matrix. i _ms, the single-time-window abundance of the i-th time window, and dat2_wd i _tms, and abundance and dat for all time windows 2_ tms; S16, through the mass-to-nucleus ratio abundance and dat 1_ wd i _ms and dat 2_ wd i _ms calculates the cosine value cos(i) of the i-th time window; Through the single time window abundance and dat1_wd i _tms and dat 2_ wd i _tms calculates the abundance ratio tms(i) for the i-th time window; Through the single time window abundance and dat1_wd i _tms and dat2_wd i _tms, and the abundance of all time windows and dat1_tms and dat2_tms are used to calculate the similarity weight w(i) of the i-th time window; S17. Calculate the similarity sml of the spectra of two viscous tobacco flavorings and fragrances based on the cosine value cos(i), abundance ratio tms(i), and similarity weight w(i) of the i-th time window, where i ranges from 1 to wn.

2. The similarity calculation method according to claim 1, wherein the diluting solvent is a mixed solvent of ethanol and water with a volume ratio of (0.2-5):1, or a mixed solvent of ethanol, water and propylene glycol with a volume ratio of (0.2-5):(0.2-5):

1.

3. The similarity calculation method according to claim 1, characterized in that, In step S14, the mass-to-nucleus ratio abundance and dat1_wd for each mass-to-nucleus ratio in the i-th time window are... i _ms is calculated as follows: For the data matrix dat1_wd of the i-th time window i Summing the data in the [2t, n] time series yields an n-dimensional vector dat1_wd. i _ms[n] represents the abundance of each mass-to-nucleus ratio in the i-th time window, and dat1_wd. i _ms; In step S14, the single-time-window abundance of the i-th time window and dat1_wd i _tms is calculated as follows: For dat1_wd i Summation of data on mass-to-nucleus ratio sequences in _ms Get dat 1_ wd i _tms represents the sum of abundance in a single time window for the i-th time window.

4. The similarity calculation method according to claim 3, characterized in that, In step S16, according to formula I, the mass-to-nucleus ratio abundance and dat1_wd are used to determine the mass-to-nucleus ratio abundance. i _ms and dat2_wd i _ms calculates the cosine value cos(i) of the i-th time window:

5. The similarity calculation method according to claim 4, characterized in that, In step S16, according to Equation II, the abundance of the single time window and dat1_wd are used. i _tms and dat2_wd i _tms calculates the abundance ratio tms(i) for the i-th time window; 6. The similarity calculation method according to claim 5, characterized in that, In step S16, according to Equation III, the abundance of the single time window and dat1_wd are used. i _tms and dat2_wd i _tms, and the abundance of all time windows and dat1_tms and dat2_tms are used to calculate the similarity weight w(i) of the i-th time window; 7. The method for calculating the similarity of viscous tobacco flavorings and fragrances according to claim 6, characterized in that, In step S17, the similarity sml of the spectra of the two viscous tobacco flavorings is calculated according to formula IV based on the cosine value cos(i), abundance ratio tms(i), and similarity weight w(i) of the i-th time window.

8. A similarity calculation device for viscous tobacco flavorings, characterized in that, include: The data acquisition and conversion unit is used to acquire LC-MS spectra of two diluted and pretreated viscous tobacco flavorings under the same conditions; The spectrum is then converted into a two-dimensional plaintext data matrix to obtain plaintext data matrices dat1[m,n] and dat2[m,n] corresponding to two viscous tobacco flavorings; the elements of the plaintext data matrix are abundance, n is the number of sequences on the mass-to-nucleus ratio sequence, and m is the number of sequences on the retention time series; The data processing unit performs noise reduction processing on the n mass-kernel ratios in the plaintext data matrix dar1[m,n] on the retained time series to obtain a noise-reduced data matrix dat1_dn[m,n]. It then performs noise reduction processing on the plaintext data matrix dat2[m,n] in the same manner to obtain a noise-reduced data matrix dat2_dn[m,n]. The noise-reduced data matrix dat1_dn[m,n] is then divided into multiple time windows on the retained time series. Let the window width of each time window be 2t, the step size of each time window be half the window width t, and the number of time windows be wn. At this time, the data matrix of the i-th time window is dat1_wd. i [2t, n], 1 ≤ i ≤ wn; then for the data matrix dat1_wd of the i-th time window. i The calculation is performed on [2t, n] to obtain the mass-to-nucleus ratio abundance and dat1_wd for each mass-to-nucleus ratio in the i-th time window. i _ms, and, obtain the single-time-window abundance and dat1_wd of the i-th time window. i _tms; calculate the sum of the abundance of each time window for wn time windows to obtain the total abundance of all time windows corresponding to the plaintext data matrix dat1[m,n], and then calculate the denoised data matrix dat2_dn[m,n] in the same way to obtain the mass-to-nucleus ratio abundance of each mass-to-nucleus ratio in the i-th time window corresponding to the plaintext data matrix, dat2_wd. i _ms, the single-time-window abundance of the i-th time window, and dat2_wd i _tms, and abundance for all time windows and dat2_tms; The similarity calculation unit is used to calculate the similarity based on the mass-to-nucleus ratio abundance and dat1_wd. i _ms and dar2_wd i The cosine value cos(i) of the i-th time window is calculated using _ms, and the abundance of the single time window and dat1_wd are used as the basis for the calculation. i _tms and dat2_wd i The abundance ratio tms(i) for the i-th time window is calculated using _tms; then, the abundance ratio for the single time window and dat1_wd are used as the basis for the comparison. i _tms and dat2_wd i _tms, and the abundance of all time windows and dat1_tms and dat2_tms are used to calculate the similarity weight w(i) of the i-th time window; finally, the similarity sml of the spectra of the two viscous tobacco flavorings is calculated based on the cosine value cos(i), abundance ratio tms(i) and similarity weight w(i) of the i-th time window, where i takes the value from 1 to wn.

9. A similarity calculation device for viscous tobacco flavorings, characterized in that, include: Memory, used to store computer programs; A processor for invoking and executing the computer program to implement the steps of the method as claimed in any one of claims 1 to 7.

10. A computer-readable storage medium comprising a software program adapted to be executed by a processor of the steps of the method as claimed in any one of claims 1 to 7.

Citation Information

Patent Citations

  • GC-QTOF detection method for aldehydes and ketones flavor components in tobaccos and tobacco products

    CN111208310A

  • Developement of metabolic biomarkers and discrimination model for determining origin of white rice

    KR1020180024293A