Method for determining nicotine content in tobacco tar based on near infrared spectrum
The spectral characteristics of the e-liquid sample were extracted through near-infrared spectroscopy technology, and the regression model was trained using the gradient hoist algorithm to solve the complexity and low efficiency of the nicotine content determination in the prior art, achieving efficient and accurate nicotine content determination.
Patent Information
- Application Number
- CN202510030881.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-08
- Publication Date
- 2025-05-06
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The prior art is difficult to quickly and accurately determine the nicotine content in e-liquid, and traditional chemical analysis methods are complex in operation, long time and high in cost, and it is difficult to extract effective features in the complexity of near-infrared spectral data.
Near-infrared spectral data of e-liquid samples were collected by a near-infrared spectrometer, spectral features such as spectral step curvature, microscopic peak polarity, frequency domain tuning fingerprint, spectral layer texture structure and spectral information complexity characteristics were extracted, and the regression model was trained using a gradient lift algorithm, and the measurement accuracy was improved through cross-verification and hyperparameter optimization.
It realizes rapid and accurate determination of the nicotine content in e-liquid, improves detection efficiency and accuracy, avoids complex chemical reagents and time-consuming analysis steps, and has the characteristics of non-destructive testing.
Smart Images

Figure CN119935943A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of cigarette oil detection, in particular to a method for measuring the nicotine content in cigarette oil based on near infrared spectroscopy. Background Art
[0002] With the rapid development of the e-cigarette market, the quality and safety of e-liquids have gradually become the focus of the industry. Nicotine is one of the main components of e-liquids, and its content directly affects the user experience of e-cigarettes and the health of consumers. Therefore, accurate and rapid determination of nicotine content in e-liquids is of great significance for production control and product quality monitoring.
[0003] At present, the traditional determination method of nicotine content in e-liquid mainly relies on chemical analysis technology, such as high performance liquid chromatography and gas chromatography. Although these methods have high accuracy, they are difficult to meet the needs of efficient, rapid and large-scale e-liquid quality testing due to their complex operation, long analysis time and high cost.
[0004] In recent years, near-infrared spectroscopy has been widely used in the fields of food, medicine, environment, etc. as a non-destructive, efficient and rapid analytical method. Near-infrared spectroscopy can directly reflect the chemical composition and molecular structure of substances, and has the advantages of no pretreatment and fast analysis process. Especially in the tobacco and e-liquid industry, near-infrared spectroscopy is used in the analysis of the components of tobacco leaves and e-liquids, and can provide an effective technical means for the determination of nicotine content. However, due to the complexity of near-infrared spectral data, how to extract effective features from a large amount of spectral data and accurately establish a nicotine content determination model is still a technical problem. Summary of the invention
[0005] Based on the above-mentioned shortcomings of the prior art, the purpose of the present invention is to provide a method for measuring the nicotine content in tobacco oil based on near infrared spectroscopy to solve the above technical problems.
[0006] To achieve the above object, the present invention provides the following technical solution: a method for determining the nicotine content in tobacco oil based on near infrared spectroscopy, comprising:
[0007] S1: Collect near-infrared spectrum data of the e-liquid sample through a near-infrared spectrometer to obtain the reflection intensity of the spectrum data within the near-infrared wavelength range;
[0008] S2: extracting spectral features according to the reflection intensity of spectral data in the near-infrared wavelength range, wherein the spectral features include spectral step curvature features, microscopic spectral peak polarity features, frequency domain tuning fingerprint features, spectral layer texture structure features and spectral information complexity features;
[0009] S3: Using the gradient boosting algorithm, the regression model is trained through the extracted spectral features to obtain a nicotine content determination model;
[0010] S4: Use cross-validation method to verify the trained measurement model, and improve the measurement accuracy of the measurement model through hyperparameter optimization;
[0011] S5: inputting the near infrared spectrum data of the tested e-liquid sample into the trained determination model to obtain the nicotine content determination value of the tested e-liquid sample.
[0012] The present invention is further configured that the processing of the e-liquid sample includes:
[0013] The e-liquid sample is prepared to a preset concentration, and when the e-liquid concentration is too high, a solvent is added to dilute it, wherein the solvent includes anhydrous ethanol;
[0014] Place the prepared e-liquid sample into a transparent container;
[0015] A transparent container containing the e-liquid sample is placed in a sample pool of a near-infrared spectrometer, and near-infrared spectrum data of the e-liquid sample is collected by the near-infrared spectrometer.
[0016] The present invention is further configured such that the calculation logic of the spectral step curvature characteristic is: Among them, C λ (t) is the spectral step curvature characteristic of the e-liquid sample t, t is the e-liquid sample number, λ is the wavelength, n is the number of spectral data points, S(λ i ,t) is the smoke oil sample t at wavelength λ i The reflection intensity at At wavelength λ i At the spectral reflection intensity S(λ i ,t) The second derivative with respect to wavelength, used to measure the curvature of the spectrum, At wavelength λ i At the spectral reflection intensity S(λ i ,t) is the first-order derivative with respect to wavelength, which is used to indicate the rate of change of the spectrum. γ1, γ2 and γ3 are exponential adjustment parameters used to adjust the sensitivity of the feature.
[0017] The present invention is further configured such that the calculation logic of the microscopic spectrum peak polarity feature is: Among them, L λ (t) is the microscopic spectral peak polarity characteristic of the e-liquid sample t, t is the e-liquid sample number, λ is the wavelength, n is the number of spectral data points, S(k,t) is the reflection intensity of the e-liquid sample t at wavelength k, and when i is 1, the wavelength k is λ1, when i is n, the wavelength k is λ n , β1 and β2 are adjustment parameters used to adjust the local peak intensity.
[0018] The present invention is further configured such that the calculation logic of the frequency domain tuning fingerprint feature is: Among them, F λ (t) is the frequency domain tuning fingerprint feature of the e-liquid sample t, t is the e-liquid sample number, λ is the wavelength, n is the number of spectral data points, S(λ i ,t) is the smoke oil sample t at wavelength λ i where δ1 is the reflection intensity at , F(·) is the Fourier transform, and δ1 and δ2 are adjustment parameters used to control the intensity and sensitivity of the frequency domain features.
[0019] The present invention is further configured such that the calculation logic of the spectral layer texture structure feature is: Among them, T λ (t) is the spectral texture structure feature of the e-liquid sample t, t is the e-liquid sample number, λ is the wavelength, n is the number of spectral data points, S(λ i ,t) is the smoke oil sample t at wavelength λ i where m is the total number of wavelengths in the spectrum, and α1 and α2 are parameters for adjusting the complexity of the texture, which are used to quantify the changes between different wavelength bands.
[0020] The present invention is further configured such that the calculation logic of the spectral information complexity feature is: Among them, H λ (t) is the spectral information complexity characteristic of the e-liquid sample t, t is the e-liquid sample number, λ is the wavelength, n is the number of spectral data points, S(λ i ,t) is the smoke oil sample t at wavelength λ i The reflection intensity at Z λ (t) is the normalization constant of near-infrared spectral data, and the calculation logic is:
[0021] The present invention is further configured that step S3 comprises:
[0022] Acquire near infrared spectral data of e-liquid samples collected by a historical near infrared spectrometer and corresponding nicotine content measurement values, extract spectral features according to the reflection intensity of the spectral data within the near infrared wavelength range, set the spectral features and the nicotine content measurement values as a data set, and divide the data set into a training set and a test set;
[0023] A regression model is constructed using a gradient boosting algorithm, hyperparameters of the regression model are initialized, the spectral features are set as inputs of the regression model, the nicotine content measurement values are set as target outputs of the regression model, and the regression model is trained using a training set, wherein the hyperparameters include a learning rate, the number of trees, the maximum depth of each tree, the number of leaf nodes, and a subsample ratio;
[0024] The regression model gradually optimizes the weights of the decision tree through multiple iterations, and finally obtains a model composed of multiple weighted trees, which can effectively fit the training data and determine the nicotine content.
[0025] The present invention is further configured that step S4 comprises:
[0026] K-fold cross validation is used to evaluate the performance of the regression model on different spectral data segmentations. Cross validation effectively detects whether the regression model is overfitting or underfitting and helps select the optimal hyperparameters.
[0027] Use grid search or Bayesian optimization for hyperparameter tuning;
[0028] Prevent overfitting by setting early stopping. The early stopping strategy stops training when the error of the validation set no longer decreases, ensuring that the regression model does not overfit on the training set.
[0029] The trained regression model is measured using the test set, and the performance of the regression model is evaluated using evaluation indicators, wherein the evaluation indicators include mean square error, root mean square error, and mean absolute error.
[0030] The present invention provides a method for determining the nicotine content in a cigarette oil based on near-infrared spectroscopy. The method comprises the following steps: collecting near-infrared spectral data of a cigarette oil sample by a near-infrared spectrometer to obtain the reflection intensity of the spectral data within the near-infrared wavelength range; extracting spectral features according to the reflection intensity of the spectral data within the near-infrared wavelength range, wherein the spectral features include spectral step curvature features, microscopic spectral peak polarity features, frequency domain tuning fingerprint features, spectral layer texture structure features and spectral information complexity features; adopting a gradient boosting algorithm to train a regression model through the extracted spectral features to obtain a nicotine content determination model; adopting a cross-validation method to verify the trained determination model, and improving the determination accuracy of the determination model through hyperparameter optimization; inputting the near-infrared spectral data of the cigarette oil sample to be tested into the trained determination model to obtain the nicotine content determination value of the cigarette oil sample to be tested, and the generated beneficial effects include:
[0031] 1. High efficiency and rapidity: Traditional methods for determining nicotine content, such as high-performance liquid chromatography or gas chromatography, are complex to operate and take a long time to analyze, and require expensive equipment and reagents. In contrast, the present invention uses near-infrared spectroscopy technology for analysis, which can complete the determination of e-liquid samples in a short time without the need for complex chemical reagents and time-consuming analysis steps, greatly improving the detection efficiency;
[0032] 2. Nondestructive testing: Near infrared spectroscopy is a nondestructive testing method that can directly analyze e-liquid samples without destroying them. It is particularly important for real-time quality control during the production process, which can reduce sample loss and improve production efficiency.
[0033] 3. High-precision determination: The present invention adopts machine learning methods such as gradient boosting algorithm, and constructs a high-precision regression model by extracting multiple spectral features from near-infrared spectral data, including spectral step curvature features, microscopic spectral peak polarity features, frequency domain tuning fingerprint features, spectral layer texture structure features and spectral information complexity features, which can accurately determine the nicotine content in the e-liquid. Through cross-validation and hyperparameter optimization, the accuracy of the determination model is further improved, ensuring that the determination results of nicotine content are more reliable and accurate.
[0034] The above description is only an overview of the technical solution of the present application. In order to more clearly understand the technical means of the present application, it can be implemented in accordance with the contents of the specification. In order to make the above and other purposes, features and advantages of the present application more obvious and easy to understand, the specific implementation methods of the present application are listed below. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for describing the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work. In the drawings:
[0036] Figure 1 The present invention is a flowchart of a method for measuring nicotine content in e-liquid based on near infrared spectroscopy, which is an exemplary embodiment of the present invention. DETAILED DESCRIPTION
[0037] The following will describe the embodiments of the present invention with reference to the accompanying drawings and preferred embodiments. Those skilled in the art can easily understand other advantages and effects of the present invention from the contents disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present invention. It should be understood that the preferred embodiments are only for illustrating the present invention, not for limiting the scope of protection of the present invention.
[0038] It should be noted that the illustrations provided in the following embodiments are only schematic illustrations of the basic concept of the present invention, and thus the drawings only show components related to the present invention rather than being drawn according to the number, shape and size of components in actual implementation. In actual implementation, the type, quantity and proportion of each component may be changed arbitrarily, and the component layout may also be more complicated.
[0039] In the following description, numerous details are discussed to provide a more thorough explanation of the embodiments of the present invention. However, it is obvious to those skilled in the art that the embodiments of the present invention can be implemented without these specific details. In other embodiments, well-known structures and devices are shown in the form of block diagrams rather than in detail to avoid making the embodiments of the present invention difficult to understand.
[0040] Methods for determining nicotine content in e-liquid based on near-infrared spectroscopy, such as Figure 1 As shown, including:
[0041] S1: Collect near-infrared spectrum data of the e-liquid sample through a near-infrared spectrometer to obtain the reflection intensity of the spectrum data within the near-infrared wavelength range;
[0042] S2: extracting spectral features according to the reflection intensity of spectral data in the near-infrared wavelength range, wherein the spectral features include spectral step curvature features, microscopic spectral peak polarity features, frequency domain tuning fingerprint features, spectral layer texture structure features and spectral information complexity features;
[0043] S3: Using the gradient boosting algorithm, the regression model is trained through the extracted spectral features to obtain a nicotine content determination model;
[0044] S4: Use cross-validation method to verify the trained measurement model, and improve the measurement accuracy of the measurement model through hyperparameter optimization;
[0045] S5: inputting the near infrared spectrum data of the tested e-liquid sample into the trained determination model to obtain the nicotine content determination value of the tested e-liquid sample.
[0046] The present invention is further configured that the processing of the e-liquid sample includes:
[0047] The e-liquid sample is prepared to a preset concentration. When the concentration of the e-liquid is too high, a solvent is added for dilution, wherein the solvent includes anhydrous ethanol. Specifically, since the nicotine content of the e-liquid sample usually has a certain volatility, directly using a high-concentration sample for near-infrared spectral analysis may cause saturation of the spectral signal, thereby affecting the accuracy of the data. Therefore, before starting the analysis, the e-liquid sample needs to be diluted to an appropriate concentration, such as 1mg / ml, and the e-liquid sample is adjusted to a concentration range suitable for spectral analysis. If the concentration of the e-liquid sample is too high, a suitable solvent needs to be added for dilution. The solvent includes anhydrous ethanol, which has good solubility and has little interference with the spectral signal in the near-infrared band, thereby avoiding the measurement results being affected by the absorption of the solvent;
[0048] Put the prepared e-liquid sample into a transparent container; specifically, the role of the transparent container is to ensure that the near-infrared light can pass through the e-liquid sample and be received by the spectrometer. Since the near-infrared spectrometer needs to analyze the spectral characteristics of the substance through the sample, choosing a transparent container can avoid the optical interference of the container itself and ensure the accurate collection of the spectral signal;
[0049] The transparent container containing the e-liquid sample is placed in the sample pool of the near-infrared spectrometer, and the near-infrared spectrometer is used to collect the near-infrared spectrum data of the e-liquid sample; specifically, the prepared transparent container containing the e-liquid sample is placed in the sample pool of the near-infrared spectrometer, and the spectrometer emits a near-infrared beam to illuminate the sample and receives the reflected light of the sample. According to the intensity and wavelength of the reflected light of the sample, the spectrometer can record the absorption and reflection characteristics of the e-liquid sample at a specific wavelength, thereby obtaining near-infrared spectrum data. These data reflect the molecular structure, chemical composition and other information of the sample, and become the basis for subsequent analysis and modeling.
[0050] The present invention is further configured such that the calculation logic of the spectral step curvature characteristic is: Among them, C λ (t) is the spectral step curvature characteristic of the e-liquid sample t, t is the e-liquid sample number, λ is the wavelength, n is the number of spectral data points, S(λ i ,t) is the smoke oil sample t at wavelength λ i The reflection intensity at At wavelength λ i At the spectral reflection intensity S(λ i ,t) The second derivative with respect to wavelength, used to measure the curvature of the spectrum, At wavelength λ i At the spectral reflection intensity S(λ i ,t) The first-order derivative with respect to wavelength is used to indicate the rate of change of the spectrum. γ1, γ2 and γ3 are exponential adjustment parameters used to adjust the sensitivity of the feature. Specifically, the spectral step curvature feature reflects the spectral characteristics of the e-liquid sample through two main spectral derivatives. The two derivatives in the calculation logic represent the curvature and the rate of change of the spectrum, respectively. The absolute value of the second-order derivative is used to measure the degree of curvature of the spectral reflection intensity when the wavelength changes. γ1 is an exponential adjustment parameter used to adjust the sensitivity of the feature to curvature. The larger this parameter is, the greater the contribution of the change in the spectral curvature region to the feature value. It is the weighted sum of the first-order derivatives, which represents the comprehensiveness of the spectral change rate. γ2 adjusts the sensitivity of the first-order derivative, and γ3 adjusts the sensitivity of the entire feature. Through multiple weighted accumulation of the first-order derivatives, the intensity evaluation of the spectral change rate is obtained; the exponential adjustment parameter γ1 is used to adjust the sensitivity of the curvature feature, and the value range is [1,2]. The exponential adjustment parameter γ2 is used to adjust the degree of influence of the change rate on the feature, and the value range is [0.5,2]. The exponential adjustment parameter γ3 is used to adjust the sensitivity of the comprehensive feature, and the value range is [0.5,2]. Through the combination of the second-order derivative and the first-order derivative, the detailed changes and global trends in the spectral data can be fully captured, making the spectral feature extraction more accurate, and providing higher prediction accuracy for the determination of nicotine content.
[0051] The present invention is further configured such that the calculation logic of the microscopic spectrum peak polarity feature is: Among them, L λ (t) is the microscopic spectral peak polarity characteristic of the e-liquid sample t, t is the e-liquid sample number, λ is the wavelength, n is the number of spectral data points, S(k,t) is the reflection intensity of the e-liquid sample t at wavelength k, and when i is 1, the wavelength k is λ1, when i is n, the wavelength k is λ n , β1 and β2 are adjustment parameters used to adjust the local peak intensity. Specifically, the microscopic spectral peak polarity feature identifies the local peaks and their polarities in the spectrum. The polarity of the local peak reflects the degree of change in the spectral signal at a specific wavelength, helping to reveal the local chemical composition and microscopic changes of the substance; Traverse all spectral data points and calculate the polarity of the microscopic spike near each point one by one. For each spectral data point λ i , with its adjacent wavelength λ i-1 and λ i+1 As a reference, calculate the polarity of this point, that is, the difference between the spectral reflection intensity at this wavelength and the minimum reflection intensity of its adjacent wavelength; by |S(k,t)-min(S(λ i-1 ,t),S(λ i+1 ,t))|Calculate the difference between the reflection intensity of each spectral point and the minimum value of its neighboring points to obtain the polarity of the peak. The size of the polarity value reflects the degree of change of the spectral characteristics near the point, which helps to identify the subtle structure and peaks in the spectrum. The adjustment parameters β1 and β2 adjust the sensitivity of the polarity difference and the contribution of the local polarity characteristics respectively, and the value range is [0.5,2]. By analyzing the polarity of tiny peaks in the spectrum, subtle changes in the chemical composition of the e-liquid sample can be identified. The introduction of the polarity characteristics of microscopic spectral peaks can capture local changes that are difficult to identify by traditional methods, and improve the perception of changes in nicotine content.
[0052] The present invention is further configured such that the calculation logic of the frequency domain tuning fingerprint feature is: Among them, F λ (t) is the frequency domain tuning fingerprint feature of the e-liquid sample t, t is the e-liquid sample number, λ is the wavelength, n is the number of spectral data points, S(λ i ,t) is the smoke oil sample t at wavelength λ i The reflection intensity at , F(·) is the Fourier transform, δ1 and δ2 are adjustment parameters used to control the intensity and sensitivity of the frequency domain features. Specifically, the frequency domain tuning fingerprint feature converts the spectral data from the time domain (wavelength domain) to the frequency domain through Fourier transform. The frequency domain can reflect the periodic change characteristics of the signal and is very effective for analyzing repetitive patterns or periodic features in the spectrum. Frequency domain features help capture some patterns that are not easy to identify in the time domain, especially the weak periodic fluctuations in the spectrum; Fourier transform F(·) is a mathematical tool to transform a signal (such as spectral data) from the time domain (wavelength domain) to the frequency domain. Through Fourier transform, complex spectral signals can be decomposed into multiple frequency components, and then the periodic structure therein can be analyzed. F(S(λ i ,t)) represents the smoke oil sample t at wavelength λ i The spectral reflection intensity S(λ i ,t) Perform Fourier transform to obtain the characteristic information in the frequency domain, and then weight the amplitude of the Fourier transform result by adjusting the parameter δ1 to enhance or suppress the influence of the frequency domain characteristics. δ1 controls the sensitivity of the frequency domain characteristic intensity; Represents the relative intensity of the frequency domain feature, and normalizes it by the ratio of the square of the Fourier transform result to the maximum reflection intensity, which helps to eliminate the absolute difference in reflection intensity, making the frequency domain feature more stable and more comparable. Then the sensitivity of the normalization process is controlled by adjusting the parameter δ2; the adjustment parameter δ1 is used to adjust the intensity of the frequency domain feature, so that the contribution of the Fourier transform amplitude to the total feature can be enlarged or reduced as needed, and the value range is [0.5,2]. The adjustment parameter δ2 is used to adjust the sensitivity of the normalized frequency domain feature to ensure that during the normalization process, the frequency domain feature can still reflect the essential difference of the spectral data, and the value range is [0.5,2]. Through Fourier transform, the present invention can effectively extract periodic changes and repetitive patterns in the spectrum. This plays an important role in the analysis of trace components in tobacco oil, and can reveal subtle fluctuations that are usually not easy to detect in the time domain, and improve the accuracy of nicotine content determination.
[0053] The present invention is further configured such that the calculation logic of the spectral layer texture structure feature is: Among them, T λ(t) is the spectral texture structure feature of the e-liquid sample t, t is the e-liquid sample number, λ is the wavelength, n is the number of spectral data points, S(λ i ,t) is the smoke oil sample t at wavelength λ i The reflection intensity at , m is the total number of wavelengths in the spectrum, and α1 and α2 are parameters for adjusting the complexity of the texture, which are used to quantify the changes between different wavelength bands. Specifically, the spectral layer texture structure feature quantifies the texture structure of the spectrum by calculating the difference between adjacent wavelength points in the spectral data, which can reveal the change pattern between the spectral wavelength bands of the e-liquid sample and reflect the characteristic changes of the sample at different wavelengths; Traverse all spectral data points, and extract the local texture information of each spectral data segment by calculating the difference between each wavelength point and its adjacent wavelength points. Measure the texture structure changes between different wavelength bands. For each wavelength point λ j , calculate the difference between it and the adjacent wavelength point. S(λ j ,t)-S(λ j+1 ,t) Calculate the difference in reflection intensity between adjacent wavelength points, which reflects the characteristic difference of the sample at these wavelengths and quantifies the changing trend of the sample in different spectral regions; adjust the parameter α1 of texture complexity to control the influence of local wavelength band differences, and the value range is [1,2]. Adjust the parameter α2 of texture complexity to control the sensitivity of the overall feature, and the value range is [1,2]. By calculating the difference between adjacent wavelength points, the texture structure features in the spectral data can be effectively extracted, which is very effective for analyzing small changes in the e-liquid samples and helps to identify small differences in the nicotine content in the e-liquid.
[0054] The present invention is further configured such that the calculation logic of the spectral information complexity feature is: Among them, H λ (t) is the spectral information complexity characteristic of the e-liquid sample t, t is the e-liquid sample number, λ is the wavelength, n is the number of spectral data points, S(λ i ,t) is the smoke oil sample t at wavelength λ i The reflection intensity at Z λ (t) is the normalization constant of near-infrared spectral data, and the calculation logic is: Specifically, the spectral information complexity feature uses the concept of information entropy to quantify the complexity of spectral data. Information entropy is a mathematical method used to measure the amount of information or uncertainty. The entropy value reflects the distribution of reflection intensity of the sample spectrum at different wavelengths. The more dispersed the information, the higher the complexity, which means that the spectral features are richer. In order to avoid the absolute reflection intensity value in the spectral data affecting the complexity calculation, the spectral data needs to be normalized. The normalization process makes the sum of the reflection intensity at all wavelengths equal to 1, thereby eliminating the absolute reflection intensity differences between different samples or under different measurement conditions, and focusing on the relative changes between wavelengths. Indicates wavelength λ i The proportion of the reflection intensity at the wavelength to the total reflection intensity reflects the information contribution of the wavelength in the spectrum. The logarithm of the ratio of reflection intensity is used to calculate the amount of information. According to information theory, the complexity of a signal is proportional to its information entropy. The logarithm of the reflection intensity represents the average uncertainty of the information. The spectral information complexity feature effectively quantifies the complexity of the spectral data through the calculation of entropy. The introduction of information entropy enables this method to reflect the diversity and complexity of the spectrum in the e-liquid sample, thereby providing more information for the prediction of nicotine content.
[0055] The present invention is further configured that step S3 comprises:
[0056] Obtain the near-infrared spectral data and corresponding nicotine content measurement values of the historical near-infrared spectrometer collected e-liquid samples, extract spectral features according to the reflection intensity of the spectral data in the near-infrared wavelength range, set the spectral features and nicotine content measurement values as a data set, and divide the data set into a training set and a test set; specifically, obtain the spectral data and the corresponding nicotine content measurement values from known historical data. Each historical sample includes two parts of information: spectral data: reflection intensity from the near-infrared spectrometer, reflection value at different wavelengths; nicotine content measurement value: the actual nicotine content obtained according to experimental chemical analysis methods or other detection methods. Extract spectral features: Extract features reflecting spectral information from the near-infrared spectral data of each e-liquid sample, including spectral step curvature features, microscopic spectral peak polarity features, frequency domain tuning fingerprint features, spectral layer texture structure features, and spectral information complexity features; pair the extracted spectral features with the nicotine content measurement values of each sample to form a data set, and divide the data set into a training set (for training regression models) and a test set (for model verification). Usually, the training set accounts for about 70%-80%, and the test set accounts for 20%-30%. This division helps evaluate the model's ability to generalize to new data;
[0057] A regression model is constructed using a gradient boosting machine algorithm, the hyperparameters of the regression model are initialized, the spectral features are set as the input of the regression model, the nicotine content measurement value is set as the target output of the regression model, and the regression model is trained using a training set. The hyperparameters include a learning rate, the number of trees, the maximum depth of each tree, the number of leaf nodes, and the sub-sample ratio. Specifically, the construction of a gradient boosting machine regression model, the gradient boosting machine is a commonly used ensemble learning algorithm, which trains multiple weak regression models (usually decision trees) through multiple iterations, and weighted averages their prediction results to form a strong regression model. The gradient boosting machine minimizes the model error by optimizing the loss function. When training the gradient boosting machine regression model, some hyperparameters need to be set: learning rate: controls the contribution of each tree to the final result. A smaller learning rate requires more trees to get a good fit, but it can usually avoid overfitting. The number of trees refers to the number of decision trees in the GBM model. More trees can usually fit the training data better, but it may also cause overfitting. The maximum depth of each tree: limiting the depth of each tree can avoid the model from overfitting complex data. The number of leaf nodes: controls the number of leaf nodes of each tree, indirectly controlling the complexity of the model. The sub-sample ratio: randomly selects a part of the samples from the training set during each training to prevent overfitting and enhance the generalization ability of the model.
[0058] The regression model gradually optimizes the weights of the decision trees through multiple iterations, and finally obtains a model composed of multiple weighted trees, which can effectively fit the training data and measure the nicotine content. Specifically, the regression model corrects the error of the previous tree by training a new decision tree in each iteration. The iterative process will gradually adjust the weight of each tree, thereby improving the prediction accuracy of the overall regression model. Finally, the model optimized through multiple iterations will contain multiple weighted decision trees, which can effectively measure on new data and output the nicotine content of the sample to be tested.
[0059] The present invention is further configured that step S4 comprises:
[0060] K-fold cross validation is used to evaluate the performance of the regression model on different spectral data segmentations. Cross validation effectively detects whether the regression model is overfitting or underfitting, and helps select the optimal hyperparameters. Specifically, K-fold cross validation is a commonly used model validation method that aims to evaluate the generalization ability of the model. The specific process is as follows: divide all training data into K subsets (usually K is 5 or 10); each time, select a subset from the K subsets as the validation set, and the remaining subsets as the training set, repeat K times, and use a different subset as the validation set each time; the comprehensive performance of the model is obtained by evaluating the average of the K validation results. The advantage of this method is that different training sets and validation sets can be used in each round of training, which effectively reduces the risk of overfitting and underfitting and provides a more robust performance evaluation; cross validation effectively detects the performance of the model under different data partitions through multiple training and validation. If the model performs extremely well in a specific partition and performs poorly in other partitions, it may indicate that the model is overfitting; on the contrary, if the model performs poorly on all partitions, it means that the model is underfitting;
[0061] Use grid search or Bayesian optimization to tune hyperparameters; specifically, the performance of regression models is often affected by hyperparameter settings. By adjusting hyperparameters, the performance of the model can be optimized. Common tuning methods include: grid search, by setting a parameter range, exhaustively enumerating all possible hyperparameter combinations to find the best hyperparameters. Although the optimal combination can be found, the amount of calculation is large and it is suitable for situations with a small parameter range; Bayesian optimization, compared to grid search, Bayesian optimization constructs a probability model (usually a Gaussian process) to find the optimal solution in a smaller parameter space. Bayesian optimization reasonably infers the "optimal" position of the parameter space and gradually adjusts the parameters, making the optimization process more efficient, especially when the hyperparameter space is large;
[0062] Prevent overfitting by setting early stopping. The early stopping strategy stops training when the error of the validation set no longer decreases, ensuring that the regression model does not overfit on the training set. Specifically, during the training process, the model may experience a period of improvement, but at some point, the model will begin to overfit on the training set, and the performance on the validation set will deteriorate. To prevent this, you can use the early stopping strategy: during the training process, check the error of the model on the validation set after a certain number of iterations (for example, every 10 rounds). If the error of the validation set does not decrease (that is, the error begins to rise) in several iterations, stop training. This ensures that the model does not overfit on the training set, thereby improving the generalization ability of the model. For regression problems, early stopping can ensure that the model stops training after achieving the optimal effect on the validation set, preventing overfitting on the training set.
[0063] The trained regression model is measured using a test set, and the performance of the regression model is evaluated by evaluation indicators, wherein the evaluation indicators include mean square error, root mean square error and mean absolute error. Specifically, after the model is trained, an independent test set is needed to evaluate the performance of the model. The test set does not participate in the training process, so it can effectively test the performance of the model in practical applications. In order to comprehensively evaluate the performance of the regression model, multiple evaluation indicators are used, mean square error: the average value of the square of the difference between the model prediction value and the actual value. A smaller mean square error indicates that the model is more accurate; root mean square error: the square root of the mean square error, used to intuitively display the size of the prediction error, the unit is consistent with the target variable; mean absolute error: the average value of the absolute value of the difference between the predicted value and the actual value. A smaller mean absolute error indicates that the prediction result of the model is more accurate. Hyperparameter tuning is performed by using K-fold cross validation, grid search or Bayesian optimization, combined with an early stopping strategy, to ensure the efficiency and accuracy of the regression model. Comprehensively evaluating the model performance through a variety of evaluation indicators can effectively improve the generalization ability and prediction accuracy of the regression model, thereby achieving accurate determination of nicotine content.
[0064] The above embodiments can be implemented in whole or in part by software, hardware, firmware or any other combination. When implemented by software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website site, computer, server or data center to another website site, computer, server or data center by wired (e.g., infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that contains one or more available media sets. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium. The semiconductor medium can be a solid-state hard disk.
[0065] It should be understood that the term "and / or" in this article is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. A and B can be singular or plural. In addition, the character " / " in this article generally indicates that the associated objects before and after are in an "or" relationship, but it may also indicate an "and / or" relationship. Please refer to the context for specific understanding.
[0066] In this application, "at least one" means one or more, and "more than one" means two or more. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, or c can be represented by: a, b, c, ab, ac, bc, or abc, where a, b, c can be single or multiple.
[0067] It should be understood that in the various embodiments of the present application, the size of the serial numbers of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.
[0068] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0069] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0070] In the several embodiments provided in the present application, it should be understood that the disclosed system can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0071] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0072] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
[0073] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application can be essentially or partly embodied in the form of a software product that contributes to the prior art. The computer software product is stored in a storage medium and includes several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the methods described in the various embodiments of the present application. The aforementioned storage media include: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks or optical disks.
[0074] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art who is familiar with the present technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.
Claims
1. A method for determining the nicotine content in tobacco oil based on near infrared spectroscopy, characterized in that: include: S1: Collect near-infrared spectrum data of the e-liquid sample through a near-infrared spectrometer to obtain the reflection intensity of the spectrum data within the near-infrared wavelength range; S2: extracting spectral features according to the reflection intensity of spectral data in the near-infrared wavelength range, wherein the spectral features include spectral step curvature features, microscopic spectral peak polarity features, frequency domain tuning fingerprint features, spectral layer texture structure features and spectral information complexity features; S3: Using the gradient boosting algorithm, the regression model is trained through the extracted spectral features to obtain a nicotine content determination model; S4: Use cross-validation method to verify the trained measurement model, and improve the measurement accuracy of the measurement model through hyperparameter optimization; S5: inputting the near infrared spectrum data of the tested e-liquid sample into the trained determination model to obtain the nicotine content determination value of the tested e-liquid sample.
2. The method for measuring the nicotine content in tobacco oil based on near infrared spectroscopy according to claim 1, characterized in that: The processing of e-liquid samples includes: The e-liquid sample is prepared to a preset concentration, and when the e-liquid concentration is too high, a solvent is added to dilute it, wherein the solvent includes anhydrous ethanol; Place the prepared e-liquid sample into a transparent container; A transparent container containing the e-liquid sample is placed in a sample pool of a near-infrared spectrometer, and near-infrared spectrum data of the e-liquid sample is collected by the near-infrared spectrometer.
3. The method for measuring the nicotine content in tobacco oil based on near infrared spectroscopy according to claim 1, characterized in that: The calculation logic of the spectral step curvature feature is: Among them, C λ (t) is the spectral step curvature characteristic of the e-liquid sample t, t is the e-liquid sample number, λ is the wavelength, n is the number of spectral data points, S(λ i ,t) is the smoke oil sample t at wavelength λ i The reflection intensity at At wavelength λ i At the spectral reflection intensity S(λ i ,t) The second derivative with respect to wavelength, used to measure the curvature of the spectrum, At wavelength λ i At the spectral reflection intensity S(λ i ,t) is the first-order derivative with respect to wavelength, which is used to indicate the rate of change of the spectrum. γ1, γ2 and γ3 are exponential adjustment parameters used to adjust the sensitivity of the feature.
4. The method for determining the nicotine content in tobacco oil based on near infrared spectroscopy according to claim 3, characterized in that: The calculation logic of the microscopic spectrum peak polarity feature is: Among them, L λ (t) is the microscopic spectral peak polarity characteristic of the e-liquid sample t, t is the e-liquid sample number, λ is the wavelength, n is the number of spectral data points, S(k,t) is the reflection intensity of the e-liquid sample t at wavelength k, and when i is 1, the wavelength k is λ1, when i is n, the wavelength k is λ n , β1 and β2 are adjustment parameters used to adjust the local peak intensity.
5. The method for measuring the nicotine content in tobacco oil based on near infrared spectroscopy according to claim 4, characterized in that: The calculation logic of the frequency domain tuning fingerprint feature is: Among them, F λ (t) is the frequency domain tuning fingerprint feature of the e-liquid sample t, t is the e-liquid sample number, λ is the wavelength, n is the number of spectral data points, S(λ i ,t) is the smoke oil sample t at wavelength λ i where δ1 is the reflection intensity at , F(·) is the Fourier transform, and δ1 and δ2 are adjustment parameters used to control the intensity and sensitivity of the frequency domain features.
6. The method for measuring the nicotine content in tobacco oil based on near infrared spectroscopy according to claim 5, characterized in that: The calculation logic of the spectral layer texture structure feature is: Among them, T λ (t) is the spectral texture structure feature of the e-liquid sample t, t is the e-liquid sample number, λ is the wavelength, n is the number of spectral data points, S(λ i ,t) is the smoke oil sample t at wavelength λ i where m is the total number of wavelengths in the spectrum, and α1 and α2 are parameters for adjusting the complexity of the texture, which are used to quantify the changes between different wavelength bands.
7. The method for measuring the nicotine content in tobacco oil based on near infrared spectroscopy according to claim 6, characterized in that: The calculation logic of the spectral information complexity feature is: Among them, H λ (t) is the spectral information complexity characteristic of the e-liquid sample t, t is the e-liquid sample number, λ is the wavelength, n is the number of spectral data points, S(λ i ,t) is the smoke oil sample t at wavelength λ i The reflection intensity at Z λ (t) is the normalization constant of near-infrared spectral data, and the calculation logic is:
8. The method for measuring the nicotine content in tobacco oil based on near infrared spectroscopy according to claim 7, characterized in that: Step S3 includes: Acquire near infrared spectral data of e-liquid samples collected by a historical near infrared spectrometer and corresponding nicotine content measurement values, extract spectral features according to the reflection intensity of the spectral data within the near infrared wavelength range, set the spectral features and the nicotine content measurement values as a data set, and divide the data set into a training set and a test set; A regression model is constructed using a gradient boosting algorithm, hyperparameters of the regression model are initialized, the spectral features are set as inputs of the regression model, the nicotine content measurement values are set as target outputs of the regression model, and the regression model is trained using a training set, wherein the hyperparameters include a learning rate, the number of trees, the maximum depth of each tree, the number of leaf nodes, and a subsample ratio; The regression model gradually optimizes the weights of the decision tree through multiple iterations, and finally obtains a model composed of multiple weighted trees, which can effectively fit the training data and determine the nicotine content.
9. The method for measuring nicotine content in tobacco oil based on near infrared spectroscopy according to claim 7, characterized in that: Step S4 includes: K-fold cross validation is used to evaluate the performance of the regression model on different spectral data segmentations. Cross validation effectively detects whether the regression model is overfitting or underfitting and helps select the optimal hyperparameters. Use grid search or Bayesian optimization for hyperparameter tuning; Prevent overfitting by setting early stopping. The early stopping strategy stops training when the error of the validation set no longer decreases, ensuring that the regression model does not overfit on the training set. The trained regression model is measured using the test set, and the performance of the regression model is evaluated using evaluation indicators, wherein the evaluation indicators include mean square error, root mean square error, and mean absolute error.