Odor concentration prediction method based on key odor substance screening and small data modeling technology

Through the screening of key odor substances and small data modeling technology, using composite odor sensory evaluation and improved correlation analysis, key odor-causing substances were screened out. Combined with the random forest regression model, the problem of poor fitting of the odor concentration prediction model was solved, and the prediction accuracy and stability were improved.

CN120748531APending Publication Date: 2025-10-03TIANJIN ACAD OF ECOLOGICAL & ENVIRONMENTAL SCI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510834290.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-20
Publication Date
2025-10-03

AI Technical Summary

Technical Problem

The existing odor concentration prediction model has insufficient sample size and poor goodness of fit, resulting in a large gap between the measurement results and the human sensory fit, especially under complex sources and changeable meteorological conditions, the response pattern fluctuates greatly.

Method used

Based on the screening of key odor substances and small data modeling technology, the author adopts composite odor sensory evaluation, odor rose diagram construction, odor characteristic index scoring, improved correlation analysis and small data modeling methods to screen out key odor-causing substances that are significantly correlated with odor attributes, and combines the random forest regression model to predict odor concentration.

Benefits of technology

The goodness of fit of odor concentration prediction was improved under small data conditions, the prediction accuracy and stability of the model were enhanced, the data volatility was reduced, and the correlation between the model and sensory evaluation was enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120748531A_ABST
    Figure CN120748531A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of odor concentration prediction, in particular to an odor concentration prediction method based on key odor substance screening and a small data modeling technology. Aiming at the problem of poor goodness of fit of an odor concentration prediction model in the prior art, the invention develops a sensory evaluation method of composite odor, establishes an odor rose diagram and realizes sensory quantitative evaluation; a peculiar smell characteristic index evaluation system and a scoring method thereof are established, by improving colinearity and false correlation of traditional correlation analysis, the incidence relation between substances and sense organs is broken through, a substance smell contribution screening technology is developed, and synchronous analysis of the substances and the sense organs is achieved; through logarithm, square or Box-Cox transformation technology processing and standardization of odor concentration sensory indexes, model selection and important parameters suitable for small data analysis are provided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of odor concentration prediction, and in particular to an odor concentration prediction method based on key odor substance screening and small data modeling technology. Background Art

[0002] Odor pollution, as a typical nuisance pollution, has received increasing attention. Odor pollution has both chemical and sensory properties. Chemical methods can detect a variety of odor components, but due to the limitations of instrument sensitivity and detection range, it is impossible to detect all components. In addition, this method cannot represent the synergistic, antagonistic and masking effects of odors between substances. Therefore, sensory analysis, as an important supplement to chemical methods, is an important evaluation method for odor pollution. Sensory analysis is completed through manual measurement and can directly reflect human olfactory perception. It mainly includes qualitative and quantitative methods such as odor concentration, odor intensity, odor description, and pleasantness. Among them, odor concentration is the only sensory control item in the "Emission Standard of Odor Pollutants" (GB 14544) and is an important indicator for odor monitoring.

[0003] With the continuous advancement of sensor technology, online monitoring technologies based on sensor arrays and olfactory simulation algorithms have emerged. Compared to manual measurement, this model can, to a certain extent, avoid errors caused by subjectivity and olfactory fatigue, and improve the timeliness of test results. However, due to factors such as sensor performance and pattern recognition technology, the application of online odor monitoring systems based on sensor arrays still has limitations. For example, the stability of online odor monitoring equipment in response to low concentrations is poor, which is mainly reflected in the balance between hardware sensitivity and noise. Although this situation can be corrected through pattern recognition improvements and algorithm corrections, when faced with complex sources and changing meteorological conditions, the single response pattern of the olfactory fitting algorithm is highly volatile, and its measurement results differ significantly from those of human sensory fitting.

[0004] Substance analysis and detection technologies based on techniques such as chromatography, mass spectrometry, and spectroscopy have systematic quality control requirements and reliable and stable measurement results. Therefore, compared to sensor technologies, substance analysis and detection have significant advantages in the field of pollution monitoring. Therefore, establishing a relationship between substance concentration and odor concentration based on the Weber-Fechner law has become a new direction for fitting odor concentrations. This method generally utilizes methods such as gas chromatography-mass spectrometry, gas chromatography, and time-of-flight mass spectrometry to obtain concentrations of pollutants of interest. Combined with measured odor concentrations, linear regression, multivariate linear regression, partial least squares regression, and neural network models are used to construct odor concentration prediction models for industries such as sewage and sludge treatment and livestock and poultry farming.

[0005] However, the goodness of fit (R 2) is less than 80%, one of the main reasons for this problem being insufficient sample size. According to the requirements of domestic and international odor concentration measurement methods (EN 13725, ASTM E679-04, and the three-point comparison odor bag method), multiple professional odorists are required to conduct odor experiments using human noses, and the test results must be statistically verified. However, some samples are highly toxic, making odor concentration measurements based on manual odor detection impossible. These reasons make it difficult to obtain odor concentration test results, which in turn affects the sample size of the true value in the prediction model, resulting in a low model fit. Summary of the Invention

[0006] To address the problem of poor goodness-of-fit of existing odor concentration prediction models due to limited sample sizes, this present invention aims to provide an odor concentration prediction method based on the screening of key odorous substances and small-data modeling techniques. By exploring the distribution characteristics of odor concentration values ​​and the correlation between odor concentration and substance concentration, and focusing on screening key odor factors and improving data quality, a small-data modeling technique based on hundreds of data points was developed for odor concentration prediction.

[0007] In order to achieve the above-mentioned object of the invention, the technical solution adopted by the present invention is:

[0008] A method for predicting odor concentration based on key odorous substances screening and small data modeling technology is proposed, which includes the following steps:

[0009] S1. Based on sensory evaluation of compound odors, an odor rose diagram is constructed to calculate the odor attributes and their intensity.

[0010] S2. Select candidate odorous substances based on the olfactory threshold, number of odorous atoms or atomic groups, saturated vapor pressure, and concentration ratio of each component of the substance;

[0011] S3. Using the odor attributes and intensities from the odor rose diagram, we improved the traditional Pearson correlation analysis by eliminating collinearity based on the variance inflation factor and controlling false correlations based on partial correlation analysis, thereby screening out key odorous substances that are significantly correlated with the odor attributes from complex components.

[0012] S4. Process and standardize the odor concentration sensory index by logarithmic, square or Box-Cox transformation techniques;

[0013] S5. Select a model suitable for odor concentration prediction and fit it. Input the types of substances selected based on odor contribution and their concentrations to predict the odor concentration.

[0014] Furthermore, in step S2, the concentration ratio of odorous substances and the number of odorous atoms and atomic groups are obtained through instrument detection and molecular structure analysis; and the random forest regression model is used to fill in the missing values ​​of olfactory threshold and saturated vapor pressure.

[0015] Furthermore, the method of step S3 specifically includes the following steps:

[0016] 1) Data Preparation

[0017] ① Substance concentration matrix X: n samples × m substances.

[0018] ②Odor attribute matrix Y: n samples × p odor attributes.

[0019] 2) Standardization of substance concentration data

[0020] In order to eliminate the difference in concentration dimension, for each substance concentration X k Perform Z-score normalization as shown in formula (1).

[0021]

[0022] Where: X k ' represents the standardized concentration of the kth substance, X k represents the concentration of the kth substance, μ k represents the sample mean of the concentration of the kth substance, σ k represents the standard deviation of the concentration of the kth substance.

[0023] 3) Data collinearity elimination

[0024] ① Establish a multiple linear regression model

[0025] A multiple linear regression model was established to predict the concentration of the kth substance using other substances, as shown in formula (2).

[0026] X k '=β0+β1X'1+β2X'2+…+β k X' k +…+β m X' m +ε (2)

[0027] Where: X k ' represents the standardized concentration of the kth substance to be analyzed, β0 represents the intercept term, β i represents the regression coefficient, m represents the number of material types, ε represents the error term, and satisfies ε~N(0, σ 2 ).

[0028] ②Calculate the coefficient of determination R k 2

[0029] After fitting the regression model using the least partial squares method, R k 2, as shown in formulas (3) to (5).

[0030]

[0031] Among them, SSE k Represents the residual sum of squares, SST k represents the total sum of squares, X′ k,i represents the standardized concentration (true value) of the kth substance in the i-th sample, represents the predicted concentration of the kth substance in the i-th sample, n represents the total number of samples, The sample mean of the standardized concentration of the k-th substance.

[0032] ③Calculate the variance inflation coefficient

[0033]

[0034] Among them, VIF k represents the variance expansion coefficient of the kth substance; R k 2 It represents the coefficient of determination when the kth substance is used as the dependent variable and regressed with other substances, which is obtained by formula (3). k If the value is greater than 10, it is considered that the substance has strong collinearity with other substances and needs to be eliminated.

[0035] 4) Calculation of partial correlation coefficient

[0036] Calculating Substance X k The partial correlation coefficient As shown in formula (7).

[0037]

[0038] in, Indicates the concentration of substance X k With Y j The raw Pearson correlation coefficient of ; Represents X k The multiple correlation coefficient with the concentration of all other substances Z is

[0039] The significance test is used to determine whether the partial correlation coefficient of a substance is statistically significant, that is, to exclude the possibility that the observed correlation is caused by random noise. The t-statistic for calculating the partial correlation coefficient is shown in formula (8).

[0040]

[0041] Among them, n represents the number of samples; q represents the variable Z dimension, that is, the number of species remaining after VIF screening minus one.

[0042] The p-value is defined as the probability of observing the current t-value given the assumption that the true correlation is 0. Based on the t-value and the degrees of freedom (df) = nq - 2, the p-value can be calculated using a t-distribution table or statistical software. If p < 0.05, the correlation is considered significant, rejecting the null hypothesis of no correlation.

[0043] 5) Filter conditions

[0044] When the partial correlation coefficient is ≥0.5, the substance is identified as having a similar odor attribute to the odor attribute Y. j Significantly associated key odor-causing substances.

[0045] Furthermore, in step S4, when the skewness is greater than 2.5 and the kurtosis is less than -1.2, the Box-Cox transformation is initiated to make the odor concentration results conform to the normal distribution as much as possible, as shown in formula (9);

[0046]

[0047] Among them, y (λ) represents the result after transformation; y represents the original data to be transformed; λ represents the parameter that determines the transformation type, which is between 0.3 and 0.7.

[0048] Furthermore, in step S5, a random forest regression, gradient boosting regression tree or Bayesian regression model is selected for fitting.

[0049] The beneficial effects of the present invention are:

[0050] In response to the problem of poor goodness of fit of odor concentration prediction models in the prior art, the present invention develops a sensory evaluation method for complex odors and establishes an odor rose diagram to achieve sensory quantitative evaluation; establishes an odor characteristic index evaluation system and a scoring method, and by improving the collinearity and false correlation of traditional correlation analysis, opens up the relationship between substances and sensory organs, develops a substance odor contribution screening technology, and achieves synchronous analysis of substances and sensory organs; processes and standardizes odor concentration sensory indicators through logarithmic, square or Box-Cox transformation technology, and proposes model selection and important parameters suitable for small data analysis. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] Figure 1 This is a technical roadmap for the method of the embodiment of the present invention;

[0052] Figure 2 It is the odor rose diagram of the rubber mixing process of the rubber products industry in the embodiment;

[0053] Figure 3 The odor rose diagram of the vulcanization process of the rubber products industry in the embodiment;

[0054] Figure 4This is a diagram of odor concentration distribution in the embodiment;

[0055] Figure 5 This is a comparison chart of the predicted results and the actual results in the embodiment. DETAILED DESCRIPTION

[0056] The specific embodiments of the present invention are described below to facilitate understanding of the present invention by those skilled in the art. However, it should be clear that the present invention is not limited to the scope of the specific embodiments. For those skilled in the art, as long as various changes are within the spirit and scope of the present invention as defined and determined by the appended claims, these changes are obvious, and all inventions and creations utilizing the concepts of the present invention are protected.

[0057] Example 1

[0058] Reference Figure 1 This embodiment takes the rubber products industry as an example to provide an odor concentration prediction method based on key odor substance screening and small data modeling technology, which specifically includes the following steps:

[0059] 1. Screening of odor contribution of odorous substances

[0060] The screening of odor characteristic substances should not only consider the chemical properties of the industry material components, but also the industry odor characteristics, and organically integrate the chemical and odor characteristics to achieve simultaneous analysis of substances and sensory organs.

[0061] This example divides the screening of odor contributions of odorous substances in the rubber products industry into three steps: First, develop a sensory evaluation method for the industry's composite odors and construct an odor rose diagram for the rubber products industry; second, establish an odor characteristic scoring method based on the substance's saturated vapor pressure, olfactory threshold, molecular structure, and content, and rank the odor characteristics of detected substances in rubber products companies; third, using the odor attributes and intensity of the rubber products industry odor rose diagram, analyze the correlation between substances with odor characteristic scores exceeding 3 and composite odor characteristics, and select substances with correlation coefficients greater than 0.5 as key odorous substances in the rubber products industry / company / production process. The specific steps are as follows:

[0062] (1) Composite odor rose diagram for rubber products industry

[0063] A composite odor rose diagram is similar to a meteorological wind rose diagram and is used to measure the attributes and intensity of industry odors. Odor attributes are located in wind direction, similar to the meteorological wind rose diagram, while odor intensity is similar to the frequency of wind speed. By creating a composite odor rose diagram, the sensory evaluation of industry odors can be quantified. The specific steps are as follows:

[0064] 1) Expert system: This system consists of 16 people who meet the requirements for odor detectors in the "Determination of Odor in Ambient Air and Exhaust Gases - Three-Point Comparison Odor Bag Method" (HJ 1262-2022).

[0065] 2) Olfactory identification index: odor intensity and odor attributes are shown in Tables 1 and 2;

[0066] 3) The expert system evaluates the odor attribute intensity level of waste gas from the production processes of 20 or more enterprises through on-site investigation and odor analysis at least twice for each enterprise. The system first sorts out the intensity test results of different odor attributes of an enterprise and takes the arithmetic mean. Then, the test results of all enterprises are summarized and the arithmetic mean is taken to obtain the industry odor attribute intensity evaluation results, and draws the odor rose diagram of the industry / production process.

[0067] Table 1 Odor intensity table

[0068]

[0069]

[0070] Table 2 20 odor words in the rubber products industry

[0071] Alcohol smell Chinese herbal flavor Camphor smell Rubber smell Plastic smell Burning rubber smell Dusty smell Glue smell Raw rice smell greasy taste nail polish smell Rotten corn / vegetable smell Glue smell Burnt smell Spicy smell Organic solvent odor sour taste Disgusting smell Pungent odor fragrance

[0072] (2) Odor characteristics scoring method

[0073] According to olfactory theory, the odor of a substance is related to factors such as its volatility, functional groups, and olfactory threshold. Therefore, this paper constructs an odor characteristic index evaluation system and scoring scheme for the rubber products industry. Using a random forest regression model, a method for imputing missing values ​​for characteristic indicators is established, enabling the scoring and ranking of the odor characteristics of substances in odor samples from the rubber industry. The specific steps are as follows:

[0074] 1) Characteristic Indicator Determination: The target variable is the odor score, with a score of "1" indicating almost no odor potential, "2" indicating weak odor potential, "3" indicating some odor potential, and "4" indicating strong odor potential. An odor characteristic index system was constructed by selecting the substance's olfactory threshold, the number of odor-producing atoms or atomic groups, the saturated vapor pressure at 20-25°C, and the substance's concentration ratio. The characteristic indicators were assigned values. The assignment scheme is shown in Table 3.

[0075] Table 3 Characteristic index assignment

[0076]

[0077] *Smelly atoms or atomic groups refer to elements from Groups 4 to 7, such as phosphorus, arsenic, sulfur, and antimony, and odor-producing functional groups, such as carbonyl, aldehyde, methanol, ester, amino, ether, carboxyl, and carbonyl.

[0078] 2) Data Collection and Processing: The concentration percentage of a substance and the number of odorous atoms and atomic clusters can be obtained through instrument detection and molecular structure analysis. However, information on the saturated vapor pressure and olfactory threshold of a substance is generally incomplete. Therefore, the present invention establishes a method for filling missing values ​​of olfactory threshold and saturated vapor pressure based on a random forest regression model.

[0079] ① Use chromatography, mass spectrometry, and ion mobility spectrometry techniques to analyze the composition and concentration of waste gas during the company's production process, query the substance's olfactory threshold and the substance's saturated vapor pressure information, and analyze the substance's molecular structure and functional group information. Score the substances with complete information according to Table 3 to obtain a "Rubber Products Industry Substance Analysis List" containing the substance name, score, olfactory threshold, number of odorous atoms or atomic groups, saturated vapor pressure, and substance concentration ratio.

[0080] ② Compared to the olfactory threshold, there are fewer missing values ​​for saturated vapor pressure, so we first fill in the missing values ​​for saturated vapor pressure. Specifically, we first fill in the missing values ​​for the olfactory threshold column with "0", and then use the randomized forest ridge regression model to fill in the missing values ​​for saturated vapor pressure. After the missing values ​​for saturated vapor pressure are filled, the same model is used to fill in the missing values ​​for olfactory threshold. Using the Python sklearn library, the code is written as follows:

[0081] A) Fill missing values ​​in the olfactory threshold column with "0"

[0082] #Library

[0083] import pandas as pd

[0084] from sklearn.impute import SimpleImputer

[0085] #Import dataset

[0086] data = pd.read_excel(r'C:\python\Rubber Products Industry Material Analysis List.xlsx')

[0087] #Create a SimpleImputer object and fill the missing values ​​of the "olfactory threshold" column with 0

[0088] imputer=SimpleImputer(missing_values=np.nan, strategy='constant', fill_value=0)

[0089] data["Olfaction threshold"] = imputer.fit_transform(data[["Olfaction threshold"]])

[0090] B) Use the random forest regression model to fill in the "saturated vapor pressure" column

[0091] # Divide the saturated vapor pressure dataset feature_cols = ["score","number of odorous atoms or groups","molecular weight","substance concentration ratio"] label_col = "saturated vapor pressure"

[0092] # Split the dataset

[0093] train_data=data[data[label_col].notnull()]

[0094] test_data=data[data[label_col].isnull()]

[0095] X_train=train_data[feature_cols]

[0096] y_train=train_data[label_col]

[0097] X_test=test_data[feature_cols]

[0098] #Train random forest model

[0099] from sklearn.ensemble import RandomForestRegressorrf=RandomForestRegressor(n_estimators=200,random_state=42)

[0100] rf.fit(X_train,y_train)

[0101] #Predict missing values

[0102] predicted_values=rf.predict(X_test)

[0103] #Fill missing values

[0104] data.loc[data[label_col].isnull(),label_col]=predicted_values ​​C) Verify the filling effect

[0105] #Import library

[0106] from sklearn.model_selection import train_test_split

[0107] from sklearn.metrics import r2_score

[0108] # Split validation set

[0109] X_train,X_val,y_train,y_val=train_test_split(X_train,y_train,test_size=0.2,random_state=42)

[0110] rf_val=RandomForestRegressor(n_estimators=100,random_state=42)

[0111] rf_val.fit(X_train,y_train)

[0112] print("Model R 2 Score: ",rf_val.score(X_val,y_val))

[0113] D) Use the random forest regression model to fill in the "Olfaction Threshold" column

[0114] #Select features

[0115] feature_cols_v2=feature_cols+[label_col]

[0116] label_col_v2="Olfaction Threshold"

[0117] #Filter out the data that needs to be predicted (data originally filled with 0)

[0118] train_data_v2=data[data[label_col_v2]! =0]

[0119] test_data_v2=data[data[label_col_v2]==0]

[0120] #Train a new model

[0121] rf_v2=RandomForestRegressor(n_estimators=200,random_state=42)

[0122] rf_v2.fit(train_data_v2[feature_cols_v2],train_data_v2[label_col_v2])

[0123] #Predict and fill

[0124] predicted_olfactory=rf_v2.predict(test_data_v2[feature_cols_v2])

[0125] data.loc[data[label_col_v2]==0,label_col_v2]=predicted_olfactory

[0126] At this point, the missing values ​​of saturated vapor pressure and olfactory threshold have been filled.

[0127] 3) Scoring of odor characteristics of detected substances: Scoring is based on the olfactory threshold, number of odor-producing atoms or atomic groups, saturated vapor pressure and proportion of substance concentration of each component of the exhaust gas. Components with an odor score exceeding 3 are selected as candidate target substances.

[0128] (3) Screening method for odor contribution of odorous substances

[0129] This paper proposes an improved correlation analysis model that improves traditional Pearson correlation analysis by eliminating collinearity based on the variance inflation factor and controlling spurious correlations based on partial correlation analysis. This model can be used to screen out the main odor-causing substances that are significantly correlated with odor properties from complex components. The specific steps are as follows:

[0130] 1) Data Preparation

[0131] ① Substance concentration matrix X: n samples × m substances.

[0132] ②Odor attribute matrix Y: n samples × p odor attributes.

[0133] 2) Standardization of substance concentration data

[0134] In order to eliminate the difference in concentration dimension, for each substance concentration X k Perform Z-score normalization as shown in formula (1).

[0135]

[0136] Where: X k ' represents the standardized concentration of the kth substance, X k represents the concentration of the kth substance, μ k represents the sample mean of the concentration of the kth substance, σ krepresents the standard deviation of the concentration of the kth substance.

[0137] 3) Data collinearity elimination

[0138] ① Establish a multiple linear regression model

[0139] A multiple linear regression model was established to predict the concentration of the kth substance using other substances, as shown in formula (2).

[0140] X k '=β0+β1X'1+β2X'2+…+β k X' k +…+β m X' m +ε (2)

[0141] Where: X k ' represents the standardized concentration of the kth substance to be analyzed, β0 represents the intercept term, β i represents the regression coefficient, m represents the number of material types, ε represents the error term, and satisfies ε~N(0, σ 2 ).

[0142] ②Calculate the coefficient of determination R k 2

[0143] After fitting the regression model using the least partial squares method, R k 2 , as shown in formulas (3) to (5).

[0144]

[0145]

[0146] Among them, SSE k Represents the residual sum of squares, SST k represents the total sum of squares, X′ k,i represents the standardized concentration (true value) of the kth substance in the i-th sample, represents the predicted concentration of the kth substance in the i-th sample, n represents the total number of samples, The sample mean of the standardized concentration of the k-th substance.

[0147] ③Calculate the variance inflation coefficient

[0148]

[0149] Among them, VIF k represents the variance expansion coefficient of the kth substance; R k 2It represents the coefficient of determination when the kth substance is used as the dependent variable and regressed with other substances, which is obtained by formula (3). k If the value is greater than 10, it is considered that the substance has strong collinearity with other substances and needs to be eliminated.

[0150] 4) Calculation of partial correlation coefficient

[0151] Calculating Substance X k The partial correlation coefficient As shown in formula (7).

[0152]

[0153] in, Indicates the concentration of substance X k With Y j The raw Pearson correlation coefficient of ; Represents X k The multiple correlation coefficient with the concentration of all other substances Z is

[0154] The significance test is used to determine whether the partial correlation coefficient of a substance is statistically significant, that is, to exclude the possibility that the observed correlation is caused by random noise. The t-statistic for calculating the partial correlation coefficient is shown in formula (8).

[0155]

[0156] Among them, n represents the number of samples; q represents the variable Z dimension, that is, the number of species remaining after VIF screening minus one.

[0157] The p-value is defined as the probability of observing the current t-value given the assumption that the true correlation is 0. Based on the t-value and the degrees of freedom (df) = nq - 2, the p-value can be calculated using a t-distribution table or statistical software. If p < 0.05, the correlation is considered significant, rejecting the null hypothesis of no correlation.

[0158] 5) Filter conditions

[0159] When the partial correlation coefficient is ≥0.5, the substance is identified as having a similar odor attribute to the odor attribute Y. j Significantly associated key odor-causing substances.

[0160] At this point, the screening of odor contributions of substances in the rubber products industry has been completed, and a list of input values ​​for the model’s independent variables has been formed.

[0161] 2. Odor concentration data processing

[0162] Odor concentration data is obtained through manual olfactory analysis. Human senses are easily affected by factors such as emotions and living environment, resulting in high volatility in the data, making it difficult to construct a good-fit model. Therefore, it is necessary to conduct in-depth analysis and preprocessing of the odor concentration data before modeling.

[0163] A comprehensive analysis of the distribution of odor concentrations in the rubber products industry was conducted, including statistical analysis of mean, median, standard deviation, skewness, and kurtosis. Because most models require normal distribution of data, data with large skewness and small kurtosis were preprocessed using logarithms and squares.

[0164] When the skewness is greater than 2.5 and the kurtosis is less than -1.2, the Box-Cox transformation is initiated to make the odor concentration results conform to the normal distribution as much as possible, as shown in formula (9).

[0165]

[0166] Among them, y (λ) represents the result after transformation; y represents the original data to be transformed; λ represents the parameter that determines the transformation type, which is between 0.3 and 0.7.

[0167] 3. Model selection

[0168] Based on the characteristics of odor concentration and substance concentration distribution in the rubber products industry, this paper proposes a model and related parameters suitable for odor concentration prediction in this industry, as shown in Table 4. The input data are the substance types and concentration values ​​after the first step of processing and the odor concentration values ​​after the second step of processing. The model is evaluated using MSE, MAPE, R 2 , MSE of 10-fold cross validation.

[0169] Table 4 Model type selection

[0170]

[0171] Example 2

[0172] Example 2 is an analysis of the actual effects of the technical solution described in Example 1.

[0173] 1. Sensory evaluation of industry complex odors and establishment of odor rose diagram for rubber products industry

[0174] An 8-member sniffing team conducted two on-site sniffing visits to 22 companies with the highest number of complaints in the rubber products industry nationwide over the past three years. They evaluated, recorded, and organized the intensity values ​​of 20 odor attributes in the companies' rubber mixing and vulcanization processes, and obtained an odor rose diagram for the rubber products industry's rubber mixing and vulcanization processes ( Figure 2-3 ).

[0175] Table 5 Odor attributes and intensity records of rubber mixing and vulcanization processes

[0176]

[0177]

[0178]

[0179] 2. Scoring of odor characteristics of main detected substances in rubber refining and vulcanization processes

[0180] 654 gas samples were collected from the rubber refining and vulcanization processes of 22 companies. All samples were analyzed using gas chromatography-mass spectrometry and gas chromatography-ion mobility spectrometry within 24 hours, resulting in the identification of 393 substances. The content, odor threshold, odor-causing elements / groups, and saturated vapor pressure of these 393 substances were compiled and summarized. A random forest regression model was used to fill in the missing odor thresholds and saturated vapor pressures for the 393 detected substances. The model was developed using Python 3.9.7, Scikit-learn 1.6.1, Numpy 1.20.3, Pandas 1.3.4, and Matplotlib 3.4.3.

[0181] According to the odor characteristic index system and scoring scheme of the rubber products industry, a list of substances with odor characteristic scores ≥ 3 was determined, totaling 57 substances, mainly including aldehydes and ketones, sulfides, alkanes and alkenes, benzene series, alcohols, phenols, and furans. Specifically, they include acetaldehyde, propionaldehyde, isovaleraldehyde, n-butyraldehyde, isobutyraldehyde, benzaldehyde, acetone, 2-butanone, methyl isobutyl ketone, cyclohexanone, limonene, trimethylamine, hexanethiol, methyl mercaptan, carbonyl sulfide, carbon disulfide, dichloromethane, chloroform, propane, isobutane, n-butane, pentane, n-hexane, cyclohexane, n-heptane, octane, dodecane, 2-methylbutane, 2-methylpentane, 3-methylpentane, 2-methylhexane, 3-methyl Hexane, methylcyclopentane, methylcyclohexane, 2,3-dimethylpentane, 2,3-dimethylbutane, 2,4-dimethylpentane, 3-methylheptane, 2,6-di-tert-butyl-p-cresol, methanol, ethanol, isopropanol, tert-butanol, 2-ethylhexanol, propylene, isobutylene, 2-methylfuran, tetrahydrofuran, benzene, styrene, toluene, ethylbenzene, p-xylene, m-xylene, o-xylene, m-ethyltoluene, and p-ethyltoluene.

[0182] Table 6 Screening results of substances with odor characteristic scores exceeding 3

[0183]

[0184]

[0185] 3. Screening results of odor contribution of substances in the rubber products industry

[0186] Based on the improved correlation analysis model, substances with VIF ≥ 10 were eliminated and partial correlation coefficients ≥ 0.5 were retained, resulting in the screening of 29 substances. The odor-contributing substances screened out were cyclohexanone, 2-methylfuran, methyl mercaptan, p-xylene, 2-methylbutane, carbon disulfide, tert-butyl alcohol, n-hexane, 2-methylhexane, 2-ethylhexanol, 3-methylheptane, p-ethyltoluene, benzaldehyde, isopropanol, 2,4-dimethylpentane, propylene, pentane, acetaldehyde, methyl isobutyl ketone, octane, cyclohexane, methanol, acetone, ethylbenzene, isovaleraldehyde, trimethylamine, dodecane, limonene, and dichloromethane.

[0187] 4. Odor concentration data processing

[0188] The odor concentration was measured using the three-point comparison odor bag method. The odor concentration range was 41 to 35481, with a mean of 2915 and a standard deviation of 5691. The data belonged to a normal right-skewed distribution, so Z-score standardization was performed ( Figure 4 ).

[0189] 5. Odor concentration prediction in rubber products industry

[0190] Three models, random forest regression, gradient boosting regression tree, and Bayesian regression, were selected for fitting. The model environment was Python 3.9.7, Scikit-learn 1.6.1, Numpy 1.20.3, Pandas 1.3.4, and Matplotlib 3.4.3. The input values ​​were the odor concentration data of rubber mixing and vulcanization in the rubber products industry, and the types and concentrations of substances selected for odor contribution. MSE, MAPE, and R were used to analyze the odor. 2 , 10-fold cross-validation MSE is used to evaluate the model. The code is as follows:

[0191] 1. Guide library

[0192] import pandas as pd

[0193] import numpy as np

[0194] import matplotlib.pyplot as plt

[0195] from sklearn.ensemble import RandomForestRegressor

[0196] from sklearn.model_selection import train_test_split

[0197] from sklearn.metrics import mean_squared_error

[0198] from sklearn.metrics import r2_score

[0199] from sklearn.metrics import mean_absolute_percentage_error

[0200] 2. Data import

[0201] odor = pd.read_excel(r"C:\Users\youra\Desktop\Odor Concentration Fitting\sample-29compounds-log.xlsx",index_col=0)

[0202] X = odor.iloc[:,0:-1]

[0203] y = odor.iloc[:,-1]

[0204] 3. Test set and training set division

[0205] X_train,X_test,y_train,y_test=train_test_split(X_selected,y,test_size=0.2,random_state=42)

[0206] 4. Model Building

[0207] (1) Random Forest Regression

[0208] model=RandomForestRegressor(n_estimators=200,random_state=42)

[0209] model.fit(X_train,y_train)

[0210] (2) Gradient Boosting Regression Tree

[0211] model=GradientBoostingRegressor(n_estimators=200, learning_rate=0.1, random_state=42)

[0212] model.fit(X_train,y_train)

[0213] (3) Bayesian regression

[0214] model = BayesianRidge()

[0215] model.fit(X_train,y_train)

[0216] 5. Model Evaluation

[0217] y_pred=model.predict(X_test)

[0218] mse=mean_squared_error(y_test,y_pred)

[0219] r2=r2_score(y_test,y_pred)

[0220] mape=mean_absolute_percentage_error(y_test,y_pred)

[0221] print(f"Test set mean square error:{mse}")

[0222] print(f"Determination coefficient (R2): {r2}")

[0223] print(f"Mean Absolute Percent Error (MAPE):{mape}")

[0224] Table 7 Evaluation results of three models

[0225] Model MSE MAPE <![CDATA[R 2 ]]> MSE of 10-fold cross validation Random Forest Regression 957 1.91 0.923 4262 Gradient Boosted Regression Tree 1017 1.65 0.913 3982 Bayesian regression 1388 4.03 0.838 5672

[0226] like Figure 5 As shown in Table 7, the model effect of random forest regression is the best. The model evaluation results in Table 7 show that the MSE and MAPE of random forest regression are the smallest, and R 2 The highest, 10-fold cross-validation MSE is at a moderate level. Figure 5 The figure shows a comparison of the odor concentration predictions obtained using the three models with the true values. The shading represents the true values ​​(blue) and their standard deviations (gray and orange). The dotted line graph shows the odor concentration predictions and their trends for the three models. Overall, the odor concentration predictions using random forest regression are closer to the true values ​​at different concentration levels. Therefore, the random forest regression model was selected as the odor concentration prediction model for the rubber products industry's mixing and vulcanization processes.

[0227] In summary, the present invention addresses the problem of poor goodness of fit of the odor concentration prediction model in the rubber products industry, develops a sensory evaluation method for the industry's composite odors, establishes an odor rose diagram for the rubber products industry, and realizes sensory quantitative evaluation of the rubber products industry; establishes an evaluation system for odor characteristic indicators in the rubber products industry and a scoring method thereof, and by improving the collinearity and false correlation of traditional correlation analysis, opens up the relationship between substances and sensory organs, develops a substance odor contribution screening technology, and realizes the synchronous analysis of substances and sensory organs; processes and standardizes the odor concentration sensory indicators through logarithmic, square or Box-Cox transformation technology, and proposes model selection and important parameters suitable for small data analysis in the rubber products industry.

[0228] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above and that the present invention can be embodied in other specific forms without departing from the spirit or essential characteristics of the present invention. Therefore, the embodiments should be considered in all respects as illustrative and non-restrictive, and the scope of the present invention is defined by the appended claims rather than the foregoing description, and it is intended that all variations that come within the meaning and range of equivalents of the claims be embraced therein.

[0229] In addition, it should be understood that although this specification is described in terms of implementation methods, not every implementation method contains only one independent technical solution. This narrative method of the specification is only for the sake of clarity. Those skilled in the art should regard the specification as a whole. The technical solutions in each embodiment can also be appropriately combined to form other implementation methods that can be understood by those skilled in the art.

Claims

1. A method for predicting odor concentration based on key odorous substance screening and small data modeling technology, characterized in that: The following steps are involved: S1. Based on sensory evaluation of compound odors, an odor rose diagram is constructed to calculate the odor attributes and their intensity. S2. Select candidate odorous substances based on the olfactory threshold, number of odorous atoms or atomic groups, saturated vapor pressure, and concentration ratio of each component of the substance; S3. Using the odor attributes and intensities from the odor rose diagram, we improved the traditional Pearson correlation analysis by eliminating collinearity based on the variance inflation factor and controlling false correlations based on partial correlation analysis, thereby screening out key odorous substances that are significantly correlated with the odor attributes from complex components. S4. Process and standardize the odor concentration sensory index by logarithmic, square or Box-Cox transformation techniques; S5. Select a model suitable for odor concentration prediction, input the types and concentrations of key odor substances screened out by odor contribution, and predict the odor concentration.

2. The odor concentration prediction method based on key odor substance screening and small data modeling technology according to claim 1 is characterized in that: In step S2, the concentration ratio of odorous substances and the number of odorous atoms and atomic groups are obtained through instrument detection and molecular structure analysis; and the random forest regression model is used to fill in the missing values ​​of olfactory threshold and saturated vapor pressure.

3. The odor concentration prediction method based on key odor substance screening and small data modeling technology according to claim 1 is characterized in that: The method of step S3 specifically includes the following steps: 1) Data Preparation ① Substance concentration matrix X: n samples × m substances; ②Odor attribute matrix Y: n samples × p odor attributes; 2) Standardization of substance concentration data In order to eliminate the difference in concentration dimension, for each substance concentration X k Perform Z-score standardization as shown in formula (1); Where: X k ' represents the standardized concentration of the kth substance, X k represents the concentration of the kth substance, μ k represents the sample mean of the concentration of the kth substance, σ k represents the standard deviation of the concentration of the kth substance; 3) Data collinearity elimination ① Establish a multiple linear regression model A multiple linear regression model was established to predict the concentration of the kth substance using other substances, as shown in formula (2); X k '=β0+β1X'1+β2X'2+…+β k X' k +…+b m X' m +e (2) Where: X k ' represents the standardized concentration of the kth substance to be analyzed, β0 represents the intercept term, β i represents the regression coefficient, m represents the number of material types, ε represents the error term, and satisfies ε~N(0, σ 2 ); ②Calculate the coefficient of determination R k 2 After fitting the regression model using the least partial squares method, R k 2 , as shown in formulas (3) to (5); Among them, SSE k Represents the residual sum of squares, SST k represents the total sum of squares, X' k,i represents the standardized concentration of the kth substance in the i-th sample, represents the predicted concentration of the kth substance in the i-th sample, n represents the total number of samples, The sample mean of the standardized concentration of the kth substance; ③Calculate the variance inflation coefficient Among them, VIF k represents the variance expansion coefficient of the kth substance; R k 2 It represents the coefficient of determination when the kth substance is used as the dependent variable and regressed with other substances, which is obtained by formula (3); when VIF k If the value is greater than 10, it is considered that the substance has strong collinearity with other substances and needs to be eliminated; 4) Calculation of partial correlation coefficient Calculating Substance X k The partial correlation coefficient As shown in formula (7); in, Indicates the concentration of substance X k With Y j The raw Pearson correlation coefficient of ; Represents X k The multiple correlation coefficient with the concentration of all other substances Z is The significance test is used to determine whether the partial correlation coefficient of a substance is statistically significant, that is, to exclude the possibility that the observed correlation is caused by random noise. The t-statistic for calculating the partial correlation coefficient is shown in formula (8). Where n represents the number of samples; q represents the variable Z dimension, that is, the number of species remaining after VIF screening minus one; The p-value is defined as the probability of observing the current t-value assuming the true correlation is 0. Based on the t-value and the degrees of freedom df = nq-2, the p-value can be calculated by looking up a t-distribution table or using statistical software. If p < 0.05, the correlation is considered significant, and the null hypothesis of "no correlation" is rejected. 5) Filter conditions When the partial correlation coefficient is ≥0.5, the substance is identified as having a similar odor attribute to the odor attribute Y. j Significantly associated key odor-causing substances.

4. The odor concentration prediction method based on key odor substance screening and small data modeling technology according to claim 1 is characterized in that: In step S4, when the skewness is greater than 2.5 and the kurtosis is less than -1.2, the Box-Cox transformation is initiated to make the odor concentration results conform to the normal distribution as much as possible, as shown in formula (9); Among them, y (λ) represents the result after transformation; y represents the original data to be transformed; λ represents the parameter that determines the transformation type, which is between 0.3 and 0.

7.

5. The odor concentration prediction method based on key odor substance screening and small data modeling technology according to claim 1 is characterized in that: In step S5, a random forest regression, gradient boosting regression tree or Bayesian regression model is selected for fitting.