Method for constructing quality sensory evaluation model of three delicacies

The quality evaluation model of "Di San Xian" (a local specialty) by using multimodal feature screening and weighted fusion solves the problems of subjectivity in traditional sensory evaluation and insufficient accuracy of single-modal detection, and realizes objective, accurate and efficient evaluation of the quality of "Di San Xian". It is suitable for quality inspection in catering and food processing enterprises.

CN121901609APending Publication Date: 2026-04-21JIANGNAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
JIANGNAN UNIV
Filing Date
2025-12-19
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

Existing technologies are insufficient to fully capture the multi-dimensional quality information of the three delicacies of Di San Xian (a type of fresh seafood), traditional linear regression models have limited fitting capabilities, single-modal detection data has poor prediction accuracy, and lacks model interpretability, thus failing to meet the rapid and standardized control requirements of industrialized food production.

Method used

A multimodal feature selection and weighted fusion method is adopted to integrate sensory evaluation data and machine detection data. Through multi-step feature selection and model training, a high-precision and interpretable quality evaluation model of "Di San Xian" (a local specialty) is constructed, including data preprocessing, multimodal feature selection, single-modal model training and weighted fusion.

Benefits of technology

It achieves objective and accurate evaluation of the quality of "Di San Xian" (a dish of stir-fried potatoes, green peppers, and eggplant), improves prediction accuracy, ensures the consistency and repeatability of evaluation, has a scientifically adapted model architecture, and features efficient and interpretable feature engineering, making it suitable for quality testing in the catering industry and food processing enterprises.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121901609A_ABST
    Figure CN121901609A_ABST
Patent Text Reader

Abstract

The invention discloses a construction method for a quality sensory evaluation model of three delicacies. The construction method comprises the following steps: acquiring a multi-source basic database; preprocessing the data; mMSFS multi-modal feature selection: screening a core feature subset from the initial features; training a single-mode model: respectively constructing and optimizing a random forest regression model or a ridge regression model according to the characteristics of four modes of an electronic nose, GC-MS, texture and nutritional indexes, and determining optimal hyper-parameters of each model through grid search; carrying out model weighted fusion: calculating a normalized fusion weight based on a decision coefficient R2 value of each single-modal model on the verification set, and carrying out weighted summation on each modal prediction result to obtain a final prediction value; and explanatory analysis: carrying out feature importance analysis and decision process visual explanation on the single-mode model and the fusion model by adopting an SHAP method. According to the method, based on multi-modal feature screening and weighted fusion, the evaluation model with the prediction precision obviously better than that of a single modal is constructed, and objective, efficient and accurate evaluation of the quality of the three delicacies of the Chinese dolichos is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a method for constructing a sensory evaluation model for the quality of "Di San Xian" (a dish of fresh potatoes, green peppers, and oysters), belonging to the field of food quality evaluation technology. Background Technology

[0002] Three Treasures of the Earth is a classic traditional Chinese stir-fried vegetable dish, and its flavor, texture, and nutritional balance are key to its quality. Currently, the industry's quality evaluation mainly relies on professional scores from trained sensory evaluators. Although sensory evaluation can comprehensively reflect consumer acceptance, its inherent subjectivity, individual differences, and instability make it difficult to meet the needs of rapid, consistent, and standardized quality control in industrialized food production, and also pose challenges to large-scale market supervision.

[0003] In existing technologies, some studies have attempted to replace some sensory evaluations with instrumental detection, such as using electronic noses to evaluate odor and texture analyzers to evaluate taste, and establishing simple statistical or shallow machine learning models for prediction. For example, headspace solid-phase microextraction combined with gas chromatography-mass spectrometry (GC-MS) combined with partial least squares (PLSR) was used to predict the growth stage attributes of jujube; GC-MS and OPLS-DA were used to analyze the effect of heating temperature on the aroma quality of Sichuan pepper chicken broth. However, these methods have significant shortcomings: First, single-modality detection data (such as odor or texture only) can only reflect one aspect of quality and cannot comprehensively capture the complex quality information of Di San Xian in terms of flavor, texture, nutrition, and other dimensions. Second, the quality formation of Di San Xian involves complex physicochemical processes, and there is often a non-linear relationship between its characteristics and sensory scores. Traditional linear regression models (such as partial least squares regression PLSR) have limited ability to fit such relationships, resulting in poor prediction accuracy. Third, the feature selection methods are singular and fail to fully consider the characteristics of different modal data and the redundancy between features, which can easily lead to overfitting or underfitting of the model and poor generalization ability. Finally, existing "black box" models lack interpretability and cannot clearly tell producers which specific indicators affect the final quality, thus making it difficult to provide targeted process optimization guidance.

[0004] Therefore, there is a need for a method to construct a sensory evaluation model for the quality of Di San Xian (a type of fresh seafood) that can integrate multi-source heterogeneous data, adapt to complex nonlinear relationships, and combine high prediction accuracy with strong model interpretability, so as to overcome the shortcomings of existing technologies and promote the intelligent development of quality control of Chinese dishes. Summary of the Invention

[0005] To solve the above problems, the present invention provides a construction method for a systematic, efficient, and interpretable sensory evaluation model of the quality of stir-fried eggplant, green pepper and potato based on multi-modal feature screening and weighted fusion, aiming to achieve objective, efficient, and accurate evaluation of the quality of stir-fried eggplant, green pepper and potato, and is applicable to quality inspection scenarios in the catering industry, food processing enterprises and market supervision. The present invention constructs an evaluation model with significantly better prediction accuracy than a single modality and with clear physical significance by deeply integrating multi-dimensional machine detection data and sensory evaluation data.

[0006] To achieve the above object, the technical solution adopted by the present invention is as follows: In a first aspect, the present invention provides a construction method for a sensory evaluation model of the quality of stir-fried eggplant, green pepper and potato, comprising the following steps: Step 1, obtain a multi-source basic database: the database includes sensory evaluation data and multi-source index data of machine detection; Step 2, data preprocessing: perform missing value filling, outlier correction, noise smoothing and sub-modal standardization processing on the multi-source index data of the machine detection in sequence to eliminate the influence of dimension and noise; Step 3, MMSFS multi-modal feature selection: adopt a multi-step screening strategy, and sequentially perform variance screening, stability screening based on Bootstrap, modal specificity screening and cross-modal redundancy removal on the data preprocessed in Step 2 to screen a core feature subset from the initial features; Step 4, single-modal model training: construct and optimize a random forest regression model or a ridge regression model respectively according to the characteristics of four modalities of electronic nose, GC-MS, texture, and nutritional index, and determine the optimal hyperparameters of each model through grid search; Step 5, model weighted fusion: calculate the normalized fusion weight based on the coefficient of determination R 2 value on the validation set of each single-modal model, and only retain the single-modal with R 2 >0 to participate in the fusion, and sum the prediction results of each modality weighted to obtain the final prediction value; Step 6, interpretability analysis: adopt the SHAP method to perform feature importance analysis and visualization interpretation of the decision-making process on the single-modal model and the fusion model.

[0007] In an embodiment of the present invention, in Step 1, the sensory evaluation data is the total score comprehensively calculated by the entropy weight method after scoring on a 0-9 scale from 5 dimensions of color, smell, taste, texture and overall acceptability; the multi-source index data of the machine detection includes texture indexes detected by a TPA texture analyzer, smell indexes jointly detected by an electronic nose and a gas chromatography-mass spectrometry GC-MS, and nutritional indexes detected.

[0008] In one embodiment of the present invention, in step 3, the core feature subset consists of 12 features, including: S1 and S8 sensor signals of the electronic nose mode; (E,E)-2,4-nonadienal, octanal, 1-octen-3-ol, and diallyl disulfide content of the GC-MS mode; potato hardness of the texture mode; and moisture content, protein content, dietary fiber content, ash content, and fat content of the nutritional index mode.

[0009] In one embodiment of the present invention, in step 4, the optimal configuration and performance result of the single-modal model training are as follows: Electronic nose modality: Random forest regressor was selected, with hyperparameter search range of max_depth=3-7, min_samples_leaf=1-3, min_samples_split=4-8, and n_estimators=200-400; GC-MS modality: Random forest regressor was selected, with hyperparameter search range of max_depth=3-7, min_samples_leaf=1-3, min_samples_split=4-8, n_estimators=200-400; Texture modality: Random forest regressors were used instead of traditional linear models, with hyperparameter search ranges of max_depth=2-5, min_samples_leaf=2-4, min_samples_split=3-6, and n_estimators=150-250; Nutritional index modality: Ridge regression model was selected, with hyperparameter search range of alpha=0.05-0.2 and fit_intercept=True / False; The model performance under optimal parameters is: R 2 =0.1344, RMSE=0.6995, MAE=0.5820.

[0010] In one embodiment of the present invention, in step 5, the weights of the model weighted fusion are based on each single-mode R. 2 The values ​​were calculated with the following weights: electronic nose model 0.32, GC-MS model 0.68, and texture and nutrient index model weighted by R. 2 If the value is not positive, the weight is 0; the prediction result of the fusion model is obtained by weighted summation according to the weights, and the performance of the final fusion model is: R 2 =0.7426, RMSE=0.3815, MAE=0.3075.

[0011] In one embodiment of the present invention, step 2 includes: Missing value handling: The K-nearest neighbor algorithm is used to fill in missing data to ensure data integrity; Outlier correction: Outliers are identified based on MAD-Z scores and replaced with the median of this feature; Noise smoothing and modal standardization: For high-noise time series data such as electronic nose and GC-MS, a weighted moving average window is used for smoothing. At the same time, in order to avoid the influence of the difference in dimensions between different modes, Z-score standardization is performed on the four types of data: electronic nose, GC-MS, texture and nutrient indicators.

[0012] In one embodiment of the present invention, step 3 includes: Variance screening: retain features with variance ≥ 0.01 and remove indices with no discriminant power. Stability screening: The frequency of each feature being selected into the model is calculated through multiple Bootstrap resampling; the differential frequency threshold is set according to the characteristics of different modal data: 0.65 for electronic nose and GC-MS, and 0.5 for texture and nutrient indicators, to retain highly stable features; Modality-specific screening: To address the strong nonlinearity of texture data, feature importance of the random forest model is used for screening; for nutrient indicators, Pearson correlation coefficient and mutual information are combined for screening. Cross-modal redundancy removal: Calculate the mutual information between features of different modalities and remove the lower priority features from highly redundant feature pairs with mutual information values ​​≥0.8; Step 3 also includes: SHAP Importance Screening: For electronic nose, GC-MS, and texture modality, the importance of features is calculated through SHAP analysis, and the core feature set covering the four modalities is finally screened.

[0013] In one embodiment of the present invention, the interpretive analysis in step 6 includes: Single-modal interpretation: Analyze the importance and influence of features in each single modality, and clarify the role of key indicators in sensory scoring under different modalities; Fusion Model Explanation: The SHAP importance analysis graph visualization tool reveals the key global impact characteristics and mechanisms of action.

[0014] Secondly, the present invention provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the method for constructing a sensory evaluation model for the quality of the "Three Delicacies of the Earth".

[0015] Thirdly, the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the steps of the method for constructing a sensory evaluation model for the quality of the "Three Delicacies of the Earth".

[0016] The beneficial effects achieved by this invention are as follows: (1) Objective and standardized evaluation: The traditional model that relies entirely on subjective human scoring is abandoned. Multi-source machine detection data is used as the model input value and the total sensory evaluation score is the output value. This ensures the objectivity and repeatability of the evaluation from the source. The Cronbach's α coefficient reaches 0.876, and the results are highly consistent, perfectly meeting the standardization requirements of industrial production.

[0017] (2) Significant improvement in prediction accuracy: Through an innovative weighted fusion strategy, complementary information from different modalities was effectively integrated. The final fusion model achieved R... 2 With an excellent performance of 0.7426, compared to the best single-modal model in the current technology (GC-MS modality, R... 2 =0.1344), R of the fusion model of this invention 2 The accuracy was improved by 4.5 times, and the RMSE decreased from 0.6995 to 0.3815, a reduction of 45.4%, achieving accurate and stable prediction of the sensory quality of the local delicacies.

[0018] (3) Scientific model architecture adaptation: Abandoning the "one-size-fits-all" modeling approach, we scientifically select adaptation algorithms such as random forest (RF) and ridge regression for different modal data physical meaning and mathematical characteristics (such as nonlinearity of texture and collinearity of nutritional indicators), and set more relaxed thresholds for texture and nutritional indicators in the feature selection stage to ensure that key quality signals are not missed, which greatly enhances the model's ability to represent different quality dimensions.

[0019] (4) Efficient and interpretable feature engineering: The proposed MMSFS multimodal feature selection strategy integrates statistical filtering, resampling stability and model-specific screening, and can efficiently extract a concise subset containing only 12 core features from high-dimensional initial features. This not only avoids the curse of dimensionality and overfitting, but also allows the model to focus on the most discriminative key indicators. Attached Figure Description

[0020] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the embodiments of the present invention will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the drawings without creative effort.

[0021] Figure 1 This is a logical flowchart of the method for constructing the sensory evaluation model for the quality of the "Three Delicacies of the Land" of the present invention.

[0022] Figure 2 To create a scatter plot that integrates model predictions and actual values.

[0023] Figure 3 This is a graph showing the importance of SHAP features in the fusion model.

[0024] Figure 4 This is a performance comparison chart of each modal model and the fusion model. Detailed Implementation

[0025] The embodiments of the technical solution of the present invention will now be described in detail with reference to the accompanying drawings. These embodiments are merely illustrative of the technical solution of the present invention and are therefore intended to limit the scope of protection of the present invention.

[0026] Unless otherwise defined, all technical and scientific terms used in this invention have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains; the terminology used in this invention is for the purpose of describing particular embodiments only and is not intended to limit the invention; the terms “comprising” and “having”, and any variations thereof, in the specification, claims and foregoing description of the drawings are intended to cover non-exclusive inclusion.

[0027] In the description of the embodiments of this invention, technical terms such as "first" and "second" are used only to distinguish different objects and should not be construed as indicating or implying relative importance or implicitly specifying the number, specific order, or primary and secondary relationship of the indicated technical features. In the description of the embodiments of this invention, "multiple" means two or more, unless otherwise explicitly defined.

[0028] In this invention, the reference to "embodiment" means that a specific feature, structure, or characteristic described in connection with an embodiment may be included in at least some embodiments of the invention. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described in this invention can be combined with other embodiments.

[0029] In the description of the embodiments of this invention, the term "and / or" is merely a description of the relationship between associated objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. Additionally, in this invention, the character " / " generally indicates that the preceding and following associated objects have an "or" relationship.

[0030] Example 1 like Figure 1 As shown in the figure, this embodiment provides a method for constructing a sensory evaluation model for the quality of "Di San Xian" (a dish of fresh, preserved vegetables, and potatoes), including the following steps: Step 1: Obtain a multi-source basic database: The database includes sensory evaluation data and machine detection multi-source index data; Step 2, Data Preprocessing: The machine-detected multi-source index data are sequentially processed with missing value filling, outlier correction, noise smoothing, and modal standardization to eliminate the influence of dimensions and noise; Step 3, MMSFS Multimodal Feature Selection: A multi-step screening strategy is adopted, and variance screening, Bootstrap-based stability screening, modality-specific screening, and cross-modal redundancy removal are performed sequentially on the data preprocessed in Step 2 to select a subset of core features from the initial features; Step 4: Single-modal model training: Based on the characteristics of the four modalities—electronic nose, GC-MS, texture, and nutrient indicators—a random forest regression model or ridge regression model is constructed and optimized respectively, and the optimal hyperparameters of each model are determined through grid search. Step 5, Model Weighted Fusion: Based on the coefficient of determination R of each single-modal model on the validation set 2 The values ​​are calculated to normalize the fusion weights, retaining only R. 2 Single modes with a value greater than 0 are included in the fusion, and the prediction results of each mode are weighted and summed to obtain the final prediction value; Step 6, Interpretive Analysis: The SHAP (SHapley Additive exPlanations) method is used to perform feature importance analysis and visualize the decision process of the unimodal model and the fusion model.

[0031] Optionally, in step 1, the sensory evaluation data is obtained by an evaluation team selected and trained according to national standards, who scores the data from five dimensions—color, odor, taste, texture, and overall acceptability—using a 0-9 scale, and then calculates the total score using the entropy weight method.

[0032] Optionally, in step 1, the machine-detected multi-source index data includes texture indexes detected by a TPA texture analyzer, odor indexes detected by a combination of an electronic nose and gas chromatography-mass spectrometry (GC-MS), and nutritional indexes detected according to national standard methods.

[0033] Optionally, in step 3, the core feature subset consists of 12 features, including: S1 and S8 sensor signals of the electronic nose mode; (E,E)-2,4-nonadienal, octanal, 1-octen-3-ol, and diallyl disulfide content of the GC-MS mode; potato hardness (T hardness) of the texture mode; and moisture content, protein content, dietary fiber content, ash content, and fat content of the nutritional index mode.

[0034] Optionally, in step 4, the optimal configuration and performance results for training the single-modal model are as follows: The GC-MS modality uses a random forest model, and the model performance under the optimal parameters is: R 2 =0.1344, RMSE=0.6995, MAE=0.5820.

[0035] Optionally, in step 5, the weights of the model weighted fusion are based on each single-mode R. 2 The values ​​were calculated with the following weights: electronic nose model 0.32, GC-MS model 0.68, and texture and nutrient index model weighted by R. 2 If the value is not positive, the weight is 0; after weighted fusion, the performance of the final fusion model is: R 2 =0.7426, RMSE=0.3815, MAE=0.3075.

[0036] Example 2 This embodiment provides a method for constructing a sensory evaluation model for the quality of "Di San Xian" (a dish made with three kinds of vegetables, fresh vegetables, and green peppers), using the programming language Python 3.8, and includes the following steps: 1. Acquisition of multi-source basic databases (1) Sensory evaluation data: 10-20 evaluators (average gender, age 20-45 years) were selected and trained according to the "General Guidelines for the Selection, Training and Management of Sensory Analysis Evaluators" (GB / T16291.1-2012). The evaluations were conducted using a 0-9 scale, with five dimensions: color, odor, taste, texture, and overall acceptability. All evaluations were conducted in an independent sensory evaluation room, with environmental conditions meeting the requirements of GB / T 13868-2009. The specific sensory evaluation forms are shown in Table 1. The consistency of the scores was verified using Cronbach's α coefficient (α > 0.8), and the overall sensory evaluation score was calculated using the entropy weight method. Cronbach's α coefficient is a reliability coefficient used to assess consistency between evaluators or evaluation items; a value greater than 0.8 is generally considered to indicate good consistency. Table 1 Sensory Evaluation Criteria for Di San Xian (Three Delicacies from the Earth)

[0037] (2) Machine testing data: Multiple samples of "Di San Xian" (a dish of potatoes, eggplant, and green peppers) covering different cooking techniques and regional flavors were collected, homogenized according to a fixed ratio of 100g potato + 100g eggplant + 25g green pepper, and then tested. Texture indicators: The TPA texture analyzer was used with a P25 probe. The pre-test, test, and post-test speeds were 2.00, 1.00, and 2.00 mm / s, respectively. The deformation was 40-50%, and the interval between two compressions was 5s. The hardness, elasticity, cohesion, and chewiness of potatoes and eggplants were recorded. Electronic nose specifications: 18 metal oxide sensors, see Table 2, incubation at 50℃ for 30 min, detection for 100 s, cleaning for 120 s; GC-MS parameters: Gas chromatography-mass spectrometry (GC-MS, J&W VF-WAXms column, programmed temperature rise from 50℃ to 250℃), screening for key volatile compounds with a matching degree ≥80% and OAV ≥1; Nutritional indicators: Moisture content (GB 5009.3-2016), crude protein content (GB 5009.5-2016), dietary fiber (GB 5009.88-2014), ash content (GB 5009.4-2016), and fat content (GB 5009.6-2016) were tested according to national standard methods.

[0038] Table 2. Display of the Response Substances and Categories of the 18 Sensors in the Electronic Nose

[0039] 2. Data Preprocessing The machine detection data was cleaned and standardized to lay the foundation for subsequent modeling. (1) Missing value handling: The K-nearest neighbor algorithm is used to fill missing data to ensure data integrity; (2) Outlier correction: Outliers are identified based on MAD-Z scores and replaced with the median of this feature; (3) Noise smoothing and modal standardization: For high-noise time series data such as electronic nose and GC-MS, a weighted moving average window is used for smoothing. At the same time, in order to avoid the influence of the difference in dimensions between different modes, Z-score standardization is performed on the four types of data: electronic nose, GC-MS, texture and nutrient indicators.

[0040] 3. MMSFS Multimodal Feature Selection To reduce dimensionality, remove redundancy, and retain key information, this invention proposes a multimodal, multi-step feature selection strategy: (1) Variance screening: retain features with variance ≥ 0.01 and remove indices with no discrimination; (2) Stability screening: The frequency of each feature being selected into the model was calculated through 120 Bootstrap resampling. Differential frequency thresholds were set according to the characteristics of different modal data: 0.65 for electronic nose and GC-MS, and 0.5 for texture and nutrient indicators, retaining high-stability features; (3) Modality-specific screening: In view of the strong nonlinearity of texture data, the feature importance of the random forest model is used for screening; for nutrient indicators, the Pearson correlation coefficient (threshold ≥ 0.15) and mutual information (threshold ≥ 0.1) are combined for screening. (4) Cross-modal redundancy removal: Calculate the mutual information between different modal features and remove the low-priority features among highly redundant feature pairs with mutual information values ​​≥0.8 (priority is determined by the modal stability threshold). (5) SHAP importance screening: For electronic nose, GC-MS, and texture modality, the importance of features is calculated by SHAP analysis, and the core feature set covering the four modalities is finally screened to ensure data comprehensiveness and eliminate redundant information.

[0041] 4. Adaptive Unimodal Model Training and Optimization For each selected modality's core features, a suitable prediction model is trained, and hyperparameters are optimized using grid search. Five-fold cross-validation is then used for evaluation. (1) Electronic nose modality: Random forest regressor was selected, and the hyperparameter search range was max_depth=3-7, min_samples_leaf=1-3, min_samples_split=4-8, n_estimators=200-400; (2) GC-MS mode: Random forest regressor was selected, and the hyperparameter search range was max_depth=3-7, min_samples_leaf=1-3, min_samples_split=4-8, n_estimators=200-400; (3) Texture mode: In order to better capture the complex relationship between texture and senses, a random forest regressor was selected instead of the traditional linear model. The hyperparameter search range was max_depth=2-5, min_samples_leaf=2-4, min_samples_split=3-6, n_estimators=150-250; (4) Nutritional index modality: Ridge regression model was selected, with hyperparameter search range of alpha=0.05-0.2 and fit_intercept=True / False.

[0042] 5. Construction of Weighted Fusion Model Single-modal models have limited predictive power. This invention employs a weighted fusion strategy to integrate the advantages of each modality: (1) Weight calculation: based on each mode R 2 The values ​​are calculated using normalized weights, retaining only R. 2 Modes with a value greater than 0 participate in the fusion, and the fused prediction value = Σ(single mode prediction value * normalized weight). (2) Fusion prediction: The prediction results of the fusion model are obtained by weighted summation according to the weights. The final performance of the fusion model is R. 2 =0.7426, RMSE=0.3815, MAE=0.3075, compared to the optimal single-mode (R 2 =0.1344), R 2 The performance improved by 4.5 times and the RMSE decreased by 45.4%, significantly outperforming the single-modality model and demonstrating the effectiveness of multi-source data fusion. 6. SHAP Model Interpretive Analysis To enhance the transparency and practical value of the model, the SHAP framework is used for explanation: (1) Single-modal interpretation: Analyze the importance and influence direction of features in each single modality, and clarify the role of key indicators in sensory scoring under different modalities; (2) Explanation of the fusion model: Through visualization tools such as the SHAP importance analysis diagram, the key global impact characteristics and mechanisms of action are revealed, providing targeted guidance for process optimization.

[0043] Example 3 This embodiment provides model building and validation: This example uses the Python 3.8 programming language.

[0044] 1. Acquisition of multi-source basic databases (1) Sensory evaluation data: Ten evaluators (5 males and 5 females, aged 20-45 years) were selected and trained according to national standards. They scored the samples according to the five dimensions of color, odor, taste, texture, and overall acceptability using a 0-9 scale, as shown in Table 1. The scores were verified by Cronbach's α coefficient (α=0.876≥0.8), and the consistency of the scores was good. The entropy weight method was used to calculate the overall score (S=0.312* odor + 0.225* texture + 0.188* taste + 0.153* color + 0.122* overall acceptability), with odor being the key evaluation indicator.

[0045] (2) Machine detection data: 68 samples of "Di San Xian" (a dish of potatoes, eggplant, and green peppers) from different restaurants with different flavors such as Sichuan, Northeast, and Cantonese cuisine were collected. After homogenization, the samples were tested according to a fixed ratio of 100g potato + 100g eggplant + 25g green pepper. The texture, electronic nose, GC-MS and nutritional indicators were measured according to the specific steps. Among them, GC-MS screened out 16 key volatile compounds (matching degree ≥80%, OAV ≥1). 2. Data Preprocessing (1) Missing value handling: The K-nearest neighbor (K=3) algorithm is used to fill missing data; (2) Outlier correction: Outliers are identified based on the MAD-Z score (threshold 3) and replaced with the median of this feature; (3) Noise smoothing and modal standardization: The electronic nose and GC-MS time series data were processed using a weighted moving average window (window length = 5); Z-score standardization was performed on the four types of data, namely electronic nose, GC-MS, texture and nutritional indicators, to eliminate the influence of dimensions.

[0046] 3. MMSFS Multimodal Feature Selection Follow these steps to complete the feature selection, filtering out 12 core features from the initial 45 features: (1) Variance screening: retain features with variance ≥ 0.01 and remove indices with no discrimination. (2) Stability screening: 120 Bootstrap resampling cycles were performed to calculate the frequency of feature selection. The threshold for the electronic nose and GC-MS modal was set to 0.65, and the threshold for the texture and nutrient index modal was set to 0.5. Features with values ​​above the threshold were retained. (3) Modality-specific screening: Texture modalities were screened using the importance of random forest features, while nutrient index modalities were screened using a combination of Pearson correlation coefficient (≥0.15) and mutual information (≥0.1); (4) Cross-modal redundancy removal: Calculate the mutual information value between features and remove redundant features with mutual information ≥ 0.8 that have lower modal priority; (5) SHAP importance screening: The importance of features was calculated by SHAP analysis for electronic nose, GC-MS, and texture modality. Finally, 12 core features were screened: covering electronic nose (S1, S8), GC-MS (E,E)-2,4-nonadienal, octanal, 1-octen-3-ol, diallyl disulfide, texture (hardness of potato, T hardness), and nutritional indicators (moisture content, protein content, dietary fiber content, ash content, fat content).

[0047] 4. Training and Optimization of Single-Modal Models The hyperparameters were optimized using grid search, and the performance was evaluated using 5-fold cross-validation. The results are as follows: (1) Electronic nose modality: Random forest regressor, with optimal parameters max_depth=5, min_samples_leaf=2, min_samples_split=5, n_estimators=300, and performance metrics: R²=0.0639, RMSE=0.7274, MAE=0.5794; (2) GC-MS modality: Random forest regressor, with optimal parameters of max_depth=5, min_samples_leaf=2, min_samples_split=6, n_estimators=300, and performance metrics of R²=0.1344, RMSE=0.6995, MAE=0.5820; (3) Texture mode: Random forest regressor, with optimal parameters max_depth=3, min_samples_leaf=3, min_samples_split=4, n_estimators=200, and performance indicators: R²=-0.0178, RMSE=0.7585, MAE=0.5784; (4) Nutritional index modality: Ridge regression model, with optimal parameters alpha=0.1 and fit_intercept=True, performance index: R 2 =-0.0807, RMSE=0.7816, MAE=0.6327.

[0048] 5. Weighted Fusion and Performance Verification (1) Weight calculation: Only R is retained 2 Electronic nose mode and GC-MS mode with values ​​>0 were used for fusion, and normalized weights were calculated: electronic nose model weight = 0.32, GC-MS model weight = 0.68; (2) Performance of the fusion model: The test set validation results are R 2 =0.7426, RMSE=0.3815, MAE=0.3075, compared to the optimal single-mode (GC-MS mode, R 2 =0.1344), R 2 It improves accuracy by 4.5 times, reduces RMSE by 45.4%, and significantly outperforms the single-modal model in prediction accuracy; (3) Visualization of prediction results: The scatter plot of the fusion model prediction value and the actual sensory evaluation value is shown in Figure 2. The scatter points are concentrated near the ideal line (y=x), indicating that the prediction results are highly consistent with the actual quality.

[0049] 6. SHAP Model Interpretive Analysis (1) Single-modal explanation: In the GC-MS modality, (E,E)-2,4-nonadienal and octanal are the most positively correlated flavor compounds, providing grassy aroma and significantly improving sensory scores; In the texture modality, the core feature is potato hardness. Too high a value will result in hard potatoes, and too low a value will result in overly mealy potatoes, both of which will reduce sensory scores. (2) Explanation of the fusion model: Through the SHAP feature importance analysis diagram of the fusion model ( Figure 3 Analysis revealed that the most important features in the whole were (E,E)-2,4-nonadienal, octanal, and 1-octen-3-ol, indicating that volatile flavor compounds are the core factors determining the sensory quality of the three delicacies of Di San Xian. This is consistent with the result that "odor is the key indicator" in sensory evaluation, and provides clear guidance for optimizing the cooking process of Di San Xian (such as controlling the generation of volatile aldehydes).

[0050] In summary, this invention provides a method for constructing a sensory evaluation model for the quality of "Di San Xian" (a traditional Chinese dish). By integrating total sensory evaluation scores with multi-source machine detection data from electronic nose, GC-MS, texture analyzer, and national standard methods, and through data preprocessing, multimodal feature screening, single-modal model training, model weighted fusion, and SHAP interpretability analysis, a sensory evaluation model for the quality of "Di San Xian" is constructed. This method solves the problems of strong subjectivity and one-sidedness of traditional sensory evaluation and single-instrument detection, achieving an objective, accurate, and efficient evaluation of the quality of "Di San Xian". Verification shows that the coefficient of determination R of the fusion model is [value missing]. 2 The R² value reached 0.7426, and the root mean square error (RMSE) was 0.3815. Compared to the optimal single-mode model, R²... 2 The accuracy was improved by approximately 4.5 times, and the RMSE was reduced by approximately 45.4%, resulting in a significant improvement in prediction precision. This invention is applicable to the quality evaluation of "Di San Xian" and similar Chinese stir-fried vegetable dishes in the catering industry, food processing enterprises, and market supervision.

[0051] Furthermore, the present invention also provides a computer device, which may include a processor, a memory, a network interface, and a database connected via a system bus. The processor of the computer device provides computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system, computer programs, and a database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The network interface of the computer device is used for communication with external terminals via a network connection. When the computer program is executed by the processor, it causes the processor to perform the steps of the sensory evaluation model construction method for the quality of the "Three Delicacies of the Land" as described in any of the above embodiments.

[0052] For the working process, working details and technical effects of the computer device provided in this embodiment, reference may be made to the embodiments of the method for constructing the sensory evaluation model of the quality of stir-fried eggplant, green pepper and potato in the foregoing text, which will not be elaborated herein.

[0053] In addition, the present invention further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the method for constructing the sensory evaluation model of the quality of stir-fried eggplant, green pepper and potato in any of the foregoing embodiments are implemented. Among them, the computer-readable storage medium refers to a carrier for storing data, which may include, but is not limited to, floppy disks, optical discs, hard disks, flash memories, USB flash drives and / or memory sticks, etc. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices.

[0054] For the working process, working details and technical effects of the computer-readable storage medium provided in this embodiment, reference may be made to the embodiments of the method for constructing the sensory evaluation model of the quality of stir-fried eggplant, green pepper and potato in the foregoing text, which will not be elaborated herein.

[0055] Those of ordinary skill in the art can understand that all or part of the processes of the methods in the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it may include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database or other medium used in the various embodiments provided in the present application may include non-volatile and / or volatile memories. The non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. The volatile memory may include random access memory (RAM) or an external cache. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).

[0056] Although the present invention has been described in detail with reference to the foregoing embodiments, for those skilled in the art, they can still modify the technical solutions described in the foregoing embodiments, or equivalently replace some of the technical features. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A method for constructing a sensory evaluation model for the quality of "Three Delicacies from the Earth" (Di San Xian), characterized in that, Includes the following steps: Step 1: Obtain a multi-source basic database: The database includes sensory evaluation data and machine detection multi-source index data; Step 2, Data Preprocessing: The machine-detected multi-source index data are sequentially processed with missing value filling, outlier correction, noise smoothing, and modal standardization to eliminate the influence of dimensions and noise; Step 3, MMSFS Multimodal Feature Selection: A multi-step screening strategy is adopted, and variance screening, Bootstrap-based stability screening, modality-specific screening, and cross-modal redundancy removal are performed sequentially on the data preprocessed in Step 2 to select a subset of core features from the initial features; Step 4: Single-modal model training: Based on the characteristics of the four modalities—electronic nose, GC-MS, texture, and nutrient indicators—a random forest regression model or ridge regression model is constructed and optimized respectively, and the optimal hyperparameters of each model are determined through grid search. Step 5, Model Weighted Fusion: Based on the coefficient of determination R of each single-modal model on the validation set 2 The values ​​are calculated to normalize the fusion weights, retaining only R. 2 Single modes with a value greater than 0 are included in the fusion, and the prediction results of each mode are weighted and summed to obtain the final prediction value; Step 6, Interpretive Analysis: The SHAP method is used to perform feature importance analysis and visualize the decision process of the single-modal model and the fusion model.

2. The method for constructing a sensory evaluation model for the quality of "Three Delicacies of the Earth" according to claim 1, characterized in that, In step 1, the sensory evaluation data is obtained by scoring five dimensions—color, odor, taste, texture, and overall acceptability—using a 0-9 scale and then calculating the total score using the entropy weight method. The machine-detected multi-source index data includes texture indexes detected by a TPA texture analyzer, odor indexes detected by a combination of an electronic nose and gas chromatography-mass spectrometry (GC-MS), and nutritional indexes.

3. The method for constructing a sensory evaluation model for the quality of "Three Delicacies of the Earth" according to claim 2, characterized in that, In step 3, the core feature subset consists of 12 features, including: S1 and S8 sensor signals of the electronic nose mode; (E,E)-2,4-nonadienal, octanal, 1-octen-3-ol, and diallyl disulfide content of the GC-MS mode; potato hardness of the texture mode; and moisture content, protein content, dietary fiber content, ash content, and fat content of the nutritional index mode.

4. The method for constructing a sensory evaluation model for the quality of "Di San Xian" (a dish of fresh vegetables, potatoes, and green peppers) according to claim 3, characterized in that, In step 4, the optimal configuration and performance results for training the single-modal model are as follows: Electronic nose modality: Random forest regressor was selected, with hyperparameter search range of max_depth=3-7, min_samples_leaf=1-3, min_samples_split=4-8, and n_estimators=200-400; GC-MS modality: Random forest regressor was selected, with hyperparameter search range of max_depth=3-7, min_samples_leaf=1-3, min_samples_split=4-8, n_estimators=200-400; Texture modality: Random forest regressors were used instead of traditional linear models, with hyperparameter search ranges of max_depth=2-5, min_samples_leaf=2-4, min_samples_split=3-6, and n_estimators=150-250; Nutritional index modality: Ridge regression model was selected, with hyperparameter search range of alpha=0.05-0.2 and fit_intercept=True / False; The model performance under optimal parameters is: R 2 =0.1344, RMSE=0.6995, MAE=0.5820.

5. The method for constructing a sensory evaluation model for the quality of "Di San Xian" (a dish of fresh vegetables, potatoes, and green peppers) according to claim 4, characterized in that, In step 5, the weights of the model weighted fusion are based on each single-mode R. 2 The values ​​were calculated with the following weights: electronic nose model 0.32, GC-MS model 0.68, and texture and nutrient index model weighted by R. 2 If the value is not positive, the weight is 0; the prediction result of the fusion model is obtained by weighted summation according to the weights, and the performance of the final fusion model is: R 2 =0.7426, RMSE=0.3815, MAE=0.3075.

6. The method for constructing a sensory evaluation model for the quality of "Di San Xian" (a local specialty) according to claim 1, characterized in that, In step 2: Missing value handling: The K-nearest neighbor algorithm is used to fill in missing data to ensure data integrity; Outlier correction: Outliers are identified based on MAD-Z scores and replaced with the median of this feature; Noise smoothing and modal standardization: For high-noise time series data such as electronic nose and GC-MS, a weighted moving average window is used for smoothing. At the same time, in order to avoid the influence of the difference in dimensions between different modes, Z-score standardization is performed on the four types of data: electronic nose, GC-MS, texture and nutrient indicators.

7. The method for constructing a sensory evaluation model for the quality of "Three Delicacies of the Earth" according to claim 1, characterized in that, In step 3: Variance screening: retain features with variance ≥ 0.01 and remove indices with no discriminant power. Stability screening: The frequency of each feature being selected into the model is calculated through multiple Bootstrap resampling; the differential frequency threshold is set according to the characteristics of different modal data: 0.65 for electronic nose and GC-MS, and 0.5 for texture and nutrient indicators, to retain highly stable features; Modality-specific screening: Considering the strong nonlinearity of texture data, feature importance of the random forest model is used for screening. For nutritional indicators, screening was conducted using Pearson correlation coefficient and mutual information. Cross-modal redundancy removal: Calculate the mutual information between features of different modalities and remove the lower priority features from highly redundant feature pairs with mutual information values ​​≥0.8; Step 3 also includes: SHAP Importance Screening: For electronic nose, GC-MS, and texture modality, the importance of features is calculated through SHAP analysis, and the core feature set covering the four modalities is finally screened.

8. The method for constructing a sensory evaluation model for the quality of "Three Delicacies of the Earth" according to claim 1, characterized in that, The interpretive analysis in step 6 includes: Single-modal interpretation: Analyze the importance and influence of features in each single modality, and clarify the role of key indicators in sensory scoring under different modalities; Fusion Model Explanation: The SHAP importance analysis graph visualization tool reveals the key global impact characteristics and mechanisms of action.

9. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of constructing a sensory evaluation model for the quality of "Di San Xian" as described in any one of claims 1-8.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the method for constructing a sensory evaluation model for the quality of "Di San Xian" as described in any one of claims 1-8.