Electronic nose feature array optimization method applied to oyster freshness detection

By employing redundancy control and stability optimization strategies, combined with mutual information and particle swarm optimization algorithms, the electronic nose feature array was optimized, thus solving the complexity problem of features and sensor arrays in oyster freshness detection and achieving efficient and accurate detection results.

CN121808340APending Publication Date: 2026-04-07SHANDONG INST OF BUSINESS & TECH +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-29
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

Existing electronic nose technology faces challenges in detecting oyster freshness due to complex gas composition and diverse detection targets. It is difficult to quickly and flexibly optimize features and sensor arrays, resulting in insufficient detection efficiency and accuracy.

Method used

A redundancy control and stability optimization strategy is adopted. Feature grouping and redundant feature screening are performed through mutual information. Combined with particle swarm optimization algorithm and support vector machine, the optimal sensor and feature combination is constructed to optimize the electronic nose feature array.

Benefits of technology

It achieves efficient and accurate oyster freshness detection in complex gas environments, reduces the number of features and sensors, and improves the stability and accuracy of the prediction model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121808340A_ABST
    Figure CN121808340A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of food detection, and particularly relates to an electronic nose feature array optimization method applied to oyster freshness detection, and the method comprises the following steps: based on collected oyster smell data, adopting a redundancy control and stability optimization strategy, firstly carrying out grouping and redundancy feature screening through mutual information, and then carrying out feature selection; secondly, evaluating feature stability by using historical data to reduce random interference, finally, completing feature and array optimization in combination with robust feature frequency, particle swarm optimization and sensor cost performance calculation, and constructing a reliable detection model through a redundancy control and stability optimization method; according to the method, the key feature screening capability and the detection reliability and accuracy are effectively improved, the problems of feature redundancy and sensor redundancy in a complex gas environment are solved, high-accuracy distinguishing of the oyster freshness is achieved, and the method is adaptive to a dynamic detection scene.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of food detection, and particularly relates to an electronic nose feature array optimization method applied to oyster freshness detection. BACKGROUND

[0002] Oysters are important economic species in marine fisheries, and are favored by global consumers due to their high nutritional and economic value. Oyster freshness is a key quality control indicator for ensuring its food quality and safety. As a core parameter in the aquaculture industry, accurate and rapid detection of oyster freshness is of great significance to food safety and resource conservation. Although traditional detection methods such as sensory evaluation, total volatile nitrogen, biogenic amines and total viable count can reflect the quality of oysters to some extent, they are usually limited by strong subjectivity, high cost, long time consumption, complex procedures and potential damage to oysters during detection.

[0003] As a kind of bionic olfactory detection technology, electronic nose has rapidly developed and been widely applied in food quality detection field due to its low cost, fast response speed and portability. Electronic nose technology simulates the human olfactory mechanism and detects the volatile compounds released by food samples through gas sensor array, providing effective data support for oyster freshness evaluation. During the process of seafood spoilage, the gases released by oysters change with the degree of spoilage, and electronic nose can capture the changes in the composition of these gases, thereby realizing rapid classification and detection of oyster freshness. However, the amount of odor data collected by electronic nose is large, and efficient feature extraction and sensor array optimization are needed to ensure and improve the reliability and accuracy of freshness prediction. The core of feature selection and array optimization is to select the most relevant features from the original data and optimize the sensor array to improve detection efficiency and performance.

[0004] Currently, electronic nose technology has been widely applied in food freshness detection, but still faces many problems in actual detection. On the one hand, the gas composition in the detection environment is complex and variable, and the target features to be detected are diverse. On the other hand, the key challenge is to dynamically construct the optimal sensor array and feature combination from the massive electronic nose data to cope with the above complex situations. In addition, the diversity of food freshness evaluation target variables in the field of food detection further increases the detection complexity. The key problem is how to quickly, flexibly and specifically optimize the features and sensor numbers according to the specific needs of different labels, which is crucial for efficient and accurate feature selection, especially in practical applications. SUMMARY

[0005] In order to overcome the problems in the prior art, the application provides an electronic nose feature array optimization method applied to oyster freshness detection.

[0006] The technical solution of the present invention to solve the above-mentioned technical problems is as follows: This invention provides a method for optimizing an electronic nose feature array for oyster freshness detection, comprising the following steps: Step 100: Collect oyster odor data under different temperature storage conditions, and extract features from the collected oyster odor data to construct the original feature dataset; Step 200: Apply a redundancy control strategy to the original feature dataset, and use mutual information to group features and filter redundant features to obtain highly correlated groups that are highly correlated with the target variable and low redundancy groups with low redundancy; merge the highly correlated groups and the low redundancy groups to obtain the filtered feature groups. Step 300: The filtered feature groups enter the stability optimization stage. The validation set corresponding to the filtered feature groups is split, and the selection frequency and relevance score of each feature are counted. The robust feature frequency score is calculated by combining the selection frequency and relevance score of each feature. Based on the robust feature frequency score, a fitness function is constructed, and the particle swarm optimization algorithm is used to perform an efficient search for the optimal feature subset to obtain the feature subset. Step 400: The feature subset enters the array optimization and evaluation stage, and the selected features and corresponding sensors are comprehensively evaluated to output the final selected optimal sensor and feature combination; Step 500: Input the features retained by the finally selected optimal sensor into the pre-built oyster freshness prediction model to predict the oyster freshness status.

[0007] Furthermore, in step 100, an electronic nose system is used to collect oyster odor data. The electronic nose system includes a gas sensor array, and features are extracted from the collected oyster odor data.

[0008] Furthermore, step 200 specifically includes: Based on the original feature dataset, the mutual information between features and the target variable is calculated to measure the correlation between each feature and the target variable; at the same time, the redundancy between features is controlled by calculating the mutual information between features. Based on the mutual information between features and the target variable, and the mutual information between features, the maximum relevance and minimum redundancy scores of each feature are calculated. The median of each maximum relevance and minimum redundancy score is used as a threshold. Features with scores higher than the threshold are classified into the high relevance group, features with scores lower than the threshold are classified into the low relevance group, and the high relevance group is retained. In the low-relevance group, the average mutual information between features is calculated, and based on the average mutual information between features, the average redundancy of features in the low-relevance group is calculated. The average redundancy of features in the low-relevance group is used as the redundancy screening threshold. Each feature in the low-relevance group is compared with the redundancy screening threshold, and features with redundancy below the redundancy screening threshold are retained, i.e., low-redundancy features. The low-redundancy features in the low-relevance group are merged with the features in the high-relevance group, i.e., the filtered feature group.

[0009] Furthermore, based on the original feature dataset, the mutual information between features and the target variable is calculated to measure the relevance of each feature to the target variable; simultaneously, redundancy between features is controlled by calculating the mutual information between features, including: ; ; In the above formula, Features x i With target variable y The correlation; Features x i Redundancy of all other features; m It is the original feature dataset; Features x i Entropy; The entropy of the target variable; Features x i With target variable y The joint entropy; Features x i With features x j The mutual information values ​​between them.

[0010] Furthermore, based on the mutual information between features and the target variable, and the mutual information among features, the scores for maximum relevance and minimum redundancy of each feature are calculated, including: ; In the above formula, The minimum redundancy score is the maximum correlation score obtained from the calculation.

[0011] Further, in step 300, the robust feature frequency score is calculated by combining the feature selection frequency and the correlation score, including: ; In the above formula, Indicates the robust feature frequency score; Representation of features xi Correlation with the target variable y; Representation of features x i Redundancy with other features; This indicates the frequency with which a feature is selected; For the addition of a very small constant; m n This indicates the filtered feature groups.

[0012] Further, in step 300, the fitness function is constructed based on the robust feature frequency score, including: Calculate the weighted average of the robust feature frequency scores for each feature to obtain the weighted robust feature frequency; A fitness function is constructed based on weighted robust feature frequencies and using the classification accuracy of the support vector classifier as the fitness evaluation index.

[0013] Furthermore, the fitness function Specifically: ; In the above formula, This represents the classification accuracy of the support vector classifier. WRFF Indicates the weighted robust feature frequency; To control and WRFF Hyperparameters for the weights of the positive composite index.

[0014] Furthermore, step 400 specifically includes: The selected features are categorized, the number of features retained for each sensor is counted, and sensors with a number of retained features greater than or equal to the average number of features are selected, while the remaining sensors are removed. Calculate the cost-performance score of the selected sensors; based on the cost-performance scores of the sensors, output the final optimal combination of sensors and features.

[0015] Furthermore, in step 500, an oyster freshness prediction model is constructed using a support vector machine classifier.

[0016] Compared with the prior art, the present invention has the following technical effects: This invention innovatively integrates redundancy control and stability optimization strategies. Using oyster odor data collected by an electronic nose as the core, it constructs a multi-stage optimization mechanism encompassing grouping, stability optimization, and array evaluation. In the grouping stage, the maximum correlation and minimum redundancy scores dynamically adapt to the complex and ever-changing target variables and the complex detection environment. The stability optimization stage quantifies odor feature redundancy through mutual information, enhances feature stability using robust feature frequencies, and employs a particle swarm optimization algorithm to optimize feature subsets, achieving adaptive retention of high-information features and precise removal of redundant data. The array optimization and evaluation stage comprehensively evaluates the selected features and their corresponding sensors, outputting the final optimal sensor and feature combination. When this combination is input into an oyster freshness prediction model built using a machine learning classifier, it can accurately distinguish the freshness status of oysters with fewer sensors and fewer feature dimensions. Attached Figure Description

[0017] To more clearly illustrate the technical solutions and advantages in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0018] Figure 1 This is a schematic diagram of the overall workflow of the present invention; Figure 2 This is a schematic diagram illustrating the collection of oyster odor data according to the present invention; Figure 3 This is a schematic diagram of the workflow for feature and array optimization and oyster freshness prediction in this invention; Figure 4 This is the result of selecting four temperature-specific datasets after feature optimization in this invention; Figure 5 This invention relates to the sensor used for performance evaluation and specific selection of the final freshness prediction model. Detailed Implementation

[0019] To further illustrate the technical means and effects adopted by the present invention to achieve its intended purpose, the specific implementation methods, structures, features, and effects of the technical solutions proposed according to the present invention are described in detail below with reference to the accompanying drawings and preferred embodiments. Specific features, structures, or characteristics in one or more embodiments may be combined in any suitable form. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.

[0020] To address the challenges of improving the screening capability, reliability, and accuracy of key features, as well as the selection and optimization of features for different detection targets under complex odor environments, this invention explores a method for selecting the most suitable and reliable features for the detection target from a vast amount of electronic nose features and optimizing the array. The study employs a redundancy control and stability optimization strategy, dividing the entire optimization process into multiple stages. First, mutual information is used for grouping and redundant feature screening to obtain feature groups that are highly correlated with the target variable and have low redundancy. Then, historical data is used to evaluate feature stability to reduce random interference. Finally, robust feature frequency, particle swarm optimization, and sensor cost-effectiveness calculations are combined to complete feature and array optimization. After screening using the above method, in the same prediction model, not only is extremely high prediction accuracy maintained, but the number of required features is also significantly reduced. Furthermore, reference suggestions are provided for subsequent array improvements, achieving high accuracy in distinguishing oyster freshness and adapting to dynamic detection targets.

[0021] In one embodiment of the present invention, reference is made to... Figures 1-3 An optimization method for electronic nose feature arrays applied to oyster freshness detection includes the following steps: Step 100: Collect oyster odor data under different temperature storage conditions, and extract features from the collected oyster odor data to construct the original feature dataset; Step 200: Apply a redundancy control strategy to the original feature dataset, and use mutual information to group features and filter redundant features to obtain highly correlated groups that are highly correlated with the target variable and low redundancy groups with low redundancy; merge the highly correlated groups and the low redundancy groups to obtain the filtered feature groups. Step 300: The filtered feature groups enter the stability optimization stage. The validation set corresponding to the filtered feature groups is split, and the selection frequency and relevance score of each feature are counted. The robust feature frequency score is calculated by combining the selection frequency and relevance score of each feature. Based on the robust feature frequency score, a fitness function is constructed, and the particle swarm optimization algorithm is used to perform an efficient search for the optimal feature subset to obtain the feature subset. Step 400: The feature subset enters the array optimization and evaluation stage, and the selected features and corresponding sensors are comprehensively evaluated to output the final selected optimal sensor and feature combination; Step 500: Input the features retained by the finally selected optimal sensor into the pre-built oyster freshness prediction model to predict the oyster freshness status.

[0022] The following is a detailed explanation of each of the above steps: Step 100: Collect oyster odor data under different temperature storage conditions, and extract features from the collected oyster odor data to construct the original feature dataset.

[0023] To ensure the freshness of the oyster samples throughout the study, all oysters were collected from a local oyster farm in Yantai, China. The electronic nose system used for odor collection in this experiment was independently developed in our laboratory. This system integrates a sensor array consisting of 10 metal-oxide-semiconductor (MOS) sensors. The models and corresponding numbers of the sensors are as follows: TGS2602 (S1), TGS2603 (S2), TGS2612 (S3), TGS2630 (S4), MQ137 (S5), MQ135 (S6), TGS2600 (S7), TGS2610 (S8), TGS2620 (S9), and TGS2611 (S10). In each odor collection experiment, approximately 250 g of fresh oyster sample was placed in a sealed gas sampling bottle. To ensure stable experimental temperature during data collection, the sampling bottle was placed in a high-low temperature test chamber. The gas sensor array was connected to the sampling bottle via two flexible conduits: one specifically for extracting the sample gas, and the other for purging the gas chamber. The airflow rate of both tubing is regulated by an air pump and a three-way valve to ensure normal circulation of sample gas and ambient air. The collection process is as follows: Figure 2 As shown.

[0024] To simulate different environmental conditions that oysters might encounter during storage, four temperature gradients were set up: 4℃, 12℃, 20℃, and 28℃. Correspondingly, the total experimental duration for each temperature condition was set to 10 days, 7 days, 3 days, and 2 days, respectively. The odor data collection process consisted of three consecutive stages: a long gas wash, a sample gas introduction, and a short gas wash. The long gas wash stage lasted 20 minutes, each sample gas collection stage lasted 300 seconds, and the short gas wash stage also lasted 300 seconds. One sample gas collection stage combined with a subsequent short gas wash stage constituted one complete odor sample. After collecting every five complete odor samples, a long gas wash was performed to refresh the system; this cycle continued until the preset total experimental duration was reached, generating a total of 2065 samples.

[0025] To initially reduce data dimensionality, this embodiment performed feature extraction on the odor information data collected by the electronic noses from four experimental groups. For electronic nose data, feature extraction is typically achieved through time-frequency signal analysis to obtain key features with discriminative power. In this embodiment, 12 types of features were extracted from the response signals of 10 sensors, covering 6 time-domain features and 6 frequency-domain features. The time-domain features include: signal response area (TD1), steady-state mean (250-350s of each response signal), peak response, differential mean, skewness, and kurtosis. The frequency-domain signals include: sum of wavelet high-frequency coefficients, sum of wavelet low-frequency coefficients, wavelet energy, Euclidean distance calculated from wavelet coefficients, DC component of Fast Fourier Transform, and amplitude of the first harmonic of Fast Fourier Transform.

[0026] After feature extraction, 3 out of every 5 samples were assigned to the validation set, and the remaining 2 were assigned to the test set. The samples were also classified according to different durations to validate and test the subsequent methods. The specific data and classification at each temperature are shown in Table 1.

[0027] Table 1. Classification of oyster characteristic data at different temperatures

[0028] Step 200: Apply a redundancy control strategy to the original feature dataset, and use mutual information to group features and filter redundant features to obtain highly correlated groups that are highly correlated with the target variable and low redundancy groups with low redundancy; merge the highly correlated groups and the low redundancy groups to obtain the filtered feature groups.

[0029] As an example, this step 200 includes the following sub-steps: Step 210: Based on the original feature dataset, calculate the mutual information between features and the target variable to measure the relevance of each feature to the target variable; simultaneously, control feature redundancy by calculating the mutual information between features, specifically including the following formula: (1); (2); In the above formula, Features x i With target variable y The correlation; Features x i Redundancy of all other features; m It is the original feature dataset; Features x i Entropy; The entropy of the target variable; Features x i With target variable y The joint entropy; Features x i With features x j The mutual information values ​​between them.

[0030] Step 220: Based on the mutual information between features and the target variable, and the mutual information between features, calculate the maximum relevance and minimum redundancy scores for each feature. Using the median of each maximum relevance and minimum redundancy score as a threshold, features with scores higher than the threshold are classified into the high relevance group, features with scores lower than the threshold are classified into the low relevance group, and the high relevance group is retained. The specific steps include the following formula: (3); (4); (5); In the above formula, The minimum redundancy score is the maximum relevance obtained from the calculation; T M This is the grouping threshold; m h The group with high correlation; m l The group with low correlation.

[0031] Step 230: In the low-relevance group, calculate the average mutual information between features; and based on the average mutual information between features, calculate the average redundancy of features in the low-relevance group; use the average redundancy of features in the low-relevance group as the redundancy screening threshold, compare each feature in the low-relevance group with the redundancy screening threshold, and retain features with redundancy lower than the redundancy screening threshold, i.e., low-redundancy features; merge the low-redundancy features in the low-relevance group with the features in the high-relevance group, i.e., the filtered feature group.

[0032] In the low correlation group m l In this process, by calculating the mutual information matrix of the features, low-redundancy features are selected, and the features of the low-redundancy group are denoted as follows: x 1 ,x 2 ,x 3 …,x ml Thus, a mutual information submatrix is ​​constructed. As shown in formula (6).

[0033] The redundancy of features in the low-correlation group is denoted as Calculate each feature within the low-correlation group x i Other m l -1 is the average mutual information of the features, used to quantify the redundancy of each feature relative to the other features, as shown in formula (7).

[0034] Higher redundancy indicates that the feature shares more information with other features and has lower independence. To filter out low-redundancy features, the average redundancy of features in low-correlation groups is calculated based on the average mutual information between features. T low , which is used as the redundancy screening threshold, as shown in formula (8).

[0035] (6); (7); (8); (9); After calculating the redundancy threshold, each feature in the low-relevance group is compared with this threshold. If the redundancy of a feature is lower than the threshold, it is considered a low-redundancy feature and retained; if the redundancy of a feature is higher than the threshold, it is considered a high-redundancy feature and discarded. After filtering out the low-redundancy features, the low-redundancy features in the low-relevance group are merged with the features in the high-relevance group to form a feature group for the next stage of filtering. m n As shown in formula (9).

[0036] Step 300: The filtered feature groups enter the stability optimization stage. The validation set corresponding to the filtered feature groups is split, and the selection frequency and relevance score of each feature are counted. The robust feature frequency score is calculated by combining the selection frequency and relevance score of each feature. Based on the robust feature frequency score, a fitness function is constructed, and the particle swarm optimization algorithm is used to perform an efficient search for the optimal feature subset to obtain the feature subset.

[0037] As an example, step 300 specifically includes the following sub-steps: Step 310: Split the validation set corresponding to the selected feature groups, and count the selection frequency of each feature, the correlation between the feature and the target variable, and the correlation between features; combine the selection frequency of each feature, the correlation between the feature and the target variable, and the correlation between features to calculate the robust feature frequency score.

[0038] The validation set corresponding to the selected feature groups is split. Specifically, the validation set is randomly divided into 5 folds and the grouping stage is repeated 10 times. The selection frequency and relevance score of each feature are calculated, and a total of 50 times are completed. The robust feature frequency score is calculated by combining the selection frequency and relevance score of each feature.

[0039] The robust feature frequency method is used to calculate the feature robustness score. The score combines the maximum relevance and minimum redundancy scores with the feature selection frequency, and is a comprehensive score that measures the high relevance, low redundancy, and strong stability of features. (10); In the above formula, RFF Indicates the robust feature frequency score; Representation of features x i With target variable y The correlation; Representation of features x i Redundancy with other features; This indicates the frequency with which a feature is selected, reflecting its inclusion in the split validation set.m n This represents the filtered feature groups; The added minimum constant is used to avoid division by zero errors. Adding 1 to the selected frequency reduces the impact of zero frequency on the calculation and ensures that the influence of frequency on feature score increases linearly, as shown in Equation (10).

[0040] Filtered feature groups m n Redundancy in Including global redundancy and local redundancy, as shown in formula (11): (11); In the above formula, α , β These represent the global redundancy and local redundancy weights, respectively, in this implementation. α and β The values ​​are 0.6 and 0.4 respectively; Representation of features x i In feature set m n The global redundancy in the formula is calculated in the same way as the redundancy in formula (2); Representation of features x i In feature set m n The local redundancy in the feature is defined as the average mutual information value between the feature and the five features with the highest correlation, and its calculation formula is shown in (12): (12); Multiplying frequency by correlation using formula (10) strengthens the importance of features that are frequently selected in multiple training sets. For example, if a feature is frequently selected in multi-fold cross-validation, its frequency value is high, and its score will increase accordingly after multiplying by correlation, making it more likely to be selected as the final feature, which helps to improve the stability of the subsequent freshness prediction model.

[0041] Step 320: After calculating the robust feature frequency score to quantify the stability and redundancy of the features, a stress function is constructed based on the robust feature frequency score and the classification accuracy of the support vector classifier. Then, the particle swarm optimization algorithm is used to perform an efficient search for the optimal feature subset to obtain the feature subset.

[0042] The particle swarm optimization algorithm is used to simulate swarm intelligence behavior and perform optimal feature subset search. By having each particle in the swarm search within the solution space and share information, the global optimum is found. Each particle represents a feature subset, and its position is represented by a binary vector. Particles continuously update their position in the search space based on their individual and global optimum solutions, and their direction of motion is determined by inertia weights, individual cognitive factors, and social cognitive factors.

[0043] In this embodiment, the classification accuracy of the support vector classifier is used as the fitness evaluation index to improve the stability and generalization ability of the final feature subset. Specific parameter settings are as follows: the inertia weight is set to 0.9 to control the influence of particle velocity; both the individual cognitive factor and the social cognitive factor are set to 2 to control the speed at which particles move towards their optimal position and the global optimal position; the number of particles is set to 30, and the number of iterations is set to 50.

[0044] Calculate the robust feature frequency score for each feature. RFF The weighted average of the scores yields the weighted robust feature frequency. WRFF That is, feature subset m n Robust Feature Frequency Score for All Features RFF The weighted average of the scores is shown in formula (13); In fitness function middle, To control With weighted robust characteristic frequency WRFF The hyperparameters of the positive composite index weights are used to balance the model's predictive power and the quality of the feature subsets. In this embodiment, Set it to 0.1, as shown in formula (14).

[0045] (13); (14); Feature subset optimized by particle swarm optimization algorithm m s Enter the array optimization and evaluation phase.

[0046] Step 400: The feature subset enters the array optimization and evaluation stage to obtain the final selected sensor and retained features.

[0047] First, the features selected in the previous stage are categorized, and the number of features retained for each sensor is counted. Sensors with a retained feature count greater than or equal to the average feature count are selected, and the remaining sensors are eliminated. Then, the cost-effectiveness score of the selected sensors is calculated. P Cost-performance ratio score P The contribution of the remaining sensors to the model performance can be quantified more intuitively, as shown in formula (15).

[0048] (15); In the above formula, R s ( x i After the above screening stages, the sensor s Chinese characteristics x i Other features x j Redundancy; m s Indicates sensor s The number of features; MI ( x i , x j Mutual information between features; m s ( m s -1) / 2 is from m s The number of all possible combinations of selecting two features from a set of features.

[0049] After completing all the above stages, the final selected sensor and features, as well as the cost-effectiveness score of the selected sensor, are output.

[0050] Step 500: Input the features retained by the finally selected optimal sensor into the oyster freshness prediction model built by the support vector machine classifier, and finally obtain a more accurate oyster freshness status with fewer sensors and features.

[0051] In the oyster freshness classification prediction of this invention, Support Vector Machine (SVM) is selected as the classification prediction model. SVM is a supervised learning method based on the maximum margin principle. Its core is to construct an optimal hyperplane to separate data points of different categories and maximize the margin between the hyperplane and the category boundary, thereby improving the model's generalization ability and classification accuracy. To ensure that the selected features perform optimally, a grid optimization method is used to fine-tune the model's hyperparameters. Specifically, the optimization targets include regularization parameters. C Kernel function types (including linear kernels and radial basis function kernels) and γ parameter.

[0052] For radial basis function kernels γ The parameter's value range is set to ['scale', 'auto', 0.00001, 0.0001, 0.001, 0.01, 0.1]. Where, when γ When the parameter is 'scale', γ The value is calculated according to formula (16); whenγ When parameter='auto', γ The value is calculated according to formula (17). X var It is the characteristic variance. n features It is the number of features; (16); (17); Recursive feature elimination is a classic feature selection method that has been widely applied and improved in various fields and tasks in recent years, especially when combined with support vector machine classifiers, demonstrating excellent feature optimization results. Recursive feature elimination optimizes the feature subset by recursively training the model and progressively removing the least important features, thereby improving the model's accuracy and generalization ability. In this invention, the recursive feature elimination method is used as a comparative experiment to evaluate the performance of the feature selection and optimization process compared to the method of this invention. The final results are shown in Table 2.

[0053] Table 2 Comparative Experiments

[0054] This invention applies to the original 120 features, and the specific selection results are as follows: Figure 4 As shown in the figure, this graph illustrates the feature selection results for four temperature-specific datasets. In the heatmap, yellow represents selected features and blue represents unselected features. Figure 4 Figures 4(a), 4(b), 4(c), and 4(d) show the feature selection results for datasets at 4℃, 12℃, 20℃, and 28℃, respectively. The number of features selected varies for each temperature dataset: 34 features for the 4℃ dataset, 29 for the 12℃ dataset, 23 for the 20℃ dataset, and 42 for the 28℃ dataset. These selected features are then sent to the array optimization and evaluation stage. In this stage, the invention categorizes the selected features by sensor and calculates the number of features retained for each sensor. Sensors with a retained feature count greater than or equal to the average feature count are selected, while the rest are filtered out, as shown in Table 2 and... Figure 5As shown in Figure 5(a), the number of sensors in the odor datasets at 4℃, 20℃, and 28℃ was reduced by 40% (from 10 to 6), and the number of sensors in the 12℃ dataset was reduced by 50% (from 10 to 5). In terms of feature selection, compared with the original 120 features, the number of optimized features was significantly reduced. The reduction rates for the four temperature datasets were 79.17%, 83.33%, 85.0%, and 75.0%, respectively, and the final number of retained features were 25, 20, 18, and 30, corresponding to the 4℃, 12℃, 20℃, and 28℃ datasets. Overall, while reducing the number of sensors and features, the optimized model improved the accuracy of the test set in most cases, and the improvement was more significant at higher storage temperatures. Specifically, as shown in Figure 5(b), the sensors selected after optimization at 4℃ were S1, S2, S4, S6, S8, and S9, among which S1 and S6 contributed the most features. According to the cost-performance ratio calculation in the array optimization method of this invention, S6 had the highest cost-performance score and S8 had the lowest. Figure 5 (c) shows the optimization results at 12℃. The selected sensors are S1, S3, S5, S6 and S9. S5 contributes the most features, and S6 has the highest cost performance score while S3 has the lowest. Figure 5 The optimized sensors at 20℃ (d) include S1, S2, S3, S5, S6, and S9. S6 not only contributes the most features but also has the highest cost-effectiveness score, while S2 has the lowest score. Finally... Figure 5 The middle (e) shows the optimization results at 28℃. The selected sensors are S2, S3, S4, S7, S8 and S9. S8 contributes the most features and has the highest cost performance score, while S9 has the lowest score.

[0055] In the field of food testing, changes in the content of volatile organic compounds (VOCs) are often used as a core analytical indicator. Existing research has confirmed the great potential of electronic nose systems in classifying and predicting VOC content in this field. This invention is also applicable to feature selection and array optimization of electronic noses, especially performing well in dynamic gas environments and scenarios with varying detection targets. To verify the flexible adaptability of this invention to changes in detection targets under complex gas environments, the target variable is replaced with the relative content of various organic volatiles during the oyster spoilage process, analyzed using gas chromatography-mass spectrometry (GC-MS), to predict the approximate content of a certain type of organic volatile in oysters at a given moment. The relative contents of various organic compounds at four storage temperatures are shown in Table 3.

[0056] Table 3. Relative contents of organic compounds at various storage temperatures obtained by gas chromatography-mass spectrometry (GC-MS) analysis.

[0057] This invention uses the maximum correlation between calculated features and labels for initial feature selection. The variation patterns of different labels significantly affect the joint distribution between labels and features, thus altering the calculated mutual information value. This change further influences subsequent feature selection and array optimization results, a particularly significant effect when processing organic compounds with irregularly varying characteristics. This invention can dynamically and rapidly acquire information under different labels, facilitating efficient feature selection for diverse monitoring and classification targets, and enabling rapid optimization of features and sensor arrays.

[0058] In this invention, the relative contents of alcohols and aldehydes are used as examples, and their relative contents are used as target variable labels. The specific classification results are shown in Table 4.

[0059] Table 4. Array optimization results of the method of the present invention with alcohols and aldehydes as target variables.

[0060] This invention uses the relative content of alcohols as the target variable to select the optimal feature set and sensor array at different temperatures. The number of selected features and sensors is adjusted accordingly with changes in storage temperature. At 4℃, 28 features and 4 sensors were selected; at 12℃, 25 features and 5 sensors were selected; at 20℃, 26 features and 5 sensors were selected; and at 28℃, the number of features and sensors increased to 46 and 7, respectively. Regarding classification accuracy, this invention, combined with a support vector classifier and the optimized feature and sensor combination at 20℃, achieved the highest accuracy of 97.46%, while the accuracy on the 4℃ test set was relatively lower at 89.33%. At 12℃ and 28℃, the accuracies were 97.10% and 92.68%, respectively. When aldehydes are used as the target variable, the selection of features and sensors varies with temperature: 19 features and 3 sensors were selected at 4℃; 24 features and 4 sensors at 12℃; 35 features and 6 sensors at 20℃; and 20 features and 3 sensors at 28℃. Regarding classification accuracy, this method, combined with an optimized support vector classifier at 20℃, achieved an accuracy of 98.31%, while the accuracy at 4℃ was 86.80%, lower than at other temperatures. The accuracies at 12℃ and 28℃ were 96.74% and 92.68%, respectively.

[0061] The above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.

Claims

1. A method for optimizing an electronic nose feature array for detecting oyster freshness, characterized in that, Includes the following steps: Step 100: Collect oyster odor data under different temperature storage conditions, extract features from the collected oyster odor data, and construct the original feature dataset; Step 200: Apply a redundancy control strategy to the original feature dataset, and use mutual information to group features and filter redundant features to obtain highly correlated groups that are highly correlated with the target variable and low redundancy groups with low redundancy; merge the highly correlated groups and the low redundancy groups to obtain the filtered feature groups. Step 300: The filtered feature groups enter the stability optimization stage. The validation set corresponding to the filtered feature groups is split, and the selection frequency and relevance score of each feature are counted. The robust feature frequency score is calculated by combining the selection frequency and relevance score of each feature. Based on the robust feature frequency score, a fitness function is constructed, and the particle swarm optimization algorithm is used to perform an efficient search for the optimal feature subset to obtain the feature subset. Step 400: The feature subset enters the array optimization and evaluation stage, and the selected features and corresponding sensors are comprehensively evaluated to output the final selected optimal sensor and feature combination; Step 500: Input the features retained by the finally selected optimal sensor into the pre-built oyster freshness prediction model to predict the oyster freshness status.

2. The method for optimizing an electronic nose feature array for oyster freshness detection according to claim 1, characterized in that, In step 100, an electronic nose system is used to collect oyster odor data. The electronic nose system includes a gas sensor array and performs feature extraction on the collected oyster odor data.

3. The method for optimizing an electronic nose feature array for oyster freshness detection according to claim 1, characterized in that, Step 200 specifically includes: Based on the original feature dataset, the mutual information between features and the target variable is calculated to measure the correlation between each feature and the target variable; at the same time, the redundancy between features is controlled by calculating the mutual information between features. Based on the mutual information between features and the target variable, and the mutual information between features, the maximum relevance and minimum redundancy scores of each feature are calculated. The median of each maximum relevance and minimum redundancy score is used as a threshold. Features with scores higher than the threshold are classified into the high relevance group, features with scores lower than the threshold are classified into the low relevance group, and the high relevance group is retained. In the low-relevance group, the average mutual information between features is calculated, and based on the average mutual information between features, the average redundancy of features in the low-relevance group is calculated. The average redundancy of features in the low-relevance group is used as the redundancy screening threshold. Each feature in the low-relevance group is compared with the redundancy screening threshold, and features with redundancy below the redundancy screening threshold are retained, i.e., low-redundancy features. The low-redundancy features in the low-relevance group are merged with the features in the high-relevance group, i.e., the filtered feature group.

4. The method for optimizing an electronic nose feature array for oyster freshness detection according to claim 3, characterized in that, Based on the original feature dataset, the mutual information between features and the target variable is calculated to measure the relevance of each feature to the target variable; Simultaneously, redundancy between features is controlled by calculating mutual information between features, including: ; ; In the above formula, Features x i With target variable y The correlation; Features x i Redundancy of all other features; m It is the original feature dataset; Features x i Entropy; The entropy of the target variable; Features x i With target variable y The joint entropy; Features x i With features x j The mutual information values ​​between them.

5. The method for optimizing an electronic nose feature array for oyster freshness detection according to claim 4, characterized in that, Based on the mutual information between features and the target variable, and the mutual information among features, the scores for maximum relevance and minimum redundancy of each feature are calculated, including: ; In the above formula, The minimum redundancy score is the maximum correlation score obtained from the calculation.

6. The method for optimizing an electronic nose feature array for oyster freshness detection according to claim 1, characterized in that, In step 300, the robust feature frequency score is calculated by combining the feature selection frequency and the correlation score, including: ; In the above formula, Indicates the robust feature frequency score; Representation of features x i With target variable y The correlation; Representation of features x i Redundancy with other features; This indicates the frequency with which a feature is selected, reflecting the inclusion of that feature in the split validation set; This is a very small constant added to avoid division by zero errors; m n This indicates the filtered feature groups.

7. The method for optimizing an electronic nose feature array for oyster freshness detection according to claim 6, characterized in that, In step 300, a fitness function is constructed based on the robust feature frequency score, including: Calculate the weighted average of the robust feature frequency scores for each feature to obtain the weighted robust feature frequency; A fitness function is constructed based on weighted robust feature frequencies and using the classification accuracy of the support vector classifier as the fitness evaluation index.

8. The method for optimizing an electronic nose feature array for oyster freshness detection according to claim 7, characterized in that, fitness function Specifically: ; In the above formula, This represents the classification accuracy of the support vector classifier. WRFF Indicates the weighted robust feature frequency; To control and WRFF Hyperparameters for the weights of the positive composite index.

9. The method for optimizing an electronic nose feature array for oyster freshness detection according to claim 1, characterized in that, Step 400 specifically includes: The selected features are categorized, the number of features retained for each sensor is counted, and sensors with a number of retained features greater than or equal to the average number of features are selected, while the remaining sensors are removed. Calculate the cost-performance score of the selected sensors; based on the cost-performance scores of the sensors, output the final optimal combination of sensors and features.

10. The method for optimizing an electronic nose feature array for oyster freshness detection according to claim 1, characterized in that, In step 500, a support vector machine classifier is used to construct an oyster freshness prediction model.