Frying oil quality detection method based on cascade feature screening optimization and ultrasonic diagnosis technology
By combining cascaded feature screening optimization with ultrasound diagnostic technology, the problem of insufficient feature utilization in frying oil quality testing has been solved, achieving high-precision, rapid, and non-destructive testing, which is applicable to the quality monitoring of frying oil for different types of oil and ingredients.
Patent Information
- Application Number
- CN202610282060.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-03-10
- Publication Date
- 2026-04-07
AI Technical Summary
Existing ultrasonic technology has problems in the quality detection of frying oil, such as insufficient utilization of features, inadequate correlation between features and physicochemical indicators, and limited model generalization ability, making it difficult to meet the needs of industrial online real-time monitoring.
A cascaded feature selection optimization method is adopted. Through ultrasonic signal preprocessing, Spearman correlation coefficient screening, consensus-stability comprehensive feature scoring index (CSFS) and machine learning model, the correlation between multidimensional acoustic features and physicochemical indicators is extracted to construct a high-precision and robust quality inspection model.
It enables rapid, non-destructive, and high-precision detection of frying oil quality, sensitively reflecting the complex degradation patterns of oils during repeated high-temperature use. It is applicable to different types of oils and ingredients, possesses good robustness and generalization ability, and supports online monitoring.
Smart Images

Figure CN121805429A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of food processing and testing technology, and in particular to a method for testing the quality of frying oil based on cascade feature screening optimization and ultrasonic diagnostic technology. Background Technology
[0002] Frying oil is widely used in food processing, and its quality directly affects food safety and product flavor. Traditional methods for assessing frying oil quality mainly rely on the determination of physicochemical indicators such as acid value, viscosity, density, polar components, saponification value, and iodine value. Although traditional physicochemical testing methods have high accuracy, they generally suffer from problems such as cumbersome sample pretreatment, long testing cycles, high consumption of chemical reagents, and difficulty in meeting the needs of industrial online real-time monitoring. In addition, commonly used indicators such as acid value and iodine value have limited sensitivity and cannot fully reflect the complex deterioration patterns of frying oil during repeated high-temperature use.
[0003] In recent years, ultrasonic diagnostic technology has been introduced into the field of frying oil quality testing due to its advantages of non-destructiveness, speed, and high sensitivity. Existing research shows that ultrasonic parameters are closely related to the physicochemical indicators of oils, such as polar components, acid value, viscosity, and total polar substances. Furthermore, during frying, the ultrasonic wave propagation speed and attenuation coefficient change significantly with variations in polar components, polymer content, and viscosity. This indicates that ultrasonic parameters can be used to reflect the degree of oil deterioration, and ultrasonic testing can serve as a rapid and non-destructive industrial quality assessment method. Patent application CN120044138A proposes an ultrasonic testing method based on stepwise temperature compensation. By constructing a multi-temperature sub-model within the 22–32℃ range and combining partial least squares and variable importance projection for feature selection, it effectively compensates for the influence of ambient temperature changes on sound wave propagation characteristics, improving the accuracy of rapeseed oil quality testing at different temperatures. However, this method still has certain limitations: First, its model construction is based on a single oil type, lacking adaptability to the acoustic differences of different vegetable oils, resulting in insufficient model versatility; second, the ultrasonic features extracted by this method are limited in dimension, mainly focusing on a few features such as sound velocity, attenuation coefficient, and frequency domain peak value, failing to fully cover the multi-level structural features of ultrasonic signals and lacking sensitivity to complex changes; in addition, variable importance projection is difficult to handle redundant relationships between features and the different requirements of different physicochemical indicators, thus affecting the stability and generalization ability of the model in high-dimensional space. Existing studies mostly rely on basic time-domain features such as ultrasonic propagation speed, attenuation coefficient, amplitude, and flight time, while insufficient mining of features in the frequency domain and time-frequency domain makes it difficult to comprehensively reflect the complex chemical and physical changes that occur in oils during high-temperature recycling.
[0004] In summary, existing methods for detecting the quality of edible oils using ultrasonic technology generally suffer from problems such as insufficient utilization of features, inadequate correlation between features and physicochemical indicators, and limited model generalization ability. Summary of the Invention
[0005] To address the shortcomings of existing technologies, this invention provides a method for detecting the quality of frying oil based on cascaded feature screening optimization and ultrasonic diagnostic technology. This method aims to overcome the drawbacks of traditional physicochemical testing methods, such as complex operation, long processing time, high cost, and poor real-time performance, while also compensating for the deficiencies of existing ultrasonic testing technology. The technical solution provided by this invention includes: A method for detecting the quality of frying oil based on cascaded feature screening optimization and ultrasound diagnostic technology includes the following steps: S1. Under constant oil temperature conditions, collect ultrasonic signals and data on multiple physicochemical indicators of frying oil; S2. The ultrasonic signal is preprocessed, and the first four echoes of the ultrasonic signal are retained as the preprocessed signal. S3. Extract acoustic features based on the preprocessed signal, and perform outlier processing on the acoustic features to obtain preprocessed features; S4. Iterate through and calculate the Spearman correlation coefficient between each of the preprocessed features and multiple of the physicochemical indicators; select the preprocessed features whose absolute value of the Spearman correlation coefficient with at least one of the physicochemical indicators is greater than a preset threshold as first-level screening features; S5. Construct a Consensus-Stability Feature Score (CSFS) to further filter the primary screening features and obtain secondary screening features; S6. Input the secondary screening features into a preset machine learning model to obtain prediction results of multiple physicochemical indicators, thereby detecting the quality of the frying oil.
[0006] Preferably, the preprocessing method in step S2 includes: S201. Extract the original ultrasonic signal from the initial position; S202. Based on the intercepted ultrasonic signal, retain the signal components with amplitudes exceeding the preset amplitude to obtain a preliminary preprocessed signal; S203. Based on the preliminary preprocessing signal, four expected time windows corresponding to the first four echoes are defined, and the first four echoes are extracted and retained from the preliminary preprocessing signal as the preprocessing signal.
[0007] Preferably, in step S3, the acoustic features include time-domain features: the maximum amplitude of each of the first four echoes, the flight time of each of the first four echoes, kurtosis, skewness, signal entropy, absolute energy, area under the time-domain spectrum curve, and average number of peaks.
[0008] Preferably, in step S3, the acoustic features include frequency domain features: spectral kurtosis, spectral skewness, area under the spectral curve, maximum spectral amplitude, frequency corresponding to the maximum amplitude, cumulative energy distribution of the first 25% of energy frequency points, cumulative energy distribution of the first 50% of energy frequency points, cumulative energy distribution of the first 75% of energy frequency points, total power spectrum, median frequency, power spectrum bandwidth, and spectral entropy.
[0009] Preferably, in step S3, the acoustic feature includes the speed of sound of the ultrasonic wave.
[0010] Preferably, the specific steps for constructing the CSFS include: S501. The primary screening features are used as independent variables and the physicochemical indicators are used as dependent variables. They are input into multiple preset machine learning models respectively. The average importance of the corresponding primary screening features in the preset machine learning models is taken as the consensus importance score C of the corresponding primary screening features. S502. Perform stratified self-sampling N times according to the combination of oil type × ingredient to obtain N data subsets; calculate the C of the primary screening feature of each data subset and sort them, and count the percentage of times each primary screening feature enters the top K positions in the N sorting, as the corresponding stability index SSI. S503. Using a nonparametric test method, calculate the significance level between each of the first-level screening features and each of the physicochemical indicators, and use it as the q value, and construct an inhibition factor (1-q). S504. Construct CSFS, and its calculation formula is as follows: ; in, Score the consensus importance of the j-th feature. As a stability index, The significance level is indicated by .
[0011] Preferably, the methods for obtaining secondary screening features include: S505. Use the CSFS to sort the first-level screening features in descending order and select F features as screening features; S506. Dummy variable encoding is performed on the oil type and ingredients, and then merged with the screening features to obtain secondary screening features.
[0012] Compared with the prior art, the present invention has the following beneficial effects: This invention enables rapid, non-destructive, and high-precision detection of the physicochemical state of oils. Compared with traditional laboratory physicochemical testing methods, the detection time can be reduced to several minutes. It comprehensively characterizes the complex degradation patterns of frying oils during repeated high-temperature use. It solves the problems of high feature redundancy in ultrasound and the different requirements of various physicochemical indicators. Through a cascaded feature screening and optimization method, it achieves a full correlation between features and physicochemical indicators. At the same time, the detection model of this invention maintains high robustness and generalization ability under different oil types and food conditions, effectively avoiding overfitting and performance fluctuations. Attached Figure Description
[0013] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0014] Figure 1 This is a schematic diagram of a frying oil quality detection method based on cascaded feature screening optimization and ultrasonic diagnostic technology, provided in a preferred embodiment of the present invention. Detailed Implementation
[0015] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0016] Example 1 like Figure 1 As shown, this embodiment provides a method for detecting the quality of frying oil based on cascaded feature screening optimization and ultrasonic diagnostic technology, including the following steps: S1. Under constant oil temperature conditions, ultrasonic signals of frying oil are collected, and data on multiple physicochemical properties of the frying oil are obtained.
[0017] To obtain stable and reliable ultrasonic signals and physicochemical index data, in some preferred embodiments, the ultrasonic pulse echo device used in this invention includes a computer, a pulse transceiver, an oscilloscope, an ultrasonic probe, a sample stage, and a water bath. The detection method is as follows: First, the fried oil sample is pretreated by filtering the oil sample (approximately 75 mL) through a 100-mesh filter to remove oil residue and impurities, and then placing it into a dedicated ultrasonic measurement sample cell to ensure that the propagation path of the sound wave in the medium is uniform. Subsequently, the sample cell is placed in a constant temperature device (water bath) to balance the oil sample temperature at 30±0.1℃ to eliminate the influence of temperature fluctuations on the ultrasonic propagation characteristics.
[0018] In some preferred embodiments, after temperature equilibrium is achieved, the ultrasonic probe is vertically inserted into the center of the sample cell, with the bottom of the probe maintained at a fixed distance from the bottom of the container to avoid interference from multiple reflections. Simultaneously, a thin temperature probe is placed outside the sound path to achieve real-time monitoring and feedback control of the oil temperature, ensuring temperature stability throughout the measurement process. After temperature stabilization, the ultrasonic pulse-echo device is activated, emitting a broadband pulse signal and receiving the pulse-echo signal from the oil sample. Multiple independent echo signals are continuously acquired to reduce random errors and improve signal repeatability and signal-to-noise ratio. After acquisition, all ultrasonic signal data is stored in a computer as raw data input for subsequent signal preprocessing and feature extraction. Physicochemical indicators, including acid value, viscosity, density, polar components, saponification value, and iodine value, are simultaneously tested on the same batch of oil samples. The results are paired with the corresponding ultrasonic signals to construct an "ultrasonic signal-physicochemical indicator" dataset, laying the foundation for subsequent modeling and analysis.
[0019] S2. The ultrasonic signal is preprocessed, and the first four echoes of the ultrasonic signal are retained as the preprocessed signal.
[0020] During the acquisition of ultrasonic signals, irrelevant signal components such as system noise and background interference are often present. Directly extracting features from the acquired signals will lead to signal feature distortion and reduced correlation, thus affecting the accuracy and stability of subsequent models. This is especially true in high-viscosity, strongly attenuating media like frying oil, where sound wave propagation is easily affected by microbubbles, interface reflections, and equipment electrical noise. Therefore, preprocessing the raw signal is essential. By removing noise, extracting effective information, and standardizing data length and amplitude range, high-quality input signals can be provided for subsequent feature extraction and modeling. Common ultrasonic signal preprocessing methods often employ a combination of overall noise reduction and signal truncation.
[0021] S3. Extract acoustic features based on the preprocessed signal, and perform outlier processing on the acoustic features to obtain preprocessed features.
[0022] The feature extraction process transforms the high-dimensional, redundant raw ultrasonic signal into key parameters that characterize changes in the physicochemical properties of oils, thereby achieving dimensionality reduction and effective characterization, and providing stable and interpretable input features for subsequent models. In the field of ultrasonic detection and analysis, various feature extraction methods for ultrasonic signals exist, including time-domain feature analysis methods (such as waveform amplitude, echo time of flight, and echo energy); frequency-domain feature analysis methods (such as dominant frequency, spectral centroid, and power spectral density); and time-frequency domain analysis methods (such as wavelet transform and short-time Fourier transform).
[0023] In some preferred embodiments, in order to eliminate extreme value interference caused by certain accidental factors during the acquisition process, avoid affecting the data distribution and the true correlation between features and physicochemical indicators, and ensure the stability of subsequent feature screening and model training, the present invention preprocesses the acoustic features. The feature preprocessing method can be the interquartile range method, or conventional technical means used by those skilled in the art according to the actual situation or site requirements.
[0024] S4. Iterate through and calculate the Spearman correlation coefficient between each of the preprocessed features and multiple of the physicochemical indicators; select the preprocessed features whose absolute value of the Spearman correlation coefficient with at least one of the physicochemical indicators is greater than a preset threshold as first-level screening features.
[0025] Spearman correlation coefficient is a nonparametric statistic used to measure the degree of monotonic correlation between two variables. The primary screening features refer to the feature set remaining after removing irrelevant features in step S4, serving as the basis for subsequent in-depth screening. Common methods for primary screening features include variance screening, univariate statistical validation, or initial screening directly based on model importance. However, these methods are difficult to specifically remove features that are not correlated with any of the physicochemical indicators. For example, a feature might be judged as "important" by the model due to data fluctuations, but in reality, it has no significant correlation with any of the physicochemical indicators of oil quality. Retaining such features would increase the complexity of subsequent models, introduce noise, or even mislead the final screening results. To address these issues, in some preferred embodiments, this invention calculates the Spearman correlation coefficient between each preprocessed feature and all physicochemical indicators, and selects the preprocessed features whose absolute value of the Spearman correlation coefficient with at least one of the physicochemical indicators is greater than a preset threshold (e.g., |ρ|≥0.30), directly focusing on the substantial correlation between the preprocessed features and oil quality. The method of this invention can reduce feature redundancy while retaining effective information, laying a reliable foundation for subsequent more refined feature selection.
[0026] S5. Construct CSFS to further filter the first-level screening features to obtain second-level screening features.
[0027] Because some features that pass the first-level screening may only show accidental correlation with physicochemical indicators under specific sample combinations, or although they are correlated in one dimension, their importance in multi-model evaluation is low, and there may even be multiple features that highly overlap and repeatedly convey the same type of information. If such features are directly used in subsequent prediction models, it will not only increase the computational complexity of the model and mask the role of core features, but also easily lead to noise in the model's fitted data, significantly reducing generalization ability and prediction accuracy. Therefore, in some preferred embodiments, this invention constructs CSFS, which refines the screening based on the first-level screening features in step S4, and mines second-level screening features that have "consensus, stability and strong correlation", providing high-quality input for the construction of subsequent prediction models.
[0028] S6. Input the secondary screening features into a preset machine learning model to obtain prediction results of multiple physicochemical indicators, thereby detecting the quality of the frying oil.
[0029] In some preferred embodiments, based on the secondary screening features obtained in step S5, the obtained secondary screening features are used as input variables and input into multiple preset machine learning models to test their predictive performance on the physicochemical properties of frying oil. The preset machine learning models include partial least squares, random forest, and extreme gradient boosting, etc., or conventional technical means used by those skilled in the art according to actual conditions or site requirements. For each physicochemical indicator, the model with the best predictive performance is selected from multiple preset machine learning models as the final prediction model corresponding to that physicochemical indicator. By obtaining the prediction results of multiple physicochemical indicators and comparing these predicted values with the quality judgment range of industry or national standards, the quality status of the frying oil can be determined, thereby detecting the quality of the frying oil.
[0030] This embodiment enables rapid, non-destructive, and high-precision detection of frying oil quality. The detection results are stable and reliable, and can sensitively reflect the changes in the physicochemical properties of oil during the frying process. The prediction model has good robustness and generalization ability, and is applicable to different oil samples and processing conditions. The overall solution is simple and environmentally friendly, and can realize online monitoring and intelligent evaluation, providing efficient and feasible technical support for food safety and quality control.
[0031] Example 2 Based on Example 1, this example is used to sample frying oil.
[0032] In some preferred embodiments, the frying oil sample of the present invention includes soybean oil and palm oil; the sampling process of the frying oil includes: taking 5 liters of oil sample into a container and heating it to 180±5°C; then adding 200±5 grams of frying ingredients; maintaining a constant temperature and frying once every 0.5 hours, frying continuously for 12 hours a day for 6 consecutive days; taking 100 ml of frying oil sample every 4 hours and storing it at -18°C, for a total of 19 samplings, obtaining a total of 19 frying oil samples. The frying ingredients can be potatoes, chicken breast, etc.
[0033] This embodiment can obtain representative samples of deteriorated oil, providing a reliable experimental basis for subsequent ultrasonic testing and model establishment. This embodiment is carried out under controlled temperature and fixed frying cycle, ensuring the repeatability and traceability of the physicochemical changes such as oil oxidation and polymerization. The use of multi-time point segmented sampling can comprehensively cover the entire process of frying oil from fresh to severely deteriorated, ensuring that the obtained samples have good stage and representativeness. At the same time, the application of different types of oil and different ingredients makes the experimental results more universal.
[0034] Example 3 Based on Example 1, this example is used to preprocess the original ultrasonic signal.
[0035] Common signal preprocessing methods, such as the single operation of combining overall noise reduction with signal truncation, directly filter the entire original signal or only truncate signal segments based on a fixed duration. These methods have significant limitations. On the one hand, the original ultrasonic signal contains invalid components such as equipment startup noise and environmental interference signals. Single filtering is insufficient to accurately separate noise from effective echoes, easily leading to echo signal amplitude attenuation or loss of detail. On the other hand, the lack of targeted extraction based on the temporal characteristics of the echoes may include weak redundant echoes from the fifth time onwards and interference signals from non-target time periods in subsequent analysis, thus affecting the accuracy of acoustic feature extraction. Furthermore, some preprocessing methods lack amplitude filtering, and the residual low-amplitude noise signals can further interfere with subsequent feature calculations and model training.
[0036] In some preferred embodiments, during the signal preprocessing stage, to reduce the impact of random noise on the signal morphology, the present invention first performs time alignment processing on multiple sets of independent pulse echo signals, and then uses a point-by-point averaging method to obtain a representative waveform, thereby effectively reducing system noise and improving signal stability and repeatability. Subsequently, the original signal is truncated into effective segments. Considering that the acquired signal may contain system excitation signals and baseline noise in the initial stage, in order to focus on analyzing the main echo information, the present invention truncates the effective signal from the initial position of the signal (e.g., 1ms) as the analysis window. Then, by setting an amplitude threshold and time window, the position of each echo is automatically identified and extracted, while background noise and weak interference are filtered out. Experimental verification shows that the first four echoes have strong amplitudes and high signal-to-noise ratios, which can fully reflect the acoustic propagation characteristics and energy attenuation law of frying oil. Therefore, the first four echoes are selected as the effective signals after preprocessing, providing high-quality input data for subsequent feature extraction and modeling.
[0037] Example 4 Based on any one of Examples 1 to 3, this example is used to extract multidimensional features that can reflect changes in the physicochemical properties of frying oil from the pre-processed ultrasonic signal, so as to achieve quantitative characterization of the deterioration process of oil quality.
[0038] To address the shortcomings of existing technologies in feature extraction for ultrasonic signals, most studies only extract a small number of basic time-domain features (such as TOF and amplitude), failing to fully explore frequency domain and time-frequency domain information. In some preferred embodiments, this invention extracts features including ultrasonic velocity. When frying oil undergoes oxidation, polymerization, and cracking reactions during repeated high-temperature use, its viscosity and density gradually increase with the formation of polar substances and macromolecular polymers. According to the principles of acoustic propagation, the increase in viscosity and density of the medium reduces its equivalent bulk modulus, causing the sound velocity to show a continuous downward trend. Therefore, ultrasonic velocity can sensitively reflect the changes in the physical properties of frying oil during the deterioration process and can serve as an important parameter for real-time monitoring of oil quality degradation. By continuously detecting changes in sound velocity, the degree of oil aging can be identified in a timely manner without damaging the sample, enabling online monitoring of quality degradation in industrial settings.
[0039] In some preferred embodiments, the present invention extracts time-domain features, including: the maximum amplitude of each of the first four echoes, the flight time of each of the first four echoes, kurtosis, skewness, signal entropy, absolute energy, area under the time-domain spectrum curve, and average number of peaks.
[0040] The first four echoes have sufficient energy and clear signals, are less susceptible to interference, and can stably transmit the absorption and scattering characteristics of the medium. The amplitude of the fifth and subsequent echoes will decrease significantly and become easily distorted, failing to reflect the true state of the oil. The changes in the properties of frying oil are a gradual process, and the sensitivity of the echoes varies at different stages. The amplitude of a single echo is easily deviated due to local fluctuations, while the maximum amplitude of the first four echoes can present the changing patterns of the medium's properties from multiple dimensions.
[0041] Among them, the maximum amplitude of the first four echoes (A1–A4) represents the energy peak of the corresponding echo, which can reflect the medium's ability to absorb and scatter ultrasonic energy. As the oil oxidizes, viscosity increases and polar components accumulate, the ultrasonic energy decays more rapidly, resulting in a more obvious decreasing trend in the echo amplitude. Therefore, this feature can effectively characterize the aging process of oil.
[0042] The flight time (T1–T4) of the first four echoes is closely related to the density and bulk modulus of the medium. It will be slightly shifted with the oxidation of oil and changes in microstructure. It is an important time-domain parameter reflecting the changes in the physical state of oil.
[0043] To further describe the waveform morphology and medium complexity, this invention uses kurtosis and skewness to quantify the changes in signal sharpness and symmetry. During the frying process, the increase of polar substances and impurities in oil will change the echo morphology, making kurtosis and skewness highly sensitive to changes in microstructure.
[0044] Signal entropy is used to measure the complexity and randomness of echoes. When the number of microbubbles or scatterers generated by oxidation increases, the entropy value increases accordingly, which can reflect the process of increasing medium complexity.
[0045] Absolute energy represents the total energy of a signal in the time domain. It is directly related to the absorption and scattering characteristics of oils and can be used to characterize the degree of energy decay.
[0046] The area under the time-domain spectrum curve is used to quantify the attenuation trend of echo energy with propagation distance, and can reflect the absorption characteristics of oils for different echo energies.
[0047] The average number of peaks is used to describe the number of abrupt changes in a signal in the time domain. An increase in the number of scatterers or microstructures will lead to an increase in the number of peaks. Therefore, this feature can be used to evaluate the uniformity of oils and changes in their scattering properties.
[0048] The time-domain features extracted by this invention can reflect the physical and chemical changes of frying oil at different aging stages from multiple aspects such as energy decay, propagation delay, waveform morphology and signal complexity.
[0049] In some preferred embodiments, the present invention extracts frequency domain features, including: spectral kurtosis, spectral skewness, area under the spectral curve, maximum spectral amplitude, frequency corresponding to the maximum amplitude, cumulative energy distribution of the first 25% of energy frequency points, cumulative energy distribution of the first 50% of energy frequency points, cumulative energy distribution of the first 75% of energy frequency points, total power spectrum, median frequency, power spectrum bandwidth, and spectral entropy.
[0050] Among them, spectral kurtosis and spectral skewness reflect the sharpness of frequency domain energy and the symmetry of its distribution, respectively. As aging of oils leads to enhanced scattering and increased structural inhomogeneity, these two indicators will change significantly, which can sensitively characterize the evolution of oils from uniform to complex structures.
[0051] The area under the spectrum curve, the maximum amplitude of the spectrum and its corresponding frequency can quantitatively describe the absorption capacity of oils to sound waves, the degree of energy attenuation and the drift of the dominant frequency position. These changes are closely related to the increase of oil viscosity, the increase of polar substances and the decrease of sound speed.
[0052] The distribution structure of spectral energy is characterized by FFT25, FFT50 and FFT75. These cumulative energy frequency points can reflect the changes in the center of gravity and distribution range of frequency domain energy, and can cover the energy migration pattern of oils from early oxidation to severe aging.
[0053] Furthermore, to quantify the degree of oil aging from the perspectives of overall energy and spectral morphology, this invention introduces total power spectrum, median frequency, and power spectrum bandwidth to describe the overall magnitude of frequency domain energy, energy centroid, and spectral expansion. All three reflect the trends of enhanced absorption, intensified scattering, and increased medium complexity in oils. Meanwhile, spectral entropy, as an indicator of spectral distribution complexity, reflects the increased randomness, non-uniformity, and structural disorder during oil degradation, and is an important characteristic for distinguishing different oxidation stages of oils.
[0054] The frequency domain features extracted by this invention can comprehensively characterize the acoustic response changes of frying oil from multiple perspectives, such as energy concentration, spectral morphology, energy distribution structure, dominant frequency position, and frequency domain complexity. Compared with traditional methods that rely on single features such as sound velocity and attenuation, the frequency domain feature system selected by this invention can more sensitively and comprehensively characterize the physicochemical degradation law of oil.
[0055] This embodiment improves the ability of ultrasonic signals to characterize changes in the quality of frying oil by extracting features from multiple dimensions, including the time domain, frequency domain, and ultrasonic parameters. Compared with traditional methods that rely solely on single features such as amplitude, this embodiment can more comprehensively reflect the physical and chemical changes that occur in oil during frying, including molecular polymerization, increase in polar components, and viscosity changes. By comprehensively utilizing these features, a good foundation can be laid for improving the accuracy and robustness of subsequent modeling and prediction.
[0056] Example 5 Based on any one of Examples 1 to 3, this example is used to construct CSFS.
[0057] Because ultrasonic features of frying oil exhibit typical sensor data characteristics such as high dimensionality, high noise, significant differences between oil types, and limited sample size, traditional feature selection methods, such as single-model importance, VIP, or correlation coefficient screening, are either easily affected by noise or ignore sample heterogeneity and statistical significance, making it difficult to guarantee the stability and generalization ability of the final feature subset across different oil types, ingredients, and data subsets. Therefore, in some preferred embodiments, this invention adopts a "consensus + stability + significance" structure in designing the feature selection method, ensuring that the final selected features simultaneously meet three core requirements: high model consensus, cross-data stability, and significant correlation with oil quality. This invention fully combines the advantages of machine learning, resampling statistics, and nonparametric tests, forming a feature selection framework with robustness and high interpretability.
[0058] The specific steps for constructing the CSFS in this embodiment include: S501. The primary screening features are used as independent variables and the physicochemical indicators are used as dependent variables. These are then input into multiple preset machine learning models. The average importance of the corresponding primary screening features in the preset machine learning models is taken as the consensus importance score C of the corresponding primary screening features.
[0059] In some preferred embodiments, firstly, the primary screening features are used as independent variables and the physicochemical indicators of the oil sample are used as dependent variables. These are then input into multiple preset machine learning models, such as random forests and extreme gradient boosting models. The models are trained to learn the mapping relationship between the features and the physicochemical indicators. After training, the importance of each primary screening feature is extracted from the multiple models. For the same primary screening feature, the average importance of its significance in the multiple models is taken. This average value is the consensus importance score C of the feature.
[0060] This embodiment integrates the evaluation results of feature importance from multiple different models, which can effectively avoid the bias caused by a single model evaluation, and make the score C more objectively reflect the consensus value of features in predicting physicochemical indicators.
[0061] S502. Perform stratified self-sampling N times according to the combination of oil type × ingredient to obtain N data subsets; calculate the C of the primary screening feature of each data subset and sort them, and count the percentage of times each primary screening feature enters the top K positions in the N sortings, as the corresponding stability index SSI.
[0062] Stratified self-sampling is a hybrid sampling method combining stratified sampling and self-sampling. Stratification refers to dividing the data into sub-layers according to preset categories, such as "oil type × ingredient" combinations, ensuring data homogeneity within each layer. Self-sampling involves independently sampling with replacement within each sub-layer and merging them into a data subset. In some preferred embodiments, this invention first constructs stratified labels based on the "oil type × ingredient" combinations in the original data, such as "soybean oil × potato, palm oil × chicken," etc., dividing samples formed by different oil types and different ingredient conditions into several layers, each layer having similar physicochemical properties and ultrasonic signal distribution. Subsequently, self-sampling is performed on each layer, sampling with replacement within each layer according to a set sampling ratio (e.g., 80%), and the samples drawn from each layer are merged to form a data subset. By repeating the above stratified self-sampling process N times (e.g., N=30), N data subsets with consistent structure but containing perturbations can be obtained. For each data subset, the consensus importance C of all first-level screening features is calculated using the method in step S501. The features are sorted from high to low according to the size of C. The set of features that enter the top K positions in each sorting is recorded. Finally, the number of times each first-level screening feature enters the top K positions in N sortings is divided by N to obtain the stability index SSI of the feature.
[0063] In this embodiment, the present invention solves the problem of heterogeneous performance of ultrasonic signals under different oil types and food conditions by evaluating the ranking frequency of features in different subsets through repeated sampling; stratified sampling ensures that the subsamples maintain a consistent structural distribution with the original data, avoids the dominance of large class samples in importance assessment, and enables SSI to effectively measure whether the importance of features is stable under data perturbation, thereby improving the repeatability and generalization of feature selection in real industrial data.
[0064] S503. Using a nonparametric test method, calculate the significance level between each of the primary screening features and each of the physicochemical indicators, and use it as the q value, and construct an inhibition factor (1-q).
[0065] First, nonparametric correlation analysis is performed on each primary screening feature with each physicochemical indicator. For example, the Spearman rank correlation coefficient is used to calculate the maximum absolute correlation coefficient of the feature among all physicochemical indicators, which serves as a statistic to measure the true strength of the correlation of the feature. Then, by randomly permuting the order of the physicochemical indicators multiple times, the statistical distribution of the feature under the assumption of no correlation is calculated, and the empirical probability q value is obtained accordingly. Multiple hypothesis testing methods, such as the Benjamini–Hochberg False Discovery Rate control (BH-FDR), are used to adjust the q value to obtain a more robust q value, which indicates whether the correlation between each feature and the physicochemical indicator is significantly higher than the random level. Finally, (1−q) is used as a significance suppression factor. The larger the q value, the more likely it is to be a random correlation, and the smaller (1−q) is, the stronger the penalty for the feature. The smaller the q value, the more significant the correlation, and the higher the score of (1−q).
[0066] This embodiment evaluates whether the importance of each feature is significantly higher than the random level, statistically eliminating spurious features caused by noise or chance. The larger the q value, the more likely the correlation between the feature and the physicochemical index is to be randomly generated. By penalizing it with (1−q), "false positive features" can be effectively avoided from entering the final model, making the feature subset more statistically reliable.
[0067] S504. Construct CSFS, and its calculation formula is as follows: ; in, Score the consensus importance of the j-th feature. Let be the stability index of the j-th feature. Let be the significance level of the j-th feature.
[0068] In some preferred embodiments, the present invention constructs a consensus importance score, stability index and inhibition factor to achieve quantitative integration of the multi-dimensional value of features, avoids the one-sidedness of a single standard, and can preferentially retain features that are "high model consensus, stable across data and significantly correlated with oil quality".
[0069] In this embodiment, the CSFS constructed by this invention is not a simple weighted sum, but rather employs a "productive triple consistency constraint," ensuring that the final retained features simultaneously meet the requirements of cross-model consistency (high C), cross-sample stability (high SSI), and statistically significant effectiveness (high 1−q). Any deficiency in any of these requirements is amplified and penalized by the multiplicative structure, thereby ensuring that the final feature subset is "small but precise," exhibiting strong robustness and generalization ability even in high-dimensional, noisy, small-sample, and highly heterogeneous data environments. This embodiment effectively improves the stability of the frying oil quality detection model in external validation, significantly reduces the risk of overfitting, and enhances feature interpretability, thus providing a more reliable modeling foundation for real-time frying oil quality monitoring in industrial scenarios.
[0070] Example 6 Based on any one of Examples 1 to 3, this example is used to obtain secondary screening features.
[0071] The methods for obtaining secondary screening features include: S505. Use the CSFS to sort the first-level screening features in descending order and select F features as screening features; In some preferred embodiments, the present invention sorts the primary screening features according to their numerical values based on the CSFS constructed in step S504, with higher scores ranking higher, indicating that the corresponding feature has a better overall value; then, the top F features are selected and determined as screening features for further model training.
[0072] S506. Dummy variable encoding is performed on the oil type and ingredients, and then merged with the screening features to obtain secondary screening features.
[0073] To transform category information into a numerical form (0-1 encoding) that the model can process, satisfying the requirements of machine learning models for numerical input while avoiding spurious numerical associations caused by direct assignment, and preserving the category differences between oil types and ingredients to enhance the model's adaptability to different scenarios, dummy variable encoding of oil types and ingredients is necessary. In some preferred embodiments, this invention first performs dummy variable encoding on the two categorical variables, "oil type" and "ingredient," such as converting oil types like "soybean oil" and "palm oil" into binary variables with 0-1 encoding, and similarly for ingredients, transforming the category information into a numerical form that the model can recognize. Then, the encoded dummy variables of oil types and ingredients are merged with the F screening features to form a comprehensive feature set containing acoustic features and category information, serving as secondary screening features.
[0074] In this embodiment, by introducing information on the types of oil and ingredients, the adaptability of features to different application scenarios (such as different combinations of oil types and ingredients) is improved. This allows the secondary screening features to not only include acoustic features directly related to oil quality, but also to incorporate scenario variables that affect changes in oil quality, providing a more comprehensive input basis for the subsequent model to cope with complex real-world scenarios.
[0075] Example 7 Based on any one of Embodiments 1 to 3, this embodiment is used to add a failure protection strategy to the preset machine learning model during external verification.
[0076] External validation is crucial for assessing a model's generalization ability, but its effectiveness is limited by the distribution of training data. When external data differs from training data—for example, if oil type or ingredient combination is not covered, or frying conditions fluctuate—secondary screening features may fail, leading to increased model detection errors. Without a failure protection strategy, the model will directly output low-precision results, potentially causing misjudgments of oil quality and threatening application safety; at worst, it may render the model unusable and waste initial investment. Therefore, implementing a failure protection strategy for external validation is essential to ensure the model's stable performance in complex scenarios.
[0077] In some preferred embodiments, the present invention provides a failure protection strategy for the network model during external validation, including: defining the relative error improvement rate of external testing. : ; in This represents the root mean square error on the external test set when the model uses preprocessed features to combine dummy variable encodings of oil type and food ingredients. This represents the root mean square error on the external test set when the model uses secondary filtered features after CSFS filtering. If the model × physicochemical index combination has a ΔRMSE ≥ 0 on external testing, the model ultimately adopts the secondary screening features obtained after CSFS screening; if ΔRMSE < 0, the model ultimately adopts the preprocessed features combined with dummy variable codes for oil type and food ingredient. In some preferred embodiments, such as when ΔRMSE ≥ 0 in external validation, the RMSE of acid value decreased by 26.6%, saponification value decreased by 41.4%, and total polar components decreased by 6.4%, indicating that the secondary screening features can improve the model's prediction accuracy in external validation, so the secondary screening features are ultimately adopted; if density and iodine value have a ΔRMSE < 0 in external validation, it indicates that the secondary screening features are ineffective on the external validation data, so in this case, the model selects the combined features of "preprocessed features + dummy variable codes for oil type and food ingredient" to ensure the model's performance.
[0078] The failure protection strategy in this embodiment improves the robustness of the model, enabling it to dynamically respond to the differences in the distribution of training data and external data. It effectively avoids the problem of decreased detection accuracy caused by CSFS screening failure, ensuring that the model can stably output reliable results in complex scenarios.
[0079] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely illustrative of the principles of the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the present invention as claimed. The scope of protection of this invention is defined by the appended claims and their equivalents.
Claims
1. A method for detecting the quality of frying oil based on cascaded feature screening optimization and ultrasonic diagnostic technology, characterized in that, Includes the following steps: S1. Under constant oil temperature conditions, collect ultrasonic signals and data on multiple physicochemical indicators of frying oil; S2. The ultrasonic signal is preprocessed, and the first four echoes of the ultrasonic signal are retained as the preprocessed signal. S3. Extract acoustic features based on the preprocessed signal, and perform outlier processing on the acoustic features to obtain preprocessed features; S4. Iterate through and calculate the Spearman correlation coefficient between each of the preprocessed features and multiple of the physicochemical indicators; select the preprocessed features whose absolute value of the Spearman correlation coefficient with at least one of the physicochemical indicators is greater than a preset threshold as first-level screening features; S5. Construct a consensus-stability comprehensive feature scoring index to further filter the primary screening features and obtain secondary screening features; S6. Input the secondary screening features into a preset machine learning model to obtain prediction results of multiple physicochemical indicators, thereby detecting the quality of the frying oil.
2. The method for detecting the quality of frying oil based on cascaded feature screening optimization and ultrasonic diagnostic technology according to claim 1, characterized in that, In step S2, the preprocessing method includes: S201. Extract the original ultrasonic signal from the initial position; S202. Based on the intercepted ultrasonic signal, retain the signal components with amplitudes exceeding the preset amplitude to obtain a preliminary preprocessed signal; S203. Based on the preliminary preprocessing signal, four expected time windows corresponding to the first four echoes are defined, and the first four echoes are extracted and retained from the preliminary preprocessing signal as preprocessing signals.
3. The method for detecting the quality of frying oil based on cascaded feature screening optimization and ultrasonic diagnostic technology according to claim 1, characterized in that, In step S3, the acoustic features include temporal features: The maximum amplitude of each of the first four echoes, the flight time of each of the first four echoes, kurtosis, skewness, signal entropy, absolute energy, area under the time-domain spectrum curve, and average number of peaks.
4. The method for detecting the quality of frying oil based on cascaded feature screening optimization and ultrasonic diagnostic technology according to claim 1, characterized in that, In step S3, the acoustic features include frequency domain features: Spectral kurtosis, spectral skewness, area under the spectral curve, maximum amplitude of the spectrum, frequency corresponding to the maximum amplitude, cumulative energy distribution of the first 25% of energy frequency points, cumulative energy distribution of the first 50% of energy frequency points, cumulative energy distribution of the first 75% of energy frequency points, total power spectrum, median frequency, power spectrum bandwidth, and spectral entropy.
5. The method for detecting the quality of frying oil based on cascaded feature screening optimization and ultrasonic diagnostic technology according to claim 1, characterized in that, In step S3, the acoustic feature includes the speed of sound of the ultrasonic wave.
6. The method for detecting the quality of frying oil based on cascaded feature screening optimization and ultrasonic diagnostic technology according to claim 1, characterized in that, The specific steps for constructing the consensus-stability comprehensive feature scoring index include: S501. The primary screening features are used as independent variables and the physicochemical indicators are used as dependent variables. They are input into multiple preset machine learning models respectively. The average importance of the corresponding primary screening features in the preset machine learning models is taken as the consensus importance score C of the corresponding primary screening features. S502. Perform stratified self-sampling N times according to the combination of oil type × ingredient to obtain N data subsets; calculate the C of the primary screening feature of each data subset and sort them, and count the percentage of times each primary screening feature enters the top K positions in the N sorting, as the corresponding stability index SSI. S503. Using a nonparametric test method, calculate the significance level between each of the first-level screening features and each of the physicochemical indicators, and use it as the q value, and construct an inhibition factor (1-q). S504. Construct a consensus-stability comprehensive characteristic scoring index, calculated using the following formula: ; in, Score the consensus importance of the j-th feature. As a stability index, The significance level is indicated by .
7. The method for detecting the quality of frying oil based on cascaded feature screening optimization and ultrasonic diagnostic technology according to claim 1, characterized in that, The methods for obtaining secondary screening features include: S505. The consensus-stability comprehensive feature scoring index is used to sort the first-level screening features in descending order, and F features are selected as screening features; S506. Dummy variable encoding is performed on the oil type and ingredients, and then merged with the screening features to obtain secondary screening features.
Citation Information
Patent Citations
Method for constructing step-by-step temperature compensation model for thermal oxidation rapeseed oil quality detection based on ultrasonic diagnosis technology
CN120044138A
Traditional Chinese medicine clinical big data storage method based on multi-mark feature selection and classification
CN109119133A
Method for constructing global temperature compensation model for thermal oxidation rapeseed oil quality detection based on ultrasonic diagnosis technology
CN120028448A
Online monitoring method and system for food frying process
CN121476408A
Method for changing over calibration curve in concentration measurement of rolling mill oil
JP1996166374A