Machine learning-based method and system for identifying the origin and breeding mode of OAV-oriented sea cucumber
By employing an OAV-guided method based on machine learning to eliminate environmental noise and identify robust sea cucumber flavor markers, the instability problem in identifying sea cucumber origin and farming methods is solved, enabling high-precision origin traceability and quality evaluation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHINESE ACAD OF INSPECTION & QUARANTINE
- Filing Date
- 2026-02-11
- Publication Date
- 2026-05-29
Smart Images

Figure CN122117129A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the interdisciplinary fields of food science, analytical chemistry and artificial intelligence, and in particular relates to a machine learning-based OAV-guided identification method and system for sea cucumber origin and farming mode. Background Technology
[0002] Sea cucumber (Apostichopus japonicus) is an important marine economic crop in my country. Due to the influence of geographical environment, water temperature, and feed, the quality of sea cucumbers varies greatly depending on the production area and farming method. Currently, the market offers various farming methods, including bottom-seeding in Dalian, net cage farming in Dalian, cofferdam farming in Dalian, and the net cage and hanging cage methods used in Fujian for farming sea cucumbers from the north to the south. Because bottom-seeded sea cucumbers are much more expensive than cage-farmed sea cucumbers, the market often sees southern sea cucumbers farmed in the north, and cage-farmed sea cucumbers not bottom-seeded. Existing technologies mainly rely on DNA barcoding, stable isotopes, or traditional non-targeted metabolomics for identification. Origin identification methods often rely on DNA barcoding, stable isotopes, mineral fingerprinting, or traditional non-targeted metabolomics full-spectrum analysis. Among these, isotope and elemental fingerprinting are easily affected by short-term fluctuations in the aquatic environment and feed changes, resulting in poor stability; while traditional full-spectrum metabolomics, although able to classify by capturing exogenous environmental pollutants, suffers from serious environmental noise problems. For example, sea cucumbers cultured in net cages often adsorb trace amounts of compounds such as styrene and phthalates released from the aquaculture facilities, while bottom-sown sea cucumbers may accumulate specific hydrocarbons from sediments. Although these exogenous environmental substances can be used to achieve highly accurate origin classification, these environmental fingerprints are abiotic metabolites, not biological metabolites of sea cucumbers, and their concentrations are extremely low, failing to reflect the true flavor and quality of sea cucumbers. This modeling logic leads to a lack of biological interpretability in the identification results, and the model is highly susceptible to interference from environmental factors such as changes in aquaculture materials, causing it to fail. Therefore, there is an urgent need for an intelligent identification method that can eliminate environmental noise, focus on key endogenous flavor compounds in sea cucumbers, and incorporate highly robust machine learning algorithms.
[0003] Based on the above problems, this invention addresses the biochemical nature of sea cucumber flavor formation and the complexity of the aquaculture environment by proposing a systematic technical solution based on biological noise reduction, robust feature screening, and interpretable modeling. By combining this solution with interpretable artificial intelligence algorithms, it provides visualized evidence for flavor traceability, enabling precise identification of the origin of sea cucumbers and specific aquaculture methods such as bottom seeding, net cages, and hanging cages. This provides a traceability technology for the deep processing, quality evaluation, and market supervision of sea cucumbers that is resistant to environmental interference, has clear biological significance, and a transparent evidence chain. It effectively solves the technical problems of existing technologies, such as the susceptibility of identification methods to environmental noise interference, lack of flavor correlation, and insufficient comprehensiveness. Through parameterized detection processes, standardized OAV noise reduction pipelines, and built-in interpretable analysis, consistent discrimination performance can be achieved across different sampling batches, different years of cultivation, and different processing methods. It can not only output origin identification results, but also simultaneously output key flavor substance fingerprints, realizing the dual functions of origin traceability and quality evaluation, providing solid technical support for anti-counterfeiting and authenticity assurance of high-end sea cucumber brands. In addition, 13 flavor feature markers based on OAV guidance have been discovered and provided, which can be used for the identification of sea cucumbers from different origins and cultivation methods. Summary of the Invention
[0004] To address the aforementioned problems in the prior art, this invention proposes a machine learning-based OAV-guided sea cucumber origin and farming model identification method and system, the method comprising: Step S1: Sample collection and grouping, specifically: collect fresh sea cucumbers from different production areas and farming methods, and prepare them into freeze-dried powder through pretreatment; Step S2: Combined chromatographic and mass spectrometric monitoring to determine the parameters of volatile organic compounds; specifically: headspace solid-phase microextraction combined with gas chromatography-quadrupole-orbit trap mass spectrometry (GC-MS) is used to detect the sample, and qualitative and quantitative data of volatile organic compounds are obtained as organic compound parameters by deconvolution and comparison with a standard library. Step S3: Odor Activity Value (OAV)-guided biological noise reduction screening to obtain key flavor variables; specifically: for different types of organic compound parameters, compare them with their corresponding olfactory thresholds to determine the odor activity value (OAV), and take substances with an odor activity value (OAV) ≥ 1 as key flavor variables, and remove non-flavor contributing substances and environmental noise with an OAV < 1. Step S4: Select core flavor biomarkers based on resampling; specifically: perform n Bootstrap resampling operations on key flavor variables, run multiple screening models in parallel during each resampling operation, and count the frequency of selection for each key flavor variable. Select the high-frequency key flavor variables as robust features of consensus and use them as core flavor biomarkers; where n is a preset value. Step S5: Construction and validation of the intelligent discrimination model; specifically: construct a random forest classification model based on core flavor markers, and test the generalization ability and anti-overfitting performance of the intelligent discrimination model through k-fold cross-validation and permutation. Step S6: Analyze the flavor formation mechanism based on SHAP interpretability; specifically: calculate the SHAP value of each core flavor marker, quantify its marginal contribution to origin identification, and generate a dependency graph to analyze the flavor formation mechanism.
[0005] Furthermore, the sea cucumber samples are mature sea cucumber samples grown in Dalian and mature sea cucumber samples grown in Fujian; the sample pretreatment in step S1 includes: crushing the sea cucumber sample and adding an internal standard solution to the headspace vial for testing; the internal standard solution is 2-methyl-3-heptanone dissolved in methanol.
[0006] Furthermore, the specific conditions for GC-MS analysis were as follows: a TG-5MS capillary column was used; helium was used as the carrier gas at a flow rate of 1.2 mL / min; the temperature program was as follows: initial temperature of 40℃ held for 2.5 min, then increased to 70℃ at 7℃ / min, then increased to 120℃ at 2℃ / min held for 1 min, and finally increased to 280℃ at 20℃ / min held for 5 min; the mass spectrometer ion source temperature was 280℃, the ionization energy was 70 eV, the scan range was m / z 35-500, and the resolution was 60,000.
[0007] Furthermore, the Bootstrap resampling time n is at least 200 times; the selection criteria for the core flavor markers are: at least two of the three algorithms are determined to be important features, and the frequency of selection in all resampling iterations is greater than 50%.
[0008] Furthermore, resampling was performed using a screening model based on OPLS-DA, random forest, and LASSO algorithms.
[0009] Furthermore, the OPLS-DA algorithm performs screening by calculating the variable projection importance (VIP) screening feature.
[0010] Furthermore, the random forest algorithm evaluates feature importance by the amount of reduction in Gini impurity.
[0011] Furthermore, the LASSO algorithm filters features by finding the minimum value of the loss function with L1 regularization.
[0012] Based on the same inventive concept, the present invention also provides a machine learning-based OAV-guided sea cucumber origin and aquaculture pattern identification system, including a processor coupled with a memory, the memory storing program instructions, and when the program instructions stored in the memory are executed by the processor, the above-mentioned machine learning-based OAV-guided sea cucumber origin and aquaculture pattern identification method is implemented.
[0013] Based on the same inventive concept, the present invention also provides a computer-readable storage medium, characterized in that it includes a program that, when run on a computer, causes the computer to execute the above-described machine learning-based OAV-guided method for identifying the origin and farming mode of sea cucumbers.
[0014] The beneficial effects of this invention include: (1) OAV-guided biological noise reduction mechanism: The odor activity value OAV is innovatively introduced as a feature filter to forcibly remove non-flavor contributing substances with OAV<1, thereby removing environmental background noise such as plastic monomers and industrial hydrocarbons from the source, and forcing the model to focus on key flavor substances produced by biological metabolic pathways such as lipid oxidation and amino acid degradation, thus achieving a leap from chemical fingerprint to sensory fingerprint. (2) Bootstrap-based integrated feature screening strategy: In response to the pain points of large individual differences in biological samples and easy overfitting of small sample data, Bootstrap resampling technology is used to construct perturbation dataset, and OPLS-DA, random forest and LASSO are integrated to perform triangular consensus screening in a differentiated manner; not only are the core markers that perform stably in multiple batches of resampling locked, but the statistical bias of a single algorithm is also effectively avoided, ensuring that the selected markers have extremely high robustness. (3) Multi-model integration and transparent evidence chain: A high-precision discrimination model based on random forest was constructed, and a permutation test was introduced to construct an empirical spatial distribution to strictly control the risk of overfitting. The built-in SHAP game theory explanation mechanism quantifies the marginal contribution of each flavor marker to the discrimination result, and generates a dependency graph to analyze how different farming models differentially regulate the accumulation of flavor substances. Attached Figure Description
[0015] The accompanying drawings, which are provided to further illustrate the invention and form part of this application, are not intended to unduly limit the invention. In the drawings: Figure 1 A schematic diagram of the architecture of the machine learning-based OAV-guided sea cucumber origin and farming mode identification system provided by the present invention.
[0016] Figure 2 This is a schematic diagram illustrating the robust feature selection process and results based on Bootstrap resampling provided by the present invention. Wherein: Figure 2(A) is a heatmap of feature selection stability; Figure 2 (B) is a line graph ranking the frequency of feature importance; Figure 2 (C) is a heatmap of feature selection stability; Figure 3 This diagram illustrates the performance evaluation and comparison of various machine learning models provided by this invention. Wherein: Figure 3 (A) Comparison of the receiver operating characteristic (ROC) curves of six models on the test set: Random Forest (RF), Support Vector Machine (SVM), K Nearest Neighbors (KNN), Logistic Regression (LR), Naive Bayes (NB), and Gradient Boosting Tree (GBDT). Figure 3 (B) is a bar chart comparing the discrimination accuracy of the above six algorithms; Figure 3 (C) is the confusion matrix of the optimal model RF on the test set; Figure 4 This is a schematic diagram illustrating the permutation test verification of the random forest model provided by the present invention.
[0017] Figure 5 This is a schematic diagram illustrating the interpretable flavor mechanism analysis based on the SHAP algorithm provided by the present invention; wherein: Figure 5 (A) is a SHAP bee colony diagram of the core flavor markers; Figure 5 (B) is a bar chart showing the average contribution of flavor markers. Detailed Implementation
[0018] The present invention will now be described in detail with reference to the accompanying drawings and specific embodiments. The illustrative embodiments and descriptions are only used to explain the present invention and are not intended to limit the present invention. As attached Figure 1 As shown, this invention proposes a machine learning-based OAV-guided identification method and system for sea cucumber origin and farming model identification. The method includes the following steps: Step S1: Sample collection and grouping, specifically: collect fresh sea cucumbers from different production areas and farming methods, and prepare them into freeze-dried powder through pretreatment; Preferred: The production areas and breeding models include at least Dalian bottom seeding DLDB, Dalian net cage DLWX, Dalian cofferdam DLWY, Fujian net cage FJWX, and Fujian hanging cage FJDL; Preferably, the sample is a sea cucumber sample, specifically a mature sea cucumber sample grown in Dalian and a mature sea cucumber sample grown in Fujian. Preferably, the sample is pretreated before step S1, specifically by crushing the sea cucumber sample and adding an internal standard solution to the headspace vial for testing; wherein the internal standard solution is 2-methyl-3-heptanone dissolved in methanol; Step S2: Combined chromatographic and mass spectrometric monitoring to determine the parameters of volatile organic compounds; specifically: headspace solid-phase microextraction combined with gas chromatography-quadrupole-orbitrap mass spectrometry (HS-SPME-GC-Q-Orbitrap-MS) is used to detect the sample, and qualitative and quantitative data of volatile organic compounds are obtained as organic compound parameters by deconvolution and comparison with the standard library; Preferred conditions for the headspace solid-phase microextraction combined with gas chromatography-quadrupole-orbit trap mass spectrometry (GC-MS) analysis are as follows: TG-5MS capillary column is used; helium is used as the carrier gas at a flow rate of 1.2 mL / min; the temperature program is: initial 40℃ held for 2.5 min, increased to 70℃ at 7℃ / min, then increased to 120℃ at 2℃ / min held for 1 min, and finally increased to 280℃ at 20℃ / min held for 5 min; the mass spectrometer ion source temperature is 280℃, the ionization energy is 70 eV, the scan range is m / z 35-500, and the resolution is 60,000. Step S3: Odor Activity Value (OAV)-guided biological noise reduction screening to obtain key flavor variables; specifically, for different types of organic compound parameters, they are compared with their corresponding olfactory thresholds to determine the odor activity value (OAV). Substances with an OAV value ≥ 1 are used as key flavor variables, while non-flavor-contributing substances and environmental noise with an OAV < 1 are eliminated. The OAV value is a dimensionless value, representing the multiple by which the compound concentration exceeds its olfactory threshold. The unit for organic compound parameters is μg / L or μg / m³. Through the odor activity value, 17 organic compound parameters covering three major aroma molecules—aldehydes, alcohols, and ketones—are identified, including (E)-2-octenal, octanal, heptanal, 1-octen-3-ol, and p-xylene. This enables robust differentiation across multiple scenarios and batches for five complex models: Dalian bottom seeding, Dalian net cages, Dalian cofferdams, Fujian net cages, and Fujian hanging cages. This achieves a leap from chemical fingerprinting to sensory fingerprinting. Preferably, the olfactory threshold is a preset value, which is the lowest concentration that can be perceived by 50% of the olfactory identification team members; Preferably, the OAV is calculated using the following formula (1); wherein, The absolute content (μg / kg) of the i-th compound in the sea cucumber sample. The olfactory threshold (μg / kg) of this compound in aqueous medium. Preferably, the organic compound is a volatile compound; Step S4: Select core flavor biomarkers based on resampling; specifically: perform n Bootstrap resampling operations on key flavor variables, run multiple screening models in parallel during each resampling operation, and count the frequency of selection for each key flavor variable. Select the high-frequency key flavor variables as robust features of consensus and use them as core flavor biomarkers; where n is a preset value. Preferably, resampling is performed using screening models based on algorithms such as OPLS-DA, Random Forest, and LASSO; further, the number of Bootstrap resampling iterations n is at least 200; the screening criteria for the core flavor markers are: at least two of the three algorithms are selected, that is, they are judged as important and robust features, and the frequency of selection in all resampling iterations is greater than 50%; Preferably, the OPLS-DA algorithm filters key flavor variables by calculating their projected importance, specifically by using the following formula (2) to calculate the projected importance of the j-th key flavor variable. ; Where p represents the total number of key flavor variables, and K represents the number of principal components in the model, which is the total number of sea cucumber classifications. The sum of squared variances of the Y variable explained by the k-th principal component represents the contribution of the sample to the k-th category of sea cucumbers. Let be the weight of the j-th key flavor on the k-th principal component; Furthermore, when At that time, the key flavor variable was identified as an important feature determined by the OPLS-DA algorithm; Preferably, the random forest algorithm assesses the importance of key flavor variables by the reduction in Gini impurity, and further, the Gini impurity is calculated using the following formula (3). Where K represents the number of importance categories, which is also the total number of sea cucumber classifications. Let be the probability that the sample belongs to the k-th category of sea cucumber; Preferably, the LASSO algorithm filters key flavor variables by finding the minimum value of the loss function with L1 regularization, and its objective function is shown in equation (4); where N is the number of samples, Y is the category label vector, that is, the sea cucumber classification number, and X is the feature matrix. For regression coefficients, The regularization parameter; the selection criterion is the regression coefficient. ; Preferred: The feature matrix X is obtained by organizing the flavor data obtained from the instrument into a matrix form and then normalizing it; specifically, the element in the a-th row and b-th column of the feature matrix indicates the normalized value of the measurement value of the a-th sea cucumber sample on the b-th flavor variable; the regression coefficient and regularization parameter are preset values or obtained automatically through optimization algorithms; As attached Figure 2 As shown in (A)-2(C), Figure 2 (A) is a heatmap of feature selection stability. The darker the color, the higher the frequency of the flavor substance being selected in 200 resampling, which intuitively shows the robustness of the marker. The statistical significance of the feature was verified by 200 random perturbations, which reduced the random error caused by small sample data. Figure 2 (B) is a line graph showing the frequency ranking of feature importance, which shows the selection frequency of the top 30 latent features under three different algorithms: VIP, Random Forest (RF), and LASSO. The dashed line in the middle indicates that the stability threshold is set at the 50th percentile. Figure 2 (C) is a Venn diagram of the consensus of the three screening algorithms, showing that the 13 core flavor markers finally determined are the intersection of multiple algorithms, that is, the consensus result; triangular consensus screening was achieved through three differentiated algorithms; not only were core markers that performed stably in multiple batches of resampling locked in, but the statistical bias of a single algorithm was also effectively avoided, ensuring that the screened markers have extremely high robustness; this invention achieves automatic discrimination and report generation by combining a conventional gas chromatography-mass spectrometry platform with standardized OAV calculation, without the need to repeatedly develop new methods for different production areas, which greatly improves detection efficiency and result reliability, and is suitable for large-scale application and continuous expansion of regulatory databases; Step S5: Construction and validation of the intelligent discrimination model; specifically: construct a random forest classification model based on core flavor markers, and test the generalization ability and anti-overfitting performance of the intelligent discrimination model through k-fold cross-validation and permutation. Preferred models include: Random Forest (RF), Support Vector Machine (SVM), K-Nearest Neighbors (KNN), Logistic Regression (LR), Naive Bayes (NB), and Gradient Boosting Tree (GBDT). See attached... Figure 3 As shown, the performance varies when using six different machine learning models; for example... Figure 3 (A) The area under the curve (AUC) of the RF model shown in (A) reaches 1.00; Figure 3 (B) is a bar chart comparing the discrimination accuracy of the above six intelligent discrimination models, showing that the RF model has the best performance; Figure 3 (C) is the confusion matrix of the best model RF on the test set. The diagonal values represent the number of correctly classified samples, showing that the prediction results of the five farming models are completely consistent with the true labels. As attached Figure 4 As shown, the permutation test validation plot of the random forest model is displayed; the histogram (gray area) shows the accuracy distribution of the spurious model constructed after 200 random shuffling of sample labels (normal distribution with low mean), and the red dashed line represents the accuracy of the model constructed based on the real labels (100%); the two are statistically significantly separated (Pvalue < 0.05), confirming that the model constructed in this invention has not overfitted and has good generalization ability.
[0019] Preferred: The core flavor markers selected based on OAV guidance are those that can be used to identify sea cucumbers from different origins and farming methods. The flavor markers are styrene, heptanal, p-xylene, 2,3-octanedione, (E,Z)-2,4-decadienal, dodecanal, 1-decanol, octanal, 1-octen-3-one, (E,Z)-2,6-nonadienal, 1-octanol, 1-octen-3-ol, and methyl hexanoate. Step S6: Analyze the flavor formation mechanism based on SHAP interpretability; specifically: calculate the SHAP value of each core flavor marker, quantify its marginal contribution to origin identification, and generate a dependency graph to analyze the flavor formation mechanism; furthermore, the SHAP value is calculated based on game theory, using the following formula (5), the SHAP value of core flavor marker i Defined as: Where F is the set of all core flavor markers, and S is a subset that does not contain core flavor marker i. For model prediction functions; This means that set S is a subset of set F after removing element i; As attached Figure 5 As shown in the analytical diagram. Figure 5 (A) is a SHAP beeswarm plot of core flavor markers. Each point in the plot represents a sample, the color of the point represents the level of the feature value, and the horizontal axis position represents the direction of its contribution to the model's judgment. Figure 5 (B) is a bar chart of the average contribution of flavor markers, which quantifies the importance ranking of each feature and shows that (E)-2-octenal, octanal and other substances have the highest contribution: By interpreting the flavor formation mechanism based on SHAP, the marginal contribution of each flavor marker to the discrimination result can be quantified, and a dependency graph can be generated to analyze how different aquaculture modes such as bottom seeding at low temperature and slow growth and net cage at high temperature and rapid growth differentiate and regulate the accumulation of flavor substances. Example 1: Collection, preparation and grouping of sea cucumber samples This embodiment details the source and preprocessing of experimental data, ensuring that the samples in the traceability model have sufficient representativeness and consistency. Sample Collection: 150 live sea cucumbers (Apostichopus japonicus) were collected from the sea areas of Dalian, Liaoning Province and Xiapu, Fujian Province, between March and May 2024. Dalian Group (Northern Model): Samples were collected from the sea areas of Wafangdian and Pulandian in Dalian, covering three typical aquaculture models: Dalian bottom seeding (DLDB, 30 cases), Dalian enclosure (DLWY, 30 cases), and Dalian net cage (DLWX, 30 cases). Southern Group (Southern Model): Samples were collected from the sea area of Xiapu, Fujian Province, all under the "Northern Sea Cucumber, Southern Aquaculture" model, covering two typical aquaculture methods: Fujian net cage (FJWX, 30 cases) and Fujian hanging cage (FJDL, 30 cases). Screening Criteria: To eliminate the influence of individual growth differences on flavor, strict screening criteria were established: adult sea cucumbers with undamaged bodies, intact spines, not in the evisceration period, and similar sizes (body weight 150±10g, body length approximately 20cm) were selected. Sample Preparation: After collection, the internal organs of the samples were quickly removed on-site and washed with seawater. They were then frozen and stored in a dry ice environment and uniformly processed by light drying to simulate the morphology of lightly dried sea cucumbers on the market. After processing, all samples were transported back to the laboratory, vacuum-packed, and stored in a -20℃ freezer for later use. Before the experiment, all lightly dried sea cucumber samples were uniformly subjected to vacuum freeze-drying. The dried samples were ground under liquid nitrogen protection and passed through a 60-mesh sieve to obtain a uniform powder, which was then sealed and placed in a desiccator for later use.
[0020] Example 2: Construction of volatile fingerprint spectrum based on HS-SPME-GC-Q-Orbitrap-MS This example describes how to obtain high-throughput volatile matter data. Instruments and reagents: A Thermo Fisher Scientific TRACE™ 1310Q Exactive GC-MS system; a Supelco DVB / CAR / PDMS (50 / 30 μm) extraction fiber; a Sigma-Aldrich C7-C30 n-alkanes mixture (for calculating the retention index KI) and 2-methyl-3-heptanone standard (internal standard). Detection method: Extraction: Weigh 0.2 g of sample powder into a headspace vial and add 6 μL of internal standard (2-methyl-3-heptanone, 0.0615 μg / μL methanol solution). After incubating the sample at 60 °C for 30 min, extract with SPME fiber for 60 min, followed by resolution at 250 °C for 3 min. Chromatographic conditions: Separation using a TG-5MS capillary column; carrier gas: helium (99.999%), flow rate: 1.2 mL / min, split ratio: 15:1. Temperature program: Initially 40℃, hold for 2.5 min; increase to 70℃ at 7℃ / min; increase to 120℃ at 2℃ / min, hold for 1 min; increase to 280℃ at 20℃ / min, hold for 5 min. Mass spectrometry conditions: EI ion source (70 eV, 280℃), transfer line temperature 250℃. Full scan mode (m / z 35-500), resolution 60,000. Data processing: Deconvolution and peak alignment were performed using TraceFinder 4.1 software. Combined with NIST 2017 library (SI>650) and retention index (LRI) comparison, a total of 112 volatile organic compounds were identified.
[0021] Example 3: Unsupervised exploration and OAV-guided biological noise reduction This embodiment demonstrates the preliminary data exploration results and the core noise reduction strategy of this invention. Preliminary Principal Component Analysis (PCA): PCA analysis was performed on 112 original volatile compounds. The results showed that although samples from different origins exhibited a certain spatial clustering trend, there was partial overlap between samples from Dalian cages (DLWX) and Fujian cages (FJWX), and some environmental pollutants (such as styrene) contributed excessively to the principal components, masking the true flavor differences. This indicates that relying solely on full-spectrum analysis is insufficient for accurate and biologically meaningful classification. OAV-guided screening: The olfactory threshold of each compound in aqueous medium was queried, and the odor activity value (OAV = concentration / threshold) was calculated. Substances with OAV < 1 were forcibly removed (including high levels of styrene, dibutyl phthalate, and other environmental background noise), retaining only 17 key flavor compounds with OAV ≥ 1. These substances are mainly aldehydes, alcohols, and ketones, directly related to the fresh aroma, fatty aroma, and earthy flavor of sea cucumber. The final 17 flavor compounds with OAV ≥ 1 are shown in the table below: Example 4: Robust Feature Biomarker Screening Based on Bootstrap Resampling This embodiment details how to identify the most stable identification indicators from key flavor compounds. Screening strategy: To address individual variability in biological samples, 200 Bootstrap resampling iterations were performed using the Python sklearn library. On the perturbed dataset generated in each resampling iteration, three feature selection algorithms were run in parallel: OPLS-DA: Calculates the projected importance of variables, requiring VIP > 1; Random Forest (RF): Calculates Gini importance, requiring Gini Importance to be in the top 20; LASSO Regression: Calculates the sparse coefficient Coef, requiring Coef ≠ 0. (See attached...) Figure 2 The characteristic stability heatmap shows the frequency of each substance being selected in 200 resampling iterations. Darker colors indicate higher stability. In the Venn diagram and consensus screening, substances deemed important in at least two algorithms and selected more than 50% of the time were selected. Ultimately, 13 core flavor markers were identified, including styrene, heptanal, p-xylene, 2,3-octanedione, (E,Z)-2,4-decadienal, dodecanal, 1-decyl alcohol, octanal, 1-octen-3-one, (E,Z)-2,6-nonadienal, 1-octanol, 1-octen-3-ol, and methyl hexanoate. These markers not only exhibit extremely high statistical significance but also demonstrate clear flavor contributions.
[0022] Example 5: Construction and Performance Evaluation of Intelligent Identification Model This embodiment demonstrates the superiority of a classification model built based on core biomarkers. Model construction: 150 samples were randomly divided into a training set (120 cases) and a test set (30 cases) in an 8:2 ratio. Based on 17 selected core biomarkers, a Random Forest classifier was constructed (number of decision trees = 100, random seed = 42). Model comparison: SVM, KNN, logistic regression, Naive Bayes, and GBDT models were constructed simultaneously for comparison. Results (e.g.) Figure 3 As shown in the figure, the random forest model achieves 100% accuracy on the test set, significantly outperforming GBDT (96.7%) and other algorithms, demonstrating its advantage in handling high-dimensional nonlinear data. (See attached figure.) Figure 3 As shown, the confusion matrix of the optimal model RF obtained from the confusion matrix analysis reveals that, among the 30 samples in the test set, the predicted labels for five types of samples—Dalian bottom seeding, Dalian cages, Dalian cofferdams, Fujian cages, and Fujian hanging cages—are completely consistent with the true labels, with no misclassifications. This demonstrates that the marker combination screened in this invention has extremely high discrimination resolution.
[0023] Example 6: Reliability verification of the model using permutation test To ensure high accuracy is not due to overfitting, rigorous statistical validation was performed. The feature matrix was kept unchanged, and the class labels (Y values) of the training set were randomly shuffled. The random forest model was retrained, and the accuracy was calculated. This process was repeated 200 times to construct an empirical empty distribution. Results analysis: see attached. Figure 4 The permutation test results show that the average accuracy of the model after randomly shuffling the labels is about 20%, which is the probability of random guessing and follows a normal distribution; while the accuracy of the model based on the real labels is 100%, which deviates far from the right side of the random distribution (P<0.005). It can be seen that the identification ability of the present invention comes from the real biological association between the marker and the place of origin, rather than data noise or overfitting.
[0024] Example 7: Explanatory Flavor Mechanism Analysis Based on SHAP As shown in Figure (5A), this invention achieves visualized source tracing and mechanism explanation; firstly, SHAP beehive graph analysis is performed: the SHAP value of each sample in the test set is calculated and a beehive graph is drawn. Each point in the graph represents a sample, the color represents the intensity of the feature value, and the X-axis position represents the direction of contribution to the model's discrimination. The results intuitively demonstrate the weight distribution of different markers in distinguishing five modes. Feature importance ranking: as shown... Figure 5 As shown in B, the SHAP bar chart quantifies the average contribution of each marker. The results show that (E)-2-octenal, octanal, and 1-octen-3-ol are the top three characteristics that contribute the most to the determination of origin.
[0025] This invention establishes a sensory-guided and data-driven method for tracing the origin and farming model of sea cucumbers, providing an objective and biologically-deep analytical method for anti-counterfeiting, market supervision, and food authenticity assessment of high-end sea cucumbers. The proposed workflow based on biological noise reduction, robust screening, and mechanism explanation can be easily applied to other marine species susceptible to environmental influences, such as abalone and scallops, providing a scalable and universal framework for molecular traceability of aquatic products. Based on the same inventive concept, the present invention also provides a machine learning-based OAV-guided sea cucumber origin and farming mode identification system, the system comprising: a detection terminal and an intelligent server; wherein: the detection terminal and the intelligent server are communicatively connected; the system is used to implement the above-mentioned machine learning-based OAV-guided sea cucumber origin and farming mode identification method; Based on the same inventive concept, the present invention also provides a machine learning-based OAV-guided sea cucumber origin and aquaculture pattern identification system, including a processor coupled with a memory, the memory storing program instructions, and when the program instructions stored in the memory are executed by the processor, the above-mentioned machine learning-based OAV-guided sea cucumber origin and aquaculture pattern identification method is implemented.
[0026] Based on the same inventive concept, the present invention also provides a computer-readable storage medium, characterized in that it includes a program that, when run on a computer, causes the computer to execute the above-described machine learning-based OAV-guided method for identifying the origin and farming mode of sea cucumbers.
[0027] A computer program (also referred to as a program, software, software application, script, or code) can be written in any form of programming language, including assembly or interpreted languages, declarative or procedural languages, and can be deployed in any form, including as a standalone program or as a module, component, subroutine, object, or other unit suitable for use in a computing environment. A computer program may, but does not necessarily, correspond to a file in a file system. A program can be stored as part of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to said program, or in multiple co-located files (e.g., a file storing one or more modules, subroutines, or code portions). A computer program can be deployed to execute on a single computer or on multiple computers located at a single site or distributed across multiple sites and interconnected by a communications network.
[0028] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0029] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0030] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0031] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0032] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the specific implementation of the present invention. Any modifications or equivalent substitutions that do not depart from the spirit and scope of the present invention should be covered within the scope of protection of the claims of the present invention.
Claims
1. A machine learning-based method for identifying the origin and farming model of sea cucumbers using OAV (Original Aspect-Guided Observation) technology, characterized in that... The method includes: Step S1: Sample collection and grouping, specifically: collect fresh sea cucumbers from different production areas and farming methods, and prepare them into freeze-dried powder through pretreatment; Step S2: Combined chromatographic and mass spectrometric monitoring to determine the parameters of volatile organic compounds; specifically: headspace solid-phase microextraction combined with gas chromatography-quadrupole-orbit trap mass spectrometry (GC-MS) is used to detect the sample, and qualitative and quantitative data of volatile organic compounds are obtained as organic compound parameters by deconvolution and comparison with a standard library. Step S3: Odor Activity Value (OAV)-guided biological noise reduction screening to obtain key flavor variables; specifically: for different types of organic compound parameters, compare them with their corresponding olfactory thresholds to determine the odor activity value (OAV), and take substances with an odor activity value (OAV) ≥ 1 as key flavor variables, and remove non-flavor contributing substances and environmental noise with an OAV < 1. Step S4: Select core flavor biomarkers based on resampling; specifically: perform n Bootstrap resampling operations on key flavor variables, run multiple screening models in parallel during each resampling operation, and count the frequency of selection for each key flavor variable. Select the high-frequency key flavor variables as robust features of consensus and use them as core flavor biomarkers; where n is a preset value. Step S5: Construction and validation of the intelligent discrimination model; specifically: construct a random forest classification model based on core flavor markers, and test the generalization ability and anti-overfitting performance of the intelligent discrimination model through k-fold cross-validation and permutation. Step S6: Analyze the flavor formation mechanism based on SHAP interpretability; specifically: calculate the SHAP value of each core flavor marker, quantify its marginal contribution to origin identification, and generate a dependency graph to analyze the flavor formation mechanism.
2. The method for identifying the origin and farming model of sea cucumber based on machine learning according to claim 1, characterized in that, The sea cucumber samples are mature sea cucumber samples grown in Dalian and mature sea cucumber samples grown in Fujian; the sample pretreatment in step S1 includes: crushing the sea cucumber samples and adding internal standard solution to the headspace vial for testing; the internal standard solution is 2-methyl-3-heptanone dissolved in methanol.
3. The method for identifying the origin and farming model of sea cucumber based on machine learning according to claim 2, characterized in that, The specific conditions for GC-MS analysis were as follows: a TG-5MS capillary column was used; helium was used as the carrier gas at a flow rate of 1.2 mL / min; the temperature program was as follows: initial temperature of 40℃ for 2.5 min, increased to 70℃ at 7℃ / min, then increased to 120℃ at 2℃ / min for 1 min, and finally increased to 280℃ at 20℃ / min for 5 min; the mass spectrometer ion source temperature was 280℃, the ionization energy was 70 eV, the scan range was m / z 35-500, and the resolution was 60,000.
4. The method for identifying the origin and farming model of sea cucumber based on machine learning according to claim 3, characterized in that, The Bootstrap resampling time n is at least 200 times; the selection criteria for the core flavor markers are: at least two of the three algorithms are determined to be important features, and the frequency of selection in all resampling iterations is greater than 50%.
5. The method for identifying the origin and farming model of sea cucumber based on machine learning according to claim 4, characterized in that, Resampling was performed using a screening model based on OPLS-DA, random forest, and LASSO algorithms.
6. The method for identifying the origin and farming model of sea cucumber based on machine learning according to claim 5, characterized in that, include: The OPLS-DA algorithm performs filtering by calculating the variable projection importance (VIP) filtering feature.
7. The method for identifying the origin and farming model of sea cucumber based on machine learning according to claim 6, characterized in that, The random forest algorithm evaluates feature importance by the amount of reduction in Gini impurity.
8. The machine learning-based OAV-guided sea cucumber origin and farming model identification system according to claim 7, characterized in that, The LASSO algorithm filters features by finding the minimum of a loss function with L1 regularization.
9. A machine learning-based OAV-guided sea cucumber origin and farming model identification system, characterized in that, The device includes a processor coupled to a memory, the memory storing program instructions, which, when executed by the processor, implement the machine learning-based OAV-guided sea cucumber origin and aquaculture mode identification method as described in any one of claims 1-8.
10. A computer-readable storage medium, characterized in that, The program, when run on a computer, causes the computer to perform the machine learning-based OAV-guided method for identifying the origin and farming pattern of sea cucumbers as described in any one of claims 1-8.