Triploid Fujian oyster producing area tracing method based on multi-omics feature fusion

By integrating multi-omics feature fusion, multi-dimensional information from inorganic elements, metabolites, and lipid molecules is combined to construct a high-accuracy and robust triploid Fujian oyster origin traceability model. This solves the shortcomings of single-omics methods in terms of accuracy and stability, achieving an origin traceability accuracy and stability of over 99%.

CN121410152APending Publication Date: 2026-01-27SANYA TROPICAL FISHERIES RES INST
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511753527.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-26
Publication Date
2026-01-27

AI Technical Summary

Technical Problem

In existing technologies, single-mathematical methods have low accuracy and poor stability when tracing the origin of triploid Fujian oysters, making it difficult to fully characterize the origin features, and the model has weak generalization ability.

Method used

A multi-omics feature fusion method was adopted to integrate multidimensional information of inorganic elements, hydrophilic metabolites and lipid molecules to construct a traceability model with high accuracy and robustness. The origin was determined by using quantitative detection data of ionomics, absolute quantitative metabolomics and macro-targeted lipidomics, combined with feature selection algorithm and machine learning classification model.

Benefits of technology

The accuracy rate of origin tracing for triploid Fujian oysters exceeded 99%, significantly better than single-mathematical methods. It possesses extremely high stability and anti-interference capabilities, and has constructed a three-dimensional traceability system that is difficult to forge, making it suitable for large-scale testing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure SMS_1
    Figure SMS_1
Patent Text Reader

Abstract

The invention belongs to the technical field of aquatic product producing area identification, and discloses a triploid Fujian oyster producing area tracing method based on multi-omics feature fusion. According to the invention, an omnibearing origin fingerprint spectrum containing an inorganic environment mark (ion group), a biological metabolism state (metabolome) and a nutrition reserve characteristic (lipidome) is constructed, and the fingerprint spectrum is difficult to counterfeit. The detection of the 26 markers required by the invention can be realized through a standardized kit and a mature mass spectrum platform, has a technical basis for large-scale detection, is suitable for being popularized in aquatic product wholesale markets, e-commerce platforms and geographical indication product authentication systems, and provides a large-scale technical means for the protection of geographical indication products.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of aquatic product origin identification technology, and more specifically, relates to a method for tracing the origin of triploid Fujian oysters based on multi-omics feature fusion. Background Technology

[0002] As an important economic shellfish, the quality and flavor of oysters are significantly affected by their growth environment (such as water quality, feed, and temperature). Oysters from different Fujian production areas vary in market price and consumer preferences. Therefore, establishing efficient and accurate methods for tracing the origin of oysters is of great significance for protecting geographical indication products, combating counterfeiting, and safeguarding consumer rights.

[0003] Existing technologies for oyster origin tracing can be categorized into several types: First, "fingerprint" technologies based on chemical composition analysis, such as elemental fingerprint analysis and isotope ratio analysis; second, technologies based on biomarkers, including lipidomics analysis, metabolomics analysis, and morphological analysis. Currently, shellfish origin tracing methods mainly include: 1. Single-analysis-based methods: such as analysis using only stable isotopes or mineral elements (ion groups). While effective to some extent, these methods are highly susceptible to dynamic environmental changes, and single-dimensional information often fails to comprehensively characterize the origin, resulting in a bottleneck in accuracy (typically between 80% and 90%); 2. Non-targeted metabolomics-based methods: although broad in coverage, they have low qualitative and quantitative accuracy and poor data reproducibility, making standardization and promotion in practical testing difficult; 3. Conventional machine learning methods: although already applied, they lack high-frequency stable feature screening for specific species (triploid Fujian oyster), resulting in weak model generalization ability. Summary of the Invention

[0004] To address the aforementioned shortcomings of existing technologies, this invention first provides a method for tracing the origin of triploid Fujian oysters based on the fusion of multi-omics features. This aims to solve the problems of insufficient accuracy and poor stability of existing single-technology traceability methods. This method integrates multidimensional information from inorganic elements, hydrophilic metabolites, and lipid molecules to construct a traceability model with extremely high accuracy (>99%) and robustness.

[0005] The objective of this invention is achieved through the following technical solution:

[0006] A method for tracing the origin of triploid Fujian oysters based on the fusion of multi-omics features is proposed. The method uses quantitative detection data of ionomics, absolute quantitative metabolomics, and macro-targeted lipidomics of triploid Fujian oysters as features, which are then processed by algorithms and machine learning classification models.

[0007] Preferably, the method for tracing the origin of triploid Fujian oysters based on multi-omics feature fusion includes the following steps:

[0008] S1. Collect oyster samples and prepare tissue homogenates;

[0009] S2. Perform multi-omics quantitative detection of tissue homogenate, absolute quantitative metabolomics and macro-targeted lipidomics to obtain data of ionomic, absolute quantitative metabolomics and macro-targeted lipidomics.

[0010] S3. Select extremely high frequency features from the multi-omics data in S2 that have been cross-validated by at least 9 feature selection algorithms to obtain a core biomarker combination containing inorganic elements, metabolites and lipid molecules.

[0011] The core marker combination includes:

[0012] Inorganic elements: Ni, Sn, Ag, Cu;

[0013] Metabolites: Proline, Trigonelline, Ursodeoxycholic acid, Imidazoleacetic acid, Methionine sulfoxide, Ornithine, Palmitoleic acid, S-Adenosylmethionine, Cysteinesulfinic acid, Cytidine-5'-diphosphate, Phenylpyruvic acid;

[0014] Lipid molecules: DG (18:0 / 18:1), FA (18:3), LPG (14:0), PA (18:0 / 18:0), PE (18:1 / 22:5), MG (22:2), PA (16:0 / 16:0), PE (14:0 / 16:1), PE (18:0 / 18:3), TG (52:6 / FA14:0), TG (54:8 / FA18:2);

[0015] S4. After normalizing the quantitative data of the marker combination described in S3, input it into the machine learning classification model for place of origin determination.

[0016] Preferably, the nine feature selection algorithms mentioned in S3 include univariate analysis of variance, LASSO regression, Boruta algorithm, VSURF, VarSelRF, RFE-RF, RFE-SVM, simulated annealing algorithm, and genetic algorithm.

[0017] Preferably, the machine learning classification model in S4 is Ridge Regression, Multinomial Regression, or Random Forest.

[0018] As a specific implementation method, in the above-mentioned method for tracing the origin of triploid Fujian oysters based on multi-omics feature fusion, the S2 ionome data are obtained through the following method:

[0019] Ionometry (ICP-MS)

[0020] Instrument: Inductively coupled plasma mass spectrometer (PerkinElmer NexION 1000G).

[0021] Pretreatment: Nitric acid-hydrogen peroxide microwave digestion method was used. 0.1 g of dry sample was weighed, 5 mL of HNO3 and 1 mL of H2O2 were added, and after microwave digestion, the volume was adjusted to 50 mL.

[0022] Detection conditions: RF power 1550 W, plasma flow rate 15 L / min, nebulizer flow rate 0.88 L / min, collision gas flow rate 4.0 mL / min. The contents of 35 elements were quantitatively determined using either the internal standard method or the external standard method.

[0023] The absolute quantitative metabolomics data described in S2 were obtained using the following methods:

[0024] Metabolomics assay (AQ700 high-throughput target)

[0025] Instrument: Ultra-high performance liquid chromatography (Agilent 1290 Infinity II) coupled with triple quadrupole mass spectrometry (ABSciex Triple Quad 6500+).

[0026] Chromatographic conditions: RP mode: ACQUITY UPLC BEH C18 column (1.7 µm, 2.1 × 150 mm); mobile phase A was water (0.1% formic acid), and mobile phase B was methanol:water = 95:5 (10 mmol / L ammonium formate).

[0027] HILIC mode: Atlantis Premier BEH Z-HILIC column (1.7 µm, 2.1 × 150 mm); mobile phase A was water:acetonitrile = 10:90 (10 mmol / L ammonium acetate), and mobile phase B was water:acetonitrile = 90:10 (10 mmol / L ammonium acetate).

[0028] Mass spectrometry conditions: ESI ion source, multiple reaction monitoring (MRM) mode. Ion source temperature 400℃, ion spray voltage +5000V / -4500V.

[0029] Quantitative method: An internal standard method was used to establish a standard curve for absolute quantification. An isotope-labeled internal standard was introduced to construct a standard curve of target metabolite concentration versus response ratio (R²>0.95). The absolute content of the metabolite (nmol / g) was calculated based on the standard curve equation and sample mass.

[0030] The data for the macro-targeted lipidome described in S2 were obtained through the following methods:

[0031] Lipomics assay (metapto-targeted lipidomics)

[0032] Instrument: Ultra-high performance liquid chromatography (Waters ACQUITY UPLC) coupled with Q-TRAP mass spectrometry (AB Sciex QTRAP4500 / 6500+).

[0033] Chromatographic conditions: Waters UPLC BEH C18 column (1.7 µm, 2.1 × 100 mm); column temperature 55 °C.

[0034] Mobile phase: Phase A is acetonitrile:water = 60:40 (10 mM ammonium acetate), Phase B is isopropanol:acetonitrile = 90:10 (10 mM ammonium acetate).

[0035] Gradient elution: elute gradient for 0-10 min at a flow rate of 0.26 mL / min.

[0036] Mass spectrometry conditions: ESI source, MRM mode scanning. Ion source temperature 600℃.

[0037] Quantitative method: Absolute quantification was performed using the isotope internal standard method. Feature matching was performed based on a self-built lipid database (VGDB) and public databases. Corresponding isotope internal standards were added for different lipid categories, and the absolute content of lipid molecules (µg / g) was calculated based on the peak area ratio of the internal standard to the target lipid.

[0038] This invention also provides the application of the above-mentioned method for tracing the origin of triploid Fujian oysters based on multi-omics feature fusion in the tracing of the origin of triploid Fujian oysters.

[0039] Compared with the prior art, the present invention has the following beneficial effects:

[0040] I. The Construction Principle of Multidimensional Complementary Spatiotemporal Fingerprints (Innovative Theoretical Foundation)

[0041] This invention is not a simple aggregation of omics, but rather a screening of three significantly complementary dimensions based on the deep mechanism of "environment-biological interaction":

[0042] Ion group (environmental fingerprint / spatial dimension): Inorganic elements are directly derived from seawater and sediments, reflecting the geochemical background and environmental pollution characteristics of the production site, and are the "spatial coordinates" of the production site.

[0043] Metabolomics (transient physiological response / current state): Hydrophilic small molecule metabolites respond rapidly to environmental stresses (such as salinity and temperature fluctuations), reflecting the real-time physiological state and acute stress response of oysters in specific habitats.

[0044] Lipidome (long-term adaptive memory / time dimension): The composition of lipids (such as membrane lipids and energy storage lipids) is affected by long-term feed structure and water temperature (such as unsaturation regulation), reflecting the long-term nutritional accumulation and chronic environmental adaptation of oysters, which constitutes the "time memory" of the place of origin.

[0045] Synergistic advantages: By integrating multi-dimensional information of "inorganic + organic" and "instantaneous + long-term", this invention constructs a three-dimensional traceability system that cannot be forged by a single means, which is the theoretical foundation for achieving an accuracy rate of >99%.

[0046] II. Significant Synergistic Effect

[0047] This invention demonstrates through comparative experiments that the performance of the multi-omics fusion model is significantly better than any single omics model.

[0048] Single metabolomics model: average accuracy approximately 82.5% (SVM); Single ionomics model: average accuracy approximately 90.3% (Multinom / RF); Single lipidomics model: average accuracy approximately 92.3% (Multinom); Dual-omics fusion (ionomics + lipids): average accuracy approximately 95.3% (Multinom); The three-omics fusion model of this invention: when using the optimized algorithm (Multinom / Ridge / RF), the average accuracy reaches 99.1% - 99.7%; even when using general algorithms (such as SVM / KNN), the accuracy can reach 97.5%.

[0049] The existing dual-omics method for oyster origin traceability has an accuracy of 95.3%. This invention uses a tri-omics method to improve the accuracy to 99.4%, an increase of 4.1 percentage points. This proves that the "transient physiological response" dimension captured by the metabolomics is a key information that cannot be replaced by the lipidomics and ionomics. The fusion of the three eliminates the blind spots of a single technology. Whether using the preferred model or the general model, its performance is significantly better than the best performance of a single omics method, achieving a qualitative leap in traceability accuracy.

[0050] The 26 markers selected by this invention are "extremely high frequency features" that have been identified in multiple feature selection methods (more than 9), exhibiting extremely high stability and universality. This avoids model overfitting due to individual differences and ensures the reliability of the method in practical applications.

[0051] This invention constructs a comprehensive geographical indication fingerprint that includes inorganic environmental imprints (ionomics), biological metabolic states (metabolics), and nutrient reserve characteristics (lipomics), making it difficult to counterfeit. The 26 biomarkers required for this invention can be detected using standardized reagent kits and a mature mass spectrometry platform, providing a technological foundation for large-scale detection. It is suitable for promotion in aquatic product wholesale markets, e-commerce platforms, and geographical indication product certification systems, offering a scalable technical means for the protection of geographical indication products. Detailed Implementation

[0052] To better illustrate the purpose, technical solution, and advantages of this invention, the invention will be further described below with reference to specific embodiments. Unless otherwise specified, the experimental methods used in the embodiments are conventional methods, and the materials and reagents used are commercially available unless otherwise specified.

[0053] Example 1

[0054] Triploid Fujian oyster samples were collected from at least four different locations in the waters off Guangdong and Fujian, China, with no fewer than 20 biological replicates from each location.

[0055] The method for tracing the origin of triploid Fujian oysters includes the following steps:

[0056] Step 1: Sample Collection and Preprocessing

[0057] Triploid Fujian oyster samples were collected and tissue homogenates were prepared.

[0058] Step 2: Multi-omics data collection

[0059] Three omics quantitative assays were performed on the same sample respectively:

[0060] 1. Ionomeascopy (ICP-MS)

[0061] Instrument: Inductively coupled plasma mass spectrometer (PerkinElmer NexION 1000G).

[0062] Pretreatment: Nitric acid-hydrogen peroxide microwave digestion method was used. 0.1 g of dry sample was weighed, 5 mL of HNO3 and 1 mL of H2O2 were added, and after microwave digestion, the volume was adjusted to 50 mL.

[0063] Detection conditions: RF power 1550 W, plasma flow rate 15 L / min, nebulizer flow rate 0.88 L / min, collision gas flow rate 4.0 mL / min. The contents of 35 elements were quantitatively determined using either the internal standard method or the external standard method.

[0064] 2. Metabolomics analysis (AQ700 high-throughput target)

[0065] Instrument: Ultra-high performance liquid chromatography (Agilent 1290 Infinity II) coupled with triple quadrupole mass spectrometry (ABSciex Triple Quad 6500+).

[0066] Chromatographic conditions: RP mode: ACQUITY UPLC BEH C18 column (1.7 µm, 2.1 × 150 mm); mobile phase A was water (0.1% formic acid), and mobile phase B was methanol:water = 95:5 (10 mmol / L ammonium formate).

[0067] HILIC mode: Atlantis Premier BEH Z-HILIC column (1.7 µm, 2.1 × 150 mm); mobile phase A was water:acetonitrile = 10:90 (10 mmol / L ammonium acetate), and mobile phase B was water:acetonitrile = 90:10 (10 mmol / L ammonium acetate).

[0068] Mass spectrometry conditions: ESI ion source, multiple reaction monitoring (MRM) mode. Ion source temperature 400℃, ion spray voltage +5000V / -4500V.

[0069] Quantitative method: An internal standard method was used to establish a standard curve for absolute quantification. An isotope-labeled internal standard was introduced to construct a standard curve of target metabolite concentration versus response ratio (R²>0.95). The absolute content of the metabolite (nmol / g) was calculated based on the standard curve equation and sample mass.

[0070] 3. Lipomics assay (metapeptide analysis)

[0071] Instrument: Ultra-high performance liquid chromatography (Waters ACQUITY UPLC) coupled with Q-TRAP mass spectrometry (AB Sciex QTRAP4500 / 6500+).

[0072] Chromatographic conditions: Waters UPLC BEH C18 column (1.7 µm, 2.1 × 100 mm); column temperature 55 °C.

[0073] Mobile phase: Phase A is acetonitrile:water = 60:40 (10 mM ammonium acetate), Phase B is isopropanol:acetonitrile = 90:10 (10 mM ammonium acetate).

[0074] Gradient elution: elute gradient for 0-10 min at a flow rate of 0.26 mL / min.

[0075] Mass spectrometry conditions: ESI source, MRM mode scanning. Ion source temperature 600℃.

[0076] Quantitative method: Absolute quantification was performed using the isotope internal standard method. Feature matching was performed based on a self-built lipid database (VGDB) and public databases. Corresponding isotope internal standards were added for different lipid categories, and the absolute content of lipid molecules (µg / g) was calculated based on the peak area ratio of the internal standard to the target lipid.

[0077] Step 3: Quantification and Screening of Characteristic Biomarkers

[0078] We employ an "Ensemble Feature Selection" strategy to identify the most stable core biomarkers from massive omics data.

[0079] 1. Initial screening and preprocessing: Perform quality control on the raw data (remove variables with zero variance), handle missing values, and standardize the data.

[0080] Missing value handling: First, remove samples or features with an excessively high proportion of missing values; for values ​​below the detection limit (0 value), fill them with half of the smallest positive value of the feature to ensure data integrity.

[0081] Data standardization: Perform log transformation and Pareto scaling on the raw data.

[0082] The preprocessing here aims to eliminate the influence of dimensions and ensure the convergence and accuracy of subsequent feature selection algorithms (such as LASSO and SVM).

[0083] 2. Multi-algorithm cross-validation: At least nine feature selection algorithms based on different principles are used for independent screening, including:

[0084] Statistical methods: univariate analysis of variance (ANOVA), univariate model screening.

[0085] Regularization method: LASSO regression (lambda.min & lambda.1se).

[0086] Tree model methods: Boruta algorithm, VSURF (Variable Selection Using Random Forests), VarSelRF.

[0087] Recursive elimination methods: RFE-RF (Recursive Feature Elimination Based on Random Forest) and RFE-SVM (Recursive Feature Elimination Based on Support Vector Machine).

[0088] Heuristic search: RF-SA / SVM-SA (Simulated Annealing Algorithm), RF-GA / SVM-GA (Genetic Algorithm).

[0089] 3. Determination of extremely high frequency features: Statistically analyze the frequency of features selected by the above algorithms, and retain only the features that are identified by more than 9 algorithms, which are defined as "extremely high frequency features".

[0090] Based on the above rigorous process, 26 core differential markers (panels) were finally identified:

[0091] Four inorganic elements: Ni (nickel), Sn (tin), Ag (silver), and Cu (copper).

[0092] Eleven metabolites: Proline, Trigonelline, Ursodeoxycholic acid, Imidazoleacetic acid, Methionine sulfoxide, Ornithine, Palmitoleic acid, S-Adenosylmethionine, Cysteinesulfinic acid, Cytidine-5'-diphosphate, and Phenylpyruvic acid.

[0093] Eleven lipid molecules: DG (18:0 / 18:1), FA (18:3), LPG (14:0), PA (18:0 / 18:0), PE (18:1 / 22:5), MG (22:2), PA (16:0 / 16:0), PE (14:0 / 16:1), PE (18:0 / 18:3), TG (52:6 / FA14:0), TG (54:8 / FA18:2).

[0094] Step 4: Data Normalization and Fusion

[0095] Feature extraction: Only quantitative data of the 26 core biomarkers selected above were extracted.

[0096] Data alignment and fusion: Based on the unique sample ID, the inner join algorithm is used to accurately align the ionomic, metabolomic, and lipidomic data and merge them into a multidimensional feature matrix to ensure that different omics data come from the same biological sample.

[0097] Secondary standardization: The merged feature matrix is ​​subjected to logarithmic transformation and Pareto standardization again.

[0098] Although preprocessing has been performed in step 3, step 4 is to construct an independent training set for the final 26 selected features. It is necessary to ensure that the feature matrix of the input model meets the requirements of the machine learning algorithm for data distribution (such as normality and homoscedasticity). Therefore, this step is an indispensable part of model construction and does not overlap with the screening and preprocessing in step 3.

[0099] Step 5: Determining the place of origin

[0100] The processed data is input into a pre-trained machine learning classification model. Based on the model performance evaluation results, the model is preferably selected from algorithms with high accuracy, including but not limited to Ridge Regression, Multinomial Regression, or Random Forest, all of which have a validation accuracy exceeding 99%; other algorithms such as SVM and KNN can also be used. The final output is the origin classification result of the sample.

[0101] Rigorous nested cross-validation (outer layer 2-fold × 10, inner layer 2-fold) was used to evaluate model performance. The model performance of the samples in this embodiment was compared across different omics datasets, and the results are shown in Table 1.

[0102] Table 1

[0103] The accuracy data mentioned above were obtained through nested cross-validation (outer layer 2 folds × 10 repetitions, for a total of 20 independent tests). Key findings: The accuracy improved by 4.1 percentage points from 95.3% with the dual-omics model to 99.4% with the tri-omics model, demonstrating that the "transient physiological response" dimension captured by the metabolomics is irreplaceable by the ionomics and lipidomics. The standard deviation of the tri-omics fusion model was only ±0.006, far lower than that of the single-omics model (metabolomics ±0.018), demonstrating extremely high predictive stability and robustness.

Claims

1. A method for tracing the origin of triploid Fujian oysters based on multi-omics feature fusion, characterized in that, The ionome, absolute quantitative metabolome, and macro-targeted lipidome data of triploid Fujian oysters were used as features and processed by algorithms and machine learning classification models.

2. The method for tracing the origin of triploid Fujian oysters based on multi-omics feature fusion according to claim 1, characterized in that, Includes the following steps: S1. Collect oyster samples and prepare tissue homogenates; S2. Perform multi-omics quantitative detection of tissue homogenate, absolute quantitative metabolomics and macro-targeted lipidomics to obtain data of ionomic, absolute quantitative metabolomics and macro-targeted lipidomics. S3. Select extremely high frequency features from the multi-omics data in S2 that have been cross-validated by at least 9 feature selection algorithms to obtain a core biomarker combination containing inorganic elements, metabolites and lipid molecules. The core marker combination includes: Inorganic elements: Ni, Sn, Ag, Cu; Metabolites: Proline, Trigonelline, Ursodeoxycholic acid, Imidazoleacetic acid, Methionine sulfoxide, Ornithine, Palmitoleic acid, S-Adenosylmethionine, Cysteinesulfinic acid, Cytidine-5'-diphosphate, Phenylpyruvic acid; Lipid molecules: DG (18:0 / 18:1), FA (18:3), LPG (14:0), PA (18:0 / 18:0), PE (18:1 / 22:5), MG (22:2), PA (16:0 / 16:0), PE (14:0 / 16:1), PE (18:0 / 18:3), TG (52:6 / FA14:0), TG (54:8 / FA18:2); S4. After normalizing the quantitative data of the marker combination described in S3, input it into the machine learning classification model for place of origin determination.

3. The method for tracing the origin of triploid Fujian oysters based on multi-omics feature fusion according to claim 1, characterized in that, The nine feature selection algorithms mentioned in S3 include univariate analysis of variance, LASSO regression, Boruta algorithm, VSURF, VarSelRF, RFE-RF, RFE-SVM, simulated annealing algorithm, and genetic algorithm.

4. The method for tracing the origin of triploid Fujian oysters based on multi-omics feature fusion according to claim 1, characterized in that, The machine learning classification model mentioned in S4 is Ridge Regression, Multinomial Regression, or Random Forest.

5. The application of the method of any one of claims 1 to 4 in the origin tracing of triploid Fujian oysters.