Metabolomics-based diagnostic markers and screening methods for Hong Kong oyster diseases
By screening 10 hepatopancreatic metabolic markers through metabolomics, and combining nuclear magnetic resonance and machine learning, a Hong Kong oyster disease diagnosis model was constructed, which solved the problem of difficulty in early detection of diseases in existing technologies and achieved efficient and accurate disease diagnosis and early warning.
Patent Information
- Application Number
- CN202310328875.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-30
- Publication Date
- 2025-09-12
- Estimated Expiration
- 2043-03-30
AI Technical Summary
Existing technologies make it difficult to accurately assess the health status of Hong Kong oysters by monitoring their physiological and biochemical indicators, resulting in difficulty in early detection and treatment of diseases, especially because their shells are hard and the symptoms of the disease are not easy to detect.
A metabolomics-based approach was used to screen out 10 hepatopancreatic metabolic markers, including glycogen, propionate, aspartic acid, lysine, arabinose, isoleucine, biotin, asparagine, fucose and nicotinamide adenine dinucleotide. A diagnostic model for Hong Kong oyster diseases was constructed by combining nuclear magnetic resonance technology and machine learning algorithms.
It has achieved rapid and accurate diagnosis of Hong Kong oyster diseases, established the foundation of a disease early warning system, improved the sensitivity and specificity of diagnosis, and is suitable for low-cost analysis of large-scale samples.
Smart Images

Figure CN116482155B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of aquatic animal disease diagnosis, and relates to metabolomics-based Hong Kong oyster disease diagnostic markers and a screening method thereof. Background Art
[0002] Hong Kong oysters (Crassostrea hongkongensis), commonly known as "white oysters" and "big oysters", are large in size, high in quality and delicious in taste. They are the most important economic shellfish species in southern production areas such as Guangdong, Guangxi, and Hainan. In addition to economic benefits, oyster farming also has important ecological benefits. As filter-feeding shellfish, oysters can effectively purify the marine environment and efficiently fix carbon in the ocean. They are called "removable carbon sinks" and are an important support for achieving the "carbon neutrality" goal. They have huge ecological development potential. Therefore, efforts to promote the development of the Hong Kong oyster industry in our province have very important economic and ecological benefits for ensuring food security, providing high-quality protein supply, conserving marine environmental resources, responding to climate crises, and achieving the "carbon neutrality" goal.
[0003] However, every year around the Qingming Festival, large-scale mortality of adult Hong Kong oysters occurs. Like other shellfish, Hong Kong oysters lack an adaptive immune system and are primarily cultured in open water bodies, making disease control through medication or water quality adjustments difficult. This severely restricts the healthy development of Hong Kong oysters. Therefore, strengthening early warning and prevention of disease outbreaks and effectively avoiding them are key technologies urgently needed for the Hong Kong oyster aquaculture industry. Identifying the symptoms of disease in Hong Kong oysters is the foundation of a disease early warning and forecasting system.
[0004] Because Hong Kong oysters are bivalves, their shells are very hard, making it difficult to observe what's going on inside. When an oyster becomes ill, the shell doesn't show any noticeable signs of illness. For example, the shell's color, shape, and texture may not look much different from those of a healthy oyster. Furthermore, even if we were able to open the oyster's shell, the symptoms inside might not be visible. This is because some pathogens can lurk in the oyster's body, making symptoms difficult to detect. Therefore, by monitoring the physiological and biochemical indicators of Hong Kong oysters, we can more accurately assess their health and promptly detect potential diseases. Currently, most research on the physiological and biochemical indicators of shellfish focuses on certain immune factors and antioxidant molecules, making it difficult to detect and treat diseases early.
[0005] In recent years, diagnostic methods based on metabolomics and artificial intelligence have been applied to early disease screening and diagnosis. Nuclear magnetic resonance (NMR)-based metabolomics offers highly reproducible and quantitative characteristics, high specificity, and the ability to simultaneously analyze metabolites qualitatively and quantitatively. Furthermore, the assay is inexpensive, making it particularly suitable for low-cost analysis of large sample volumes. However, there is currently no metabolomics-based diagnostic technology for Hong Kong oyster diseases. Summary of the Invention
[0006] In order to overcome the deficiencies of the prior art, the present invention aims to provide a metabolomics-based diagnostic marker for Hong Kong oyster diseases.
[0007] The purpose of the present invention is achieved through the following technical solutions:
[0008] A diagnostic marker combination for Hong Kong oyster diseases, comprising glycogen, propionic acid, aspartic acid, lysine, arabinose, isoleucine, biotin, asparagine, fucose, and nicotinamide adenine dinucleotide.
[0009] The present invention also provides the use of combination A as a diagnostic marker for Hong Kong oyster diseases, wherein the combination A is glycogen, propionic acid, aspartic acid, lysine, arabinose, isoleucine, biotin, asparagine, fucose, and nicotinamide adenine dinucleotide.
[0010] The present invention also provides a method for screening diagnostic markers of Hong Kong oyster diseases based on metabolomics, comprising the following steps:
[0011] S1. Collect hepatopancreas samples from Hong Kong oysters during mass mortality periods and normal periods at different times and locations as analytical samples;
[0012] S2. Perform metabolomics analysis on each analyzed sample using nuclear magnetic resonance technology to obtain the original metabolic fingerprint of each hepatopancreas sample;
[0013] S3. Use the R package Rnmr1D to perform chromatographic processing on the original metabolic fingerprints of the diseased and healthy hepatopancreatic samples, generating a two-dimensional matrix with metabolite information per row and the analyzed sample per column. Perform metabolite peak identification, peak area integration, and absolute quantification on the two-dimensional matrix for further machine learning.
[0014] S4. Use nine machine learning algorithms in the R package tidymodels to learn the two-dimensional matrix data of S3. Randomly use 3 / 4 of the above-mentioned diseased and healthy control hepatopancreatic sample data as the training set and 1 / 4 as the test set for learning. Randomly loop and iterate dozens of times. Statistically compare the average values of the accuracy, F value, Kappa value, Precision value, Recall value, and ROC_AUC value of different models on the test set to obtain the model with the best performance.
[0015] S5. Based on the optimal model obtained above, feature screening was performed using the Python packages ShapRFECV, Shap_hypertune, and BorutaShap to obtain differential metabolite combinations, and the differential metabolites obtained by the intersection of the three algorithms were selected as the combination;
[0016] S6. The datasets of differential metabolite combinations screened above were respectively trained using the nine machine learning algorithms of S4, with 3 / 4 of the dataset used as the training set and 1 / 4 as the test set. The training set was randomly iterated dozens of times. The average values of the accuracy, F value, Kappa value, Recall value, and ROC_AUC value of different metabolite combinations on different model test sets were statistically compared to finally obtain the relatively optimal metabolite combination.
[0017] S7. By comparing the performance of different models with the optimal metabolite combination, it was found that the model based on SVM performed the best. The confusion matrix confirmed that the metabolomic data of diseased and healthy Hong Kong oysters could be effectively classified.
[0018] S8. Perform differential analysis and correlation evaluation on the screened metabolites;
[0019] S9. Finally, based on the dataset of the optimal metabolite combination, a classification model was constructed using machine learning SVM to obtain an oyster disease diagnosis model.
[0020] Preferably, in the above method, the 9 machine algorithms described in S4 include: decision tree, support vector machine, extreme gradient boosting tree, neural network, random forest, multivariate linear regression, least absolute shrinkage and selection operator, logistic regression, and orthogonal partial least squares discriminant analysis.
[0021] Compared with the prior art, the present invention has the following beneficial effects:
[0022] This invention discloses metabolomics-based diagnostic markers for Hong Kong oyster diseases and their screening methods. The diagnostic markers comprise a combination of 10 hepatopancreatic tissue metabolite markers. Using nuclear magnetic resonance (NMR) technology, the metabolomics analysis of diseased Hong Kong oyster hepatopancreatic tissue was performed. Using artificial intelligence data analysis, differential metabolites were identified between diseased and healthy Hong Kong oysters. The method validates the diagnostic capabilities of the markers for Hong Kong oyster diseases, digitizes the phenotypic symptoms of Hong Kong oyster disease, and lays the foundation for the establishment of an early warning system for Hong Kong oyster diseases.
[0023] The diagnostic marker screening method of the present invention is highly operable, the model construction method is simple, and the resulting diagnostic model is effective, highly sensitive, and good specificity, making it suitable for diagnosing diseases of Hong Kong oysters. The present invention enables rapid diagnosis of diseases of Hong Kong oysters and provides a method for diagnosing other shellfish diseases, thus having great value for use and promotion. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] Figure 1 This is a histogram of all evaluation indicators of the test set using all metabolites as the input data set;
[0025] Figure 2This is a bar chart of all evaluation indicators of the test set using 10 metabolites as the input data set;
[0026] Figure 3 The confusion matrix is obtained using 10 metabolites as input data set;
[0027] Figure 4 This is the correlation analysis diagram of 10 metabolites;
[0028] Figure 5 Figure 4 is a differential analysis diagram of 10 metabolites. DETAILED DESCRIPTION
[0029] To better illustrate the purpose, technical solutions and advantages of the present invention, the present invention will be further described below with reference to the accompanying drawings and examples. In the examples, the experimental methods used are conventional methods unless otherwise specified, and the materials and reagents used are all commercially available unless otherwise specified.
[0030] Example 1 Screening of disease diagnostic markers
[0031] 1. Collect hepatopancreas samples from Hong Kong oysters at different times and locations during mass mortality periods and normal periods as analytical samples;
[0032] 2. Perform metabolomics analysis on each sample using nuclear magnetic resonance technology to obtain the original metabolic fingerprint of each hepatopancreas sample;
[0033] 3. Use the R package Rnmr1D to perform chromatographic processing on the original metabolic fingerprints of diseased and healthy hepatopancreatic samples, generating a two-dimensional matrix with metabolite information per row and the analyzed sample per column. Metabolite peak identification, peak area integration, and absolute quantification were performed on the two-dimensional matrix for further machine learning.
[0034] 4. Nine machine learning algorithms in the R package tidymodels were used, including decision tree (DT), support vector machine (SVM), extreme gradient boosting (XGBoost), multilayer perceptron (MLP), random forest (RF), multiple linear regression (MLR), least absolute shrinkage and selection operator (LASSO), logistic regression (LOG), and orthogonal partial least squares discriminant analysis (OPLS-DA). The two-dimensional matrix data from step 3 were learned. Three-quarters of the hepatopancreatic sample data from the diseased and healthy controls were randomly used as the training set, and one-quarter was used as the test set. The data were randomly iterated 40 times. The average values of the accuracy, F value, Kappa value, precision value, recall value, and ROC_AUC value of the different models on the test set were statistically compared. Figure 1 ), and found through comparison that the model based on XGBoost has the best performance and can effectively classify the metabolomics data of diseased and healthy Hong Kong oysters;
[0035] 5. Based on the XGBoost model obtained above, feature screening was performed using Python packages ShapRFECV, Shap_hypertune, BorutaShap, etc., and 16, 19, and 24 differential metabolite combinations were obtained, respectively. It was also found that the intersection of the three algorithms was 10 differential metabolite combinations.
[0036] 6. The datasets of 16, 19 and 24 metabolite combinations screened out above were respectively learned by the above 9 machine learning algorithms, of which 3 / 4 of the dataset was used as the training set and 1 / 4 as the test set, and the dataset was randomly iterated 40 times. By statistically comparing the average values of the accuracy, F value, Kappa value, Recall value and ROC_AUC value of different metabolite combinations on different model test sets, it was finally concluded that the number of relatively optimal metabolite combinations was 10. The 10 metabolites were glycogen, propionate, aspartic acid, lysine, arabinose, isoleucine, biotin, asparagine, fucose and nicotinamide adenine dinucleotide.
[0037] 7. At the same time, by comparing the performance of different models of 10 metabolite combinations ( Figure 2 ), and found that the model based on SVM had the best performance, with the average values of accuracy and area under the ROC curve (AUC) both above 0.9; through the confusion matrix, it was clear that the metabolomics data of diseased and healthy Hong Kong oysters could be effectively classified ( Figure 3 ).
[0038] 8. Perform differential analysis and correlation evaluation on the screened metabolites. Figure 4 As shown in Figure 2, the concentrations of 10 metabolites were significantly different between diseased and healthy oysters. Figure 5 ) showed that the correlation coefficients among the 10 metabolites were very low, indicating that the lower the repeated information among the screened markers, the simpler the metabolite combination tends to be and the more information it contains.
[0039] 9. Finally, based on the dataset of 10 metabolite combinations, a classification model was constructed using machine learning support machine learning (SVM) to obtain an oyster disease diagnosis model. Subsequently, by measuring the concentrations of 10 metabolites in the oyster hepatopancreas—glycogen, propionate, aspartic acid, lysine, arabinose, isoleucine, biotin, asparagine, fucose, and nicotinamide adenine dinucleotide—and inputting them into the oyster disease diagnosis model, it was determined whether the oyster was diseased.
[0040] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit the scope of protection of the present invention. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the essence and scope of the technical solutions of the present invention.
Claims
1. Use of combination A in the preparation of a product for the diagnosis of Hong Kong oyster diseases, characterized in that: The combination A is glycogen, propionic acid, aspartic acid, lysine, arabinose, isoleucine, biotin, asparagine, fucose and nicotinamide adenine dinucleotide.
Citation Information
Patent Citations
HIV clinical plan
CN109689866A
Metabolomics-based pancreatic cancer diagnosis marker as well as screening method and application thereof
CN110646554A
Method for distinguishing oysters from different producing areas based on quantitative nuclear magnetic resonance metabolomics
CN114894831A