Chinese sturgeon parent fish quality model system established based on noninvasive sample data
By constructing a health assessment model for Chinese sturgeons based on machine learning and using blood biochemical index data for evaluation, the problem of insufficient accuracy and real-time accuracy of Chinese sturgeons health assessment in the existing technology is solved, and a comprehensive and accurate assessment and timely warning of health status is achieved.
Patent Information
- Application Number
- CN202510045293.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-13
- Publication Date
- 2025-06-06
AI Technical Summary
The existing technology is difficult to achieve accurate and real-time assessment of the health status of Chinese sturgeons. The traditional methods are highly subjective and difficult to meet the high requirements for evaluation by modern aquaculture.
By collecting blood biochemical index data of sturgeon, after preprocessing, a health assessment model is constructed using machine learning models (such as logistic regression, random forest, support vector machine, etc.), train and cross-validation to obtain the optimal model, and finally a health score is performed based on the model output.
A comprehensive and accurate assessment of the health status of Chinese sturgeons has been achieved, which can promptly warn of health problems, thereby improving the health level of sturgeons.
Smart Images

Figure CN120108508A_ABST
Abstract
Description
Technical Field
[0001] The invention relates to the technical field of aquaculture, and in particular to a method for evaluating the health of Chinese sturgeon broodstock based on blood data. Background Art
[0002] The Chinese sturgeon is a national first-class protected animal. In recent years, due to the intensification of human activities, the habitat and spawning conditions of the Chinese sturgeon have further deteriorated, and natural reproduction activities have shown a discontinuous trend, and the continuation of the species faces severe challenges. Breeding Chinese sturgeon in an artificial environment is an important protection measure. At present, the health of artificially cultured Chinese sturgeon broodstock varies. Some broodstock have problems such as poor growth and development, gonadal degeneration, and severe aging. There is an urgent need to accurately assess the health of broodstock. Traditional health assessment methods for Chinese sturgeon broodstock mainly rely on manual observation and empirical judgment. This method is highly subjective and cannot meet the high requirements of modern aquaculture for assessment accuracy and real-time performance. In recent years, with the development of science and technology, data mining and machine learning technologies have made significant progress in the biomedical field, providing effective tools for complex biological data analysis. However, there is still a lack of systematic and standardized health assessment models for special farmed fish such as the Chinese sturgeon, making it difficult to achieve efficient monitoring and early warning. Summary of the invention
[0003] In view of the defects of the prior art, the present invention aims to provide a Chinese sturgeon broodstock health model system based on non-invasive sample data. The system scientifically integrates the blood biochemical index data of Chinese sturgeon broodstock, constructs a machine learning-based health assessment model and health score, and comprehensively evaluates the health level of Chinese sturgeon broodstock.
[0004] In order to achieve the above object, the present invention provides a method for evaluating the health of Chinese sturgeon broodstock based on blood data, comprising:
[0005] (1) Collect blood samples of Chinese sturgeon broodstock to construct a biochemical index data set, perform preprocessing, and establish reference value ranges for biochemical indexes;
[0006] (2) selecting a machine learning model to construct a Chinese sturgeon broodstock health model, inputting the preprocessed biochemical indicator data set into the Chinese sturgeon broodstock health model for training, and using cross-validation to obtain the optimal Chinese sturgeon broodstock health model;
[0007] (3) The blood data of the Chinese sturgeon broodstock to be evaluated is input into the optimal Chinese sturgeon broodstock health model, and a health score is calculated based on its output to complete the health evaluation of the Chinese sturgeon broodstock.
[0008] Further, the preprocessing includes:
[0009] Data cleaning: processing missing values and outliers in the biochemical index data set by MICE method or deletion method;
[0010] Data conversion: using StandardScaler to standardize the biochemical index data set to obtain a standardized format;
[0011] Data segmentation: dividing the biochemical indicator data set into a training set and a validation set.
[0012] Furthermore, the reference value range of the biochemical index is calculated by using the percentile method.
[0013] Furthermore, the machine learning models selected in step (2) include logistic regression, random forest, support vector machine, BP neural network, and K nearest neighbor.
[0014] Furthermore, the model training in step (2) is specifically as follows:
[0015] Divide the preprocessed biochemical index data set into k parts, and perform cross-validation on the selected machine learning model; wherein k is the number of selected machine learning models;
[0016] The grid search method is used to tune hyperparameters during the machine learning model training, and L1 or L2 regularization is implemented to prevent overfitting;
[0017] The accuracy, precision, recall, F1 score and AUC of the machine learning model are compared, and the optimal machine learning model is finally selected as the healthy model.
[0018] Furthermore, the health score in step (3) is specifically:
[0019] Score=AB*log(odds)
[0020] odds = p / (1-p)
[0021] Among them: Score is the health score, which is used to evaluate the health of Chinese sturgeon broodstock; A is a constant term, which represents the starting score without any additional information; B is a coefficient, which is used to adjust the rate of change of the score; p is the probability of an unhealthy state, which is the output of the optimal Chinese sturgeon broodstock health model.
[0022] The present invention also provides a Chinese sturgeon broodstock health assessment system based on blood data, comprising:
[0023] The data collection module is used to collect blood from Chinese sturgeon broodstock and construct a biochemical indicator data set;
[0024] The data preprocessing module is used to preprocess the collected biochemical index data set and establish the reference value range of the biochemical index;
[0025] Model building module, used to select machine learning models to build a Chinese sturgeon broodstock health model, including logistic regression, random forest, support vector machine, BP neural network, and K nearest neighbor;
[0026] A model training module, used for inputting the pre-processed biochemical indicator data set into the Chinese sturgeon broodstock health model for training, and adopting cross-validation to obtain the optimal Chinese sturgeon broodstock health model;
[0027] The health scoring module is used to perform health scoring based on the output of the optimal Chinese sturgeon broodstock health model and complete the health assessment of Chinese sturgeon broodstock;
[0028] The model deployment and optimization module is used to deploy the optimal machine learning model and health score in the Chinese sturgeon breeding base to realize the health monitoring of Chinese sturgeon broodstock; and regularly update and retrain according to new data and feedback.
[0029] Beneficial effects of the present invention:
[0030] The present invention conducts health modeling based on the biochemical indicators of the blood of Chinese sturgeon broodstock. The model system can comprehensively monitor the health status of Chinese sturgeon broodstock and timely warn of health problems, which is of great significance to improving the health level of broodstock. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] Figure 1 The present invention is a flowchart of a method for evaluating the health of Chinese sturgeon broodstock based on blood data according to an embodiment of the present invention.
[0032] Figure 2 1 is the ROC curve under different machine learning models of the embodiments of the present invention. DETAILED DESCRIPTION
[0033] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.
[0034] like Figure 1 As shown, the embodiment of the present invention provides a method for evaluating the health of Chinese sturgeon broodstock based on blood data, comprising the following steps:
[0035] S101. Collect blood from Chinese sturgeon broodstock to construct a biochemical index data set, perform preprocessing, and establish a reference value range for biochemical indexes.
[0036] The embodiment of the present invention collects blood biochemical indicators of Chinese sturgeon broodstock through a Chinese sturgeon breeding base. The collection process is as follows:
[0037] (1) Capturing broodstock: After catching the Chinese sturgeon, immediately transfer it to a pre-prepared stretcher and turnover water tank. After it stabilizes, collect samples and cover its eyes with a wet towel to help it calm down quickly.
[0038] (2) Blood sample collection and processing: When collecting blood, place the fish belly up and immerse the head in water to maintain normal breathing, and at the same time raise the tail moderately to keep it out of the water. Use a syringe to collect blood at the base of the anal fin. The collected blood is placed at room temperature and protected from light for 4 hours. After the blood is stratified, it is centrifuged (4,000 rpm, 10 minutes, 4°C). The serum is then drawn out, transferred to a centrifuge tube, and stored in a -80°C ultra-low temperature refrigerator for subsequent determination of blood biochemical indicators.
[0039] (3) Blood biochemical indicators: Liver function indicators include alanine aminotransferase, aspartate aminotransferase, alkaline phosphatase, total protein, albumin and globulin; muscle and heart function indicators include creatine kinase and lactate dehydrogenase; metabolic function indicators include urea, glucose, triglycerides, total cholesterol, high-density lipoprotein cholesterol and low-density lipoprotein cholesterol; inflammation and immune response indicator is C-reactive protein; electrolyte balance indicators include osmotic pressure, potassium, sodium, chloride and calcium; acid-base balance indicator is pH value; hormone level indicators include estradiol, testosterone and 11-ketotestosterone.
[0040] The collected biochemical indicator data sets are preprocessed, including: data cleaning, data conversion, and data segmentation.
[0041] Data cleaning: Missing values and outliers in the biochemical index data set were processed by MICE (Multiple Imputation by Chained Equations) or deletion method.
[0042] Data conversion: The biochemical indicator data set was standardized using StandardScaler to obtain a unified standardized format.
[0043] That is, each feature of the data is scaled to zero mean and unit variance. Specifically, it is standardized by the following steps:
[0044] (1) Calculate the mean: For each feature, calculate its mean; (2) Calculate the standard deviation: For each feature, calculate its standard deviation; (3) Standardize: For each value of each feature, subtract the mean and divide by the standard deviation.
[0045] Data Split: The data is randomly split into a training set (70%) and a validation set (30%). The training set is used to train the model, and the test set is used to finally evaluate the model performance.
[0046] Reference range of blood biochemical indexes: 95% percentile range was selected as the reference value range. The calculation formula used the 2.5th percentile (P2.5) and the 97.5th percentile (P97.5) as the upper and lower limits. The results are shown in Table 1.
[0047] Table 1
[0048] index Reference value range ALT 4.00—43.93U / L Aspartate aminotransferase 50.00—285.95 U / L Lactate dehydrogenase 393.00—2195.68U / L Alkaline phosphatase 21.00—96.95U / L Creatine kinase 226.00—3719.55U / L Total Protein 19.51—55.30g / L albumin 9.30—31.60g / L globulin 9.70—28.70g / L Urea 0.01—0.70mmol / L glucose 0.07—3.18mmol / L Triglycerides 0.30—8.06mmol / L Total cholesterol 0.65—5.62mmol / L HDL cholesterol 0.17—2.22mmol / L LDL cholesterol 0.03—1.72mmol / L C-reactive protein 0.00—34.01mg / L Osmotic pressure 253.00—313.00mmol / kg Potassium 1.20—3.25mmol / L sodium 118.50—188.39mmol / L chlorine 105.80—159.20mmol / L calcium 1.04—1.47mmol / L pH 7.39—7.95mmol / L Estradiol 0.29—220.09 pg / ml Testosterone 1.03—433.76 pg / ml 11-Ketotestosterone 9.13—213.20 pg / ml
[0049] S102, selecting a machine learning model to construct a Chinese sturgeon broodstock health model, inputting the preprocessed biochemical indicator data set data into the Chinese sturgeon broodstock health model for training, and using cross-validation to obtain the optimal Chinese sturgeon broodstock health model.
[0050] First, a Chinese sturgeon broodstock health model was constructed based on machine learning models such as logistic regression, random forest, support vector machine, BP neural network, and K nearest neighbor, etc. Its input is biochemical indicator data, and the output is the probability of unhealthy state of Chinese sturgeon broodstock.
[0051] Among them, logistic regression is mainly used to solve binary classification problems; support vector machines are suitable for processing linear and nonlinear classification problems and processing high-dimensional data; random forests process the interaction between high-dimensional data and features; BP neural networks are suitable for multi-layer neural networks, based on the gradient descent method, and their information processing capabilities come from multiple combinations of simple nonlinear functions; K nearest neighbors perform classification by measuring the distance between different feature values.
[0052] Then the training set of the preprocessed biochemical index data set is divided into k parts, where k is the number of selected machine learning models (here k=5). The selected machine learning model is trained, and the cross-validation technique is used to evaluate the performance of the model on different subsets to ensure the generalization ability of the model.
[0053] Machine learning models use grid search or random search methods to tune hyperparameters to improve performance, implement L1 or L2 regularization to prevent overfitting, and improve the robustness of the model.
[0054] Model evaluation indicators include:
[0055] (1) Accuracy: The ratio of the number of correctly predicted samples to the total number of samples.
[0056] (2) Recall: The ratio of the number of samples correctly predicted as positive to the number of actual positive samples.
[0057] (3) Precision: The ratio of the number of samples correctly predicted as positive to the number of samples predicted as positive.
[0058] (4) F1 score: the harmonic mean of precision and recall.
[0059] (5) AUC is a numerical representation of the area under the ROC curve, which is used to measure the overall performance of the classification model.
[0060] The evaluation results are shown in Table 2 and Figure 2 shown.
[0061] Table 2
[0062]
[0063] It can be seen from the above table that the support vector machine has the best prediction effect, so the support vector machine is used as the health model of Chinese sturgeon broodstock.
[0064] S103, input the blood data of the Chinese sturgeon broodstock to be evaluated into the optimal Chinese sturgeon broodstock health model, perform health scoring based on the output, and complete the health evaluation of the Chinese sturgeon broodstock.
[0065] Establish a health score for Chinese sturgeon broodstock, the scoring formula is:
[0066] Score=AB*log(odds)
[0067] odds = p / (1-p)
[0068] Among them: Score is the health score, which is used to evaluate the health of Chinese sturgeon broodstock; A is a constant term, A = Score 0 +B*log(odds 0 ), Score 0 is the benchmark score; B is a coefficient used to adjust the rate of change of the score, which is related to PDO (Points to Double the Odds), B = PDO / log(2); p is the probability of unhealthy state, which is the output of the optimal Chinese sturgeon broodstock health model; odds is the odds, which is the ratio of unhealthy state to healthy state; Log(odds): takes the natural logarithm of the odds, which is used to convert the odds into a score.
[0069] In the embodiment of the present invention, PDO=5; Score0=60; odds0=1.2840 is set, and Score=61.8037-7.2135*log(odds) is obtained. The health scores of the four Chinese sturgeon broodstock are shown in Table 3.
[0070] Table 3
[0071] index Sample 1 Sample 2 Sample 3 Sample 4 Alanine aminotransferase (U / L) 12.00 14.00 18.00 21.00 Aspartate aminotransferase (U / L) 89.00 106.00 80.00 78.00 Lactate dehydrogenase (U / L) 1276.00 307.00 860.00 588.00 Alkaline phosphatase (U / L) 91.00 44.00 32.00 24.00 Creatine kinase (U / L) 871.00 1372.00 1058.00 1229.00 Total protein (g / L) 30.60 14.80 42.20 50.90 Albumin (g / L) 13.90 6.30 22.80 29.70 Globulin (g / L) 16.70 8.50 19.40 21.20 Urea (mmol / L) 0.28 0.33 0.09 0.19 Glucose (mmol / L) 0.93 0.73 1.10 2.22 Triglyceride (mmol / L) 1.01 0.63 2.17 5.07 Total cholesterol (mmol / L) 2.21 0.66 2.43 3.23 High density cholesterol (mmol / L) 0.39 0.00 1.09 1.26 Low density cholesterol (mmol / L) 0.70 0.10 0.65 0.64 C-reactive protein (mg / L) 16.69 0.00 6.08 21.38 Osmotic pressure (mmol / kg) 266.00 268.00 264.00 294.00 Potassium (mmol / L) 1.94 2.22 1.89 1.83 Sodium (mmol / L) 141.70 157.80 148.50 143.30 Chlorine (mmol / L) 119.20 108.30 116.60 118.90 Calcium (mmol / L) 1.33 1.27 1.36 1.22 pH value (no unit) 7.64 7.85 7.57 7.49 Estradiol (pg / ml) 1.17 40.03 61.02 34.73 Testosterone (pg / ml) 10.02 17.44 302.25 50.02 11-Ketotestosterone (pg / ml) 9.56 19.55 111.75 92.32 Quality Rating 85.05 60.12 91.73 64.51
[0072] The health model is deployed in the Chinese sturgeon breeding base and health scores are performed to achieve health assessment of broodstock. At the same time, a regular physical examination mechanism is established to evaluate the accuracy and stability of the model in practical applications, and a feedback mechanism is established based on the actual prediction results and the observed health status of Chinese sturgeon broodstock to regularly update and retrain the model.
[0073] The embodiment of the present invention also provides a Chinese sturgeon broodstock health assessment system based on blood data, based on the above method, including:
[0074] The data acquisition module is used to collect the blood of Chinese sturgeon broodstock and construct a biochemical indicator data set.
[0075] The data preprocessing module is used to preprocess the collected biochemical indicator data set and establish the reference value range of biochemical indicators.
[0076] The model building module is used to select machine learning models to build a health model for Chinese sturgeon broodstock, including logistic regression, random forest, support vector machine, BP neural network, and K nearest neighbor.
[0077] The model training module is used to input the preprocessed biochemical indicator data set into the Chinese sturgeon broodstock health model for training, and adopt cross-validation to obtain the optimal Chinese sturgeon broodstock health model.
[0078] The health scoring module is used to perform health scoring based on the output of the optimal Chinese sturgeon broodstock health model and complete the health assessment of Chinese sturgeon broodstock.
[0079] The model deployment and optimization module is used to deploy the optimal machine learning model and health score in the Chinese sturgeon breeding base to realize the health monitoring of Chinese sturgeon broodstock; and regularly update and retrain according to new data and feedback.
[0080] The above is only a preferred embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any technician familiar with the technical field can make equivalent replacements or changes according to the technical solution and inventive concept of the present invention within the technical scope disclosed by the present invention, which should be covered by the protection scope of the present invention.
Claims
1. A method for evaluating the health of Chinese sturgeon broodstock based on blood data, characterized in that: The steps include: (1) Collect blood samples of Chinese sturgeon broodstock to construct a biochemical index data set, perform preprocessing, and establish reference value ranges for biochemical indexes; (2) selecting a machine learning model to construct a Chinese sturgeon broodstock health model, inputting the preprocessed biochemical indicator data set into the Chinese sturgeon broodstock health model for training, and using cross-validation to obtain the optimal Chinese sturgeon broodstock health model; (3) The blood data of the Chinese sturgeon broodstock to be evaluated is input into the optimal Chinese sturgeon broodstock health model, and a health score is calculated based on its output to complete the health evaluation of the Chinese sturgeon broodstock.
2. The method for evaluating the health of Chinese sturgeon broodstock based on blood data according to claim 1, characterized in that: The pre-processing comprises: Data cleaning: processing missing values and outliers in the biochemical index data set by MICE method or deletion method; Data conversion: using StandardScaler to standardize the biochemical index data set to obtain a standardized format; Data segmentation: dividing the biochemical indicator data set into a training set and a validation set.
3. The method for evaluating the health of Chinese sturgeon broodstock based on blood data according to claim 1, characterized in that: The reference value range of the biochemical indexes: the reference value range is calculated using the percentile method.
4. The method for evaluating the health of Chinese sturgeon broodstock based on blood data according to claim 1, characterized in that: The machine learning models selected in step (2) include logistic regression, random forest, support vector machine, BP neural network, and K nearest neighbor.
5. The method for evaluating the health of Chinese sturgeon broodstock based on blood data according to claim 1, characterized in that: The model training in step (2) is specifically as follows: Divide the preprocessed biochemical index data set into k parts, and perform cross-validation on the selected machine learning model; wherein k is the number of selected machine learning models; The grid search method is used to tune hyperparameters during the machine learning model training, and L1 or L2 regularization is implemented to prevent overfitting; The machine learning model is adopted to compare the accuracy, precision, recall, F1 score and AUC, and finally the optimal machine learning model is selected as the healthy model.
6. The method for evaluating the health of Chinese sturgeon broodstock based on blood data according to claim 1, characterized in that: The health score in step (3) is specifically: Score = AB * log (odds) odds = p / (1-p) Among them: Score is the health score, which is used to evaluate the health of Chinese sturgeon broodstock; A is a constant term, which represents the starting score without any additional information; B is a coefficient, which is used to adjust the rate of change of the score; p is the probability of an unhealthy state, which is the output of the optimal Chinese sturgeon broodstock health model.
7. A Chinese sturgeon broodstock health assessment system based on blood data, characterized in that: include: The data collection module is used to collect blood from Chinese sturgeon broodstock and construct a biochemical indicator data set; The data preprocessing module is used to preprocess the collected biochemical index data set and establish the reference value range of the biochemical index; Model building module, used to select machine learning models to build a Chinese sturgeon broodstock health model, including logistic regression, random forest, support vector machine, BP neural network, and K nearest neighbor; A model training module, used for inputting the pre-processed biochemical indicator data set into the Chinese sturgeon broodstock health model for training, and adopting cross-validation to obtain the optimal Chinese sturgeon broodstock health model; The health scoring module is used to perform health scoring based on the output of the optimal Chinese sturgeon broodstock health model and complete the health assessment of Chinese sturgeon broodstock; The model deployment and optimization module is used to deploy the optimal machine learning model and health score in the Chinese sturgeon breeding base to realize the health monitoring of Chinese sturgeon broodstock; and regularly update and retrain according to new data and feedback.