Intelligent disease early warning system and storage medium for aquaculture environment
By building an intelligent disease early warning system for aquaculture environments and utilizing machine learning technology and multi-source data collection, we have solved the problems of delayed disease monitoring and inaccurate early warning in aquaculture, achieved real-time and accurate disease early warning and personalized prevention and control, and improved aquaculture efficiency and safety.
Patent Information
- Application Number
- CN202511029685.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-25
- Publication Date
- 2025-09-16
- Estimated Expiration
- 2045-07-25
AI Technical Summary
The current aquaculture sector lacks effective disease monitoring and early warning methods. Traditional methods have delayed responses and low accuracy, making it difficult to fully cover the various environmental factors that affect the health of aquatic animals. In addition, early warning information is not transmitted in a timely manner, making disease prevention and control difficult.
Establish an intelligent disease early warning system for aquaculture environment, use machine learning technology to build a correlation model between environmental factors and disease occurrence, and combine multi-source heterogeneous environmental data collection, edge computing, Internet of Things communication, machine learning prediction models and application service layers to achieve real-time and accurate disease early warning.
It realizes comprehensive monitoring and real-time early warning of the breeding process, improves the accuracy and efficiency of disease prevention and control, provides precise and efficient disease prevention and control measures, ensures that early warning information is delivered in time and generates personalized prevention and control plans.
Smart Images

Figure CN120526554B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of aquaculture automation control, and in particular relates to an intelligent disease early warning system for an aquaculture environment and a storage medium. Background Art
[0002] Aquatic products are an important source of high-quality protein for humans. In recent years, with growing consumer demand, my country's aquaculture industry has entered a period of rapid development. According to statistics, my country's total aquatic product output reached 64.5 million tons in 2020, firmly ranking first in the world. Farmed aquatic product output accounts for 77% of the total aquatic product output. Intensive aquaculture methods, such as ponds and factory-scale recirculating water systems, are highly susceptible to pathogenic microorganisms due to high stocking densities and poor water flow.
[0003] The current aquaculture sector still lacks effective means for disease monitoring and early warning. Traditional practices rely primarily on regular manual testing of key water quality indicators, combined with experience to determine disease risks. However, this approach has significant shortcomings, such as delayed response, untimely warnings, and low accuracy. Specifically, the technical challenges currently facing aquatic disease prevention and control mainly include incomplete monitoring of environmental factors. Numerous environmental factors affect the health of aquatic animals, ranging from conventional indicators such as water temperature and pH to unconventional indicators such as turbidity and heavy metal content. Traditional manual detection methods often fail to fully cover these factors, resulting in inaccurate and incomplete monitoring results.
[0004] Furthermore, the interactions between environmental factors are extremely complex. Research has shown that these factors exhibit complex nonlinear correlations, meaning that changes in a single indicator do not necessarily lead to disease. The key lies in identifying the specific combinations of factors that trigger disease, which places high demands on early warning algorithms. Furthermore, existing early warning models lack generalizability. The key combinations of disease-causing factors vary significantly across different farmed species and geographical environments. Traditional empirical models lack universal applicability and are limited in scope, making them inadequate for meeting the demands of actual farming operations. Finally, early warning information is not delivered in a timely manner. Currently, the aquaculture sector lacks robust information management tools, preventing the automatic transmission of environmental monitoring data and timely delivery of disease warning information to frontline staff. This severely impacts early warning effectiveness and hinders effective disease prevention and control.
[0005] Therefore, this field needs to develop an innovative intelligent disease early warning system and storage medium for aquaculture environments, in order to provide new ideas and new solutions for aquatic disease prevention and control from the perspective of the aquaculture environment. Summary of the Invention
[0006] The purpose of the present invention is to provide an intelligent disease early warning system and storage medium for aquaculture environments. The system innovatively establishes a correlation model between environmental factors and disease occurrence, and uses machine learning technology to achieve an early warning accuracy of more than 85%. It can realize comprehensive monitoring and real-time early warning of the breeding process, and support accurate and efficient disease prevention and control; thereby solving the problems of low accuracy and delayed response of traditional early warning methods, and significantly improving the effect of aquaculture disease prevention and control.
[0007] To achieve the above-mentioned objectives, the present invention provides an intelligent disease early warning system for aquaculture environment, including a data acquisition layer, a data transmission layer, a processing platform layer and an application service layer; the data acquisition layer includes an environmental monitoring unit, a water quality sensor group and a portable detection device, which are used to realize the automatic collection of multi-source heterogeneous environmental data, and can comprehensively cover various environmental factors that affect the health of aquatic animals; the data transmission layer includes a wireless transmission module and an edge computing node, the wireless transmission module is connected to the data acquisition layer via a high-speed data line, and the edge computing node is used for data preprocessing and transmission; the data acquisition layer is connected to the data transmission layer via a bus interface. The data transmission layer aggregates the data collected by each node to the cloud through the Internet of Things communication technology. At the same time, it is equipped with an edge computing node to perform preliminary cleaning and compression at the data source to reduce the network load.
[0008] The processing platform layer includes a machine learning prediction model and an environmental factor correlation analysis model. The processing platform layer is connected to the data transmission layer via a wireless network to provide intelligent early warning decisions. The early warning accuracy of the machine learning prediction model is over 85%.
[0009] The environmental factor association analysis model uses mathematical statistics and association rule mining technology to characterize the quantitative influence mechanism between different environmental factors and screen out the key factor combination that has a significant impact on the disease. On this basis, the machine learning prediction model integrates multiple algorithms such as convolutional neural networks, long short-term memory networks, random forests, etc. to make accurate predictions on the probability of disease occurrence in the future. At the same time, the processing platform layer also has a model evaluation and optimization unit, which evaluates the model effect through cross-validation and other technologies. When the warning accuracy rate does not meet the standard, the model retraining process is automatically started to continuously improve the system's prediction performance. Among them, the environmental factor association analysis model is used to establish the correlation between environmental factors and disease occurrence, and pass the analysis results to the machine learning prediction model for disease warning.
[0010] The application service layer directly serves end-user farmers, providing user-friendly and visual information services. The application service layer includes an early warning information push module and a prevention and control plan generation module. The application service layer is connected to the processing platform layer via a data interface to enable accurate early warning information push and prevention and control plan generation. The early warning information push module receives early warning signals output by the processing platform, classifies them according to risk level, and promptly delivers them to users through multiple channels such as app push, text messages, and voice calls. The prevention and control plan generation module connects to the internal knowledge base, matches disease characteristics, and automatically generates specific measures such as drug dosage and environmental control. It also instantiates them based on the scale of the farming industry, directly guiding users in implementing prevention and control measures. A user feedback channel is also provided to collect user experience evaluations, forming a closed loop.
[0011] Preferably, the environmental monitoring unit includes a temperature sensor, a humidity sensor, and a light sensor, covering the monitoring of air parameters above the water surface; wherein the temperature sensor has a measurement accuracy of ±0.1°C and a range of -10 to 50°C; the humidity sensor has a measurement accuracy of ±2%RH and a range of 0 to 100%RH; the light sensor has a measurement accuracy of ±3% and a range of 0 to 200,000 Lux;
[0012] The water quality sensor group includes a pH sensor, a dissolved oxygen sensor, and an ammonia nitrogen sensor, responsible for monitoring key underwater water quality parameters. The pH sensor has a measurement accuracy of ±0.01 and a range of 0 to 14 pH. The dissolved oxygen sensor has a measurement accuracy of ±0.1 mg / L and a range of 0 to 20 mg / L. The ammonia nitrogen sensor has a measurement accuracy of ±0.1 mg / L and a range of 0 to 100 mg / L.
[0013] The portable testing device is a handheld multi-parameter water quality analyzer with COD, turbidity, and heavy metal ion detection functions. It can be carried by staff to the aquaculture pond to quickly test turbidity, COD and other indicators. The handheld multi-parameter water quality analyzer is connected to the data acquisition layer through a waterproof connector.
[0014] Taking into account the environmental differences at different depths and locations, the sensors in the environmental monitoring unit are evenly installed at an angle of 120° 2 meters above the aquaculture water body, and the water quality sensor group is evenly installed at an angle of 60° at a depth of 0.5 meters in the aquaculture water body. The sensors in the data acquisition layer are connected to the data acquisition module of the data acquisition layer through the RS485 bus, and the sampling frequency is 1 time / minute.
[0015] Preferably, the wireless transmission module includes a 5G communication unit and a WiFi communication unit; wherein, the 5G communication unit supports SA and NSA dual modes, with an uplink rate of 1Gbps and a downlink rate of 3Gbps; the WiFi communication unit supports WiFi6 protocol, with a transmission rate of 2.4Gbps; the superposition of multiple communication standards ensures the stability of data backhaul.
[0016] The edge computing node is equipped with a 3.2GHz quad-core processor, 8GB of RAM, and a 256GB solid-state drive. It has data cleaning, outlier processing, and data compression functions, with a data processing capacity of ≥1,000 records per second. The edge computing node has a data caching mechanism that locally stores more than 72 hours of monitoring data in the event of a network interruption, and automatically uploads the cached data to the processing platform layer after the network is restored.
[0017] The wireless transmission module is connected to the data acquisition layer through a Gigabit Ethernet interface, with a data transmission delay of less than 10ms. Fiber optic communication is used between the edge computing node and the wireless transmission module, with a transmission bandwidth of 10Gbps; the overall structure constitutes a low-latency, highly reliable, and highly complementary data transmission architecture.
[0018] Preferably, the machine learning prediction model adopts a combination of deep learning and ensemble learning, which includes a data preprocessing module, a feature extraction module, a model training module and a prediction output module; wherein the data preprocessing module standardizes and normalizes the collected environmental and water quality data; the feature extraction module uses a convolutional neural network to extract time series features and uses an attention mechanism to highlight the influence weight of key environmental factors; the model training module combines LSTM and RandomForest algorithms, uses historical data for model training and parameter optimization, and the training data set contains no less than 10,000 historical records; the prediction output module generates an early warning signal based on the model prediction results and combined with a confidence threshold;
[0019] The environmental factor association analysis model has the following key points in its implementation:
[0020] (1) For continuous variables such as temperature, pH, oxygen, and ammonia nitrogen, the quantile binning method is used for discretization. At the same time, the binning interval is adaptively adjusted in combination with the maximum information entropy criterion to ensure the optimization of feature expression.
[0021] (2) For each monitoring indicator, slice processing is performed according to the time window (e.g., 24 hours); by extracting multiple statistics such as mean, variance, and peak, the time series characteristics of the variable can be fully characterized;
[0022] (3) In the frequent pattern mining stage, in addition to using the Apriori algorithm, a time series pattern mining algorithm is also introduced; this algorithm fully considers the previous and next correlations between indicators; the generated association rules are screened using the chi-square test to remove false positive results;
[0023] (4) When constructing the environmental factor weight matrix, weights are given from the perspectives of data-driven and expert knowledge. In terms of data-driven, the entropy method is used to calculate information gain; in terms of expert knowledge, the analytic hierarchy process (AHP) is used for subjective weighting. Finally, the two are integrated through weighted averaging and other methods.
[0024] The analysis process runs every 12 hours, and the generated feature matrix is passed to the prediction module in JSON format. The entire correlation analysis unit is deployed as a microservice, with the ability to scale independently and support deep coupling with the machine learning platform.
[0025] Specifically, the environmental factor association analysis model is based on the Apriori algorithm and the grey correlation analysis method, and the association analysis between environmental factors and disease occurrence is achieved through the following steps:
[0026] a. Bin the environmental factor data and establish a candidate set of association rules;
[0027] b. Calculate support and confidence to filter out strong association rules;
[0028] c. Use grey correlation to calculate the influence of each environmental factor on the occurrence of the disease;
[0029] d. Generate an environmental factor weight matrix and input it into the machine learning prediction model to improve the accuracy of early warning;
[0030] The processing platform layer also includes a model evaluation and optimization module, which evaluates the performance of the model through cross-validation and confusion matrix analysis. When the warning accuracy is lower than 85%, the model retraining mechanism is automatically triggered to improve the model performance through parameter tuning and sample expansion.
[0031] Preferably, the specific implementation of the machine learning prediction model is:
[0032] A. Data preprocessing module processes data through the following steps;
[0033] a. Use mean filling method to handle missing values and use 3σ criterion to remove outliers;
[0034] b. Perform Min-Max normalization on continuous data and one-hot encoding on discrete data;
[0035] c. Divide the dataset into training set and test set in a ratio of 8:2;
[0036] To improve robustness, we prioritized data-driven methods such as KNN for missing value filling, and non-parametric detection models such as Isolation Forest for outlier detection. These methods ensured data quality and laid a solid foundation for subsequent feature extraction and model training.
[0037] B. The convolutional neural network of the feature extraction module consists of three convolutional layers, with kernel sizes of 3×3, 5×5, and 7×7, respectively, and a ReLU activation function. The attention mechanism uses a multi-head self-attention structure with 8 heads to capture feature dependencies at different time scales.
[0038] The feature extraction module utilizes a multi-scale one-dimensional convolutional neural network. By setting convolution kernels with different receptive field sizes, the network can adaptively extract local and global features from time series data. Furthermore, the attention mechanism utilizes a multi-head self-attention mechanism. This multi-head design further enhances the diversity of feature selection, enabling the model to more accurately capture key information in the data.
[0039] C. The LSTM network combined with the model training module contains two hidden layers, each with 128 nodes and a time step of 24. The Random Forest model contains 100 decision trees, with a maximum depth of 10 and a minimum number of leaf node samples of 5. The outputs of the two models are integrated using the stacking method, using LightGBM as the secondary learner.
[0040] The model training module uses LSTM as its core framework, effectively controlling overfitting through methods such as Dropout and L1 / L2 regularization. Furthermore, the model incorporates the integration capabilities of tree models such as Random Forest, further reducing model variance and improving model stability and accuracy. Furthermore, the model incorporates transfer learning, leveraging historical data from similar farms to aid modeling and enhance generalization.
[0041] The prediction output module uses the soft voting method to fuse the model prediction results, sets the confidence threshold to 0.8, and triggers an early warning when the prediction probability exceeds the threshold; at the same time, it calculates the precision, recall rate and F1 value based on the confusion matrix, and feeds the evaluation results back to the model training module for continuous optimization.
[0042] In terms of loss functions, in addition to the commonly used cross-entropy, business metrics such as precision and recall are also incorporated for optimization. Furthermore, Focal Loss is introduced to alleviate data imbalance. For hyperparameter tuning, heuristic search strategies such as Bayesian optimization are employed. The optimization space covers multiple aspects, including the number of network layers, number of units, learning rate, and regularization term, ensuring optimal model performance.
[0043] The key paths of machine learning prediction models are solidified in the ONNX format, enabling cross-platform deployment. The prediction service API adopts the RESTful standard and automatically generates call documentation via Swagger, facilitating user invocation and integration. The service boasts a throughput of up to 100 queries per second and a P99 latency of less than 100ms, fully meeting real-time early warning requirements.
[0044] Preferably, the environmental factor association analysis model adopts a multidimensional analysis method, including the following steps:
[0045] Step S1, environmental factor identification: Based on the principal component analysis method, key factor data that have a significant impact on the occurrence of the disease are screened from monitoring data such as temperature, pH value, dissolved oxygen, and ammonia nitrogen. The screening standard is that the cumulative contribution rate reaches more than 85%;
[0046] Step S2, data preprocessing: standardize the screened environmental factor data and use the sliding window method to calculate the temporal variation characteristics of each factor, with a window size of 24 hours and a step size of 1 hour;
[0047] Step S3, association rule mining: combining the Apriori algorithm and the time series association rule mining method, setting the minimum support to 0.2 and the minimum confidence to 0.6, mining the association rules between the combination of environmental factors and the occurrence of diseases, and verifying the significance of the rules based on the chi-square test;
[0048] Step S4, factor weight calculation: using a combination of analytic hierarchy process and entropy weight method to establish an environmental factor weight evaluation system, and calculate the weight of each factor's impact on different types of diseases;
[0049] Step S5, model integration: encapsulate the association rules and weight information of environmental factors into feature vectors, and pass them to the machine learning prediction model in JSON format to improve the accuracy and real-time performance of early warnings;
[0050] The environmental factor association analysis model updates the analysis results every 12 hours and evaluates the model performance through cross-validation. The verification accuracy must reach more than 80%.
[0051] Preferably, the environmental factor association analysis model is implemented by the following steps:
[0052] a. Use the equal frequency binning method to divide the continuous environmental factor data into 10 intervals, and use the optimal binning algorithm based on information entropy to optimize the intervals;
[0053] b. Set the minimum support threshold to 0.3 and the minimum confidence threshold to 0.7 in the Apriori algorithm, and generate frequent itemsets through iteration;
[0054] c. When calculating the grey correlation degree, temperature is selected as the reference sequence, the resolution coefficient is taken as 0.5, and the correlation coefficient of each environmental factor is obtained;
[0055] d. Normalize the correlation coefficient to generate a 10×10 weight matrix;
[0056] The model evaluation and optimization module includes the following functions: 10-fold cross-validation is used to evaluate model performance and calculate accuracy, precision, recall, and F1 value; confusion matrix is used to analyze the warning effect of different types of diseases and calculate the warning success rate of each type of disease; when the warning accuracy of any type of disease is less than 85%, the model optimization mechanism is triggered;
[0057] The model optimization mechanism is implemented through the following steps:
[0058] a. Use grid search to optimize the hyperparameters of LSTM and Random Forest. The parameter search range includes learning rate [0.001, 0.1], number of hidden layer nodes [64, 256], and number of decision trees [50, 200].
[0059] b. Use the SMOTE algorithm to oversample minority class samples and balance the data set distribution;
[0060] c. Use samples in the newly added verified_label_samples tag library for incremental learning. verified_label_samples is updated once a month, and the sample size is no less than 1,000.
[0061] Cross-validation was used to evaluate the model's generalization performance, and a confusion matrix was used to analyze the characteristics of misclassified samples. When the prediction accuracy for any disease type fell below 85%, retraining strategies such as hyperparameter optimization, sample rebalancing, and incremental learning were triggered to ensure that algorithm performance improved with actual application.
[0062] Preferably, the warning information push module includes an information classification unit, a multi-channel push unit and a user feedback unit; wherein the information classification unit divides warning information into three levels according to the degree of urgency: red, orange and yellow. Red warnings must be handled within 5 minutes, orange warnings within 30 minutes, and yellow warnings within 2 hours; the multi-channel push unit supports four notification methods: SMS, APP push, email and voice call, ensuring that the warning information delivery rate reaches 99.9%; the user feedback unit records user processing results and satisfaction evaluation for continuous system improvement;
[0063] The prevention and control plan generation module includes a plan knowledge base, an intelligent decision-making unit, and a plan update unit. The plan knowledge base stores no less than 1,000 standard prevention and control plans, covering 20 common aquaculture diseases. The intelligent decision-making unit uses a decision tree algorithm to automatically match the optimal prevention and control plan based on current environmental parameters and disease type, and optimizes plan parameters according to the scale of aquaculture. The plan update unit evaluates the prevention and control effects monthly and updates the plan library based on the evaluation results.
[0064] The prevention and control plan generation module, based on a knowledge graph, stores a library of standard prevention and control plans. This knowledge base is organized as "disease-plan-effect" triplets, managed using a graph database, and supports complex semantic search. Upon receiving a disease alert, the system matches the three most relevant candidate plans from the library and provides detailed guidance from multiple perspectives, including drug ratios, environmental control, and feeding management. Utilizing algorithms such as decision trees and hierarchical analysis, the system dynamically optimizes plan parameters based on the current aquaculture environment and scale, improving operability.
[0065] The application service layer interacts with the processing platform layer for data through the REST API interface. The interface response time is less than 100ms, and a load balancing mechanism is provided to support no less than 1,000 concurrent users.
[0066] Preferably, the solution knowledge base adopts a graph database storage structure, which specifically includes disease nodes storing disease characteristics, disease occurrence patterns, and degree of harm; specific measures include but are not limited to prevention and control solution nodes storing drug use, environmental regulation, and biological control; the connection between contents includes relationship edges storing solution application conditions and historical application effects; the knowledge base adopts knowledge graph technology, and the knowledge graph is implemented through Neo4j, supporting complex semantic queries and associative reasoning;
[0067] The intelligent decision-making unit generates a prevention and control plan through the following steps:
[0068] a. Use the C4.5 decision tree algorithm to build a solution selection model, set the decision tree depth to 5, and the minimum number of sample splits to 50;
[0069] b. Based on the current environmental parameters, farming scale and disease type, retrieve the top three candidate solutions with the highest matching degree from the knowledge base;
[0070] c. Use the analytic hierarchy process to calculate the comprehensive score of each plan, taking into account the three dimensions of prevention and control effectiveness, economic cost, and operational difficulty;
[0071] d. Select the solution with the highest score and optimize key parameters through genetic algorithm;
[0072] The implementation process of the program update unit is as follows: monthly feedback on the application of the prevention and control program is collected, with indicators including but not limited to cure rate, recovery time and recurrence rate; the fuzzy comprehensive evaluation method is used to quantitatively evaluate the effectiveness of the program, and the weights of the evaluation indicators are determined by the expert scoring method; when the program score is lower than 0.8, the optimization process is triggered, and the program is improved through case reasoning technology; the optimized program must undergo more than 3 experimental verifications, and after verification, it is updated to the program knowledge base.
[0073] As the system is applied in real-world settings, data on prevention and control effectiveness continues to accumulate. The solution update unit evaluates target farms monthly on metrics such as cure rates and recovery times. For underperforming solutions (e.g., with a score <0.8), a case-based reasoning-based improvement process is automatically initiated. Optimized solutions undergo repeated experimental verification before they are officially released, ensuring continuous iteration of the knowledge base.
[0074] The collaborative working mechanism of the intelligent disease early warning system for aquaculture environment is:
[0075] Distributed message queues are used between layers for data transmission. Apache Kafka is used as the message middleware, and more than three broker nodes are configured to ensure high availability. Message retention is set to 7 days. The data collection layer encapsulates monitoring data in JSON format and pushes it to the message queue. The data transmission layer subscribes to the messages, pre-processes them, and then forwards them to the processing platform layer.
[0076] The system's fault-tolerance mechanisms include: deploying a Zookeeper cluster for service registration and discovery, automatically switching services when a node fails; using a master-slave hot backup mechanism to synchronize data between the backup server and the master server; setting a scheduled data snapshot with an interval of one hour, and retaining the last seven days of backups;
[0077] The system's security measures include: using SSL / TLS encryption protocols to protect data transmission security; implementing a JWT-based identity authentication mechanism; setting firewall rules to restrict external access; storing sensitive data with AES-256 encryption; and establishing an operation log audit mechanism to record all key operations.
[0078] The system's scalability design includes: adopting a microservice architecture to decouple various functional modules; using Docker container technology to achieve service encapsulation; orchestrating and scheduling containers through Kubernetes; supporting horizontal expansion and dynamically adjusting the number of service instances based on load conditions; and reserving standardized interfaces to support the access of new sensors and algorithms.
[0079] The present invention also provides a computer-readable storage medium having computer program instructions stored thereon; when the computer program instructions are executed by a processor, the above-mentioned intelligent disease early warning system for aquaculture environment is implemented.
[0080] The present invention uses the above-mentioned intelligent disease early warning system for aquaculture environment and storage medium, and the beneficial effects are as follows:
[0081] (1) The present invention is an intelligent disease early warning system for aquaculture environments based on the Internet of Things and artificial intelligence technologies, namely, the development of an intelligent and precise aquaculture environment monitoring and disease early warning system; this system in the present invention is intended to comprehensively improve the efficiency and accuracy of aquaculture management. The system will use multi-parameter sensors to dynamically monitor various elements in the aquaculture environment. This means that key indicators such as water quality, temperature, and dissolved oxygen content will be tracked in real time, providing farmers with comprehensive environmental data. By constructing a heterogeneous sensor network in which the environment, water quality, and equipment are interconnected, the present invention achieves comprehensive perception of the aquaculture environment. This network can integrate data from multiple modalities, providing a rich and solid foundation for subsequent intelligent analysis.
[0082] (2) In terms of data analysis, the system of the present invention will use intelligent algorithms such as machine learning to conduct in-depth mining of the collected data; through these algorithms, the system can accurately predict the risk of epidemics and generate early warning information that can be interpreted and used to guide prevention and control; this will enable farmers to take measures in advance and effectively avoid the outbreak of epidemics.
[0083] (3) The system of the present invention uses mobile Internet technology to realize the automatic collection and uploading of monitoring data and the timely push of early warning information; no matter where the breeder is, he can understand the breeding environment conditions in real time through mobile devices such as mobile phones and obtain early warning information in a timely manner.
[0084] (4) The system of the present invention also has the ability of self-learning and self-optimization; in the machine learning pipeline, the present invention incorporates functions such as model evaluation, hyperparameter optimization, and incremental learning, forming a self-improving and self-evolving closed-loop training mechanism. This mechanism enables the accuracy of the early warning model to continue to improve with the accumulation of application practice, thereby ensuring the accuracy and reliability of the early warning and providing more reliable support for farmers.
[0085] (5) In terms of modeling, this paper innovatively proposes a collaborative modeling paradigm of "environmental factor association analysis + machine learning prediction." By quantitatively characterizing the interactions between environmental factors, the constructed early warning model is not only environmentally adaptable but also pathologically interpretable, thereby enhancing the model's practicality and credibility.
[0086] (6) At the application level, the present invention adopts the strategy of "information classification + multi-channel access" to maximize the early warning efficiency; at the same time, by exposing cloud service capabilities through standardized interfaces, third-party systems can be easily integrated and secondary developed, further expanding the application scope of the present invention.
[0087] (7) The present invention follows a microservices architecture design, with loose coupling and strong cohesion between components, and has elastic expansion, fault isolation, and hot-swappable capabilities. Combined with containerized deployment, the system can be quickly replicated and applied on a large scale in different aquaculture scenarios.
[0088] (8) This invention addresses the critical need for aquaculture disease prevention and control, providing a systematic solution across all aspects, from monitoring, analysis, early warning, to decision-making. Through data-driven intelligent technology, this invention addresses the many shortcomings of traditional empirical judgments, providing a new approach and powerful impetus for ensuring aquaculture safety and promoting high-quality development of the fishery industry.
[0089] The technical solution of the present invention is further described in detail below through the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0090] Figure 1 Schematic diagram of the overall architecture of the intelligent disease early warning system for aquaculture environment and storage medium embodiment system of the present invention;
[0091] Figure 2 Schematic diagram of hardware deployment of the intelligent disease early warning system for aquaculture environment and storage medium embodiment system of the present invention;
[0092] Figure 3 Schematic diagram of the software architecture of the intelligent disease early warning system for aquaculture environment and storage medium embodiment system of the present invention;
[0093] Figure 4 A training flow chart of a machine learning prediction model for an intelligent disease early warning system for aquaculture environments and a storage medium embodiment of the present invention;
[0094] Figure 5 This is a principle block diagram of an environmental factor correlation analysis model for an intelligent disease early warning system for aquaculture environments and a storage medium embodiment of the present invention;
[0095] Figure 6 This is a flowchart of the application service layer warning information push of the intelligent disease early warning system for aquaculture environment and storage medium embodiment of the present invention;
[0096] Figure 7 Generate a flow chart for the application service layer prevention and control solution of the intelligent disease early warning system for aquaculture environment and storage medium embodiment of the present invention;
[0097] Figure 8 Schematic diagram of a typical application scenario of the intelligent disease early warning system for aquaculture environment and storage medium embodiment system of the present invention. DETAILED DESCRIPTION
[0098] The technical solution of the present invention is further described below with reference to the accompanying drawings and embodiments.
[0099] Unless otherwise defined, technical or scientific terms used in the present invention shall have the same meaning as commonly understood by one of ordinary skill in the art to which the present invention belongs.
[0100] Example
[0101] like Figure 1-Figure 3 As shown in FIG, the intelligent disease early warning system for aquaculture environment includes a data acquisition layer, a data transmission layer, a processing platform layer and an application service layer.
[0102] The data collection layer includes environmental monitoring units, water quality sensor groups and portable detection equipment. The main function of this layer is to achieve full-factor perception of the aquaculture environment.
[0103] The environmental monitoring unit includes a temperature sensor, a humidity sensor, and a light sensor, located above the aquaculture water surface. In this embodiment, the temperature and humidity sensors use a high-precision digital DHT22. The temperature sensor has a measurement accuracy of ±0.1°C and a range of -10°C to 50°C. The humidity sensor has a measurement accuracy of ±2%RH and a range of 0 to 100%RH. The light sensor uses a photoresistor, with a measurement accuracy of ±3% and a range of 0 to 200,000 Lux.
[0104] Each sensor is connected to the data logger via an RS485 bus, powered by the bus. The 485 interface has terminal resistors, allowing for a communication distance of up to 1,200 meters. The sampling cycle is one minute, and the sensors are arranged in a triangle, spaced 5-10 meters apart, depending on the size of the aquaculture water surface.
[0105] The water quality sensor set, submerged 0.5 meters underwater, includes a pH sensor, a dissolved oxygen sensor, and an ammonia nitrogen sensor. The pH electrode in this example uses the E-201-C composite glass electrode, with a measurement range of 0 to 14 pH and an accuracy of ±0.01. The dissolved oxygen electrode uses a fluorescence method, model JPBJ-608, with a measurement range of 0 to 20 mg / L and an accuracy of ±0.05 mg / L. The ammonia nitrogen electrode is an electrochemical sensor, model NH3-SE-G006, with a measurement range of 0.1 to 1000 mg / L.
[0106] The electrode signals are extracted via waterproof cables, processed by a signal conditioning circuit, and then connected to a data logger. To ensure representative data, the sensors are evenly spaced in a plum blossom pattern along the diagonal of the pond, with spacing of 1-2 meters to avoid blind spots and pollution sources.
[0107] Portable testing equipment is a handheld multi-parameter water quality analyzer. As a supplement to fixed sensors, it is carried by staff and allows for detailed local testing at suspected abnormalities. It features over 10 functional probes, including COD, turbidity, and heavy metal ion detection. It also has built-in data storage and wireless transmission units, enabling flexible and maneuverable operation.
[0108] The technical specifications of each probe in this example are as follows: turbidity 0-1000 NTU, accuracy ±2%. COD 0-1000 mg / L, accuracy ±5%. Heavy metal ion 0.001-10 mg / L, accuracy better than ±1%. The instrument has an IP67 protection rating and can operate continuously for at least 8 hours.
[0109] The data transmission layer includes wireless transmission modules and edge computing nodes. This layer is responsible for achieving reliable backhaul of monitoring data from the breeding site to the cloud.
[0110] In this embodiment, the wireless transmission module uses dual-mode 5G / 4G and WiFi6 communication, providing mutual backup. The 5G module supports both SA and NSA networking, with uplink rates exceeding 500Mbps. The WiFi6 module integrates MIMO technology, with access rates exceeding 1Gbps. The antenna gain is 10dBi, with a coverage radius of 1km. The communication protocol uses MQTT, with a payload of less than 256 bytes, low overhead, and strong real-time performance. Transmission latency is less than 10ms, and packet loss rate is less than 0.1%. All wireless links utilize AES-256 encryption to ensure data security.
[0111] In this embodiment, the edge computing node acts as the link between cloud and edge collaboration, performing preliminary processing at the source of data. The hardware is a high-performance industrial computer equipped with an 8-core 2.5GHz industrial-grade CPU, 16GB ECC memory, a 512GB SSD, and a Gigabit Ethernet card.
[0112] The operating system uses Linux tailored for IoT scenarios, with low-latency kernel mode enabled. Nodes are pre-installed with common algorithm modules such as data cleaning, compression, and encryption, and are containerized for plug-and-play operation.
[0113] When offline, edge nodes can utilize local caching to store over 72 hours of monitoring data. Once communication is restored, an incremental upload mechanism is triggered to ensure data integrity. The node management platform provides remote monitoring and configuration updates, enabling self-healing and one-click upgrades.
[0114] The processing platform layer serves as the brain of the system in this embodiment, including a machine learning prediction model and an environmental factor correlation analysis model. This layer integrates AI capabilities such as environmental factor correlation analysis and machine learning prediction to achieve intelligent epidemic warning.
[0115] The purpose of the environmental factor association analysis model is to identify key combinations of factors that trigger disease from complex environmental parameters, providing interpretable features to support predictive models. The model is based on association rule mining and incorporates techniques such as causal inference and graphical modeling to strengthen the causal and stable nature of the results.
[0116] First, data exploration is conducted on each monitoring indicator, constant features are removed, and continuous variables are discretized into bins. The binning process uses maximum entropy segmentation, which can adaptively adjust the bin width based on the data distribution, overcoming the subjective shortcomings of traditional equal-width binning.
[0117] On this basis, we used the improved Apriori algorithm and FP-Growth algorithm to mine frequent itemsets. The support, confidence, and lift thresholds were set to 0.01, 0.6, and 2, respectively. We also conducted a statistical test of the significance of the rules, eliminating spurious associations with P > 0.05.
[0118] Considering the lag effects of environmental factors, the model uses a sliding window approach to calculate support, incorporating the time dimension to extract dynamic associations among variables. The window length is 24 hours, with a sliding step of 1 hour, reflecting the cumulative effects on a diurnal scale.
[0119] After identifying strong association rules, the model further constructs a directed acyclic graph (DAG) to reveal inter-factor dependencies. It also introduces the d-separation criterion to eliminate confounding effects and purify the core driving factors. Ultimately, an "environmental parameter network" is formed, capturing the full picture of factor interactions.
[0120] Finally, the model calculates the importance weights of each environmental factor based on network topology and association strength. This weighting takes into account both data-driven and expert knowledge. First, conditional entropy is used to determine the information gain of the variables, which serves as objective weights. Domain experts then subjectively score key factors, forming a hierarchical analysis model. This weighted fusion of these two factors yields an environmental risk profile that is both a priori and data-driven.
[0121] The entire process is encapsulated as microservices and deeply integrated with the data warehouse system. Batch processing frequency can be flexibly configured, such as updating every 12 hours. The model pipeline is built using the bubble framework, enabling visual monitoring and performance profiling throughout the entire process.
[0122] Building on correlation analysis, the machine learning prediction model further leverages machine learning techniques to quantitatively predict the probability of aquatic disease occurrence within a specific time window. This model uses multidimensional time series data as input, selects a suitable deep learning framework as the main network, and integrates it with other models to build an integrated prediction system.
[0123] During the data preprocessing phase, adaptive smoothing strategies are employed to address different types of noise. KNN interpolation is used for occasional missing values, while matrix completion algorithms are used for continuous missing values. Unsupervised models such as One-Class SVM are used for outlier detection to overcome the subjectivity of threshold selection.
[0124] During feature engineering, in addition to the risk factor combinations output by correlation analysis, we also incorporate higher-order features such as temporal statistics (such as mean and variance) and wavelet transform coefficients. We also perform autocorrelation tests on each feature to eliminate redundant information. The number of selected features is limited to 20 to avoid the curse of dimensionality.
[0125] Considering the long-range dependencies of aquaculture time series, the forecasting model uses a long short-term memory (LSTM) network as its backbone. Environmental parameters are used as LSTM inputs, with each time step corresponding to one hour and 72 steps expanded to reflect the cumulative effects over a three-day timeframe.
[0126] The LSTM model is stacked horizontally in 2-3 layers, with 128 memory cells per layer. These layers are connected using an attention mechanism to highlight the impact of key time points. Five-fold cross-validation is used to adjust the number of layers to achieve a bias-variance balance. L1 and L2 regularization terms of 0.001-0.01 are also introduced to control overfitting.
[0127] While LSTM is highly capable of capturing time series features, it lacks the ability to model feature interactions. To address this deficiency, the model incorporates tree models such as random forests at the second level and integrates them with the LSTM results through stacking. The tree model hyperparameter tuning range is: number of trees 50-200, maximum depth 5-20, and minimum number of node samples 2-10.
[0128] During training, the batch size was set to 72, and the memory usage was less than 10GB. The loss function used cross-entropy, while also incorporating classification metrics such as sensitivity and precision to guide the model's focus on edge cases. Training was automatically terminated if the test set accuracy did not improve for five consecutive epochs to prevent hallucination fitting.
[0129] The model management platform provides self-service services, enabling one-click initiation of model training, evaluation, and release. Version management seamlessly integrates with the container environment, ensuring traceability of iterations. The inference API utilizes a RESTful interface, achieving a single-call latency of less than 50ms and supporting loads exceeding 500 queries per second.
[0130] like Figure 5As shown in the figure, a high-density sensor network covering 10 indicators, including meteorological and water quality, has been deployed in a smart shrimp pond. The sensors collect and transmit data every 10 minutes, accumulating tens of thousands of time series records. The task is to analyze the correlation between different combinations of environmental factors and shrimp disease occurrence and predict the probability of disease risk within the next hour.
[0131] (1) Data preprocessing:
[0132] Raw data contains noise and anomalies, such as outliers caused by probe failures and missing data caused by network delays. The system uses methods such as thresholding and the 3-Sigma principle to adaptively identify different types of anomalies.
[0133] For occasional random missing values, we used the nearest neighbor (KNN) interpolation. Taking into account the spatiotemporal autocorrelation of environmental data, we considered data at adjacent time steps and spatial locations as "neighbors," and used DTW dynamic time warping as the distance metric. For continuous missing values, we used a matrix completion algorithm instead, leveraging the potential relationships of other complete variables to estimate the distribution of missing values.
[0134] Outlier processing is a two-step process. First, a one-class support vector machine (SVM) is used to construct an anomaly detector. This detector fits a tight boundary around the normal data and identifies the few points that deviate from this boundary as suspected anomalies. Further identification is then performed based on the physical context. For example, a sudden rise in dissolved oxygen may be due to equipment drift and requires manual review.
[0135] (2) Association rule mining:
[0136] Treating the multiple indicators in each monitoring record as a "basket of items," and the range of values for each indicator as a different "item," this creates an association rule mining problem. The goal is to identify strong correlation patterns between the combined characteristics of environmental factors and the occurrence of shrimp diseases.
[0137] First, the numerical range of each indicator is binned and mapped to discrete attributes. This binning method uses the maximum entropy principle, selecting the split point with the highest information entropy gain. Continuous variables, such as water temperature, are categorized into three levels: high, medium, and low. Discrete variables, such as rainfall, are categorized into two categories: present and absent.
[0138] Next, we use the improved Apriori algorithm to find frequent itemsets. Traditional Apriori uses a layer-by-layer iterative algorithm to generate candidate sets, which is computationally expensive when the number of items is large. This system instead uses the FP-Growth algorithm, which directly finds frequent items using the FP-Tree data structure. This algorithm is more efficient for high-dimensional sparse itemsets. The support, confidence, and lift thresholds are set to 0.01, 0.6, and 2, respectively.
[0139] Taking into account the temporal nature of environmental factors, a sliding time window is incorporated into the generation of frequent items to extract patterns of combinations across the factor's historical intervals. For example, when exploring the relationship between hypoxia and disease onset, not only the current moment is considered, but also the dissolved oxygen sequence from the previous 1 to 24 hours. If all values are low, it indicates persistent hypoxia and a higher risk. The window span is set to 24 hours, with a sliding step of 1 hour, to reflect the cumulative effects of the diurnal time scale.
[0140] Finally, strong association rules were selected from the frequent item set using the Lift metric. A chi-square test was used to evaluate significance, and false rules with a P value greater than 0.05 were removed. For example, if the rule "High temperature (>32°C) and low oxygen (<3mg / L) for three consecutive days increases the probability of shrimp disease by 5 times" was found with a confidence level of 0.8 and a lift of 10, it would pass the significance test and be included in the association rule library.
[0141] (3) Verification of causal relationship:
[0142] Frequently co-occurring environmental combinations and shrimp diseases are correlated, but not necessarily causal. Many associations may be spurious. For example, while rising air and water temperatures both contribute to disease, they share a common cause: increased solar radiation. Therefore, correlation is not causation.
[0143] Therefore, the system further employs the g-formula to verify the causal pathway from environmental combinations to shrimp diseases through reverse experiments. The basic principle is to first learn the time-dependent structure (CBN) of environmental factors, and then use do-calculus operations to calculate the changing trend of disease risk when a factor takes a specific value.
[0144] Take temperature and bacterial diseases as an example. The system fits a temporal causal model between the two from the data: for every 1°C increase in temperature, the incidence rate increases by 1%. Then, assuming temperatures of 15°C, 20°C, and so on, and plugging them into the causal model, it finds that the risk of disease increases with rising temperatures. This confirms the causal relationship between temperature and disease.
[0145] Through causal testing, spurious correlations are filtered out from massive association rules, identifying the key drivers most likely to trigger the disease. These factors provide highly targeted input features for subsequent predictive models.
[0146] (4) Build a machine learning prediction model:
[0147] Using the environmental factor-shrimp disease causal chain, a nonlinear mapping from monitoring indicators to future shrimp disease risks is constructed to form a disease early warning model. This system uses a time series classification model, and its main modules are as follows: Figure 4 shown.
[0148] During the feature engineering phase, based on the aforementioned correlation factors, statistical features (mean, variance, kurtosis, etc.) and frequency domain features (Fourier transform) were derived to enhance information representation. Furthermore, to overcome the imbalance problem (shrimp disease samples are far fewer than normal samples), SMOTE oversampling was used to increase the proportion of small-category samples.
[0149] Considering the long time lag between environmental factors and disease onset, the prediction model uses an LSTM network to learn temporal dependencies. Hourly monitoring data from the past three days is fed into the network to predict the probability of disease onset within the next hour. Weighted cross-entropy was used as the loss factor, and the hyperparameter search space was: 1-3 LSTM layers, 16-256 units, dropout rate 0-0.5, and learning rate 1e-5-1e-2. Due to the small sample size, 5-fold cross-validation was used.
[0150] To fully utilize features beyond time series, we built a stacking ensemble of LSTM and Random Forest models. First, we fed the hidden state of the LSTM at the last time step (the semantic representation of the environment) into the RF. Then, we took the weighted average of the LSTM output (probability of disease onset) and the RF output (complementary information). We optimized the ensemble weights using grid search with a step size of 0.1.
[0151] In a model comparison experiment, six common classifiers, including Naive Bayes (NB), Support Vector Machine (SVM), LSTM, CNN, RF, and XGBoost, were fitted to the training set and their hyperparameters were uniformly optimized to obtain performance upper bounds. The results showed that the LSTM+RF ensemble model significantly outperformed the other single models in terms of AUC (0.91), F1 (0.85), and FPR (0.05), approaching or even exceeding the level of human expert judgment.
[0152] In terms of interpretability, the system uses common methods such as LIME to analyze the feature dimensions that the model focuses on. It found that temperature, dissolved oxygen, and pH ranked among the top three weighted factors. Furthermore, the Integrated Gradients method was used to reveal the nonlinear relationship between each feature and the predicted probability. For example, when the temperature rises to 28 degrees Celsius, the risk increases sharply. When the dissolved oxygen drops below 3 mg / L, the hazard increases sharply. This is highly consistent with expert experience.
[0153] The model is implemented using TensorFlow 2.0, leveraging the Keras high-level API for rapid iteration. The inference service is packaged in the ONNX format, wrapped as a RESTful API via Flask, and uses NVIDIA TensorRT for device-side acceleration, achieving a single call latency of less than 10ms.
[0154] The application service layer includes an early warning information push module and a prevention and control plan generation module. This layer is directly aimed at farmers and provides simple and easy-to-use intelligent early warning decision-making services.
[0155] The early warning information push module is responsible for obtaining the epidemic warning signals of the processing platform in real time and reaching users through multiple channels such as APP notifications, text messages, voice calls, etc. The push process is as follows: Figure 6 shown.
[0156] When the epidemic risk probability generated by the processing platform per minute exceeds the warning threshold (such as 0.8), and the correlation analysis shows that the current combination of environmental factors matches the high-risk pattern, it is judged to be a valid warning. At this time, the message is processed by the classification module and divided into three levels according to the degree of risk and urgency: Level I (red), Level II (orange), and Level III (yellow), with corresponding risk thresholds of 0.95, 0.9, and 0.8 respectively.
[0157] Level I alerts require action within 5 minutes, Level II alerts within 30 minutes, and Level III alerts require follow-up within 2 hours. Alert information is automatically pushed to the responsible person via app push (delay < 3 seconds), SMS (success rate > 95%), and voice call (connection rate > 90%). If the user does not read the message within 10 minutes, the system will automatically resend it until manually confirmed or the maximum number of retries (5) is reached.
[0158] Early warning information is presented using the "Six Whys" principle, concisely informing farmers of what has occurred (what), the possible causes (why), the necessary measures (how), the timeframe for completion (when), who is in charge (who), and the progress of the incident (progress). Data visualization tools such as trend charts are also used to enhance the credibility and operability of early warnings.
[0159] The system also provides user feedback channels. It not only supports manual feedback but also integrates environmental monitoring data as supporting evidence. If objective evidence, such as the disappearance of abnormal environmental indicators or the elimination of equipment faults, is provided, the system will automatically disable or lower the warning level, streamlining manual operations to the greatest extent possible.
[0160] The underlying service utilizes high-availability message queues like RabbitMQ to ensure reliable message delivery. Elastic scaling is achieved through a microservices registry, enabling cluster throughput of up to 100,000 requests per second. Service development adheres to the Open API specification, offering self-describing interfaces, convenient invocation, and support for further development.
[0161] In addition to providing timely warnings, the prevention and control plan generation module should also guide users to take targeted measures to curb the spread of the epidemic. Figure 7 As shown in the figure, the system has designed an intelligent recommendation engine for prevention and control plans, which runs in coordination with the early warning service.
[0162] The recommendation engine, centered around a knowledge graph, accumulates structured disease prevention and control knowledge. The knowledge base is represented by triples of "disease-solution-effect," and uses Neo4j as its backend, naturally supporting graph queries. In this example, the system stores over 1,000 prevention and control solutions, covering 20 common aquaculture diseases, including bacterial, viral, and parasitic diseases.
[0163] When an epidemic alert is received, the system first compares the disease pattern and, in conjunction with context such as the species and breeding cycle, searches the knowledge base to initially screen out five candidate solutions. Each solution is then comprehensively evaluated based on its effectiveness, cost, and implementation difficulty. Effectiveness focuses on the reduction in morbidity and mortality, referencing local historical and regional averages. Costs are calculated by calculating the input-output ratio of factors such as drugs and labor. Implementation difficulty takes into account process complexity and required personnel skills. Decision optimization algorithms such as hierarchical analysis are used to automatically select the best solution based on multiple criteria.
[0164] Next, the system deeply customizes the optimal solution to suit the current farm situation. First, it leverages real-time data from the farming environment through simulation and reinforcement learning to dynamically optimize key process parameters such as drug dosage and timing. Second, it personalizes process steps and operating instructions based on farm staffing and user preferences.
[0165] Customized plans are distributed to edge nodes via a remote configuration center, scheduling farming equipment for execution. The execution process can be visually tracked through video surveillance and personnel trajectories to ensure plan implementation.
[0166] At the same time, the system continuously tracks the effectiveness of the plan implementation. Using feedback data such as environmental restoration, reduced disease losses, and improved aquaculture efficiency, it conducts cost-benefit analysis and quantitatively assesses the plan's universality through causal inference and other methods. The results of this assessment are promptly fed back into the knowledge base to refine the plan.
[0167] The knowledge operations team can also manage the updating, editing, reviewing, and publishing of the knowledge graph through a visual workbench. Based on hot trends in aquaculture issues, they can proactively collect the experience of aquaculture experts to enrich the breadth and depth of their solutions.
[0168] The solution recommendation service is developed using Java Stack, following domain-driven design to decouple business logic from technical implementation. An SDK package is also provided for user-friendly deployment. JMeter-tested service performance shows an average response time of less than 1 second, meeting the requirements of interactive intelligent decision-making.
[0169] To ensure the coordination of the aforementioned "monitoring-analysis-warning-decision-making" processes, the system in this embodiment adopts the following designs in terms of communication, storage, computing, and security:
[0170] (1) Communication mechanism:
[0171] Since it involves a large number of heterogeneous data sources and different functional modules, the system adopts a service-oriented architecture (SOA) to achieve component decoupling and standard interoperability.
[0172] In terms of technology selection, lightweight IoT protocols such as MQTT are used between the data collection and transport layers to reduce bandwidth usage. Apache Kafka is used to build the edge-to-cloud data channel, enabling flexible scalability. High-performance RPC frameworks such as gRPC are used between microservices in the cloud to ensure millisecond-level call execution.
[0173] Regarding communication security, the system enforces that all API calls must go through the API Gateway, implementing unified authentication, authorization, flow control, and metering and billing. HTTPS is enabled by default for external interfaces and supports national encryption algorithms. Mutual authentication (mTLS) is enabled for internal inter-service communication. Furthermore, to ensure physical isolation, the system has dedicated virtual private clouds for core components, with VPNs and bastion hosts used to manage access.
[0174] (2) Storage mechanism:
[0175] To balance the real-time nature of data aggregation and offline batch processing of analysis, the system adopts the Lambda architecture and builds a data lake with both a speed layer and a batch layer.
[0176] The speed layer is built on Apache Druid. Real-time data from the edge is imported into Druid via Kafka. Using core mechanisms such as columnar storage and inverted indexing, it supports petabyte-scale OLAP queries in milliseconds. This provides a solid foundation for services such as real-time early warning and anomaly detection.
[0177] The batch layer uses the Hadoop system, using tools like Sqoop to regularly import incremental data from Kafka into HDFS daily. Based on this, Hive, Spark, and other big data processing components are used to conduct periodic full-data correlation analysis and risk factor mining.
[0178] The Speed Layer and Batch Layer are transparent to business systems. The system utilizes integrated analytical engines like Presto to mask heterogeneous storage differences and implement unified SQL calculations. OLAP middleware like Kylin is used to build subject-oriented multidimensional data marts, further simplifying calculation views.
[0179] In addition, to address unexpected situations like device failures, the data lake also has a built-in disaster recovery mechanism. The system leverages HDFS's inherent multi-replication strategy to ensure triple-copy fault tolerance for batch data. For real-time data, regular snapshots (e.g., every five minutes) are taken, and cold backups are provided using distributed block storage services like Alibaba Cloud ESSD, achieving an RPO of less than five minutes.
[0180] (3) Resource scheduling and task scheduling:
[0181] In order to make full use of heterogeneous computing resources, the system adopts cloud-native technology stacks such as Kubernetes (k8s) to achieve elastic resource scheduling, task flow orchestration, etc.
[0182] AI workloads such as training and inference are managed using GPU clusters, with GPU devices managed using the Kubernetes Device Plugin mechanism. By extending the Kubernetes Default Scheduler, affinity matching between PODs and GPU card types is achieved, with support for preemptive scheduling. The entire cluster uses virtual nodes (Virtual Kubelet) to connect to Alibaba Cloud Elastic Container Instances (ECI), breaking through the limitations of physical cluster scale.
[0183] For offline batch processing and workflow orchestration, the system uses Argo Workflow. Based on Kubernetes Custom Resource Definitions (CRDs), it defines DAG pipelines as native Kubernetes resources. At the scheduling level, it integrates with developed schedulers such as Volcano and Yunikorn, ensuring dedicated resources for tasks while enabling finer-grained priority and fairness control.
[0184] In addition, the system fully leverages the Prometheus monitoring system and the Istio service mesh to achieve full-stack observability. Monitoring indicators cover four dimensions: system, cluster, service, and instance, providing a digital foundation for fault diagnosis, capacity planning, and cost optimization.
[0185] (4) Data privacy protection:
[0186] Because the system involves sensitive data such as aquaculture production, it prioritizes data privacy and security, and is designed in strict compliance with Level 3 security protection requirements. SSL / TLS encryption is mandatory for transmission. Externally exposed APIs require callers to obtain OAuth 2.0 authorization. For sensitive data, homomorphic encryption methods such as Fully Homomorphic Encryption (FullyHomomorphic Encryption) ensure agnosticism throughout the entire transmission, storage, and computation process. Furthermore, each piece of data is immutably watermarked and timestamped upon generation, ensuring data traceability.
[0187] At the storage layer, all petabyte-level production data in the system is encrypted at rest. Keys are managed by a dedicated KMS system, and authorization adheres strictly to the principle of least privilege. Access control is implemented down to the field level, using RBAC. Log auditing covers the entire DDL / DML / DCL operation chain, providing instant alerts for suspicious behavior.
[0188] Furthermore, in response to new regulations such as the Cybersecurity Law, the system is actively embracing privacy-preserving computing. By building privacy-preserving computing frameworks such as Multi-Party Computing (MPC) and Federated Learning (FL) on Kubernetes, the system ensures that all participants' original data remains locally accessible while sharing data modeling and analysis, mitigating data aggregation risks and providing a technical foundation for cross-organizational collaboration.
[0189] The system in this embodiment is applied to a shrimp breeding base. The breeding base has a pond area of 500 mu, an annual output of 200 tons of whiteleg shrimp, and an output value of 20 million yuan, which is a certain scale.
[0190] like Figure 8 As shown in the figure, the technical architecture of the system in this scenario is as follows:
[0191] Data Perception Layer: Three types of sensors are deployed to cover key environmental factors, such as pond atmosphere and water quality. Meteorological sensors monitor temperature, humidity, and light intensity, and are located at the four corners of the ponds, 2 meters above the water surface, with a sampling period of 5 minutes. Water quality sensors monitor nine parameters, including dissolved oxygen, pH, and ammonia nitrogen. Four sensors are placed diagonally in each pond, fixed 0.5 meters underwater, with a sampling period of 10 minutes. A portable multi-parameter water quality meter primarily monitors COD, heavy metals, and other parameters. Handheld by workers, it is used four times daily.
[0192] Data backhaul utilizes NB-IoT, with a signal sector radius of 5 kilometers, power consumption less than 500mW, and a battery life of one year. The gateway device is Alibaba Cloud IoT Edge, equipped with edge capabilities such as MQTT messaging and function computing, enabling local access and real-time cleaning. The cloud pipeline utilizes the open-source Apache StreamSets framework, supporting distributed stream processing.
[0193] The data center uses Alibaba Cloud DataWorks to build an enterprise-level data bus, connecting the entire chain from data integration, storage, computing, quality, security, and services. Real-time data is imported into MaxCompute online, suitable for monitoring scenarios with high timeliness requirements. Offline data is imported into MaxCompute offline daily on T+1 for periodic large-scale modeling and analysis. The computing engine, primarily based on PAI, provides a visual workbench for the entire process, from data preprocessing to model development, evaluation, and deployment.
[0194] The application layer focuses on two key scenarios: intelligent early warning and intelligent decision-making. The early warning service, based on a correlation analysis model, identifies high-risk environmental combinations for shrimp diseases. For example, three consecutive days of high temperatures (>32°C) combined with low dissolved oxygen (<4 mg / L) can increase the incidence of shrimp diseases fivefold. If real-time environmental data matches this pattern, a Level 1 (red) alert is immediately issued, notifying farmers to take swift countermeasures.
[0195] The decision-making service is built on a massive, structured knowledge base of disease prevention and control. The knowledge base contains over 500,000 entries, covering over 50 common shrimp diseases and detailing over 300 specific prevention and control measures. The system intelligently matches disease characteristics and provides treatment plans, including medication, feeding, and disinfection. It also instantiates parameters (such as medication dosage and stocking density) based on the scale of the aquaculture operation. Farmers simply follow the system's instructions with a single click, significantly reducing the decision-making threshold.
[0196] The service uses Angular 10 to build a farming management dashboard, integrating real-time environmental monitoring, disease risk warnings, and recommended treatment plans, with a mobile version available. In the year since the system went live, the environmental data collection rate has reached 99%, alarm response time has been reduced by 86%, and the survival rate of farmed animals has increased by 5 percentage points, generating direct economic value of 3 million yuan.
[0197] Taking vitiligo syndrome as an example, the specific application effect of the system in this embodiment is demonstrated:
[0198] At 2:00 PM on a certain day in June 2020, the dissolved oxygen probe reading in Pond 1 plummeted to 3.2 mg / L, and the pH also dropped to 6.5, persisting for two hours. Systematic association analysis revealed that this combination had previously led to shrimp disease outbreaks. Furthermore, a machine learning model predicted a 92% probability of disease within the next hour.
[0199] Based on the aforementioned triggering conditions, the system automatically generated a Level 1 alert at 2:05 PM and sent it to the pool owner via the aquaculture app, phone, and text message. It also provided recommendations: immediately activate the aerator, arrange for manual water mixing and bottom modification, and apply quicklime throughout the pool every four hours at a dosage of 20g / m2.
[0200] The farm staff followed the advice and recorded their progress in an app. The system then collected real-time environmental data and evaluated the effectiveness of the treatment. By 10 PM that evening, the dissolved oxygen level had returned to above 7 mg / L and the pH had returned to 7.5. The system determined the risk had been resolved and automatically downgraded the warning signal. The entire process took less than 10 hours. Later, statistics showed that this disease prevention and control effort avoided approximately 500,000 yuan in economic losses.
[0201] In implementing the technical solution of this embodiment, those skilled in the art can clearly understand that this embodiment can be implemented by software plus general hardware tuning, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment of this application or certain parts of the embodiments.
[0202] Therefore, the present invention adopts the above-mentioned intelligent disease early warning system and storage medium for aquaculture environment. The system innovatively establishes a correlation model between environmental factors and disease occurrence, and uses machine learning technology to achieve an early warning accuracy of more than 85%. It can realize comprehensive monitoring and real-time early warning of the breeding process, and support accurate and efficient disease prevention and control; thereby solving the problems of low accuracy and delayed response of traditional early warning methods, and significantly improving the effect of aquaculture disease prevention and control.
[0203] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit the same. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that they can still modify or replace the technical solutions of the present invention with equivalents, and these modifications or equivalent replacements cannot cause the modified technical solutions to deviate from the spirit and scope of the technical solutions of the present invention.
Claims
1. Intelligent disease early warning system for aquaculture environment, characterized by: It includes a data acquisition layer, a data transmission layer, a processing platform layer, and an application service layer; the data acquisition layer includes an environmental monitoring unit, a water quality sensor group, a portable detection device, and a data acquisition module; the data transmission layer includes a wireless transmission module and an edge computing node, and the wireless transmission module is connected to the data acquisition layer via a high-speed data line; the processing platform layer includes a machine learning prediction model and an environmental factor correlation analysis model, and the processing platform layer is connected to the data transmission layer via a wireless network; the application service layer includes an early warning information push module and a prevention and control plan generation module, and the application service layer and the processing platform layer are connected via a data interface; The machine learning prediction model uses a combination of deep learning and ensemble learning, and includes a data preprocessing module, a feature extraction module, a model training module, and a prediction output module. The data preprocessing module standardizes and normalizes the collected environmental and water quality data. The feature extraction module uses a convolutional neural network to extract time series features and employs an attention mechanism to highlight the influence of key environmental factors. The model training module combines LSTM and Random Forest algorithms, using historical data for model training and parameter optimization. The training data set contains no less than 10,000 historical records. The prediction output module generates warning signals based on the model prediction results and confidence thresholds. The environmental factor association analysis model is based on the Apriori algorithm and the grey correlation analysis method. The following steps are used to perform the association analysis between environmental factors and disease occurrence: a. Bin the environmental factor data and establish a candidate set of association rules; b. Calculate support and confidence to filter out strong association rules; c. Use grey correlation to calculate the influence of each environmental factor on the occurrence of the disease; d. Generate an environmental factor weight matrix and input it into the machine learning prediction model to improve the accuracy of early warning; The processing platform layer also includes a model evaluation and optimization module, which evaluates the performance of the model through cross-validation and confusion matrix analysis. When the warning accuracy is lower than 85%, the model retraining mechanism is automatically triggered to improve the model performance through parameter tuning and sample expansion.
2. The intelligent disease early warning system for aquaculture environment according to claim 1 is characterized by: The environmental monitoring unit includes a temperature sensor, a humidity sensor, and a light sensor. The temperature sensor has a measurement accuracy of ±0.1°C and a range of -10°C to 50°C. The humidity sensor has a measurement accuracy of ±2%RH and a range of 0 to 100%RH. The light sensor has a measurement accuracy of ±3% and a range of 0 to 200,000 Lux. The water quality sensor group includes a pH sensor, a dissolved oxygen sensor, and an ammonia nitrogen sensor; the pH sensor has a measurement accuracy of ±0.01 and a range of 0 to 14 pH; the dissolved oxygen sensor has a measurement accuracy of ±0.1 mg / L and a range of 0 to 20 mg / L; and the ammonia nitrogen sensor has a measurement accuracy of ±0.1 mg / L and a range of 0 to 100 mg / L. The portable detection device is a handheld multi-parameter water quality analyzer, which has COD, turbidity and heavy metal ion detection functions. The handheld multi-parameter water quality analyzer is connected to the data acquisition layer through a waterproof connector; The sensors in the environmental monitoring unit are evenly installed at an angle of 120° 2 meters above the aquaculture water body, and the water quality sensor group is evenly installed at a depth of 0.5 meters at an angle of 60°. The sensors in the data acquisition layer are connected to the data acquisition module of the data acquisition layer through the RS485 bus, and the sampling frequency is 1 time / minute.
3. The intelligent disease early warning system for aquaculture environment according to claim 1 is characterized by: The wireless transmission module includes a 5G communication unit and a WiFi communication unit. The 5G communication unit supports both SA and NSA modes, with an uplink rate of 1Gbps and a downlink rate of 3Gbps. The WiFi communication unit supports the WiFi6 protocol with a transmission rate of 2.4Gbps. The edge computing node is equipped with a 3.2GHz quad-core processor, 8GB of RAM, and a 256GB solid-state drive. It has data cleaning, outlier processing, and data compression functions, with a data processing capacity of ≥1,000 records per second. The edge computing node has a data caching mechanism that locally stores 72 hours of monitoring data in the event of a network interruption. After the network is restored, the cached data is automatically uploaded to the processing platform layer to achieve breakpoint resumption. The wireless transmission module is connected to the data acquisition layer through a Gigabit Ethernet interface, with a data transmission delay of less than 10ms. Fiber optic communication is used between the edge computing node and the wireless transmission module, with a transmission bandwidth of 10Gbps.
4. The intelligent disease early warning system for aquaculture environment according to claim 1 is characterized by: The specific implementation of the machine learning prediction model is: A. Data preprocessing module processes data through the following steps; a. Use mean filling method to handle missing values and use 3σ criterion to remove outliers; b. Perform Min-Max normalization on continuous data and one-hot encoding on discrete data; c. Divide the dataset into training set and test set in a ratio of 8:2; B. The convolutional neural network of the feature extraction module consists of three convolutional layers, with kernel sizes of 3×3, 5×5, and 7×7, respectively, and a ReLU activation function. The attention mechanism uses a multi-head self-attention structure with 8 heads to capture feature dependencies at different time scales. C. The LSTM network combined with the model training module contains two hidden layers, each with 128 nodes and a time step of 24. The Random Forest model contains 100 decision trees, with a maximum depth of 10 and a minimum number of leaf node samples of 5. The output results of the LSTM and Random Forest models are integrated using the stacking method, using LightGBM as the secondary learner. The prediction output module uses the soft voting method to fuse the model prediction results, sets the confidence threshold to 0.8, and triggers an early warning when the prediction probability exceeds the threshold; at the same time, it calculates the precision, recall rate and F1 value based on the confusion matrix, and feeds the evaluation results back to the model training module for continuous optimization.
5. The intelligent disease early warning system for aquaculture environment according to claim 1 is characterized by: The environmental factor association analysis model adopts a multi-dimensional analysis method, including the following steps: Step S1, Identification of Environmental Factors: Based on the principal component analysis method, key factor data with significant impact on disease occurrence are screened from the temperature, pH value, dissolved oxygen, and ammonia nitrogen monitoring data. The screening criterion is that the cumulative contribution rate is greater than 85%; Step S2, data preprocessing: standardize the screened environmental factor data and use the sliding window method to calculate the temporal variation characteristics of each factor, with a window size of 24 hours and a step size of 1 hour; Step S3, association rule mining: combining the Apriori algorithm and the time series association rule mining method, setting the minimum support to 0.2 and the minimum confidence to 0.6, mining the association rules between the combination of environmental factors and the occurrence of diseases, and verifying the significance of the rules based on the chi-square test; Step S4, factor weight calculation: using a combination of analytic hierarchy process and entropy weight method to establish an environmental factor weight evaluation system, and calculate the weight of each factor's impact on different types of diseases; Step S5, model integration: encapsulate the association rules and weight information of environmental factors into feature vectors, and pass them to the machine learning prediction model in JSON format; The environmental factor association analysis model updates the analysis results every 12 hours, and evaluates the model performance through cross-validation. The verification accuracy must be greater than 80%.
6. The intelligent disease early warning system for aquaculture environment according to claim 5, characterized in that: The environmental factor association analysis model is implemented through the following steps: a. Use the equal frequency binning method to divide the continuous environmental factor data into 10 intervals, and use the optimal binning algorithm based on information entropy to optimize the intervals; b. Set the minimum support threshold to 0.3 and the minimum confidence threshold to 0.7 in the Apriori algorithm, and generate frequent itemsets through iteration; c. When calculating the grey correlation degree, temperature is selected as the reference sequence, the resolution coefficient is taken as 0.5, and the correlation coefficient of each environmental factor is obtained; d. Normalize the correlation coefficient to generate a 10×10 weight matrix; The model evaluation and optimization module includes the following functions: 10-fold cross-validation is used to evaluate model performance and calculate accuracy, precision, recall, and F1 value; confusion matrix is used to analyze the warning effect of different types of diseases and calculate the warning success rate of each type of disease; when the warning accuracy of any type of disease is less than 85%, the model optimization mechanism is triggered; The model optimization mechanism is implemented through the following steps: a. Use grid search to optimize the hyperparameters of the LSTM and Random Forest algorithms. The parameter search range includes the learning rate [0.001, 0.1], the number of hidden layer nodes [64, 256], and the number of decision trees [50, 200]. b. Use the SMOTE algorithm to oversample minority class samples and balance the data set distribution; c. Use samples in the newly added verified_label_samples tag library for incremental learning. verified_label_samples is updated once a month, and the sample size is ≥ 1000.
7. The intelligent disease early warning system for aquaculture environment according to claim 1 is characterized by: The early warning information push module includes an information classification unit, a multi-channel push unit, and a user feedback unit. The information classification unit classifies early warning information into three levels of urgency: red, orange, and yellow. Red warnings are handled within 5 minutes, orange warnings within 30 minutes, and yellow warnings within 2 hours. The multi-channel push unit supports four notification methods: SMS, app push, email, and voice call, ensuring a 99.9% early warning information delivery rate. The user feedback unit records user processing results and satisfaction ratings. The prevention and control plan generation module includes a plan knowledge base, an intelligent decision-making unit, and a plan update unit. The plan knowledge base stores no less than 1,000 standard prevention and control plans, covering 20 common aquaculture diseases. The intelligent decision-making unit uses a decision tree algorithm to automatically match the optimal prevention and control plan based on current environmental parameters and disease type, and optimizes plan parameters according to the scale of aquaculture. The plan update unit evaluates the prevention and control effects monthly and updates the plan library based on the evaluation results. The application service layer interacts with the processing platform layer for data through the REST API interface. The interface response time is less than 100ms, and a load balancing mechanism is provided to support no less than 1,000 concurrent users.
8. The intelligent disease early warning system for aquaculture environment according to claim 7, characterized in that: The solution knowledge base uses a graph database storage structure, specifically including disease nodes that store disease characteristics, disease patterns, and severity of damage. Its specific measures include, but are not limited to, prevention and control solution nodes that store drug use, environmental regulation, and biological control. The connection between content includes relationship edges that store solution applicability conditions and historical application results. The solution knowledge base uses knowledge graph technology and is implemented through Neo4j, supporting complex semantic queries and associative reasoning. The intelligent decision-making unit generates a prevention and control plan through the following steps: a. Use the C4.5 decision tree algorithm to build a solution selection model, set the decision tree depth to 5, and the minimum number of sample splits to 50; b. Based on the current environmental parameters, farming scale and disease type, retrieve the top three candidate solutions with the highest matching degree from the knowledge base; c. Use the analytic hierarchy process to calculate the comprehensive score of each plan, taking into account the three dimensions of prevention and control effectiveness, economic cost, and operational difficulty; d. Select the solution with the highest score and optimize key parameters through genetic algorithm; The implementation process of the program update unit is as follows: monthly feedback on the application of the prevention and control program is collected, with indicators including but not limited to cure rate, recovery time and recurrence rate; the fuzzy comprehensive evaluation method is used to quantitatively evaluate the effectiveness of the program, and the weights of the evaluation indicators are determined by the expert scoring method; when the program score is lower than 0.8, the optimization process is triggered, and the program is improved through case reasoning technology; the optimized program must undergo more than 3 experimental verifications, and after verification, it is updated to the program knowledge base.
9. A computer-readable storage medium, characterized in that Computer program instructions are stored on the computer-readable storage medium; when the computer program instructions are executed by the processor, the intelligent disease early warning system for aquaculture environment as described in any one of claims 1 to 8 is implemented.
Citation Information
Patent Citations
Intelligent aquaculture monitoring and early warning system based on artificial intelligence
CN116228031A
Ecological environment health intelligent assessment and management system based on PSR and entropy weight dynamic self-adaption
CN120088113A