Animal husbandry epidemic disease early warning method and system based on big data analysis
By combining space-time transmission models, multi-source data fusion warning models, gene-environment correlation models and lightweight deep learning models in the animal husbandry epidemic early warning system, the problems of epidemic spread path prediction error, limited computing resources and low efficiency in the deep learning model in the existing technology are solved, and accurate early warning of the epidemic and real-time assessment of the healthy status of the animal husbandry are achieved, providing a scientific basis for epidemic prevention and control.
Patent Information
- Application Number
- CN202510275455.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-10
- Publication Date
- 2025-06-27
AI Technical Summary
The existing technology is difficult to accurately capture the epidemic spread path in complex geographical environments. The computing resources of the gene-environment correlation model are limited, and the operation efficiency of the deep learning model is low, resulting in low epidemic monitoring efficiency.
By obtaining the breeding farm distribution data and transportation network data in the geographical information system, the preliminary epidemic spread path is generated in combination with the spatiotemporal communication model, and dynamically adjust it based on the real-time updated data. The multi-source data fusion warning model is used to predict high-risk areas, and the correlation between gene fragments and environmental parameters is identified through the gene-environmental association model, and the risk of adaptive mutations of strains is judged. At the same time, a lightweight deep learning model is deployed at the edge computing nodes, and live analyzing animal husbandry behavior and voiceprint data, assessing animal husbandry health status, and integrating mutation risk analysis and health assessment results to generate a comprehensive epidemic warning report.
Accurate prediction of the spread path of the epidemic has been achieved, the accuracy of prediction of high-risk areas has been improved, and the adaptive mutation risk of the strain can be quickly identified. Through real-time analysis of animal husbandry behavior and voiceprint data, the health status of animal husbandry is evaluated, and a comprehensive epidemic warning report is generated to provide a scientific basis for epidemic prevention and control.
Smart Images

Figure CN120221122A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of livestock disease early warning, and particularly relates to a livestock disease early warning method and system based on big data analysis. Background Art
[0002] In recent years, with the development of technologies such as big data, the Internet of Things, and artificial intelligence, significant progress has been made in livestock disease early warning systems. Current systems can already use Geographic Information System (GIS) technology to monitor the health status, vital signs, and behavior data of livestock in real time, and use big data analysis technology to predict and identify epidemics. In addition, gene-environment association models are used to analyze the relationship between pathogen genomic data and environmental parameters to identify potential adaptive mutation risks. In terms of lightweight edge intelligence, edge computing nodes combined with deep learning models (such as LSTM and CNN) are deployed for real-time analysis of abnormal livestock behaviors and voiceprint classification, further improving the efficiency of disease monitoring.
[0003] However, there are still some deficiencies in the existing technologies. First, it is difficult for GIS technology to accurately capture the epidemic spread path in complex geographical environments, and due to the uneven distribution of farms, the dynamic changes in the transportation network, and the lag in data updates, there are large errors in predicting the epidemic spread path. Second, when dealing with huge pathogen genomic data and constantly changing environmental parameters, the gene-environment association model faces problems such as limited computing resources and low comparison efficiency, and it is difficult to quickly identify the adaptive mutation risks of virus strains. Finally, in terms of lightweight edge intelligence, the computing power and storage space of edge devices are limited, resulting in low operating efficiency of deep learning models and it is difficult to reduce the computational complexity and storage requirements while ensuring the accuracy of the models. Summary of the Invention
[0004] To solve the above technical problems, the present invention proposes a livestock disease early warning method and system based on big data analysis to solve the problems existing in the above-mentioned prior arts.
[0005] To achieve the above object, in the first aspect, the present invention provides a livestock disease early warning method based on big data analysis, including:
[0006] Obtain the farm distribution data and transportation network data in the Geographic Information System, and generate a preliminary epidemic spread path in combination with the spatio-temporal propagation model;
[0007] Adjust the preliminary epidemic spread path according to the real-time updated farm density data and the dynamic change information of the transportation network to obtain a corrected spread path;
[0008] For the corrected spread path, adopt a multi-source data fusion early warning model to integrate the surrounding farm density and transportation network data to predict high-risk areas;
[0009] Obtain pathogen genomic data and environmental parameter data, compare them through a gene-environment association model, and identify the correlation between gene fragments and environmental parameters;
[0010] Based on the correlation between the gene fragments and environmental parameters, determine whether there is a risk of strain adaptive mutation, and generate a mutation risk analysis result;
[0011] Adopt a lightweight deep learning model, deploy the LSTM algorithm on edge computing nodes, and analyze livestock behavior data in real time to identify abnormal behavior patterns;
[0012] Deploy a CNN model through edge computing nodes to classify livestock voiceprint data and obtain a voiceprint classification result;
[0013] Combine the abnormal behavior pattern and the voiceprint classification result to generate an assessment report on the health status of livestock; integrate the mutation risk analysis result with the assessment report on the health status of livestock to generate a comprehensive epidemic warning report.
[0014] Preferably, the process of generating a preliminary epidemic spread path includes:
[0015] Obtain the distribution data of farms in the geographic information system, and extract the geographical location information of the farms;
[0016] Obtain traffic network data, and extract the structural characteristics and connection relationships of the traffic network;
[0017] Combine the farm distribution data and the traffic network data to construct a spatial association matrix;
[0018] Adopt a preset spatio-temporal propagation model, input the spatial association matrix, and calculate the contact frequency between farms;
[0019] According to the contact frequency and the preset infection rate parameter, calculate the epidemic spread speed;
[0020] Adopt a path generation algorithm, and generate a preliminary epidemic spread path based on the epidemic spread speed and the traffic network structure.
[0021] Preferably, the process of obtaining a corrected spread path includes:
[0022] Obtain the real-time updated farm density data, extract the geographical location change information of the farms, and obtain the updated farm distribution data;
[0023] Obtain the traffic network dynamic change data, extract the updated information of the traffic network structural characteristics, and obtain the updated traffic network data;
[0024] Construct an updated spatial association matrix by combining the updated farm distribution data and traffic network data;
[0025] Adopt a preset spatio-temporal propagation model, input the updated spatial association matrix, and calculate the dynamic contact frequency between farms;
[0026] Calculate the corrected epidemic spread speed according to the dynamic contact frequency and the preset infection rate parameter;
[0027] Adopt a path generation algorithm, combine the corrected epidemic spread speed and the updated traffic network structure, and generate a corrected epidemic spread path.
[0028] Preferably, the steps of predicting high-risk areas include:
[0029] Obtain the corrected epidemic spread path data, extract the farm location information involved in the path, and obtain the target farm distribution data;
[0030] According to the target farm distribution data, combine the surrounding farm density data, calculate the farm density value in the area, and obtain the updated density distribution data;
[0031] Obtain the dynamic change data of the traffic network, extract the traffic network structure information related to the target farm, and obtain the updated traffic network data;
[0032] Adopt a multi-source data fusion algorithm, combine the updated density distribution data and traffic network data, calculate the regional risk index, and obtain the risk distribution data;
[0033] According to the preset warning threshold, if the regional risk index exceeds the threshold, then determine that the area is a high-risk area, and obtain a list of high-risk areas;
[0034] Adopt a machine learning classification model, input the risk distribution data and the list of high-risk areas, optimize the regional risk prediction result, and obtain the final high-risk area prediction data.
[0035] Preferably, the process of identifying the correlation between gene fragments and environmental parameters includes:
[0036] Obtain the pathogen genome data, extract the gene fragment characteristic values, and obtain the genome characteristic data set;
[0037] Obtain the environmental parameter data, extract the environmental parameter characteristic points, and obtain the environmental parameter characteristic data set;
[0038] Adopt a gene-environment association model, input the genome characteristic data set and the environmental parameter characteristic data set, analyze the correlation between gene fragments and environmental parameters, and obtain the association characteristic data set.
[0039] Preferably, the process of generating the mutation risk analysis result includes:
[0040] According to a preset correlation threshold, if the correlation feature value exceeds the threshold, it is determined that the gene fragment has a strong correlation with the environmental parameter, and a strongly correlated gene data set is generated;
[0041] Using a support vector machine classification model, inputting the strongly correlated gene data set, optimizing the correlation features between the gene fragment and the environmental parameter, and generating an optimized correlation feature data set;
[0042] According to the optimized correlation feature data set, adjusting the gene-environment correlation model parameters to generate the final correlation model parameters;
[0043] Using the final correlation model parameters, inputting new genomic data and environmental parameter data, predicting the correlation features between the new gene fragment and the environmental parameter, and generating a predicted correlation feature data set;
[0044] Based on the predicted correlation feature data set, if the mutation probability of the gene fragment exceeds the preset risk threshold, it is determined that there is an adaptive mutation risk, and a mutation risk analysis result is generated.
[0045] Preferably, the steps of identifying abnormal behavior patterns include:
[0046] Obtaining livestock behavior data, cleaning and preprocessing the data to generate a standardized behavior data set;
[0047] Using a lightweight LSTM model, inputting the standardized behavior data set, extracting behavior features, and generating a behavior feature data set;
[0048] According to the behavior feature data set, the behavior feature values are calculated in real time through an edge computing node to generate a real-time feature value data set;
[0049] If the real-time feature value deviates from the preset normal range, it is determined that there is abnormal behavior, and an abnormal behavior marker data set is generated.
[0050] Preferably, the process of obtaining the voiceprint classification result includes:
[0051] Obtaining livestock voiceprint data, and using a preset noise reduction algorithm to remove environmental noise to generate noise-reduced voiceprint data;
[0052] Performing standardization processing on the noise-reduced voiceprint data to generate a standardized voiceprint data set;
[0053] Using a lightweight CNN model, inputting the standardized voiceprint data set, extracting voiceprint features, and generating a voiceprint feature data set;
[0054] Based on the voiceprint feature dataset, the voiceprint features are classified in real time through edge computing nodes to generate voiceprint classification results.
[0055] Preferably, the process of generating the comprehensive epidemic warning report includes:
[0056] Obtain the mutation risk analysis data, calculate the mutation risk value using a preset risk assessment model, and generate a mutation risk dataset;
[0057] Extract health status indicators based on the livestock health status assessment report to generate a health status dataset;
[0058] Fuse the mutation risk dataset with the health status dataset, and use the decision tree algorithm to evaluate the epidemic risk level to generate a comprehensive epidemic warning report.
[0059] In a second aspect, the present invention also provides a livestock disease warning system based on big data analysis, including:
[0060] A preliminary path generation module for obtaining farm distribution data and traffic network data in a geographic information system, and combining with a spatio-temporal propagation model to generate a preliminary epidemic diffusion path;
[0061] A path correction module for adjusting the preliminary epidemic diffusion path according to the real-time updated farm density data and traffic network dynamic change information to obtain a corrected diffusion path;
[0062] A risk prediction module for using a multi-source data fusion warning model for the corrected diffusion path to integrate the surrounding farm density and traffic network data to predict high-risk areas;
[0063] A gene-environment analysis module for obtaining pathogen genome data and environmental parameter data, and comparing them through a gene-environment association model to identify the correlation between gene fragments and environmental parameters;
[0064] A mutation risk judgment module for judging whether there is a risk of strain adaptive mutation according to the correlation between the gene fragment and the environmental parameter, and generating a mutation risk analysis result;
[0065] A behavior analysis module for using a lightweight deep learning model to deploy the LSTM algorithm at an edge computing node to analyze livestock behavior data in real time and identify abnormal behavior patterns;
[0066] A voiceprint classification module for classifying livestock voiceprint data through a CNN model deployed at an edge computing node to obtain voiceprint classification results;
[0067] The comprehensive warning module is used to generate an evaluation report on the health status of livestock by combining the abnormal behavior patterns and voiceprint classification results; integrate the mutation risk analysis results with the evaluation report on the health status of livestock to generate a comprehensive epidemic warning report.
[0068] Compared with the prior art, the present invention has the following advantages and technical effects:
[0069] The present invention discloses a method for warning livestock diseases based on big data analysis. First, using the distribution of farms and traffic network data of the geographic information system, combined with the spatio-temporal propagation model to generate a preliminary epidemic diffusion path, and dynamically adjust according to the real-time updated data; secondly, adopt a multi-source data fusion warning model to predict high-risk areas; judge the risk of strain mutation by analyzing the correlation between the pathogen genome and environmental parameters; further, deploy a deep learning model at the edge computing node to analyze livestock behavior and voiceprint data in real time to evaluate the health status of livestock; finally, integrate the mutation risk analysis and livestock health assessment results to generate a comprehensive epidemic warning report. The present invention realizes the accurate warning of animal epidemics through multi-dimensional data analysis and edge intelligent computing, providing a scientific basis for epidemic prevention and control. BRIEF DESCRIPTION OF THE DRAWINGS
[0070] The drawings constituting a part of this application are used to provide a further understanding of this application. The schematic embodiments of this application and their descriptions are used to explain this application and do not constitute an improper limitation to this application. In the drawings:
[0071] Figure 1 is the flowchart of the method of the embodiment of the present invention;
[0072] Figure 2 is the schematic diagram of the system of the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0073] It should be noted that, without conflict, the embodiments in this application and the features in the embodiments can be combined with each other. The following will refer to the drawings and combine the embodiments to detail this application.
[0074] It should be noted that the steps shown in the flowchart of the drawings can be executed in a computer system such as a set of computer executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.
[0075] Embodiment 1
[0076] As Figure 1 shown, this embodiment provides a method for warning livestock diseases based on big data analysis, including:
[0077] S101. Obtain farm distribution data and transportation network data in the geographic information system, and generate a preliminary epidemic spread path based on the preset space-time propagation model.
[0078] Obtain farm distribution data from the geographic information system and extract the geographical location information of the farms. Obtain transportation network data and extract the structural characteristics and connection relationships of the transportation network. Combine farm distribution data and transportation network data to construct a spatial association matrix. Use the preset spatiotemporal propagation model, input the spatial association matrix, and calculate the contact frequency between farms. Calculate the epidemic spread rate based on the contact frequency and the preset infection rate parameters. Use the path generation algorithm to generate a preliminary epidemic spread path based on the epidemic spread rate and transportation network structure. According to the generated epidemic spread path, adjust the spatiotemporal propagation model parameters and optimize the path prediction results.
[0079] Specifically, the acquisition of farm distribution data in the geographic information system involves the spatial positioning information of the farm, which usually includes latitude and longitude coordinates, administrative divisions and addresses. For example, a large chicken farm is located in the suburbs of a city, with longitude and latitude of 30 degrees and 15 minutes north latitude and 120 degrees and 20 minutes east longitude, which can be obtained through satellite remote sensing or field mapping. Attribute information such as breeding scale, species, and feeding methods also needs to be collected, which has an important impact on the risk assessment of epidemic transmission.
[0080] Traffic network data mainly includes road grade, traffic capacity and connection relationship. Taking a certain county as an example, national roads, provincial roads and rural roads form a multi-level road network. Road grade and traffic capacity directly affect the speed of disease transmission. The daily traffic volume of trunk roads can reach tens of thousands of vehicles, while rural roads may only have hundreds of vehicles. Road network structural characteristics such as node degree and centrality reflect the connectivity and importance of different regions.
[0081] The construction of the spatial association matrix requires comprehensive consideration of the physical distance and traffic accessibility between farms. For example, if two farms ten kilometers apart are connected by a high-grade road, their spatial association will be higher than if they are connected by a rural road. Matrix elements can be expressed as time distance or weighted distance. The calculation of contact frequency between farms needs to consider factors such as the flow of transport vehicles and personnel. For example, there are five round trips between two farms every day for feed transport vehicles and two exchanges between technicians. These are all potential routes for the spread of the disease. Combined with the preset infection rate parameters, such as a 20% probability of infection from a single contact, the spread rate of the epidemic can be estimated.
[0082] The generation of the epidemic spread path needs to consider spatio-temporal transmission characteristics. Taking the spread of avian influenza as an example, the virus can be transmitted through the air and is affected by the wind direction. It can also be mechanically transmitted through transportation vehicles. After an epidemic occurs in a farm, farms within a radius of three kilometers around it will be directly threatened, while the indirect transmission through the transportation network takes longer. After the initial path is generated, the model parameters can be adjusted according to the actual transmission situation. For example, the influence range of air transmission can be adjusted from three kilometers to five kilometers, or the time parameter of transportation transmission can be adjusted. Environmental variables such as seasonal factors and weather conditions also need to be considered during the path optimization process. For example, in the cold season, the virus activity increases and the transmission speed may accelerate. In rainy weather, due to the decrease in road traffic capacity, the risk of mechanical transmission may decrease. These factors will all affect the prediction of the final epidemic spread path.
[0083] S102. Adjust the preliminary epidemic spread path according to the real-time updated farm density data and the dynamic change information of the transportation network to obtain the corrected spread path.
[0084] Obtain the real-time updated farm density data, extract the geographical location change information of the farms to obtain the updated farm distribution data. Obtain the dynamic change data of the transportation network, extract the updated information of the transportation network structure characteristics to obtain the updated transportation network data. Combine the updated farm distribution data and the transportation network data to construct the updated spatial association matrix. Use the preset spatio-temporal transmission model, input the updated spatial association matrix, and calculate the dynamic contact frequency between farms. According to the dynamic contact frequency and the preset infection rate parameter, calculate the corrected epidemic spread speed. Use the path generation algorithm, combine the corrected epidemic spread speed and the updated transportation network structure to generate the corrected epidemic spread path. According to the corrected epidemic spread path, adjust the spatio-temporal transmission model parameters to optimize the path prediction result.
[0085] Specifically, the dynamic changes in farm density data need to consider factors such as newly added farms, closed farms, and adjusted breeding scales. For example, in a certain area, there were originally ten pig farms with a scale of more than 10,000 heads. Through real-time monitoring, it is found that two new large-scale farms are added, and one is temporarily closed due to environmental protection requirements. These changes will directly affect the assessment of the risk of disease transmission.
[0086] The dynamic changes in the transportation network include information such as changes in road traffic capacity, newly added transportation lines, and seasonal traffic restrictions. For example, during the Spring Festival travel rush, the average daily traffic volume of an inter-provincial road increased from 5,000 vehicles on weekdays to 8,000 vehicles, and a new provincial road was opened. These changes will affect the actual connection strength between farms.
[0087] The update of the spatial correlation matrix needs to comprehensively consider factors such as geographical distance, transportation convenience, and transportation frequency. For example, for two farms that are fifty kilometers apart, if a new highway is built in the middle, the spatio-temporal distance may be shortened from the original two hours to forty minutes, resulting in an increase in the risk coefficient of disease transmission.
[0088] The calculation of the dynamic contact frequency needs to combine information such as the disinfection records of transport vehicles and vehicle trajectory data. For example, if there are ten transport vehicles coming and going to a certain farm every week, and through the analysis of the transport trajectories, it is found that three of them have visited multiple farms within a week. In this case, the weight of the contact frequency needs to be increased.
[0089] The calculation of the corrected epidemic spread speed needs to consider environmental variables such as seasonal factors and weather conditions. For example, in the cold season, the survival time of the virus is extended and the transmission risk increases, so the basic spread speed needs to be increased, while in hot and dry weather conditions, it may need to be decreased.
[0090] The path generation algorithm needs to consider the diversity and probability distribution of the transmission chain. For example, starting from an infection source, it may spread to other farms through multiple channels such as direct contact, vehicle transportation, and personnel flow at the same time. Each transmission chain has its occurrence probability and transmission speed. The adjustment of the spatio-temporal transmission model parameters should be based on the verification of historical epidemic data. For example, by comparing the predicted transmission path with the actual epidemic development trajectory, it is found that the model has a high prediction accuracy for short-distance transmission, while there is an underestimation for long-distance cross-regional transmission, and the distance attenuation parameter needs to be adjusted accordingly.
[0091] S103. For the corrected diffusion path, adopt a multi-source data fusion early warning model, integrate the density of surrounding farms and traffic network data, and predict high-risk areas.
[0092] Obtain the corrected epidemic spread path data, extract the location information of the farms involved in the path to obtain the target farm distribution data. According to the target farm distribution data, combine the density data of surrounding farms, calculate the density value of farms in the region to obtain the updated density distribution data. Obtain the dynamic change data of the traffic network, extract the traffic network structure information related to the target farm to obtain the updated traffic network data. Adopt a multi-source data fusion algorithm, combine the updated density distribution data and traffic network data, calculate the regional risk index to obtain the risk distribution data. According to the preset early warning threshold, if the regional risk index exceeds the threshold, then determine that the region is a high-risk area to obtain a list of high-risk areas. Adopt a machine learning classification model, input the risk distribution data and the list of high-risk areas, optimize the regional risk prediction result to obtain the final high-risk area prediction data. According to the final high-risk area prediction data, adjust the parameters of the multi-source data fusion algorithm to optimize the prediction accuracy of the early warning model.
[0093] Specifically, extracting the information of farm locations from the corrected epidemic spread path is essentially to locate specific risk sources. For example, there are multiple farm locations in a certain poultry farming area. According to the historical epidemic data, avian influenza has occurred in this area. Based on the spread path data, three of these farms can be locked as key monitoring points. The geographical coordinates and surrounding environmental characteristics of these farms need to be recorded in detail.
[0094] The analysis of the density of surrounding farms is crucial for assessing the risk of epidemic spread. For example, in an area with a radius of five kilometers, if the number of farms exceeds ten and the average inventory of each farm reaches more than five thousand, this high-density distribution will significantly increase the risk of disease transmission. The density distribution data needs to be updated at any time to reflect the dynamic changes in the farming scale.
[0095] The information of the traffic network structure includes elements such as road grades and traffic capacities. For example, there is a complex road network composed of provincial roads and rural roads around a high-risk farm, with a daily traffic flow of three thousand vehicles, and the proportion of vehicles transporting poultry reaches 15%. All these data will affect the risk assessment results.
[0096] Multi-source data fusion is a process of integrating information from different dimensions such as density distribution and traffic data. For example, the density of farms in a certain area is eight per square kilometer, the daily traffic flow is two thousand vehicles, and it is located downstream of a river. The risk index calculated by integrating these factors is 0.85.
[0097] The setting of the warning threshold needs to consider multiple factors. For example, setting the risk index of 0.75 as the warning threshold. When the density of farms in a certain area exceeds the standard and the traffic flow surges, resulting in the risk index reaching 0.8, the system will automatically list this area in the high-risk list. The machine learning model continuously optimizes its prediction ability through training data. For example, inputting historical data such as the epidemic records, changes in farm distribution, and traffic flow in a certain area in the past three years, the model can predict the risk level of this area in different seasons. When the prediction result shows that the risk index of this area may exceed 0.8 during the spring migration period, the system will issue a warning in advance. The parameter adjustment of the warning model is a continuous optimization process. By comparing the warning results with the actual epidemic occurrence situation, the risk assessment parameters are continuously corrected. For example, if it is found that although the density of farms in a certain area is high, due to the perfect biosecurity measures, the actual incidence rate is low, the weight of the density factor can be adjusted accordingly.
[0098] S104. Obtain pathogen genome data and environmental parameter data, and compare them through a gene-environment association model to identify the correlation between gene fragments and environmental parameters.
[0099] Obtain pathogen genomic data, extract gene fragment characteristic values, and obtain a genomic characteristic data set. Obtain environmental parameter data, extract environmental parameter characteristic points, and obtain an environmental parameter characteristic data set. Use a gene-environment association model, input the genomic characteristic data set and the environmental parameter characteristic data set, analyze the correlation between gene fragments and environmental parameters, and obtain an association characteristic data set. According to a preset association threshold, if the association characteristic value exceeds the threshold, it is determined that the gene fragment has a strong correlation with the environmental parameter, and a strong association gene data set is obtained. Use a machine learning classification model, input the strong association gene data set, optimize the association characteristics between gene fragments and environmental parameters, and obtain an optimized association characteristic data set. According to the optimized association characteristic data set, adjust the parameters of the gene-environment association model, optimize the prediction accuracy of the association model, and obtain the final association model parameters. Use the final association model parameters, input new genomic data and environmental parameter data, predict the association characteristics between new gene fragments and environmental parameters, and obtain a predicted association characteristic data set.
[0100] Specifically, the acquisition of genomic data is usually achieved through high-throughput sequencing technology. The length of the nucleotide sequence obtained after sequencing is usually between several hundred and several thousand base pairs. Taking the avian influenza virus as an example, its genomic characteristic values include the hemagglutinin and neuraminidase gene sequences, which can reflect the pathogenicity and transmission ability of the virus.
[0101] Environmental parameters involve multiple dimensions such as temperature, humidity, and air pressure. For example, in South China, when the temperature is between 20 and 30 degrees Celsius and the relative humidity is between 70% and 80%, the transmission risk of the avian influenza virus is relatively high. Gene fragment feature extraction can use sequence alignment methods to compare the target sequence with the standard sequence and identify key mutation sites. Environmental parameter characteristic points can be identified through spatio-temporal clustering methods. For example, clustering meteorological data within a year by season to obtain typical combinations of environmental parameters.
[0102] In correlation analysis, the canonical correlation analysis method can be used to calculate the correlation coefficient between gene characteristics and environmental parameters. When the correlation coefficient exceeds 0.8, a significant correlation can be considered. Taking Streptococcus suis as an example, there is an obvious correlation between its virulence genes and environmental temperature. When the environmental temperature is within the range of 25 to 30 degrees Celsius, the expression level of virulence genes increases significantly. Through a machine learning classification model such as a support vector machine, this association characteristic can be further optimized. The model training data includes known gene expression profiles and corresponding environmental parameter records, and the optimal model parameters are determined through cross-validation. The tuning process of the association model needs to consider multiple indicators such as prediction accuracy, sensitivity, and specificity. Taking the foot-and-mouth disease virus as an example, the prediction accuracy of the association between its structural protein coding gene and environmental humidity can reach 85%, which provides an important basis for epidemic warning. The final prediction model can input new sequence data and environmental parameters to predict potential epidemic risks.
[0103] In the practical application of association analysis, the impact of seasonal variations also needs to be considered. For example, in summer, the gene expression patterns of certain pathogens change significantly, and this change is closely related to environmental factors such as increased temperature and precipitation. By establishing a dynamic prediction model, the impact brought by such seasonal variations can be captured more accurately. The model can dynamically adjust the warning threshold according to the environmental parameter characteristics of different seasons, improving the accuracy of prediction.
[0104] In practice, the association analysis between the genome and environmental parameters also needs to consider geospatial factors. For example, in mountainous areas, the change in altitude leads to gradient changes in temperature and air pressure, and such changes may affect the gene expression and transmission characteristics of pathogens. By integrating geographical information system data, the impact of environmental factors on the genetic characteristics of pathogens can be analyzed more comprehensively.
[0105] S105. According to the relevance between gene fragments and environmental parameters, determine whether there is a risk of strain adaptive mutation, and generate a mutation risk analysis result.
[0106] Obtain pathogen genome data and environmental parameter data, extract gene fragment feature values and environmental parameter feature points, and generate a genome feature data set and an environmental parameter feature data set. Adopt a gene-environment association model, input the genome feature data set and the environmental parameter feature data set, analyze the relevance between gene fragments and environmental parameters, and generate an association feature data set. According to a preset association threshold, if the association feature value exceeds the threshold, it is determined that the gene fragment and the environmental parameter have a strong relevance, and a strong association gene data set is generated. Adopt a support vector machine classification model, input the strong association gene data set, optimize the association features between gene fragments and environmental parameters, and generate an optimized association feature data set. According to the optimized association feature data set, adjust the parameters of the gene-environment association model to generate final association model parameters. Adopt the final association model parameters, input new genome data and environmental parameter data, predict the association features between new gene fragments and environmental parameters, and generate a predicted association feature data set. Based on the predicted association feature data set, if the gene fragment mutation probability exceeds the preset risk threshold, it is determined that there is a risk of adaptive mutation, and a mutation risk analysis result is generated.
[0107] Specifically, genome data can obtain the nucleotide sequence information of pathogen samples through a high-throughput sequencer, and environmental parameter data includes physical and chemical conditions such as temperature, humidity, and pH. Taking the influenza virus as an example, from the genome sequences of virus strains collected from different regions, nucleic acid sequence fragments are extracted as feature values, such as the amino acid sequence composition of the hemagglutinin gene fragment. The feature points of environmental parameters include key indicators such as seasonal temperature change curves and relative humidity distributions.
[0108] The gene-environment association model analyzes the correlation between gene fragments and environmental factors through statistical methods. For example, to analyze the relationship between the influenza virus hemagglutinin gene and temperature, the Pearson correlation coefficient between gene sequence variations and temperature changes can be calculated. If the correlation coefficient exceeds a preset threshold (such as 0.8), it is considered that there is a significant association between the gene fragment and temperature. The support vector machine classification model can further optimize the association features. Taking the influenza virus as an example, the gene sequence variations of hemagglutinin and environmental parameters such as temperature and humidity are used as feature vectors and input into the model. Through the kernel function, they are mapped into a high-dimensional space to find the optimal classification hyperplane, so as to identify a more accurate gene-environment association pattern. The adjustment of the association model parameters is based on the results of cross-validation. For example, by adjusting the kernel function parameters, regularization coefficients, etc., the prediction accuracy of the model on the validation set can reach the optimal. For the influenza virus, it may be found that certain hemagglutinin gene mutations are highly correlated with the low-temperature environment, which provides a basis for predicting the environmental adaptability of the virus. When predicting the association features, the newly obtained virus genome data and the current environmental parameters are input into the model. If it is found that the mutation probability at a certain site of the hemagglutinin gene exceeds the risk threshold (such as 0.6) under low-temperature conditions, it indicates that the virus may undergo adaptive evolution and enhance its transmission ability in the low-temperature environment. This analysis method is not only applicable to livestock influenza viruses, but also can be extended to other livestock pathogens. Such as the association analysis between the spike protein gene of coronaviruses and humidity, or the relationship study between the drug-resistant genes of Mycobacterium tuberculosis and antibiotic concentrations. By systematically analyzing the gene-environment association features, the evolutionary trends of pathogens can be predicted, providing a scientific basis for disease prevention and control. The association analysis between genomic features and environmental parameters can reveal the adaptation mechanism of pathogens to the environment, help understand the disease transmission law, and provide an important reference for public health decision-making.
[0109] S106. Adopt a lightweight deep learning model, deploy the LSTM algorithm on the edge computing node, and analyze livestock behavior data in real time to identify abnormal behavior patterns.
[0110] Obtain livestock behavior data, clean and preprocess the data to generate a standardized behavior data set. Adopt a lightweight LSTM model, input the standardized behavior data set, extract behavior features, and generate a behavior feature data set. According to the behavior feature data set, calculate the behavior feature values in real time through the edge computing node to generate a real-time feature value data set. If the real-time feature value deviates from the preset normal range, it is determined that there is abnormal behavior, and an abnormal behavior marker data set is generated. Use the abnormal behavior marker data set to train the LSTM model, optimize the ability to identify abnormal behavior, and generate optimized model parameters. According to the optimized model parameters, input real-time behavior data, predict the probability of abnormal behavior, and generate an abnormal probability data set. Based on the abnormal probability data set, if the abnormal probability exceeds the preset threshold, generate a behavior abnormal warning message.
[0111] Specifically, in this embodiment, the livestock is taken as an example of pigs. The livestock behavior data collection of pigs involves multi-dimensional indicators such as activity trajectories, feeding and drinking frequencies, and resting postures. Through the infrared cameras and sensor networks installed in the pigsty, the movement trajectories and behavioral characteristics of pigs can be recorded all day long. There may be noises such as equipment failures and environmental interferences in the original data, and data cleaning and standardization processing are required. For example, the collected pig activity trajectory data is uniformly converted into position information in a standard coordinate system, and the feeding behavior data is normalized into frequency statistics within a fixed time period every day.
[0112] The lightweight long short-term memory network model can effectively extract the temporal characteristics of pig behavior. By dividing the standardized data into time windows, the normal activity patterns of pigs can be identified. For example, healthy pigs usually have a higher activity level in the early morning and evening, while they mainly rest at noon. The model can learn this circadian rhythm feature to form a behavioral pattern feature vector. The edge computing node is deployed on-site in the pigsty and can process sensor data in real time and extract feature values. Taking the movement speed of pigs as an example, if the instantaneous speed of a certain pig suddenly rises to more than three times the normal value, or the activity level within multiple consecutive time windows is significantly lower than the group average level, it may indicate abnormal behavior. By accumulating enough abnormal behavior marked samples, the recognition ability of the model can be continuously optimized. For example, it is found that some pigs will show early signs such as a decrease in food intake and slow movement before getting sick, and these characteristics can be important learning targets for the model. After the model parameters are optimized, similar abnormal behavior patterns can be captured more accurately.
[0113] In practical applications, the system continuously monitors the behavioral characteristics of each pig. When the probability of abnormal behavior of a certain pig continuously exceeds the warning threshold of 85% within ten minutes, the system will immediately push a warning message to the breeding personnel. This timely warning mechanism can help the breeding personnel discover problem pigs as early as possible, take corresponding prevention and control measures, effectively reduce the risk of disease transmission, and improve the health level and production efficiency of livestock.
[0114] S107. Deploy a CNN model through an edge computing node to classify livestock voiceprint data and obtain a voiceprint classification result.
[0115] Obtain livestock voiceprint data, use a preset noise reduction algorithm to remove environmental noise, and generate noise-reduced voiceprint data. For the noise-reduced voiceprint data, perform normalization processing to generate a normalized voiceprint dataset. Use a lightweight CNN model, input the normalized voiceprint dataset, extract voiceprint features, and generate a voiceprint feature dataset. According to the voiceprint feature dataset, use an edge computing node to classify voiceprint features in real time and generate a voiceprint classification result. If the voiceprint classification result deviates from the preset normal range, it is determined that there is a voiceprint anomaly, and an abnormal voiceprint marker dataset is generated. Use the abnormal voiceprint marker dataset to train the CNN model, optimize the voiceprint anomaly classification ability, and generate optimized model parameters. According to the optimized model parameters, input real-time voiceprint data, predict the voiceprint anomaly probability, and generate an anomaly probability dataset.
[0116] Specifically, the voiceprint data collection is the group voice information collected by arranging a microphone array in the pigsty. Through an omnidirectional acoustic acquisition device, various calls made by livestock can be effectively recorded. The environmental noise mainly comes from the operating sounds of equipment such as environmental control equipment and ventilation systems. The adaptive filtering algorithm can effectively reduce these interferences. For the noise-reduced voiceprint, through amplitude normalization and duration normalization processing, the voice data collected at different times is unified to the same scale range.
[0117] The voiceprint feature extraction uses a deep convolutional neural network, and the network structure includes multiple convolutional layers and pooling layers. Perform time-frequency domain conversion on the input voiceprint data to generate a spectrogram, and then extract voice features through convolution operations. For example, when livestock shows a stress response, the calls it makes will present a specific energy distribution pattern in the frequency spectrum, and these features can be effectively captured through the convolutional layer. The edge computing node processes voiceprint data in real time and obtains the classification result of the voiceprint through forward propagation calculation. The voiceprint classification result includes two categories: normal and abnormal. Normal voiceprints are characterized by regular and stable frequency and energy distribution. Abnormal voiceprints, on the other hand, are characterized by sudden high-frequency screams or continuous low-frequency moans. When an abnormal voiceprint is detected, the system will mark and store it for subsequent model optimization. For example, during the feeding process, livestock will make abnormal calls when they are hungry, sick, or frightened, and these voice features are significantly different from those in the normal state. During the model training process, the network parameters are continuously adjusted using the marked abnormal voiceprint data. Optimize the convolutional kernel parameters through the backpropagation algorithm to improve the recognition accuracy of the model for abnormal voiceprints. The optimized model can more accurately capture the subtle changes in voice features.
[0118] In practical applications, the system analyzes the real-time collected voiceprint data and calculates the probability value of it belonging to the abnormal category. When the abnormal probability exceeds the preset threshold, such as 0.8, the system will issue a warning message in a timely manner to remind the breeding personnel to pay attention to the livestock status. The voiceprint analysis technology has important application value in the field of intelligent breeding. Through the real-time monitoring of voice characteristics, the abnormal conditions of livestock can be quickly detected, providing a basis for timely intervention. This non-contact monitoring method not only does not affect the normal activities of livestock but also can work all-weather, improving the intelligent level of breeding management.
[0119] S108. Combine the abnormal behavior pattern and the voiceprint classification result to generate an assessment report on the health status of livestock.
[0120] Obtain the livestock voiceprint data, use a preset noise reduction algorithm to remove environmental noise, and generate the noise-reduced voiceprint data. For the noise-reduced voiceprint data, perform normalization processing to generate a normalized voiceprint data set. Use a lightweight CNN model, input the normalized voiceprint data set, extract voiceprint features, and generate a voiceprint feature data set. According to the voiceprint feature data set, classify the voiceprint features in real time through an edge computing node to generate a voiceprint classification result. If the voiceprint classification result deviates from the preset normal range, it is determined that there is a voiceprint abnormality, and an abnormal voiceprint marked data set is generated. Use the abnormal voiceprint marked data set to train the CNN model, optimize the voiceprint abnormal classification ability, and generate optimized model parameters. According to the optimized model parameters, input the real-time voiceprint data, predict the voiceprint abnormal probability, and generate an abnormal probability data set.
[0121] Specifically, the collection of livestock voiceprint data needs to consider environmental factors, including the noise of ventilation equipment inside the pig house, the calls of other animals and other interference. The noise reduction algorithm can use the wavelet transform method to decompose the signal into different frequency bands to identify and remove the characteristic frequency bands of environmental noise. For example, the fixed frequency noise generated when the ventilation equipment is running can be filtered by setting a threshold to retain the main frequency components of the pig's call. The standardization of voiceprint data involves amplitude normalization and duration unification. By normalizing the energy of the collected sound signal, the voiceprint data collected at different times and locations can be made comparable. The duration can be unified by using a fixed window interception method, such as cutting each sound sample to a uniform length of three seconds. The design of a lightweight convolutional neural network model needs to balance accuracy and computing resource consumption. A deep separable convolution structure can be used to reduce the number of model parameters while maintaining feature extraction capabilities. The network structure can contain three to four convolution layers, each layer uses a smaller convolution kernel size, such as three by three, and cooperates with the maximum pooling layer to reduce the size of the feature map. Hardware resource limitations need to be considered when deploying edge computing nodes. The model can be quantized into an eight-bit fixed-point format to reduce computing and storage overhead. On edge devices such as the Raspberry Pi, the time to process a three-second voiceprint data can be controlled within one hundred milliseconds, meeting real-time requirements. Voiceprint anomaly detection is based on statistical features and pattern recognition. Normal pig calls have relatively stable frequency distribution and energy characteristics, while calls when sick or frightened will have obvious deviations. The normal range of multiple dimensions such as frequency characteristics and energy distribution can be set, and samples outside the range are marked as abnormal samples. During model optimization, transfer learning methods can be used. First, pre-train the model with a large amount of normal voiceprint data, and then fine-tune the model parameters with a small number of abnormal samples. This can improve the model's ability to identify new abnormal types. In practice, it is found that using 5,000 normal samples for pre-training and then fine-tuning with 100 abnormal samples can increase the accuracy of anomaly detection by 15%. The reliability evaluation of the prediction results requires the integration of multiple indicators. The similarity between the voiceprint features and the normal pattern can be calculated, and the abnormal probability score can be generated by combining the continuity analysis in the time series. When the anomaly is detected multiple times in a row and the probability exceeds 80%, the abnormal alarm is triggered to reduce the false alarm rate.
[0122] S109. Integrate the mutation risk analysis results with the livestock health status assessment report to generate a comprehensive epidemic warning report.
[0123] Obtain mutation risk analysis data, calculate the mutation risk value using a preset risk assessment model, and generate a mutation risk data set. Extract health status indicators according to the livestock health status assessment report to generate a health status data set. Fuse the mutation risk data set with the health status data set, use the decision tree algorithm to evaluate the epidemic risk level, and generate a comprehensive epidemic warning report. According to the comprehensive epidemic warning report, use the support vector machine algorithm to analyze the epidemic development trend and generate an epidemic trend prediction result. If the epidemic trend prediction result is higher than the preset threshold, use the neural network model to generate prevention and control strategy suggestions and generate a prevention and control strategy data set. According to the prevention and control strategy data set, use an automated script to deploy prevention and control measures and generate prevention and control measure execution results. According to the prevention and control measure execution results, update the comprehensive epidemic warning report to generate the latest epidemic warning report.
[0124] Specifically, mutation risk analysis requires continuous collection of environmental data and livestock behavior data. Environmental data includes temperature, humidity, ammonia concentration, etc., while behavior data includes activity intensity, feeding frequency, voiceprint characteristics, etc. By establishing a mutation risk assessment model based on the genetic algorithm, weighted analysis can be performed on different indicators. For example, when the temperature exceeds the normal range and the livestock activity frequency decreases significantly, the mutation risk value will increase accordingly. The key indicators in the health status assessment report include the body temperature change trend, feed intake fluctuation, voiceprint abnormality frequency, etc. Through data standardization processing, indicators of different dimensions are unified into the range of zero to one. For example, individuals with a body temperature change amplitude exceeding 0.5 degrees and a duration exceeding four hours are marked as high-risk samples.
[0125] The data fusion process uses a multi-layer decision tree model, taking the mutation risk value and health status indicators as input features. Multiple thresholds are set for the branch nodes of the decision tree. For example, when the mutation risk value is greater than 0.7 and the number of abnormal health indicators exceeds three, it is determined as a high-risk level. The risk level of each pigsty is scored on a five-point scale, and a score exceeding four triggers an alarm. Epidemic trend prediction uses the support vector machine algorithm, and the input features include time series data such as historical risk level changes and changes in the number of abnormal individuals. Feature mapping is performed through the radial basis kernel function to predict the risk trend for the next seven days. When the prediction result shows that the risk level continues to rise and the slope exceeds the preset threshold, the generation of prevention and control strategies is initiated.
[0126] The prevention and control strategy generation uses a deep neural network model, which is pre-trained on historical prevention and control case data. According to the current epidemic characteristics, it generates multi-dimensional prevention and control suggestions including isolation plans, disinfection plans, treatment plans, etc. For example, in the case of a high incidence of respiratory diseases, it is recommended to increase the disinfection frequency to three times a day and increase the ventilation time. The deployment of prevention and control measures uses an automated control system to convert the prevention and control strategies into specific execution instructions. For example, automatically adjust the operating parameters of ventilation equipment, control the startup time of disinfection equipment, and send instructions for dividing isolation areas. The system collects execution result data every two hours, including the disinfection coverage area, the number of ventilation and air changes, etc.
[0127] The early warning report update mechanism is based on real-time data analysis. It compares the latest execution effects of prevention and control measures with the expected goals and dynamically adjusts the risk level assessment criteria. When the execution effect fails to meet the expectations, the system automatically increases the sampling frequency and early warning sensitivity to ensure timely detection of potential risks.
[0128] Embodiment 2
[0129] As Figure 2 shown, based on the same inventive concept, this embodiment also provides a livestock disease early warning system based on big data analysis, mainly including:
[0130] A preliminary path generation module, which is used to obtain the farm distribution data and traffic network data in the geographic information system, and combine with the spatio-temporal propagation model to generate a preliminary epidemic diffusion path;
[0131] A path correction module, which is used to adjust the preliminary epidemic diffusion path according to the real-time updated farm density data and the dynamic change information of the traffic network to obtain a corrected diffusion path;
[0132] A risk prediction module, which is used to adopt a multi-source data fusion early warning model for the corrected diffusion path, integrate the surrounding farm density and traffic network data, and predict high-risk areas;
[0133] A gene-environment analysis module, which is used to obtain pathogen genome data and environmental parameter data, and compare them through a gene-environment association model to identify the correlation between gene fragments and environmental parameters;
[0134] A mutation risk judgment module, which is used to judge whether there is a risk of strain adaptive mutation according to the correlation between the gene fragment and the environmental parameter, and generate a mutation risk analysis result;
[0135] A behavior analysis module, which is used to adopt a lightweight deep learning model, deploy the LSTM algorithm at the edge computing node, and analyze livestock behavior data in real time to identify abnormal behavior patterns;
[0136] A voiceprint classification module, which is used to deploy a CNN model through an edge computing node, classify livestock voiceprint data, and obtain a voiceprint classification result;
[0137] A comprehensive early warning module, which is used to combine the abnormal behavior pattern and the voiceprint classification result to generate an evaluation report on the health status of livestock; integrate the mutation risk analysis result with the evaluation report on the health status of livestock to generate a comprehensive epidemic early warning report.
[0138] The livestock disease early warning system based on big data analysis provided in this embodiment has all the advantages of the livestock disease early warning method based on big data analysis provided in Embodiment 1.
[0139] The above is only a preferred specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed in the present application should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A livestock disease early warning method based on big data analysis, characterized in that: The following steps are involved: Obtain farm distribution data and transportation network data from the geographic information system, and combine them with the spatiotemporal transmission model to generate a preliminary epidemic spread path; According to the real-time updated farm density data and traffic network dynamic change information, the preliminary epidemic spread path is adjusted to obtain a revised spread path; According to the revised diffusion path, a multi-source data fusion early warning model is used to integrate the surrounding farm density and transportation network data to predict high-risk areas; Obtain pathogen genome data and environmental parameter data, compare them through gene-environment association models, and identify the correlation between gene fragments and environmental parameters; According to the correlation between the gene fragment and the environmental parameters, whether there is a risk of adaptive mutation of the strain is determined, and a mutation risk analysis result is generated; Adopt lightweight deep learning models and deploy LSTM algorithms on edge computing nodes to analyze livestock behavior data in real time and identify abnormal behavior patterns; Deploy CNN models through edge computing nodes to classify livestock voiceprint data and obtain voiceprint classification results; Combining the abnormal behavior pattern with the voiceprint classification results, generating a livestock health status assessment report; The mutation risk analysis results are integrated with the livestock health status assessment report to generate a comprehensive epidemic warning report.
2. The method according to claim 1, characterized in that The process of generating a preliminary epidemic spread path includes: Obtain farm distribution data from the geographic information system and extract the geographical location information of the farms; Obtain traffic network data and extract the structural characteristics and connection relationships of the traffic network; Combining the farm distribution data and the transportation network data, constructing a spatial association matrix; Using a preset spatiotemporal propagation model, inputting the spatial association matrix, and calculating the contact frequency between farms; Calculate the epidemic spread rate based on the contact frequency and the preset infection rate parameters; A path generation algorithm is used to generate a preliminary epidemic spread path based on the epidemic spread speed and transportation network structure.
3. The method according to claim 1, characterized in that The process of obtaining the corrected diffusion path includes: Obtain real-time updated farm density data, extract the geographical location change information of the farms, and obtain updated farm distribution data; Obtain dynamic change data of the traffic network, extract updated information of the traffic network structure characteristics, and obtain updated traffic network data; Combining the updated farm distribution data and transportation network data, constructing an updated spatial association matrix; Using a preset spatiotemporal propagation model, inputting the updated spatial association matrix, and calculating the dynamic contact frequency between farms; Calculate the corrected epidemic spread rate according to the dynamic contact frequency and the preset infection rate parameter; A path generation algorithm is used to generate a revised epidemic spread path in combination with the revised epidemic spread speed and the updated traffic network structure.
4. The method according to claim 1, characterized in that The steps to predict high-risk areas include: Obtain the corrected epidemic spread path data, extract the location information of the farms involved in the path, and obtain the target farm distribution data; According to the target farm distribution data, combined with the surrounding farm density data, the farm density value in the area is calculated to obtain updated density distribution data; Obtain dynamic change data of the traffic network, extract traffic network structure information related to the target farm, and obtain updated traffic network data; Using a multi-source data fusion algorithm, combining the updated density distribution data and traffic network data, calculating the regional risk index, and obtaining risk distribution data; According to the preset warning threshold, if the regional risk index exceeds the threshold, the region is determined to be a high-risk region, and a list of high-risk regions is obtained; A machine learning classification model is used to input the risk distribution data and a list of high-risk areas, optimize the regional risk prediction results, and obtain the final high-risk area prediction data.
5. The method according to claim 1, characterized in that The process of identifying associations between gene fragments and environmental parameters involves: Obtain pathogen genome data, extract gene fragment feature values, and obtain a genome feature data set; Acquire environmental parameter data, extract environmental parameter feature points, and obtain an environmental parameter feature data set; The gene-environment association model is adopted, the genome feature data set and the environmental parameter feature data set are input, the association between the gene fragments and the environmental parameters is analyzed, and the association feature data set is obtained.
6. The method according to claim 1, characterized in that The process of generating mutation risk analysis results includes: According to a preset correlation threshold, if the correlation feature value exceeds the threshold, it is determined that the gene fragment has a strong correlation with the environmental parameter, and a strongly correlated gene data set is generated; Using a support vector machine classification model, inputting the strongly associated gene data set, optimizing the association features between gene fragments and environmental parameters, and generating an optimized association feature data set; According to the optimized association feature data set, adjusting the gene-environment association model parameters to generate final association model parameters; Using the final association model parameters, inputting new genome data and environmental parameter data, predicting the association characteristics of new gene fragments and environmental parameters, and generating a predicted association characteristic data set; Based on the predicted association feature data set, if the probability of gene segment mutation exceeds a preset risk threshold, it is determined that there is an adaptive mutation risk, and a mutation risk analysis result is generated.
7. The method according to claim 1, characterized in that Steps to identifying unusual patterns of behavior include: Obtain livestock behavior data, clean and preprocess the data, and generate standardized behavior data sets; Using a lightweight LSTM model, inputting the standardized behavior data set, extracting behavior features, and generating a behavior feature data set; According to the behavior feature data set, the behavior feature value is calculated in real time by the edge computing node to generate a real-time feature value data set; If the real-time characteristic value deviates from the preset normal range, it is determined that there is abnormal behavior and an abnormal behavior marking data set is generated.
8. The method according to claim 1, characterized in that The process of obtaining voiceprint classification results includes: Acquire livestock voiceprint data, use a preset noise reduction algorithm to remove environmental noise, and generate noise-reduced voiceprint data; Performing standardization processing on the denoised voiceprint data to generate a standardized voiceprint data set; Using a lightweight CNN model, inputting the standardized voiceprint dataset, extracting voiceprint features, and generating a voiceprint feature dataset; According to the voiceprint feature data set, the voiceprint features are classified in real time through the edge computing node to generate a voiceprint classification result.
9. The method according to claim 1, characterized in that: The process of generating a comprehensive epidemic warning report includes: Obtaining the mutation risk analysis data, calculating the mutation risk value using a preset risk assessment model, and generating a mutation risk data set; Extracting health status indicators and generating a health status data set based on the livestock health status assessment report; The mutation risk dataset is fused with the health status dataset, and a decision tree algorithm is used to evaluate the epidemic risk level to generate a comprehensive epidemic warning report.
10. A livestock disease early warning system based on big data analysis, characterized in that: include: The preliminary path generation module is used to obtain farm distribution data and transportation network data in the geographic information system, and combine it with the spatiotemporal propagation model to generate a preliminary epidemic spread path; A path correction module is used to adjust the preliminary epidemic diffusion path according to the real-time updated farm density data and the dynamic change information of the traffic network to obtain a corrected diffusion path; A risk prediction module is used to predict high-risk areas based on the corrected diffusion path by using a multi-source data fusion early warning model and integrating surrounding farm density and transportation network data; The gene-environment analysis module is used to obtain pathogen genome data and environmental parameter data, compare them through the gene-environment association model, and identify the correlation between gene fragments and environmental parameters; A mutation risk judgment module is used to judge whether there is a risk of adaptive mutation of the strain based on the correlation between the gene fragment and the environmental parameters, and generate a mutation risk analysis result; Behavior analysis module, which uses a lightweight deep learning model and deploys LSTM algorithm on edge computing nodes to analyze livestock behavior data in real time and identify abnormal behavior patterns; The voiceprint classification module is used to deploy the CNN model through edge computing nodes, classify livestock voiceprint data, and obtain voiceprint classification results; A comprehensive early warning module, which is used to generate a livestock health status assessment report by combining the abnormal behavior pattern with the voiceprint classification results; The mutation risk analysis results are integrated with the livestock health status assessment report to generate a comprehensive epidemic warning report.
Citation Information
Cited By
Pig epidemic prevention and control flow regulation system and method based on big data
CN120674103A
Epidemic disease prevention and control method and system based on continuous monitoring of feeding environment
CN121279786A
Wild animal epidemic disease monitoring, prevention and control method and system based on artificial intelligence
CN121281867A
Animal disease intelligent early warning method and system based on cloud side-end cooperation
CN121354933A
Multi-source data fusion-based fowl adenovirus time sequence early warning method and system
CN122314458A