Pig farm breeding management system and method based on industrial big data

By constructing a multi-layered adaptive anomaly detection and causal analysis pig farm management system, the problems of delayed disease early warning and insufficient multi-source data fusion in existing technologies have been solved. This has enabled accurate early warning of pig herd health status and dynamic resource scheduling, thereby improving breeding production efficiency and animal welfare.

CN121707495APending Publication Date: 2026-03-20CHONGQING ANIMAL HUSBANDRY TECH EXTENSION STATION
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511864999.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-11
Publication Date
2026-03-20

AI Technical Summary

Technical Problem

Existing pig farm management systems lag behind in disease early warning, lack sufficient integration of multi-source data, and struggle to provide effective intervention during the subclinical phase, thus limiting the accuracy and foresight of management decisions.

Method used

A pig farm management system based on industrial big data was constructed, including a multi-source heterogeneous data acquisition module, a data fusion and feature engineering module, a subclinical anomaly detection module, a causal inference and root cause tracing module, and a dynamic resource optimization and scheduling module. Through multi-layer adaptive anomaly detection and causal analysis, accurate early warning and dynamic resource scheduling are achieved.

Benefits of technology

It enables very early and high-precision early warning of the health status of pig herds, and can keenly capture subtle abnormal patterns that are difficult to detect by traditional methods, thereby reducing the risk of disease outbreaks and improving the efficiency of breeding production and animal welfare.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121707495A_ABST
    Figure CN121707495A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of industrial big data and intelligent breeding, and particularly discloses a pig farm breeding management system and method based on industrial big data. The system comprises a multi-source heterogeneous data acquisition module, a data fusion and feature engineering module, a subclinical period anomaly detection module, a causal inference and root cause tracing module and a dynamic resource optimization scheduling module. High-order features are extracted through multi-source data fusion, a double-layer adaptive model is adopted to realize sub-clinical period anomaly detection, then root causes are deduced and positioned in combination with causality, and finally an accurate resource scheduling instruction is generated based on multi-objective optimization. According to the invention, extremely early warning and accurate intervention of the health state of the swinery can be realized, the disease outbreak risk is reduced, and the breeding benefit is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of industrial big data and intelligent breeding technology, specifically relating to a pig farm breeding management system and method based on industrial big data. Background Technology

[0002] In modern animal husbandry, especially in large-scale pig farming, achieving refined and intelligent management of the production process is the core objective of improving breeding efficiency and ensuring biosecurity and animal welfare. This involves the collection, analysis, and decision support of massive amounts of data from multiple dimensions, including the breeding environment, animal behavior, and physiological health. Among these, data-driven health early warning and production optimization technologies have become key development directions.

[0003] The pig farm management system based on industrial big data aims to integrate real-time and historical data generated by IoT sensors, automated feeding, and environmental monitoring systems to build a data model. This model enables early warning of pig health status, accurate assessment of growth performance, and optimal allocation of breeding resources. Its fundamental goal is to transform the traditional management model, which relies on manual experience, into a predictable and intervention-oriented digital management model through data fusion and intelligent analysis.

[0004] Existing technologies typically build disease prediction models based on historical epidemic data. These models heavily rely on data from past outbreaks for training and struggle to effectively capture subtle physiological abnormalities in pigs during the subclinical phase. Because early warning signals are delayed, alerts are often triggered only when disease symptoms are obvious and widespread transmission has begun. This leads to lagging control measures and an inability to intervene effectively before infection spreads, resulting in significant economic losses.

[0005] Existing systems have shortcomings in integrating and analyzing the correlation between environmental, feeding, and individual health data, making it difficult to achieve deep fusion and collaborative early warning of cross-dimensional data. This limits the accuracy and foresight of management decisions. Therefore, how to achieve very early and high-precision early warning of pig herd health status and perform dynamic resource scheduling based on multi-source data fusion has become an urgent technical challenge to be solved. Summary of the Invention

[0006] The purpose of this invention is to provide a pig farm breeding management system and method based on industrial big data, so as to solve the problems of delayed disease early warning, insufficient integration of multi-source data, and inability to effectively intervene in the subclinical stage in the existing technology.

[0007] To achieve the above objectives, this invention provides a pig farm management system based on industrial big data. The system includes a multi-source heterogeneous data acquisition module, a data fusion and feature engineering module, a subclinical anomaly detection module, a causal inference and root cause tracing module, and a dynamic resource optimization and scheduling module.

[0008] The multi-source heterogeneous data acquisition module is used to collect raw data streams in real time from various sensors and automated equipment distributed within the physical space of the pig farm, covering environmental dimensions, individual behavior dimensions, physiological indicator dimensions, and production operation dimensions.

[0009] The environmental data specifically includes temperature, humidity, ammonia concentration, carbon dioxide concentration, light intensity, and wind speed in each pigsty unit.

[0010] Individual behavioral data is acquired through inertial measurement units and ultra-wideband positioning modules deployed in the ear tags or collars of pigs. Specifically, it includes the daily movement trajectory, activity level, feeding and drinking frequency, lying time, and social interaction distance of individual pigs.

[0011] Physiological data are acquired through non-contact or implantable sensors, including individual pigs' body surface temperature, respiratory rate, heart rate variability, and specific acoustic characteristics.

[0012] Production operation data is obtained through automated feeding systems, environmental control systems, and manual data entry terminals. Specifically, it includes the precise amount of feed for each pig, feed composition, vaccination records, regrouping records, and operating parameters of ventilation and temperature control equipment in the pigsty.

[0013] The data fusion and feature engineering module is connected to the multi-source heterogeneous data acquisition module. It is used to perform time synchronization, missing value imputation and outlier cleaning on the received raw multi-dimensional data stream, and to construct an aquaculture data cube under a unified spatiotemporal benchmark.

[0014] Furthermore, the module extracts three types of high-order features from the cleaned data cube using a sliding time window mechanism.

[0015] The first category is individual time-series dynamic characteristics, specifically including the mean, variance, autocorrelation coefficient, and trend slope based on first-order difference of the core physiological indicators of individual pigs within a continuous 7-day time window.

[0016] The second category is the spatial distribution characteristics of the group, which specifically includes the real-time spatial aggregation index of all pigs in the same pig house unit, the average distance between individuals calculated based on movement trajectories, and the distribution entropy of activity hotspot areas.

[0017] The third category is cross-dimensional correlation features, which are obtained by calculating the time-lag cross-correlation function between the environmental parameter sequence and the group's average behavioral index sequence. Specifically, this includes the peak value of the correlation coefficient between the change in ammonia concentration and the change in the average lying time of pigs, and its corresponding time lag.

[0018] The subclinical anomaly detection module is connected to the data fusion and feature engineering module. It is used to build and run a two-layer adaptive anomaly detection model based on the extracted high-order features to achieve very early and fine-grained early warning of the health status of pig herds.

[0019] The core of this module lies in its cascaded detection architecture. The first layer is a population baseline deviation detection layer based on multivariate statistical process control.

[0020] This layer establishes a dynamically updated multivariate feature baseline model for each pig house unit. This baseline model is defined by the multivariate Gaussian distribution parameters of the high-order features of the current healthy pig population in the pig house.

[0021] This layer calculates the Mahalanobis distance of each pig's feature vector relative to the dynamic baseline of its pig house unit in real time, and continuously compares this distance with the control upper limit obtained based on historical health data statistics.

[0022] When the Mahalanobis distance of any individual pig continuously exceeds the set control limit for three consecutive detection cycles, the first layer of detection triggers a primary anomaly marker and pushes the individual and its associated feature vector to the second layer of detection.

[0023] The second layer is an individual anomaly pattern recognition layer based on a deep temporal convolutional network.

[0024] This layer receives primary anomalous individuals from the first layer and their complete high-order feature sequences within the most recent 14-day time window.

[0025] This layer incorporates a deep temporal convolutional network model pre-trained on a large amount of historical data labeled with the onset points of subclinical abnormalities.

[0026] The network model takes the aforementioned 14-day feature sequence as input and outputs an anomaly confidence score between 0 and 1, which represents the degree of matching between the pattern contained in the input sequence and the historical subclinical abnormal pattern.

[0027] When the anomaly confidence score is greater than the preset threshold of 0.85, the second layer of detection confirms the subclinical abnormal event and generates a formatted warning signal containing the abnormal individual identifier, the anomaly confidence score, the anomaly pattern type code, and the first deviation time stamp.

[0028] The causal inference and root cause tracing module is connected to the subclinical abnormality detection module and the data fusion and feature engineering module. It is used to automatically start the targeted causal analysis process after receiving the subclinical abnormality warning signal in order to determine the potential root cause most likely to cause the abnormality.

[0029] The process first focuses on the individual under warning and constructs a local causal discovery dataset that includes the individual's own historical data, data from other individuals in the same dormitory during the same period, and environmental and equipment data within the dormitory.

[0030] Furthermore, this module applies a constraint-based causal discovery algorithm, specifically the PC algorithm, to search for and construct a locally causal directed acyclic graph on the dataset.

[0031] In this diagram, nodes represent different data variables, and edges represent potential causal relationships between variables.

[0032] This module identifies the set of precursor variables with the strongest causal effect by analyzing the incoming edges of the directed acyclic graph that point to the core physiological or behavioral abnormality nodes of the abnormal individual.

[0033] Finally, the module outputs a root cause analysis report, which lists potential root causes in descending order of causal effect strength. Root cause types include persistent exceedance of specific environmental parameters, recent changes in feeding formula, close contact with specific infected individuals, or missing immunization records.

[0034] The dynamic resource optimization and scheduling module is connected to the subclinical abnormality detection module and the causal inference and root cause tracing module. It is used to generate and execute precise resource intervention and scheduling instructions based on early warning signals and root cause analysis results.

[0035] This module incorporates a decision engine based on a multi-objective optimization model. The engine's objective function simultaneously minimizes the risk of disease transmission, maximizes the overall welfare of the pig herd, and keeps resource adjustment costs within budget.

[0036] Decision variables include isolation and transfer plans for individuals under warning, environmental parameter adjustment plans for abnormal pig houses, preventive feeding or medication plans for associated pigs, and allocation plans for relevant human resources.

[0037] Each time a new warning signal is received, the module invokes the decision engine to solve for the Pareto optimal scheduling scheme by combining the current pigpen vacancy status, equipment availability, drug inventory, and staff scheduling in real time.

[0038] Subsequently, the module decomposes the solution into a series of specific control instructions with timing and execution object labels, and sends them to the corresponding automated actuators, including intelligent isolation gate controllers, environmental control units, precision feeding stations, and mobile inspection terminals, through standard industrial communication protocols.

[0039] As one embodiment of the present invention, the update mechanism of the dynamic baseline model in the population baseline deviation detection layer based on multivariate statistical process control is as follows:

[0040] Every day at midnight, the system selects data from all individual pigs that have not been marked by any anomaly detection layer in the past 7 days, recalculates the sample mean vector and covariance matrix of their higher-order features, and uses an exponential smoothing method to fuse the newly calculated statistics with the statistics of the old baseline model. The smoothing coefficient is set to 0.1 to achieve a gradual adaptive update of the baseline, thereby tracking the normal feature drift of the pig herd caused by growth stage or seasonal changes.

[0041] As one embodiment of the present invention, the structure of the deep temporal convolutional network model includes 4 temporal convolutional blocks and 2 fully connected layers.

[0042] Each temporal convolutional block consists of a one-dimensional convolutional layer, a batch normalization layer, and a modified linear unit activation function in sequence.

[0043] The kernel sizes of the four convolutional layers are 7, 5, 3, and 3, respectively, and the number of filters are 64, 128, 256, and 256, respectively.

[0044] The network ultimately maps the temporal features into a fixed-length vector through a global average pooling layer, and then outputs the anomaly confidence score through two fully connected layers with 128 and 1 neurons respectively.

[0045] The training loss function for this network is a weighted binary cross-entropy, where the weights of subclinical abnormal samples are set to 5 times the weights of normal samples.

[0046] As one embodiment of the present invention, the execution process of the constraint-based causal discovery algorithm (PC algorithm) is as follows:

[0047] We first assume that there are undirected edges between all variables as the initial graph.

[0048] Secondly, given condition sets of different orders, conditional independence is tested for each pair of variables, and edges that do not meet the independence conditions are gradually removed.

[0049] The conditional independence test used was a partial correlation test based on the Gaussian hypothesis, with a significance level set at 0.01.

[0050] Finally, based on the V-structure rule and the direction propagation rule, the causal direction of the remaining undirected edges is determined on the obtained skeleton graph, thus forming a locally causal directed acyclic graph.

[0051] As one embodiment of the present invention, the mathematical expression of the multi-objective optimization model includes three objectives and four types of constraints.

[0052] The first objective is to minimize the risk of disease transmission, which is quantified by weighted summation of the estimated infectivity of early warning individuals and their highly associated contacts identified in the causal graph.

[0053] The second objective is to maximize the overall welfare of the pig herd, which is quantified by a comprehensive score that ensures the environmental comfort and feeding satisfaction of non-abnormal pigs.

[0054] The third objective is to minimize resource adjustment costs, including equipment energy consumption, drug consumption, and labor costs.

[0055] The four types of constraints are pig house capacity constraints, maximum equipment load constraints, drug inventory constraints, and personnel skill matching constraints.

[0056] The model is solved using a non-dominated sorting genetic algorithm with an elitist strategy, with a population size of 100 and a maximum number of generations of evolution of 200.

[0057] As one embodiment of the present invention, the system further includes a digital twin simulation and strategy evaluation module.

[0058] This module is invoked after the dynamic resource optimization and scheduling module generates the preliminary scheduling plan and before it is actually issued and executed.

[0059] This module constructs a high-fidelity virtual simulation environment based on the accurate 3D model of the current pig farm and the status data of all entities.

[0060] This module injects the initial scheduling plan into the simulation environment and simulates the dynamics of the pig herd, environmental changes, and disease transmission process over the next 72 hours at an acceleration rate.

[0061] This module outputs key assessment metrics, including the predicted curve of changes in the number of abnormal individuals, the average stress level of the population, and the total amount of resources consumed.

[0062] The decision engine of the dynamic resource optimization scheduling module will receive this simulation evaluation result. If the evaluation index does not reach the preset safety and efficiency thresholds, the decision engine will readjust the optimization weights or search space and iteratively generate a new scheduling scheme until it passes the simulation evaluation.

[0063] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0064] 1. This invention systematically solves the core contradiction of delayed health early warning and extensive decision-making in pig farming by constructing a complete technical closed loop from multi-source data fusion, high-order feature extraction, two-layer adaptive anomaly detection, causal root cause tracing to dynamic resource scheduling.

[0065] 2. The system extracts individual temporal dynamics, group spatial distribution, and cross-dimensional correlation features through the data fusion and feature engineering modules, laying a data foundation for the identification of deep anomaly patterns.

[0066] 3. The subclinical abnormality detection module adopts a two-layer architecture that combines population deviation detection based on multivariate statistical process control with individual pattern recognition based on deep temporal convolutional networks. It realizes progressive analysis from population baseline screening to individual specific pattern confirmation, and can keenly capture subtle abnormal patterns that are difficult to detect by traditional methods before the appearance of clinical symptoms, thus significantly advancing the warning time.

[0067] 4. The causal inference and root cause tracing module introduces a causal discovery algorithm, upgrading correlation analysis to causal analysis. This can accurately locate the root cause most likely to cause the abnormality from many related factors, providing a clear target for subsequent intervention and avoiding misjudgments that may occur based on correlation.

[0068] 5. The dynamic resource optimization and scheduling module, based on multi-objective optimization and digital twin simulation, transforms early warning information and root cause analysis into a set of executable scheduling instructions that consider the balance of risk, welfare and cost. This realizes automated closed-loop management from perception to decision-making to execution, improving the accuracy of resource utilization and the timeliness of emergency response.

[0069] 6. The various modules of the entire system work closely together to form a new paradigm of intelligent aquaculture management characterized by foresight, precision and adaptability, which reduces the risk of disease outbreaks and improves the efficiency of aquaculture production and the level of animal welfare. Attached Figure Description

[0070] Figure 1 This is a schematic diagram of the overall technical solution architecture of the pig farm breeding management system based on industrial big data proposed in this invention;

[0071] Figure 2 This is a schematic diagram of the core principle framework of the two-layer adaptive anomaly detection module in the subclinical phase of this invention.

[0072] Figure 3 This is a flowchart illustrating the logical process of constructing aquaculture data cubes and extracting high-order features in the data fusion and feature engineering module of this invention.

[0073] Figure 4 This is a flowchart illustrating the logical flow of the causal inference and root cause tracing module in this invention for targeted causal analysis.

[0074] Figure 5 This is a schematic diagram of the decision-making and execution framework of the dynamic resource optimization and scheduling module in this invention, which is based on multi-objective optimization and digital twin simulation. Detailed Implementation

[0075] This invention provides a pig farm management system based on industrial big data. Please refer to the appendix. Figures 1 to 5This system constructs a closed-loop technical architecture encompassing data perception, intelligent analysis, and precise execution. Its core lies in achieving very early warning and precise intervention of pig herd health status through the collaborative operation of multiple modules. Specifically, the system includes a multi-source heterogeneous data acquisition module, a data fusion and feature engineering module, a subclinical anomaly detection module, a causal inference and root cause tracing module, and a dynamic resource optimization and scheduling module. These modules are connected and interact according to a strict logical sequence and data flow relationship, collectively forming an intelligent management entity with adaptive and self-optimizing capabilities.

[0076] The multi-source heterogeneous data acquisition module acts as the sensory nerve ending of the system, responsible for capturing raw data streams in real time and continuously from every corner of the pig farm's physical space.

[0077] The deployment of this module covers four major data domains: environmental dimension, individual behavior dimension, physiological indicator dimension, and production operation dimension.

[0078] In terms of the environment, the module integrates sensor arrays deployed on the ceiling and side walls inside each pigsty unit.

[0079] The array includes a high-precision digital temperature and humidity sensor, an electrochemical ammonia and carbon dioxide concentration sensor, a silicon photodiode light intensity sensor, and a hot-wire anemometer.

[0080] These sensors operate at a sampling frequency of once per second and, through wired or wireless IoT protocols, aggregate real-time readings of temperature, humidity, ammonia concentration, carbon dioxide concentration, light intensity, and wind speed from each pigsty unit to a regional gateway.

[0081] After the regional gateway performs initial packaging and timestamping of the data, it uploads it to the system data center via the industrial Ethernet backbone network.

[0082] At the individual behavior level, this module relies on smart ear tags or collars worn by each pig in the pen. The smart device integrates a nine-axis inertial measurement unit and an ultra-wideband positioning module.

[0083] The inertial measurement unit continuously acquires triaxial acceleration, triaxial angular velocity and triaxial magnetometer data at a frequency of 50 Hz.

[0084] The raw data is preprocessed by the built-in microprocessor on the device and a threshold-based activity recognition algorithm is used to calculate and output the instantaneous activity intensity, posture angle and behavior classification results of individual pigs in real time. The behavior classification includes basic states such as walking, running, eating, drinking and lying down.

[0085] The ultra-wideband positioning module communicates with at least four positioning base stations deployed in the pigsty. By measuring the signal arrival time difference, it calculates the two-dimensional plane coordinates of individual pigs at an update rate of 10 times per second, with an accuracy of up to 0.3 meters.

[0086] Behavioral and location data are integrated within smart devices to generate data packets containing timestamps, individual identifiers, behavioral status, activity levels, and precise location coordinates. These data packets are then periodically reported to the data aggregation node within the pigsty via Bluetooth Low Energy or a dedicated radio frequency channel.

[0087] In terms of physiological indicators, this module employs a sensing strategy that combines non-contact and minimally invasive implantation. For body surface temperature and respiratory rate, an infrared thermal imaging camera and millimeter-wave radar are installed above the pigs' resting area.

[0088] Infrared thermal imaging cameras capture thermal images of pigs' bodies at a rate of one frame per minute, and image processing algorithms extract temperature values ​​for specific areas such as the base of the ears and eyes.

[0089] Millimeter-wave radar transmits frequency-modulated continuous waves and analyzes the phase changes of the reflected signals to detect the micro-movements in the pig's chest cavity in a non-contact manner, thereby calculating the respiratory rate with an accuracy of up to 2 times per minute.

[0090] For deeper physiological parameters such as heart rate variability, a biocompatible encapsulated electrocardiogram (ECG) sensor can be implanted in the core breeding population or specific monitoring individuals. This sensor collects ECG signals at a sampling rate of 256 Hz and wirelessly transmits the data to an in-house receiver via subcutaneous near-field communication.

[0091] Directional microphone arrays are deployed at key locations in the pigsty to continuously collect ambient sounds. Through voiceprint recognition and event detection algorithms, specific acoustic feature events such as coughing, sneezing, and abnormal grunting of pigs are separated and identified.

[0092] In terms of production operations, this module is deeply integrated with the existing automation system of the pig farm.

[0093] The automated feeding system records the feed trough number, feeding time, feed formula code, and actual consumption amount obtained through weighing sensors each time feed is dispensed, and links this data with the feeding records of individual pigs wearing RFID ear tags.

[0094] The environmental control system provides real-time operating status, set parameters, and actual power data for each device, including fans, wet curtains, heaters, and air inlets.

[0095] The system provides a dedicated manual data entry terminal for farmers or veterinarians to record critical events that cannot be automatically collected, including vaccination records, regrouping records, treatment records, and external observation notes for each pig.

[0096] All production operation data is accompanied by precise timestamps and operator identifiers, and is imported into the system through a standard data interface.

[0097] The data fusion and feature engineering module is connected to the multi-source heterogeneous data acquisition module and is the core hub for the system to perform data governance and knowledge extraction.

[0098] Please refer to the attached document. Figure 3 This module first performs a rigorous data preprocessing pipeline on the incoming raw multi-dimensional data stream.

[0099] The first step is time synchronization. Because data acquisition frequencies and transmission delays vary from source to source, the module maintains a high-precision network time protocol clock server.

[0100] Before entering the processing queue, the timestamps of all incoming data are uniformly calibrated to the server's time base, with the error controlled within the millisecond level.

[0101] For non-equal interval data, cubic spline interpolation is used to resample them into a uniform equal interval time series, with the basic time resolution set to 1 minute.

[0102] The second step is missing value imputation and outlier cleaning.

[0103] The module incorporates a multi-layered processing strategy based on data characteristics.

[0104] For environmental sensor data, if there is a short-term missing data, linear interpolation or forward filling is used;

[0105] If the data is missing for more than 10 minutes, it will be marked as invalid and a device check alarm will be triggered.

[0106] For pig behavior and physiological data, since individuals may temporarily leave the monitoring area, a method of imputation based on the mean of data from the same period of individuals is used.

[0107] Outlier cleaning employs a dual filtering approach based on statistical distribution and business rules.

[0108] For example, if a pig's body temperature reading is less than 35 degrees Celsius or greater than 41 degrees Celsius for three consecutive cycles, it is considered a sensor malfunction and is discarded.

[0109] If a pig consumes more than 5% of its body weight in a single feeding, it is considered an abnormal data point and requires verification in conjunction with video recordings.

[0110] After preprocessing, the module constructs an aquaculture data cube under a unified spatiotemporal reference.

[0111] This data cube is a multidimensional array structure with three core dimensions: time, space, and metrics. The time dimension is aggregated by day, week, and month, with the smallest granularity being the minute.

[0112] The spatial dimension is organized according to the pig farm, building, pen, and individual level.

[0113] The indicator dimensions encompass all the cleaned original indicators and the subsequently derived feature indicators.

[0114] The data cube is stored in a time-series database and supports efficient multi-dimensional slicing, dicing, and roll-up / drill-down operations.

[0115] Furthermore, this module dynamically extracts three types of high-order features from the cleaned data cube using a sliding time window mechanism to serve advanced analysis.

[0116] The step size of the sliding window is set to 1 day, and the window length is dynamically adjusted according to the feature type.

[0117] The first category is individual time-series dynamic characteristics. For each pig, the module selects its core physiological and behavioral indicators, including body surface temperature, respiratory rate, daily activity level, and feeding duration, and calculates them within a continuous 7-day time window.

[0118] The extracted features include not only the mean and variance of the sequences within the window, but also focus on characterizing their dynamic evolution patterns.

[0119] Specifically, its autocorrelation coefficient is calculated, with lag order ranging from 1 to 3, to quantify the short-term memory of the sequence itself.

[0120] Simultaneously, the trend slope based on first-order difference is calculated, i.e., a linear fit is performed on the sequence within the window, and the resulting slope value is used to determine whether the indicator is in an upward, downward, or stable trend. For the indicator sequence within the window... Its first-order difference sequence is trend slope Fitting by least squares method get, Indicator Series The Middle The index value corresponding to each position ( The range of values ​​for is 1, 2, ..., T). Indicator Series The location index of the data points in the middle, It is the intercept term of the least squares fitted line, and it is the constant term in the fitting formula.

[0121] The second category is the spatial distribution characteristics of the population.

[0122] For each pigsty unit, the module calculates the real-time spatial distribution index of all pigs on site based on the snapshot data every minute.

[0123] The spatial clustering index is obtained by calculating the reciprocal of the average area of ​​all triangles after the Delaunay triangulation. The larger the value, the more densely the pigs are clustered.

[0124] The average distance between individuals is calculated by directly taking the mean of the Euclidean distances between all pairs of pigs.

[0125] The calculation of the activity hotspot distribution entropy first involves dividing the pigsty into a grid, counting the percentage of time each pig spends in each grid, and then calculating the Shannon entropy of this probability distribution. A high entropy value indicates that the pigs' activities are dispersed, while a low entropy value indicates that the activities are concentrated in a few areas.

[0126] The third category is cross-dimensional correlation features. These features aim to capture the potential correlation between environmental factors and group behavior. The module selects key environmental parameter sequences, such as ammonia concentration and temperature, and sequences of average group behavioral indicators, such as average lying time and average activity level. By calculating the time-lag cross-correlation function between these two sequences, it seeks the time-lag point with the strongest correlation. Specifically, for the environmental sequence... and behavioral sequence In time lag Cross-correlation coefficients under The calculation formula is:

[0127] ;

[0128] The total number of data points for the environmental sequence and the behavioral sequence. For environmental sequence The mean, Behavioral sequence The mean, For behavioral sequences In the The specific value of the time. The module's calculation time delay ranges from -12 hours to +12 hours. Record its peak value and the corresponding time delay For example, it might be found that changes in ammonia concentration lag behind changes in the average lying time of pigs by 6 hours, with a negative correlation between the two, and the peak correlation coefficient reaching -0.7. Such characteristics provide a quantitative basis for understanding the lagged effects of environmental stress on group behavior.

[0129] The subclinical abnormality detection module is connected to the data fusion and feature engineering module and is the core of the intelligent early warning system.

[0130] Please refer to the attached document. Figure 2This module builds and runs a two-layer adaptive anomaly detection model, whose design philosophy is to first perform broad-spectrum population baseline screening and then perform in-depth individual pattern identification.

[0131] The first layer is a population baseline deviation detection layer based on multivariate statistical process control.

[0132] This layer maintains a dynamically updated multivariate feature baseline model for each pig house unit.

[0133] The baseline model is essentially a multivariate Gaussian distribution, and its parameters are the mean vector. With covariance matrix It is defined by the higher-order characteristics exhibited by the current healthy pig population in the pigsty.

[0134] Healthy pigs are identified as those that have not been flagged by any abnormal detection layer in the past 72 hours.

[0135] The module initiates the baseline update process at 2 AM daily.

[0136] During the update, the system selects high-order feature data from all healthy individuals over the past 7 days and calculates their sample mean vector. With the sample covariance matrix .

[0137] To avoid drastic changes in the baseline due to daily data fluctuations, an exponential smoothing method is used to merge the new statistics with the old baseline parameters. The smoothing update formula is as follows:

[0138] ;

[0139] ;

[0140] This is the updated baseline mean vector after exponential smoothing fusion; This is the vector of the old baseline mean before the update; This is the updated baseline covariance matrix after exponential smoothing fusion; The old baseline covariance matrix is ​​shown before the update; the smoothing coefficient of 0.1 ensures that the baseline can progressively track the normal characteristic drift of the pig herd due to growth and seasonal changes, while maintaining sufficient sensitivity to sudden anomalies.

[0141] During the real-time detection phase, the module calculates the current high-order feature vector x for each pig in the pigsty every 4 hours.

[0142] Then, the eigenvector is calculated relative to the dynamic baseline of its pigsty unit. Mahalanobis distance :

[0143] ;

[0144] Mahalanobis distance takes into account the correlation between features and can more accurately measure the degree to which an individual deviates from the group norm.

[0145] The system calculates the statistical distribution of Mahalanobis distance based on 90 days of historical health data and sets its 95th percentile as the upper limit of control.

[0146] The real-time detection logic is as follows:

[0147] The first layer of detection is triggered when the Mahalanobis distance of any individual pig is continuously greater than the control limit for three consecutive detection cycles.

[0148] At this point, the system generates a primary anomaly marker, which includes the anomaly individual identifier, trigger time, Mahalanobis distance value of the deviation, and the main characteristic items that caused the deviation.

[0149] Subsequently, the individual and its complete high-order feature sequence within the most recent 14-day time window were packaged into a data packet and pushed to the second-layer detection for in-depth analysis.

[0150] The second layer is an individual anomaly pattern recognition layer based on a deep temporal convolutional network.

[0151] This layer receives primary exception data packets from the first layer.

[0152] Its core is a pre-trained deep temporal convolutional network model, which is specifically designed to identify subtle temporal patterns associated with subclinical diseases.

[0153] The network model has a carefully designed structure, consisting of four temporal convolutional blocks and two fully connected layers.

[0154] Each temporal convolutional block consists of a one-dimensional convolutional layer, a batch normalization layer, and a modified linear unit activation function in sequence.

[0155] The kernel sizes of the four convolutional layers are 7, 5, 3, and 3, respectively. This design allows the network to first capture patterns over a longer time span and then gradually focus on more refined short-term fluctuations.

[0156] The number of filters is 64, 128, 256, and 256 respectively. The gradually increasing number of feature maps enables the network to learn more complex feature representations.

[0157] The network ultimately maps the temporal features into fixed-length vectors through a global average pooling layer, which are then processed by two fully connected layers. Finally, the fully connected layers use the Sigmoid activation function to output a scalar between 0 and 1, which is the anomaly confidence score.

[0158] The network was trained on a large amount of historical data labeled with the onset points of subclinical abnormalities. The loss function used during training was weighted binary cross-entropy, in which samples retrospectively diagnosed as subclinical abnormalities by veterinarians were given higher weights, five times the weight of normal samples. This forced the network to pay more attention to a few but critical abnormal patterns.

[0159] During the inference phase, when the network calculates an anomaly confidence score greater than the preset threshold of 0.85 for the input 14-day feature sequence, the second layer detection confirms the subclinical abnormal event.

[0160] At this point, the system generates a formatted advanced warning signal.

[0161] This signal not only includes the identifier of the abnormal individual and the confidence level of the abnormality, but also the abnormal pattern type code. The code points to the typical abnormal category clustered by the activation patterns of the intermediate layer of the network, such as the latency pattern of the respiratory system and the latency pattern of the digestive system.

[0162] At the same time, the signal records the first deviation timestamp, that is, the time when the first layer of detection was first triggered, providing a basis for tracing the origin of the anomaly.

[0163] The causal inference and root cause tracing module is connected to the subclinical abnormality detection module and the data fusion and feature engineering module.

[0164] This module is automatically activated upon receiving a pre-existing abnormality warning signal in the subclinical phase. Its mission is to locate the most likely root cause of the abnormality from a massive number of related factors, thus advancing the analysis from "what" to "why".

[0165] Please refer to the attached document. Figure 4 This module initiates the targeted causal analysis process.

[0166] The first step in the process is to build a local causal discovery dataset.

[0167] The module focuses on the individual being warned, defining the spatiotemporal scope of the analysis, typically the pigsty unit where the individual resides, with a time range from 7 days before the anomaly occurred to the present.

[0168] The dataset includes three types of data: all historical time-series data of the individual being alerted;

[0169] High-order trait data and environmental response data of all other individuals in the same pig house at the same time period;

[0170] All environmental sensor data and equipment operation log data within the pigsty.

[0171] All variables were standardized to eliminate the influence of dimensions.

[0172] The second step of the process is to apply a constraint-based causal discovery algorithm, specifically the PC algorithm in this embodiment, to search for and construct a locally causal directed acyclic graph on the dataset.

[0173] The execution of the PC algorithm begins with a completely undirected graph, where each node represents a variable in the dataset.

[0174] The core of the algorithm is to gradually remove non-existent edges from the graph through a series of conditional independence checks. The specific process is as follows:

[0175] First, examine the independence of each pair of variables under the zero-order condition set, and then remove the edges between independent variables.

[0176] Next, test the independence of the remaining variable pairs under the first-order condition set, and so on, gradually increasing the order of the condition set.

[0177] The conditional independence test used in this system is a partial correlation test based on the Gaussian hypothesis, with a significance level set at 0.01.

[0178] After all unnecessary edges are removed, a skeleton graph describing the dependencies between variables is obtained.

[0179] The third step in the process is to determine the causal direction. On the obtained skeleton diagram, orientation is determined according to the V-structure rules:

[0180] For a triple of the form A, B, and C, if A is connected to B, B is connected to C, but A is not connected to C, and B is not in any condition set of A and C, then the orientation is A to B and C to B. Subsequently, the direction propagation rule is applied to avoid the formation of new V-structures or cycles, thereby determining the direction for the remaining undirected edges, ultimately forming a locally causal directed acyclic graph.

[0181] In this directed acyclic graph, incoming edges pointing to core abnormal nodes of the individual under warning, such as nodes with abnormally elevated respiratory rate or abnormally decreased activity level, represent potential causal paths.

[0182] The module quantifies the contribution of each precursor variable to the outlier node by calculating the strength of the causal effect, for example, by using regression coefficients or the average causal effect calculated by the intervention.

[0183] Finally, the module outputs a structured root cause analysis report. The report lists the identified potential root causes in descending order of causal effect strength.

[0184] Root cause types are categorized into several main types:

[0185] Specific environmental parameters continued to exceed the standard;

[0186] Feeding formula has recently been changed;

[0187] Close contact with a specific infected individual, determined by the duration and distance of contact calculated based on ultra-wideband positioning data;

[0188] The individual had missing immunization records, and a system comparison of the immunization schedule revealed that the individual had not received a certain vaccine as planned.

[0189] The dynamic resource optimization and scheduling module, connected to the subclinical abnormality detection module and the causal inference and root cause tracing module, is the final key link in translating analytical decisions into practical actions.

[0190] Please refer to the attached document. Figure 5 This module incorporates a decision engine based on a multi-objective optimization model and introduces digital twin simulation for strategy pre-evaluation to ensure the scientific nature and robustness of the scheduling scheme.

[0191] The core of the decision engine is a multi-objective optimization model.

[0192] The model contains three objectives that need to be optimized simultaneously and four types of hard constraints that must be satisfied.

[0193] The first objective is to minimize the risk of disease transmission. This risk is quantified by a transmission risk index, which is a weighted sum of the estimated infectivity of the alerted individual and the highly associated contacts identified in the causal graph.

[0194] The estimated infectivity is determined by the individual's outlier confidence score, the pathogen transmission characteristics model, and the intensity of contact.

[0195] The second objective is to maximize the overall welfare of the pig herd.

[0196] This level is quantified by a welfare scoring function that takes into account both the environmental comfort and feeding satisfaction of non-abnormal pigs.

[0197] Environmental comfort is calculated based on temperature and humidity indices and ammonia concentration, while feeding satisfaction is calculated based on the ratio of actual feed intake to standard requirements.

[0198] The third objective is to minimize resource adjustment costs. This is an economic objective, including the additional energy consumption incurred due to adjusting environmental equipment, the cost of preventative medication or nutritional supplements, and the labor costs incurred due to additional manpower deployment.

[0199] The four types of constraints are as follows:

[0200] Pigsty capacity constraints, meaning that after isolation and relocation, the number of pigs in each pen is less than the design limit;

[0201] Maximum load constraints on equipment, such as maximum fan speed and maximum heater power;

[0202] Drug inventory constraints mean that the planned drug usage in the scheduling plan is less than the current warehouse inventory.

[0203] Personnel skill matching constraint, that is, the assigned inspection or treatment tasks must be performed by personnel with the corresponding qualifications.

[0204] This multi-objective optimization model is solved using a non-dominated sorting genetic algorithm with an elitist strategy.

[0205] The algorithm initializes a random population of size 100, with each individual representing a complete scheduling scheme code.

[0206] Through iterative evolution via selection, crossover, and mutation, the maximum number of generations is 200.

[0207] The algorithm ultimately outputs a Pareto optimal solution set, and the decision engine then selects the final scheduling scheme from the solution set based on the current management strategy preferences, such as whether to focus more on risk control or cost control.

[0208] Before the actual implementation of the plan, the system calls the digital twin simulation and strategy evaluation module for a pre-run.

[0209] This module constructs a high-fidelity virtual simulation environment based on a precise 3D building information model of the pig farm, the real-time status attributes of each pig, and the performance parameters of all equipment.

[0210] The scheduling scheme generated by the decision engine, including isolation instructions, environmental setting adjustments, and feeding scheme changes, is injected into the simulation environment as input.

[0211] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.

[0212] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A pig farm management system based on industrial big data, characterized in that, include: The multi-source heterogeneous data acquisition module is used to collect raw data streams in real time from various sensors and automated equipment distributed in the physical space of the pig farm, covering environmental dimensions, individual behavior dimensions, physiological indicator dimensions, and production operation dimensions, namely environmental dimension data, individual behavior dimension data, physiological indicator dimension data, and production operation dimension data. The data fusion and feature engineering module is connected to the multi-source heterogeneous data acquisition module. It is used to perform time synchronization, missing value imputation and outlier cleaning on the received raw multi-dimensional data stream, and to construct an aquaculture data cube under a unified spatiotemporal benchmark. The subclinical anomaly detection module is connected to the data fusion and feature engineering module and is used to construct and run a two-layer adaptive anomaly detection model based on the extracted high-order features. The causal inference and root cause tracing module is connected to the subclinical abnormality detection module and the data fusion and feature engineering module, and is used to automatically start the targeted causal analysis process after receiving the subclinical abnormality warning signal; The dynamic resource optimization and scheduling module is connected to the subclinical abnormality detection module and the causal inference and root cause tracing module. It is used to generate and execute precise resource intervention and scheduling instructions based on the early warning signal and the root cause analysis results.

2. The pig farm management system based on industrial big data according to claim 1, characterized in that, The environmental data includes temperature, humidity, ammonia concentration, carbon dioxide concentration, light intensity, and wind speed in each pigsty unit. The individual behavior dimension data is obtained through inertial measurement units and ultra-wideband positioning modules deployed in the ear tags or collars of pigs, including the daily movement trajectory, activity level, feeding and drinking frequency, lying time, and social interaction distance of individual pigs. The physiological indicators data are acquired through non-contact or implantable sensors, including the individual pig's body surface temperature, respiratory rate, heart rate variability, and specific acoustic characteristics. The production operation data is obtained through automated feeding systems, environmental control systems, and manual input terminals, including the precise feeding amount per pig, feed composition, vaccination records, regrouping records, and operating parameters of ventilation and temperature control equipment in the pigsty.

3. The pig farm management system based on industrial big data according to claim 2, characterized in that, The data fusion and feature engineering module is also used to extract three types of high-order features from the cleaned data cube through a sliding time window mechanism; The first category is individual time-series dynamic characteristics, including the mean, variance, autocorrelation coefficient, and trend slope based on first-order difference of the core physiological indicators of individual pigs within a continuous 7-day time window; The second category is the spatial distribution characteristics of the group, including the real-time spatial aggregation index of all pigs in the same pig house unit, the average distance between individuals calculated based on movement trajectories, and the distribution entropy of activity hotspot areas. The third category is cross-dimensional correlation features, which are obtained by calculating the time-lag cross-correlation function between the environmental parameter sequence and the group's average behavioral index sequence. This includes the peak correlation coefficient of ammonia concentration change lagging behind the change in the average lying time of pigs and its corresponding time lag.

4. The pig farm management system based on industrial big data according to claim 3, characterized in that, The dual-layer adaptive anomaly detection model includes a first layer based on multivariate statistical process control for population baseline deviation detection and a second layer based on deep temporal convolutional networks for individual anomaly pattern recognition. The first layer establishes a dynamically updated multivariate feature baseline model for each pig house unit. This baseline model is defined by the multivariate Gaussian distribution parameters of the high-order features of the current healthy pig population in the pig house. The first layer calculates the Mahalanobis distance of the feature vector of each pig relative to the dynamic baseline of its pig house unit in real time, and continuously compares this distance with the control upper limit obtained based on historical health data statistics; When the Mahalanobis distance of any individual pig is continuously greater than the set control limit for three consecutive detection cycles, the first layer triggers a primary anomaly marker and pushes the individual and its associated feature vector to the second layer. The second layer receives primary anomalous individuals from the first layer and their complete high-order feature sequences within the most recent 14-day time window; The second layer incorporates a deep temporal convolutional network model pre-trained on a large amount of historical data labeled with the onset points of subclinical abnormalities; The deep temporal convolutional network model takes the aforementioned 14-day feature sequence as input and outputs an anomaly confidence score between 0 and 1. When the abnormal confidence score is greater than a preset threshold, the second layer confirms the subclinical abnormal event and generates a formatted warning signal containing the abnormal individual identifier, abnormal confidence score, abnormal pattern type code, and first deviation time stamp.

5. A pig farm management system based on industrial big data according to claim 4, characterized in that, The targeted causal analysis process first focuses on the individual under warning and constructs a local causal discovery dataset that includes the individual's own historical data, concurrent data of other individuals in the same dormitory, and environmental and equipment data within the dormitory. The causal inference and root cause tracing module applies a constraint-based causal discovery algorithm to search for and construct a local causal directed acyclic graph on the dataset. The causal inference and root cause tracing module analyzes the incoming edges of the directed acyclic graph pointing to the core physiological or behavioral abnormality nodes of the abnormal individual, identifies the set of precursor variables with the strongest causal effect, and outputs a root cause analysis report that lists potential root causes in descending order of causal effect strength.

6. A pig farm management system based on industrial big data according to claim 5, characterized in that, The dynamic resource optimization and scheduling module embeds a decision engine based on a multi-objective optimization model; The optimization objective function of the decision engine simultaneously minimizes the risk of disease transmission, maximizes the overall welfare of the pig herd, and keeps resource adjustment costs within budget. The decision variables include the isolation and transfer plan for individuals under warning, the environmental parameter adjustment plan for abnormal pig houses, the preventive feeding or medication plan for related pigs, and the allocation plan for relevant human resources. Upon receiving a new warning signal, the dynamic resource optimization scheduling module invokes the decision engine to solve for the Pareto optimal scheduling scheme by combining the current vacancy status of pigsties, equipment availability, drug inventory, and real-time constraints of personnel scheduling. The scheme is then decomposed into a series of specific control instructions with timing and execution object labels, which are then sent to the corresponding automated actuators via standard industrial communication protocols.

7. A pig farm management system based on industrial big data according to claim 6, characterized in that, The update mechanism of the dynamic baseline model in the population baseline deviation detection layer based on multivariate statistical process control is as follows: Every day at midnight, the system selects data from all individual pigs that have not been marked by any anomaly detection layer in the past 7 days, recalculates the sample mean vector and covariance matrix of their higher-order features, and uses the exponential smoothing method to fuse the newly calculated statistics with the statistics of the old baseline model, thereby achieving a gradual adaptive update of the baseline.

8. A pig farm management system based on industrial big data according to claim 7, characterized in that, The constraint-based causal discovery algorithm is the PC algorithm, and its execution process is as follows: First, we assume that there are undirected edges between all variables as the initial graph; Secondly, given condition sets of different orders, conditional independence is tested for each pair of variables, and edges that do not meet the independence conditions are gradually removed. The conditional independence test method used was the partial correlation test based on the Gaussian hypothesis; Finally, based on the V-structure rule and the direction propagation rule, the causal direction of the remaining undirected edges is determined on the obtained skeleton graph, thus forming a locally causal directed acyclic graph.

9. A pig farm management system based on industrial big data according to claim 8, characterized in that, The system also includes a digital twin simulation and strategy evaluation module; The digital twin simulation and strategy evaluation module is invoked after the dynamic resource optimization and scheduling module generates a preliminary scheduling plan and before it is actually issued and executed. The digital twin simulation and strategy evaluation module constructs a high-fidelity virtual simulation environment based on the accurate 3D model of the current pig farm and the status data of all entities. The digital twin simulation and strategy evaluation module injects the preliminary scheduling scheme into the simulation environment and simulates the dynamics of the pig herd, environmental changes and disease transmission process in the future at an acceleration rate. The digital twin simulation and strategy evaluation module outputs key evaluation indicators, including the predicted curve of the number of abnormal individuals, the average stress level of the group, and the total amount of resources consumed.

10. A pig farm breeding management method based on industrial big data, characterized in that, The pig farm management system based on industrial big data, as described in any one of claims 1 to 9, is used to implement pig farm management.

Citation Information

Cited By

  • Breeding robot operation data traceability and operation analysis system and method

    CN122048395A