LSTM-CNN pig attitude classification method based on attitude sensor data
Through the LSTM-CNN pig pose classification method based on attitude sensor data, the problem of low manual observation efficiency and difficult to identify occlusion problems in pig health monitoring is solved, and accurate recognition and efficient monitoring of pig poses are achieved.
Patent Information
- Application Number
- CN202510303819.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-14
- Publication Date
- 2025-06-27
AI Technical Summary
The prior art relies on manual observation in pig health monitoring, resulting in untimely abnormal discovery and low work efficiency. In the scene of gathering pigs, the computer vision-based pose recognition model is difficult to accurately identify due to occlusion problems.
The LSTM-CNN pig pose classification method based on attitude sensor data is adopted, and the acceleration and gyroscope data of pigs are collected through attitude sensors, combined with sliding window feature extraction algorithm and time domain analysis method, statistical features are extracted, and posture recognition is performed through deep learning models.
It realizes accurate identification of four postures of pigs: standing, lying, eating and physical movement, avoids occlusion problems, improves identification accuracy and work efficiency, and is suitable for complex environments in group breeding scenarios.
Smart Images

Figure CN120220233A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of computer science and artificial intelligence, and particularly relates to a method for classifying pig postures based on attitude sensor data using LSTM-CNN. Background Art
[0002] The health of pigs is one of the important challenges faced by the livestock breeding industry. Previously, the means of monitoring pig health mainly relied on manual observation, which could lead to untimely detection of pig abnormalities, low work efficiency. The behavioral performance of pigs in the pig farm is a key indicator reflecting their welfare and health status, and has a direct impact on the economic benefits of the pig farm. In animal behavior monitoring technology, compared with traditional manual monitoring, using a small and light six-axis attitude sensor to collect pig posture data, and then using deep learning technology to build a model for pig behavior classification training and recognition, demonstrates the advantages of low cost and high efficiency, can continuously provide valuable behavior information, provide strong support for pig farm management, help improve the automated monitoring level and personnel work efficiency of breeding enterprises, and reduce the breeding costs of enterprises. Summary of the Invention
[0003] The purpose of the present invention is to provide a method for classifying pig postures based on attitude sensor data using LSTM-CNN to achieve posture classification and recognition of pigs in free stalls. The specific scheme is as follows:
[0004] S1. Collect the motion data of the attitude sensor of the pig, including acceleration and gyroscope readings, and establish a pig attitude sensor data set;
[0005] S2. Based on the sliding window feature extraction algorithm, use time domain analysis method to perform separate feature extraction on each dimension of features in the data set to obtain the final feature data set;
[0006] S3. Standardize the attitude sensor feature data set to eliminate the influence brought by different sensor dimensions and ensure the consistency of data in model training;
[0007] S4. Construct a deep learning model that combines a convolutional neural network (CNN) and a long short-term memory network (LSTM). CNN is used to extract local features in time series data, while LSTM is used to capture long-term dependencies in time series;
[0008] S5. Divide the training set and the test set from the pig posture recognition attitude sensor feature data set, use the training set to train the dual-stream LSTM-CNN pig posture recognition model, use the test set to test the model performance, and finally screen the best performance model by adjusting the parameters.
[0009] Preferably, the specific process of step S1 is as follows:
[0010] S11. Fix the device under the pig's neck with a flexible band using an attitude sensor, facilitating adjustment of the flexible band according to actual conditions at any time to ensure that the sensor is always in the correct position.
[0011] S12. Fix the camera above the pigsty to capture video images of the pig's movement from a top view.
[0012] Preferably, the specific process of step S2 is as follows:
[0013] S21. Determine the size and step length of the sliding window. The size of the sliding window is determined according to the time scale of the pig's behavior characteristics, and the step length is set according to the data acquisition frequency and window size.
[0014] S22. Apply the sliding window to each dimension of the sensor data and calculate the features within the window, including the maximum value, minimum value, average value, standard deviation, skewness, and kurtosis.
[0015] Preferably, the specific process of step S3 is as follows:
[0016] Use the Z-score normalization method for the feature dataset, calculate the mean and standard deviation of the data, and transform each data point using the following formula:
[0017]
[0018] where x is the original data point and z is the transformed data point. Ensure that the scales of different features are consistent to improve the performance and stability of the machine learning model.
[0019] Preferably, the specific process of step S4 is as follows:
[0020] At the input end of the model, design an input layer to receive the feature dataset processed by Z-score normalization, ensuring that the data format matches the expected input of the model. The model combines a convolutional neural network (CNN) and a long short-term memory network (LSTM). The CNN part consists of multiple convolutional layers, batch normalization layers, activation function layers, and pooling layers, which are specifically used to extract local features from time series data. These two sub-models are combined through appropriate connection layers to form a two-stream structure for in-depth feature integration. A fully connected layer is added between the CNN and LSTM sub-models to achieve feature integration. To solve the problem of gradient flow in deep network training, residual connections are also added to the model to improve training stability.
[0021] Preferably, the specific process of step S5 is as follows:
[0022] S51. Divide the pose dataset extracted by the sliding window feature into a training set and a test set, where the training dataset accounts for 80% and the test set accounts for 20% to test the model performance.
[0023] S52. According to the size of the divided training set, adjust the number of training epochs, the number of data for each batch training, and the learning rate hyperparameter of model training to achieve the optimal model performance.
[0024] The present invention discloses an LSTM-CNN pig pose classification method based on pose sensor data, which realizes the accurate recognition of four poses of pigs, namely standing, lying, eating, and body movement, in a free-stall scenario through innovative technical means. The method first uses a pose sensor to collect the acceleration and gyroscope data of pigs, and combines the video images captured by a camera for data annotation to ensure the accuracy and integrity of the data. Compared with traditional vision methods, this sensor data-based method is not affected by pig occlusion or light changes and is suitable for complex environments in group farming scenarios. Then, the data is processed by a sliding window feature extraction algorithm and a time domain analysis method to extract statistical features that can characterize the dynamic changes of pig poses, such as maximum value, minimum value, and average value, thereby improving the model's recognition ability for rapidly changing behaviors. In addition, Z-score normalization processing is used to eliminate the differences in the dimensions of different sensor data, ensuring the consistency and stability of the data in model training. In terms of model construction, the method combines a convolutional neural network and a long short-term memory network. The convolutional neural network mines local features of the time series through convolutional layers, while the long short-term memory network captures long-term dependencies to jointly achieve a comprehensive understanding of the dynamic changes of pig poses. At the same time, residual connections are added to the model to optimize the training stability of the deep network and further improve the classification performance. After reasonable division of the dataset and fine-tuning of hyperparameters, the model with the best performance on the test set is selected to ensure its high reliability and practicality in actual applications. This method is not only innovative in technology but also can significantly reduce the monitoring costs of breeding enterprises, improve the automation level and work efficiency, and is not restricted by environmental factors, suitable for diverse breeding scenarios, showing broad application prospects and promotion potential. Description of the Drawings
[0025] Figure 1 It is a flowchart of an LSTM-CNN pig pose classification method based on pose sensor data according to the present invention;
[0026] Figure 2 It is a network structure diagram of a deep learning model in an LSTM-CNN pig pose classification method according to the specific embodiment of the present invention. Detailed Embodiments
[0027] To make the above objects and features of the present invention more obvious and understandable, the present invention will be further described in detail below in conjunction with embodiments and drawings. However, the embodiments of the present invention are not limited thereto. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention.
[0028] As an effective method for behavior detection, pose estimation has attracted extensive attention in the field of animal health and welfare detection in recent years. By monitoring the changes in the individual behaviors of pigs over time, it will contribute to the early detection of diseases and allow earlier and more effective interventions. However, in the actual breeding environment, pose estimation of pigs faces major challenges, mainly due to the occlusion problems caused by the aggregation of pigs in groups, which often makes it difficult for computer vision-based pose recognition models to accurately identify due to the lack of necessary semantic information. In contrast, the data collected by the present invention using pose sensors is not affected by occlusion, thus avoiding the inaccurate recognition problems caused thereby.
[0029] As Figure 1 and Figure 2 shown, the present invention provides an LSTM-CNN pig pose classification method based on pose sensor data, which can solve the occlusion problem in the scenario of pigs gathering in groups. The present invention first uses pose sensors to collect the joint angle data of pigs, and these data include the spatial position and movement trajectory of the pig's neck. Since the sensor data does not depend on visual images, it is not affected by the mutual occlusion between pigs, ensuring the integrity and accuracy of the data, and can realize the discrimination of the pose and behavior of pigs.
[0030] To achieve the above object, this embodiment will be described by including the following parts: data set collection, data set preparation (i.e., preprocessing), model construction, model training and tuning, as follows:
[0031] (1) Acquisition environment.
[0032] The experimental site of this embodiment is the Xiangjiang Pig Farm in Guigang, Yangxiang Co., Ltd. Pose sensor and video data from March 19th to 22nd, 2024 were collected as the basic data for research. The actual scenarios of each pen are as Figure 1 shown. The main equipment used for data collection includes cameras and pose sensors. Among them, the cameras used to shoot the behavior videos of pigs are installed above the pig pens. The pose sensor selects a nine-axis WIFI communication IoT remote accelerometer gyroscope pose angle sensor, model WT901WIFI, which integrates a three-axis accelerometer and a three-axis gyroscope inside.
[0033] (2) Data collection.
[0034] In this embodiment, the data for the study was collected from a pigsty in Yangxiang Xiangjiang Pig Farm, Guigang. Given the large number of pigs in the same pigpen, to improve the accuracy of data collection and reduce the possible stress response of the pigs, we randomly selected two pigs from each pigpen and placed them in separate pigpens for observation. During the experiment, we used elastic bands to fix the sensors on the necks of the pigs. At the same time, we installed a camera on the wall in front of the pigpen to record the behavior process of the pigs. To reduce the impact of personnel flow on the behavior of the pigs, after the installation and debugging of the equipment were completed, the experimenters left the pigsty and observed and recorded the behavior activities of the pigs remotely through the camera. The recorded video files were then viewed through the memory card. The observation of the pigs was usually divided into three stages: morning, noon, and evening in a day. To avoid long-term stress responses in the pigs, we changed the pigs for observation every day. During the transition of each observation stage, the experimenters would enter the pigsty to check whether the pigs had stress responses and replenish the feed to ensure the normal physiological activities of the pigs. The experiment conducted observational recordings of different pigs in four pigpens for four days, thus providing the basic data for the construction of the subsequent dataset.
[0035] (2) Construction of the original dataset.
[0036] In this embodiment, a total of 191 MB of original sensor data and 15.2 GB of video data were collected during data collection. To classify and label the original sensor data, this study adopted the method of manual annotation. During the process of manually annotating the postures of the pigs, using the video data as a reference, the posture sensor data was classified and labeled, and the research focused on four common postures of the pigs. The posture names and descriptions are shown in Table 1 below:
[0037] Table 1: Table of posture name descriptions.
[0038]
[0039] (3) Preprocessing of the dataset.
[0040] In this embodiment, since the data collected by the attitude sensor is of a time-series nature, the sensor can sample 10 valid data per second in total. Each attitude data contains a total of six elements: acceleration X, Y, Z and angular velocity X, Y, Z. For the same type of attitude, the data shows significant variability in both time and space. The variability in time is reflected in the fact that the duration for which the pig maintains each attitude is different. Some attitudes may change rapidly, while other attitudes may last longer, resulting in a large difference in the attitude samples in the time dimension. The variability in space indicates that even for the same type of attitude, the amplitude shown may vary. Therefore, even for the same type of attitude, the data waveform diagrams will not be exactly the same. Directly using these original time-series data for model training may encounter many challenges. Therefore, feature extraction is performed on the original data in order to construct a more abstract feature set to improve the training effect of the model. We use the sliding time window method to segment the original data set. Each time the window slides, the recognition algorithm performs feature extraction once. Denote the i-th attitude data as x i (a i X ,a i Y ,a i Z ,g i X ,g i Y ,g i Z ) For each window, that is, for every 10 data, feature extraction is performed on each dimension of the data respectively. Three statistical feature quantities are selected, including the maximum value, the minimum value, and the average value as the newly added feature values. Among them, the maximum value and the minimum value reflect the change intensity of the attitude, while the average value reflects the overall situation and trend of the attitude. As shown in equations (1) - (3)
[0041]
[0042] Taking the data of one window as the operation unit, each time the window slides, feature extraction is performed once. Each piece of data after feature extraction is a d-dimensional feature vector, denoted as f k =(f k,1 ,f k,2 ,...,f k,d ) The feature set used to train the behavior recognition model can be expressed as {(f1, y1), …, (f k , y k ), …}, where y k is the true behavior category y k ∈ {standing, lying prone, eating, body movement}.
[0043] Use the Z-score normalization method for the feature dataset, calculate the mean and standard deviation of the data, transform each data point, and use the following formula
[0044]
[0045] where x is the original data point and z is the transformed data point. Ensure that the scales of different features are consistent to improve the performance and stability of the machine learning model.
[0046] (4) Pose classification model.
[0047] The LSTM-CNN network proposed in this embodiment has three stages. After the input passes through the feature-extracted data, it needs to go through an initial stage containing 4 one-dimensional convolutional layers and pooling layers to deeply mine and refine the spatio-temporal features in the sequence; subsequently, use the residual structure to consolidate and transmit key information; then, capture and analyze the long-term dependent temporal dynamics through a double-layer LSTM network; finally, complete the detailed classification of the pig group pose with the help of a multi-layer fully connected network.
[0048] More specifically, as follows:[[]]END]]
[0049] Convolutional initial feature extraction stage: Responsible for extracting features from the original input data. This sequence includes two groups of convolutional layers, each containing two one-dimensional convolutional layers. The number of channels in the convolutional layers increases from low to high as [32, 64, 128, 256]. The convolutional kernel size is set to 11, the stride is 1, and appropriate padding is used to keep the input and output sizes consistent. After each group of convolutions, batch normalization is followed to improve training stability, and the LeakyReLU activation function is used to enhance non-linear expression. Max pooling is used for further downsampling between groups, and finally a Dropout layer is added to improve the generalization ability of the model. When the input signal where N is the batch size, C in is the number of input channels, L in is the length of the input sequence. The convolutional kernel set where j = 1, 2,..., C out indicates that there are C out output channels, and each filter w j is a C in ×K matrix used to extract specific types of features from the input. For the output its calculation can be given by the following formula:[[]]END]]
[0050]
[0051] where Y n,j,l represents the value of the nth sample, the jth channel, and the lth position in the output sequence. X n,c,lS+k is the value at the corresponding position and channel in the input sequence, considering the stride S. wj,c,k is the weight corresponding to the input channel c at position k in the j-th output channel filter. b j is the bias term of the j-th output channel.
[0052] LSTM (Long Short-Term Memory) Temporal Pattern Capturing Stage: In this embodiment, the LSTM layer adopts two stacked layers, each with 256 hidden units, which increases the expressive power of the model and enables the model to learn deeper temporal features and more complex sequence patterns. In the early stage, the model extracts local spatio-temporal features of the input data (such as statistical features of acceleration and angular velocity) through convolutional units. These features are then used as the input of the LSTM layer. Through a complex gating mechanism, the LSTM not only retains and updates these local features, but more importantly, it can capture and integrate the long-term dependencies of these features evolving over time, which is crucial for understanding the behavioral patterns of the pig group's posture changing continuously over time. Therefore, the LSTM layer plays a key role in refining and encoding deep temporal patterns from time series data in this embodiment, laying a solid foundation for subsequent classification tasks. The LSTM unit achieves this function through its unique gating mechanism, including the forget gate f t , the input gate i t , the output gate o t and the candidate state update gate c' t . The operation of these gates follows the following core formulas:
[0053] f t = σ(W f ·[h t-1 , x t +b f ) (6)
[0054] i t = σ(W i ·[h t-1 , x t +b i ) (7)
[0055] c' t = g t (8)
[0056] c t = f t ⊙ c t-1 + i t ⊙ c' t (9)
[0057] o t = σ(W o ·[h t-1 , x t +bo ) (10)
[0058] h t = o t ⊙ tanh(c t ) (11)
[0059] Classification decision stage: Finally, through a series of designed Dense Units, the network transforms the high-level features extracted in the previous stage into specific pose classifications. These Dense Units gradually reduce the dimensionality, from 256 dimensions to 128 dimensions, then to 64 dimensions, and finally map to the actual class outputs. Each step is accompanied by a LeakyReLU activation function to introduce non-linearity to achieve the final classification decision. When the input vector is The weight matrix is The bias vector is Then the output of the Dense Unit can be calculated by the following formula:
[0060] y = σ(Wx + b) (12)
[0061] The above are only the preferred embodiments of the present invention, and do not impose any form of limitation on the present invention. Based on the technical essence of the present invention, any simple modifications, equivalent replacements, and improvements made to the above embodiments within the spirit and principles of the present invention still fall within the protection scope of the technical solution of the present invention.
Claims
1. A LSTM-CNN pig posture classification method based on posture sensor data, characterized in that: The following steps are involved: S1, collect the pig's posture sensor motion data, including acceleration and gyroscope readings, and establish a pig posture sensor data set; S2, based on the sliding window feature extraction algorithm, use the time domain analysis method to extract each dimension of the data set separately to obtain the final feature data set; S3, standardize the attitude sensor feature data set to eliminate the impact of different sensor dimensions and ensure the consistency of data in model training; S4. Build a deep learning model that combines convolutional neural networks and long short-term memory networks. Convolutional neural networks are used to extract local features in time series data, while long short-term memory networks are used to capture long-term dependencies in time series. S5. Divide the pig posture recognition posture sensor feature data set into a training set and a test set, use the training set to train the two-stream LSTM-CNN pig posture recognition model, use the test set to test the model performance, and finally select the best performance model by adjusting the parameters.
2. The LSTM-CNN pig posture classification method based on posture sensor data according to claim 1 is characterized in that: The specific process of step S1 is as follows: S11. Fix the posture sensor with an elastic band just below the pig's neck, so that the elastic band can be adjusted at any time according to the actual situation to ensure that the sensor is always in the correct position; S12, fixing the camera above the pig pen, and shooting the video images of the pigs' movements from a bird's-eye view.
3. The LSTM-CNN pig posture classification method based on posture sensor data according to claim 1 is characterized in that: The specific process of step S2 is as follows: S21, determining the size and step length of the sliding window, the size of the sliding window is determined according to the time scale of the pig's behavior characteristics, and the step length is set according to the frequency of data collection and the window size; S22. Apply a sliding window to each dimension of sensor data and calculate the features within the window, including maximum value, minimum value, mean value, standard deviation, skewness, and kurtosis.
4. The LSTM-CNN pig posture classification method based on posture sensor data according to claim 1 is characterized in that: The specific process of step S3 is as follows: Use the Z-score standardization method on the feature data set, calculate the mean and standard deviation of the data, and transform each data point using the following formula: Here, x is the original data point and z is the transformed data point, ensuring the consistency of scales of different features, thereby improving the performance and stability of the machine learning model.
5. The LSTM-CNN pig posture classification method based on posture sensor data according to claim 1 is characterized in that: The specific process of step S4 is as follows: At the input end of the model, the input layer is designed to receive the feature data set that has been normalized by Z-score to ensure that the data format matches the expected input of the model. The model integrates convolutional neural networks and long short-term memory networks. The convolutional neural network part consists of multiple convolutional layers, batch normalization layers, activation function layers, and pooling layers, which are specifically used to extract local features from time series data. The two sub-models are combined through appropriate connection layers to form a dual-stream structure to achieve deep integration of features. A fully connected layer is added between the convolutional neural network and the long short-term memory network sub-models to achieve feature integration. In order to solve the gradient flow problem in deep network training, residual connections are also added to the model to improve training stability.
6. The LSTM-CNN pig posture classification method based on posture sensor data according to claim 1 is characterized in that: The specific process of step S5 is as follows: S51, dividing the posture data set extracted by the sliding window feature into a training set and a test set, wherein the training data set accounts for 80% and the test set accounts for 20% to test the model performance; S52. According to the size of the training set, adjust the number of model training rounds, the amount of data for each training batch, and the learning rate hyperparameters to achieve optimal model performance.