A method for detecting abnormal behaviors of pigs

Through the improved Yolov5n network model and dual-stream convolutional autoencoder network, combined with K-means clustering and binary classifier, general detection of pig abnormal behavior is realized, solving the problems of data set distribution imbalance and training data dependence in the existing methods, and achieving efficient and accurate abnormal behavior recognition.

CN115359511BActive Publication Date: 2025-05-30SOUTH CHINA AGRICULTURAL UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210934696.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-04
Publication Date
2025-05-30
Estimated Expiration
2042-08-04

AI Technical Summary

Technical Problem

The existing pig abnormal behavior detection methods cannot realize general pig abnormal behavior detection, and there are problems with unbalanced distribution of data sets and dependence on manual annotation of training data.

Method used

The improved Yolov5n network model is used for object detection and cropping, combined with an end-to-end trainable object-centric dual-stream convolutional autoencoder network to extract the appearance and motion feature vectors of pigs, and unsupervised anomaly behavior detection is achieved through K-means clustering and binary classifiers.

Benefits of technology

Effectively detect abnormal behavior of pigs in complex occlusion pig farm environment, making up for the lack of training data in the existing methods, and can accurately identify abnormal behavior of pigs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115359511B_ABST
    Figure CN115359511B_ABST
Patent Text Reader

Abstract

The present invention provides a method for detecting abnormal behaviors of pigs, comprising the following steps: S1: Extracting images frame by frame from the real-time collected pig life videos; S2: Using an improved Yolov5n model to perform object detection and cropping on the extracted images of each frame to obtain the target screenshots of each pig in each frame image; S3: Extracting feature vectors through a two-stream convolutional autoencoder; S4: Using K-means and a classification algorithm to cluster and classify the feature vectors; S5: Obtaining the classification scores of all targets in the current frame through a classifier, and combining all the classification scores to form an abnormal prediction map; S6: Performing Gaussian filtering time series smoothing on the abnormal prediction map, and recording the highest classification score as the abnormal score of the current frame image; S7: Judging whether the abnormal score of the current frame image is a positive number; if so, there is no abnormal behavior, otherwise, there is. The method for detecting abnormal behaviors of pigs provided by the present invention solves the problem that the current abnormal detection methods cannot achieve general detection of abnormal behaviors of pigs.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of artificial intelligence, and more specifically, to a method for detecting abnormal behaviors of pigs. Background Art

[0002] In the livestock industry, especially in pig farms with a closed breeding environment, infectious diseases among animals cause serious damage to their well-being, are prone to fatal infections, and cause huge economic losses to farmers. In addition to providing a good living environment for the pig herd, a necessary condition for realizing welfare pig farming is to continuously monitor animal behaviors to detect abnormalities as early as possible for timely diagnosis and treatment, so as to maximize benefits.

[0003] The behaviors of pigs reflect the welfare status and social interactions of animals, and are important bases for analyzing the health status of pigs and healthy breeding management. Close interactions among pigs may have a negative impact on the health of pigs and reduce animal welfare. For example, both male and female pigs will exhibit mounting behaviors, especially during the estrus period, which is usually manifested as a pig putting its two front hooves on the body or head of another pig, and the other pig either lies still or quickly dodges, resulting in bruises, lameness, and leg fractures. These injuries will cause serious economic losses to the livestock industry. Therefore, by timely monitoring different abnormal behaviors of live pigs, the abnormal conditions of live pigs can be evaluated, thereby preventing pig diseases or preventing the spread of diseases, and improving the welfare level of pig farming.

[0004] In recent years, due to the remarkable performance of deep learning in the field of anomaly detection, a large number of studies have adopted methods based on neural networks. However, realizing the monitoring of pig behaviors in a closed pig farm breeding environment poses a huge challenge to computer vision. For example, confusion between different pigs due to visual similarity, sudden movements caused by aggressive behaviors of pigs, frequent occlusions, pigs crowding together, etc. The training effect of anomaly behavior detection under supervised learning is easily affected by the unbalanced distribution of video surveillance data sets, and its performance largely depends on the availability and quality of manually annotated training data sets, and is not suitable for anomaly behavior detection based on videos. Therefore, a large number of researchers have started from the perspective of unsupervised learning and proposed an anomaly behavior detection method that does not require training with labeled data sets and adapts to the video data itself, and has achieved good results on various data sets. However, the existing unsupervised methods mainly design targeted algorithms for specific abnormal behaviors, such as aggressive, tail-biting, mounting and other abnormal behaviors for identification. One of their disadvantages is that only one dedicated algorithm can be designed for detecting one abnormal behavior, and it is impossible to achieve general detection of abnormal behaviors of pigs. Summary of the Invention

[0005] In order to overcome the technical defect that the current anomaly detection methods cannot achieve general detection of abnormal behaviors of pigs, the present invention provides a method for detecting abnormal behaviors of pigs.

[0006] To solve the above technical problems, the technical solution of the present invention is as follows:

[0007] A method for detecting abnormal behaviors of pigs, comprising the following steps:

[0008] S1: Real-time collect the living videos of pigs, and extract images frame by frame from the living videos of pigs;

[0009] S2: Use an improved Yolov5n network model to perform object detection and cropping on the extracted images of each frame, and respectively obtain the target screenshots of each pig in each frame of the image;

[0010] The improved Yolov5n network model is: after the 4th, 6th, and 8th layers of the backbone feature extraction network of the existing Yolov5n network model, add channel attention modules, and splice the channel attention modules with the upsampling layers of the 18th, 22nd, and 26th layers of the neck network, and add a C3 layer and a channel attention module after the 11th layer of the backbone feature extraction network;

[0011] S3: Construct an end-to-end trainable object-centered two-stream convolutional autoencoder network to extract the appearance feature vectors and motion feature vectors of each pig in the target screenshots, and perform feature fusion to form the feature vector corresponding to the frame;

[0012] The two-stream convolutional autoencoder network is trained only with images of normal pig behaviors;

[0013] S4: Use the K-means clustering algorithm to cluster the fused feature vectors, and input the results into a binary classifier for training to obtain a trained classifier;

[0014] S5: In each frame of the image, obtain the classification scores of all target screenshots in the current frame of the image through the classifier, and combine all the classification scores to form an abnormal prediction map of the current frame of the image;

[0015] S6: Perform Gaussian filtering time series smoothing on the abnormal prediction map of the current frame of the image, and record the highest classification score as the abnormal score of the current frame of the image;

[0016] S7: Judge whether the abnormal score of the current frame of the image is a positive number;

[0017] If so, there is no abnormal behavior of the pigs in the current frame of the image;

[0018] If not, there is an abnormal behavior of the pigs in the current frame of the image.

[0019] In the above solution, an improved Yolov5n network model is used to perform object detection and cropping on each frame of image to obtain the target screenshots of each pig in each frame of image. All pigs can still be effectively detected in the actual pig farm environment with complex occlusion. Then, a classifier trained only with images of normal pig behaviors is used to classify the target screenshots to obtain classification scores, and then the abnormal scores of each frame of image are obtained. Finally, pig abnormal behavior detection is realized according to the abnormal scores, making up for the lack of real abnormal behavior training data and being able to accurately identify the abnormal behaviors of pigs.

[0020] Preferably, the channel attention module includes compression, excitation, and scaling operations; among them,

[0021] The compression operation is: using global average pooling to compress the dimension H*W*C of the original feature layer into 1*1*C;

[0022] The excitation operation is: using two fully connected layers to fuse the feature map information of each feature channel, and then using the Sigmoid function to normalize the weights;

[0023] The scaling operation is: mapping the weights output after the excitation operation to a set of weights of feature channels, and then multiplying and weighting them with the features of the original feature map to achieve feature recalibration of the original features in the channel dimension.

[0024] Preferably, the improvement of the Yolov5n network model also includes adding a 64-fold downsampling detection layer to make the scale of the output feature map 20×20.

[0025] Preferably, in step S5, a target screenshot is selected from the current frame of image, the feature vector of the selected target screenshot is extracted through step S3 and clustered into k clusters, and then the clustering results are respectively input into k classifiers to obtain k classification scores, and the highest classification score is selected as the abnormal score of the selected target screenshot. Repeat this step until the abnormal classification scores of all target screenshots in the current frame of image are obtained.

[0026] Preferably, the classifier is a binary classifier, and the definition of the i-th binary classifier is as follows:

[0027]

[0028] Among them, w j represents the weight vector, b represents the bias value, x represents the sample input into the binary classifier, x can be classified as a normal sample or an abnormal sample, and x j represents the j-th element of the sample, and m represents the dimension of x.

[0029] Preferably, k binary classifiers are trained through the following steps:

[0030] A1: Select the images of normal pig behaviors from the pig life videos as training images;

[0031] A2: Use the improved Yolov5n network model to perform object detection and cropping on the training images, and obtain the target screenshots of each pig in the training images respectively;

[0032] A3: Convert the target screenshots into grayscale images, and obtain the corresponding frame difference images by subtracting the pixel values from the adjacent frame images of the training images;

[0033] A4: Use the grayscale frame images and grayscale frame difference images obtained in step A3 as the inputs of the appearance sub-network and the action sub-network in the object-centered pig abnormal behavior convolutional autoencoder network respectively, and extract the appearance feature vectors and action feature vectors of each pig in the target screenshots through this network;

[0034] The autoencoder network includes an appearance sub-network and an action sub-network. The appearance sub-network is used to extract appearance feature vectors from the target screenshots, and the action sub-network is used to extract action feature vectors from the frame difference images;

[0035] A5: Fuse the appearance feature vectors and action feature vectors to obtain the fused feature vectors of the training images;

[0036] A6: Perform k-means clustering on the fused feature vectors to obtain the clustering result clusters i, i = 1, 2,...., k;

[0037] A7: Input the clustering results into k binary classifiers to obtain k trained binary classifiers.

[0038] Preferably, both the appearance sub-network and the action sub-network include an attention module and a memory module; where

[0039] The calculation formula of the attention module is:

[0040]

[0041]

[0042] u t,t′ = a(s t-1 ,h t′ )

[0043] where, c t represents the context vector at time t, T represents the total time length, α t,t′ represents the attention weight in the neighborhood of t' at time t, h t′ represents the hidden unit output at the t' -th moment, α represents the attention weight, u t,t′ represents the output score in the neighborhood of t' at time t, ut,k represents the output score of the k-neighborhood at time t, s t-1 represents the hidden state at time t-1;

[0044] The memory module includes M memory items p m , m = 1, …, M, for recording various prototype feature patterns of normal pig behavior data;

[0045] For each query mapping By weighting the corresponding weights of the memory item p m perform weighted averaging to read the memory item and obtain the feature

[0046]

[0047]

[0048] where represents the weight of the memory item p m′ and p m′ represents the m'-th memory item;

[0049] The update formula for the memory item is:

[0050]

[0051]

[0052]

[0053] where ← represents the update operation, f represents the L2 norm, and v t ′k,m represents the reconstruction of the matching probability value , represents the query index set of the memory module.

[0054] Preferably, when updating the memory item, if the weighted score ε of the t-th frame image t is greater than the preset threshold, the t-th frame image is regarded as an abnormal frame and is not used to update the memory item;

[0055] The weighted score ε is calculated by the following formula t :

[0056]

[0057]

[0058] where represents the weight value of the feature, Represents a certain feature within the t-neighborhood, I t Represents the feature at the t-th moment, where i and j represent spatial indices.

[0059] Preferably, the loss function of the autoencoder is:

[0060]

[0061] where, is the reconstruction error, is the feature compactness loss function, is the feature separation loss function, is a hyperparameter.

[0062] Preferably,

[0063] the reconstruction error is:

[0064]

[0065] The feature compactness loss function is:

[0066]

[0067]

[0068] The feature separation loss function is:

[0069]

[0070]

[0071] where, T represents the total time, t represents the time index, k represents the index of the query mapping, K represents the total number of query mappings, represents a certain feature within the t-neighborhood, I t represents the feature at the t-th moment, p p represents the query mapping of the nearest memory item, and p is the index of the nearest item of the query mapping of, represents the weight of the m-th memory item, m represents the memory item index, M represents the total number of memory items, p n represents the query mapping of the second-nearest memory item.

[0072] Compared with the prior art, the beneficial effects of the technical solution of the present invention are:

[0073] The present invention provides a method for detecting abnormal behaviors of pigs. By using an improved Yolov5n network model to perform object detection and cropping on each frame of image, target screenshots of each pig in each frame of image can be obtained, and all pigs can still be effectively detected in the actual pig farm environment with complex occlusion. Then, an end-to-end trainable object-centered two-stream convolutional autoencoder network is constructed, which is trained only with videos of normal pig behaviors, only focuses on the pig objects existing in the scene, does not require manual extraction of image features, and can accurately locate the abnormalities in each frame, and determine the scale and duration of the occurrence of pig abnormal behaviors. At the same time, a memory module and a memory update strategy with the ability to learn and store the "prototype" features of normal pig behaviors are proposed; then, an unsupervised binary classification method is used to solve the problem of pig abnormal behavior detection. The feature vectors learned by the autoencoder network are clustered, and the obtained clustering results are used to train a classifier. By using the trained classifier to classify the target screenshots to obtain classification scores, and then obtaining the abnormal scores of each frame of image, and finally realizing the detection of pig abnormal behaviors according to the abnormal scores, which makes up for the lack of training data for real abnormal behaviors and can accurately identify the abnormal behaviors of pigs. BRIEF DESCRIPTION OF THE DRAWINGS

[0074] Figure 1 It is a flowchart of the implementation steps of the technical solution of the present invention;

[0075] Figure 2 It is a schematic structural diagram of the improved Yolov5n network model in the present invention;

[0076] Figure 3 It is a schematic flowchart of obtaining a frame difference map in the present invention;

[0077] Figure 4 It is a schematic diagram of the object-centered pig abnormal behavior detection network in the present invention;

[0078] Figure 5 It is a schematic structural diagram of a sub-network of the autoencoder network in the present invention;

[0079] Figure 6 It is a schematic diagram of memory item reading in the present invention;

[0080] Figure 7 It is a schematic diagram of memory item update in the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0081] The drawings are only for illustrative purposes and should not be construed as a limitation of this patent;

[0082] For better illustration of this embodiment, some components in the drawings are omitted, enlarged or reduced, and do not represent the dimensions of the actual product;

[0083] For those skilled in the art, it is understandable that some well-known structures and their descriptions in the drawings may be omitted.

[0084] The technical solution of the present invention will be further described below in conjunction with the drawings and embodiments.

[0085] Embodiment 1

[0086] As Figure 1 shown, a method for detecting abnormal behaviors of pigs includes the following steps:

[0087] S1: Real-time collect the living videos of pigs, and extract images frame by frame from the living videos of pigs;

[0088] S2: Use an improved Yolov5n network model to perform object detection and cropping on the extracted images of each frame, and respectively obtain the target screenshots of each pig in each frame image;

[0089] The improved Yolov5n network model is: add channel attention modules after the 4th, 6th, and 8th layers of the backbone feature extraction network of the existing Yolov5n network model, splice the channel attention modules with the upsampling layers of the 18th, 22nd, and 26th layers of the neck network, and add a C3 layer and a channel attention module after the 11th layer of the backbone feature extraction network;

[0090] S3: Construct an end-to-end trainable object-centered two-stream convolutional autoencoder network to extract the appearance feature vectors and motion feature vectors of each pig in the target screenshots, and perform feature fusion to form the feature vectors of the corresponding frame;

[0091] The two-stream convolutional autoencoder network is trained only with images of normal pig behaviors;

[0092] S4: Use the K-means clustering algorithm to cluster the fused feature vectors, and input the results into a binary classifier for training to obtain a trained classifier;

[0093] S5: In each frame image, obtain the classification scores of all target screenshots in the current frame image through the classifier, and combine all the classification scores to form the abnormal prediction map of the current frame image;

[0094] S6: Perform Gaussian filtering time series smoothing on the abnormal prediction map of the current frame image, and record the highest classification score as the abnormal score of the current frame image;

[0095] S7: Judge whether the abnormal score of the current frame image is a positive number;

[0096] If so, there is no abnormal behavior of the pigs in the current frame image;

[0097] Otherwise, there are abnormal behaviors of the pigs in the current frame image.

[0098] In the specific implementation process, an improved Yolov5n network model is used to perform object detection and cropping on each frame of image to obtain the target screenshots of each pig in each frame of image. All pigs can still be effectively detected in the actual pig farm environment with complex occlusion. Then, the appearance and motion feature vectors of each pig in the target screenshots are extracted by inputting them into a two-stream convolutional autoencoder network, and the fused feature vectors are clustered. The obtained clustering results are used to train a classifier. The target screenshots are classified by using the trained classifier to obtain classification scores, and then the abnormal scores of each frame of image are obtained. Finally, the abnormal behaviors of pigs are detected according to the abnormal scores, making up for the lack of real abnormal behavior training data and being able to accurately identify the abnormal behaviors of pigs.

[0099] Embodiment 2

[0100] A method for detecting abnormal behaviors of pigs includes the following steps:

[0101] S1: Real-time collect the living videos of pigs and extract images frame by frame from the living videos of pigs;

[0102] S2: Use an improved Yolov5n network model to perform object detection and cropping on the extracted frames of images to obtain the target screenshots of each pig in each frame of image respectively;

[0103] As Figure 2 shown, the improved Yolov5n network model is: add channel attention modules after the 4th, 6th, and 8th layers of the backbone feature extraction network of the existing Yolov5n network model, splice the channel attention modules with the upsampling layers of the 18th, 22nd, and 26th layers of the neck network, and add a C3 layer and a channel attention module after the 11th layer of the backbone feature extraction network;

[0104] In actual implementation, aiming at the problems of the existing Yolov5n network model such as inaccurate bounding box positioning, difficulty in distinguishing overlapping objects, and poor robustness, this embodiment adds an SE-Net channel attention module to the Backbone of the backbone feature extraction network for improvement, establishing a feature map through the interaction between convolutional network channels, enabling the network model to automatically learn global feature information and highlight useful feature information, while suppressing other less important feature information, making the network model more focused on the purpose of training occluded objects.

[0105] More specifically, the channel attention module includes compression, excitation, and scaling operations; among them,

[0106] The compression operation is: using global average pooling to compress the dimension H*W*C of the original feature layer into 1*1*C;

[0107] The excitation operation is as follows: use two fully connected layers to fuse the feature map information of each feature channel, and then use the Sigmoid function to normalize the weights;

[0108] The scaling operation is as follows: map the weights output after the excitation operation to the weights of a set of feature channels, and then multiply and weight them with the features of the original feature map to achieve feature recalibration of the original features in the channel dimension.

[0109] More specifically, the improvement of the Yolov5n network model also includes adding a 64-fold downsampling detection layer to make the scale of the output feature map 20×20.

[0110] In the specific implementation process, based on the original 3 detection layers with different scales (40x40, 80x80, 160x160) in the Yolov5n network model, add a detection layer with an ultra-small scale (20x20), that is, after 8-fold, 16-fold, and 32-fold downsampling, add a 64-fold downsampling detection layer, so as to obtain the feature map of the 20x20 scale detection layer, further deepen the network depth, enable the network model to extract higher-level semantic information, and the information is also more abundant, thereby enhancing the multi-scale learning ability of the model in complex scenarios and improving the detection performance of the model.

[0111] S3: Construct an end-to-end trainable object-centered two-stream convolutional autoencoder network to extract the appearance feature vectors and motion feature vectors of each pig in the target screenshot, and perform feature fusion to form the feature vector of the corresponding frame;

[0112] The described two-stream convolutional autoencoder network is trained only with images of normal pig behaviors;

[0113] S4: Use the K-means clustering algorithm to cluster the fused feature vectors, and input the results into a binary classifier for training to obtain a trained classifier;

[0114] S5: In each frame of image, obtain the classification scores of all target screenshots in the current frame image through the classifier, and combine all the classification scores to form the anomaly prediction map of the current frame image;

[0115] More specifically, in step S5, select a target screenshot from the current frame image, extract the feature vector of the selected target screenshot through step S3 and cluster it into k clusters, then input the clustering results into k classifiers respectively to obtain k classification scores, and select the highest classification score as the anomaly score of the selected target screenshot. Repeat this step until the anomaly classification scores of all target screenshots in the current frame image are obtained.

[0116] More specifically, the classifier is a binary classifier, and the definition of the i-th binary classifier is as follows:

[0117]

[0118] where w j represents the weight vector, b represents the bias value, x represents the sample input to the binary classifier, x can be classified as a normal sample or an abnormal sample, x j represents the j-th element of the sample, and m represents the dimension of x.

[0119] More specifically, k binary classifiers are trained through the following steps:

[0120] A1: Select images of normal pig behaviors from pig life videos as training images;

[0121] In actual implementation, perform masking processing on images containing multiple pig pens. Add a masking layer on the original image with the inventory of pig pens as the boundary to cover the pigs in other pens, so as to construct a training dataset with multiple scenarios (different pig houses, different numbers of pigs, different degrees of occlusion, different lighting, different pig body sizes, etc.), and manually label the pigs in the images;

[0122] A2: Use the improved Yolov5n network model to perform object detection and cropping on the training images to obtain the target screenshots of each pig in the training images respectively;

[0123] In actual implementation, since the preset anchor boxes of the existing Yolov5n network model are mainly for the coco dataset (a public dataset provided by the Microsoft team), which is completely different from the aspect ratio of the annotation boxes of the training dataset in this embodiment (the maximum aspect ratio of the annotation boxes in the coco dataset reaches more than 1:8, and the maximum aspect ratio of the training dataset in this embodiment is 5.6), therefore, this embodiment generates 12 anchor boxes as the initial boxes of the model through the K-means algorithm clustering, and then randomly mutates the width and height of the anchor boxes 100,000 times using the Genetic Alogirthm genetic algorithm. If the effect after mutation is better, then use the result after mutation as the final result of the anchor box, otherwise skip it.

[0124] A3: Convert the target screenshots into grayscale images, and obtain the corresponding frame difference images by subtracting the pixel values from the adjacent frame images of the training images, as Figure 3 shown;

[0125] A4: Use the grayscale frame images and grayscale frame difference images obtained in step A3 as the inputs of the appearance sub-network and the action sub-network in the object-centered convolutional autoencoder network for pig abnormal behavior respectively, and extract the appearance feature vectors and action feature vectors through this network, as Figure 4 shown;

[0126] The autoencoder network includes an appearance sub-network and an action sub-network. The appearance sub-network is used to extract appearance feature vectors from the target screenshots, and the action sub-network is used to extract action feature vectors from the frame difference images;

[0127] More specifically, both the appearance sub-network and the action sub-network include an attention module and a memory module; where

[0128] The calculation formula of the attention module is:

[0129]

[0130]

[0131] u t,t′ = a(s t-1 , h t′ )

[0132] where c t represents the context vector at time t, T represents the total time length, α t,t′ represents the attention weight of the neighborhood of t' at time t, h t′ represents the output of the hidden unit at the t'-th moment, α represents the attention weight, u t,t′ represents the output score of the neighborhood of t' at time t, u t,k represents the output score of the neighborhood of k at time t, s t-1 represents the hidden state at time t-1;

[0133] In the specific implementation process, the autoencoder network consists of two sub-networks to form a two-stream structure, namely an appearance sub-network and an action sub-network. Both sub-networks include an encoder, a memory module, and a decoder. The encoder and the decoder both include a spatial convolutional layer, three convolutional LSTM layers (ConvLSTM), three attention modules, and two max pooling layers (MaxPool), as Figure 5 shown.

[0134] By constructing an end-to-end trainable object-centered two-stream convolutional autoencoder network, it only focuses on the pig objects existing in the scene, does not require manual extraction of image features, and can accurately locate the anomalies in each frame, determine the scale and duration of the abnormal behavior of the pigs, with the technical advantages of time-saving, high efficiency, high accuracy, and high robustness.

[0135] The memory module includes M memory items p m , m = 1,..., M, which are used to record various prototype feature patterns of the normal behavior data of the pigs;

[0136] such asFigure 6 As shown Figure 6 where C represents calculating the cosine similarity between the two, S represents the softmax function, and W represents the weighted average; to read the memory item, by calculating each query mapping and all memory items p m the cosine similarity between them, a two-dimensional graph of size M×K is obtained, and then the softmax function is applied vertically to obtain the read matching probability

[0137]

[0138] For each query mapping by performing a weighted average on the memory item p with the corresponding weight m to read the memory item and obtain the feature

[0139]

[0140] where represents the weight of the memory item p m′ and p m′ represents the m'-th memory item;

[0141] As Figure 7 shown Figure 7 where C represents calculating the cosine similarity between the two, S represents the softmax function, W represents the weighted average, and n represents the maximum normalization; for the update operation, for each memory item p m , by calculating and then selecting the query mapping m closest to p and then using the query index set to update the memory item, and the update formula for the memory item is:

[0142]

[0143]

[0144]

[0145] where ← represents the update operation, f represents the L2 norm, and v t ′k,m represents the reconstruction of the write matching probability value A5: Fuse the appearance feature vector and the action feature vector to obtain the fused feature vector of the training image;

[0146]

[0147] ​A6: Perform k-means clustering on the fused feature vectors to obtain the clustering result cluster i, where i = 1, 2,...., k;

[0148] A7: Input the clustering result into k binary classifiers to obtain k trained binary classifiers.

[0149] In the specific implementation process, a context is constructed through K-means clustering. In this context, a subset of normal samples is equivalent to pseudo-abnormal samples relative to another subset, thus making up for the lack of real abnormal samples. K-means clustering clusters normal samples into k clusters, and each cluster represents a certain normal behavior of pigs, which is different from the behaviors represented by other clusters. That is, from the perspective of a given cluster i, samples belonging to other clusters (from the dataset {1, 2,...., k}\i [\ represents except i]) can be regarded as abnormal samples.

[0150] S6: Perform Gaussian filtering temporal smoothing on the abnormal prediction map of the current frame image, and record the highest classification score as the abnormal score of the current frame image;

[0151] S7: Determine whether the abnormal score of the current frame image is positive;

[0152] If so, there is no abnormal behavior of the pigs in the current frame image;

[0153] If not, there is abnormal behavior of the pigs in the current frame image.

[0154] Embodiment 3

[0155] A method for detecting abnormal behaviors of pigs includes the following steps:

[0156] S1: Real-time collect the living videos of pigs and extract images frame by frame from the living videos of pigs;

[0157] S2: Use the improved Yolov5n network model to perform object detection and cropping on the extracted frame images to obtain the target screenshots of each pig in each frame image;

[0158] The improved Yolov5n network model is: add channel attention modules after the 4th, 6th, and 8th layers of the backbone feature extraction network of the existing Yolov5n network model, splice the channel attention modules with the upsampling layers of the 18th, 22nd, and 26th layers of the neck network, and add a C3 layer and a channel attention module after the 11th layer of the backbone feature extraction network;

[0159] More specifically, the channel attention module includes compression, excitation, and scaling operations; among them,

[0160] The compression operation is as follows: using global average pooling to compress the dimension H*W*C of the original feature layer into 1*1*C;

[0161] The excitation operation is as follows: using two fully connected layers to fuse the feature map information of each feature channel, and then using the Sigmoid function to normalize the weights;

[0162] The scaling operation is as follows: mapping the weights output after the excitation operation to the weights of a group of feature channels, and then multiplying and weighting them with the features of the original feature map to achieve feature recalibration of the original features in the channel dimension.

[0163] More specifically, the improvement of the Yolov5n network model also includes adding a 64-fold downsampling detection layer to make the scale of the output feature map 20×20.

[0164] S3: Construct an end-to-end trainable object-centered two-stream convolutional autoencoder network to extract the appearance feature vectors and motion feature vectors of each pig in the target screenshot, and perform feature fusion to form the feature vector of the corresponding frame;

[0165] The two-stream convolutional autoencoder network is trained only with images of normal pig behaviors;

[0166] S4: Use the K-means clustering algorithm to cluster the fused feature vectors, and input the results into a binary classifier for training to obtain a trained classifier; S5: In each frame of the image, obtain the classification scores of all target screenshots in the current frame image through the classifier, and combine all the classification scores to form the anomaly prediction map of the current frame image;

[0167] More specifically, in step S5, select a target screenshot from the current frame image, extract the feature vector of the selected target screenshot through step S3 and cluster it into k clusters, then input the clustering results into k classifiers respectively to obtain k classification scores, and select the highest classification score as the anomaly score of the selected target screenshot. Repeat this step until the anomaly classification scores of all target screenshots in the current frame image are obtained.

[0168] More specifically, the classifier is a binary classifier, and the definition of the i-th binary classifier is as follows:

[0169]

[0170] where, w j represents the weight vector, b represents the bias value, x represents the sample input to the binary classifier, x can be classified as a normal sample or an abnormal sample, x j represents the j-th element of the sample, and m represents the dimension of x.

[0171] More specifically, k binary classifiers are trained through the following steps:

[0172] A1: Select images of normal pig behaviors from pig life videos as training images;

[0173] A2: Use the improved Yolov5n network model to perform object detection and cropping on the training images, and obtain the target screenshots of each pig in the training images respectively;

[0174] A3: Convert the target screenshots into grayscale images, and obtain the corresponding frame difference images by subtracting the pixel values from the adjacent frame images of the training images;

[0175] A4: Use the grayscale frame images and grayscale frame difference images obtained in step A3 as the inputs of the appearance sub-network and the action sub-network in the object-centered convolutional autoencoder network for abnormal pig behaviors respectively, and extract the appearance feature vectors and action feature vectors of each pig in the target screenshots through this network;

[0176] The autoencoder network includes an appearance sub-network and an action sub-network. The appearance sub-network is used to extract appearance feature vectors from the target screenshots, and the action sub-network is used to extract action feature vectors from the frame difference images;

[0177] More specifically, both the appearance sub-network and the action sub-network include an attention module and a memory module; among them,

[0178] The calculation formula of the attention module is:

[0179]

[0180]

[0181] u t,t′ = a(s t-1 , h t′ )

[0182] where c t represents the context vector at time t, T represents the total time length, α t,t′ represents the attention weight in the neighborhood of t' at time t, h t′ represents the output of the hidden unit at the t' -th moment, α represents the attention weight, u t,t′ represents the output score in the neighborhood of t' at time t, u t,k represents the output score in the neighborhood of k at time t, s t-1 represents the hidden state at time t - 1;

[0183] The memory module includes M memory items p m , m = 1, …, M, which are used to record various prototype feature patterns of normal pig behavior data;

[0184] For each query mapping Read the memory item by weighted averaging the memory item p with the corresponding weight and obtain the feature m where,

[0185]

[0186]

[0187] wherein, represents the weight of the memory item p m′ and p m′ represents the m'-th memory item;

[0188] The update formula of the memory item is:

[0189]

[0190]

[0191]

[0192] wherein, ← represents the update operation, f represents the L2 norm, and v t ′k,m represents the reconstruction of the matching probability value and represents the query index set of the memory memory module.

[0193] More specifically, when updating the memory item, if the weighted score ε of the t-th frame image t is greater than the preset threshold, the t-th frame image is regarded as an abnormal frame and the abnormal frame is not used to update the memory item;

[0194] The weighted score ε is calculated by the following formula t :

[0195]

[0196]

[0197] wherein, represents the weight value of the feature, represents a certain feature within the t neighborhood, and I t represents the feature at the t-th moment, and i, j represent the spatial indices.

[0198] In the specific implementation process, when both normal and abnormal samples exist, to prevent the memory from recording the features of abnormal pig samples, a weighted rule scoring score is used to measure the abnormality of video frames, and the memory item is updated only when the frame is determined to be normal.

[0199] More specifically, the loss function of the autoencoder is:

[0200]

[0201] where, is the reconstruction error, is the feature compactness loss function, is the feature separation loss function, is a hyperparameter.

[0202] More specifically,

[0203] the reconstruction error is:

[0204]

[0205] the feature compactness loss function is:

[0206]

[0207]

[0208] the feature separation loss function is:

[0209]

[0210]

[0211] where, T represents the total time, t represents the time index, l represents the index of the query mapping, K represents the total number of query mappings, represents a certain feature within the t neighborhood, I t represents the feature at the t-th moment, p p represents the query mapping of the nearest memory item, p is the index of the nearest item of the query mapping represents the weight of the m-th memory item, m represents the memory item index, M represents the total number of memory items, p n represents the query mapping of the second-nearest memory item.

[0212] In the specific implementation process, the memory memory module of the autoencoder is trained through the feature compactness function and the feature classification loss function, so as to ensure the diversity and discriminability of the memory items.

[0213] A5: Fuse the appearance feature vector and the action feature vector to obtain the fused feature vector of the training image;

[0214] A6: Perform k-means clustering on the fused feature vector to obtain the clustering result cluster i, where i = 1, 2,...., k;

[0215] A7: Input the clustering result into k binary classifiers to obtain k trained binary classifiers.

[0216] S6: Perform Gaussian filtering temporal smoothing on the anomaly prediction map of the current frame image, and record the highest classification score as the anomaly score of the current frame image;

[0217] S7: Determine whether the anomaly score of the current frame image is positive;

[0218] If so, there is no abnormal behavior of the pigs in the current frame image;

[0219] If not, there is abnormal behavior of the pigs in the current frame image.

[0220] Obviously, the above embodiments of the present invention are merely examples for clearly explaining the present invention, rather than limiting the implementation manners of the present invention. For those of ordinary skill in the art, other different forms of changes or variations can be made based on the above description. It is not necessary and impossible to list all the implementation manners here. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included in the protection scope of the claims of the present invention.

Claims

1. A method for detecting abnormal behaviors of pigs, characterized in that, it includes the following steps: S1: Real-time collect the living videos of pigs, and extract images frame by frame from the living videos of pigs; S2: Use the improved Yolov5n network model to perform object detection and cropping on the extracted images of each frame, and respectively obtain the target screenshots of each pig in each frame of the image; The improved Yolov5n network model is: add channel attention modules after the 4th, 6th, and 8th layers of the backbone feature extraction network of the existing Yolov5n network model, splice the channel attention modules with the upsampling layers of the 18th, 22nd, and 26th layers of the neck network, and add a C3 layer and a channel attention module after the 11th layer of the backbone feature extraction network; S3: Construct an end-to-end trainable and object-centered two-stream convolutional autoencoder network for extracting the appearance feature vectors and motion feature vectors of each pig in the target screenshots, and perform feature fusion to form the feature vectors of the corresponding frame; The two-stream convolutional autoencoder network includes an appearance sub-network and an action sub-network, and both sub-networks include an attention-based convolutional LSTM module and a memory module; wherein, The calculation formula of the attention module is: Among them, represents the context vector at the moment, represents the total time length, represents at time t the attention weight of the neighborhood, represents the output of the hidden unit at the moment, represents the attention weight, represents t at the moment the output score of the neighborhood, represents t at the moment k the output score of the neighborhood, represents the hidden state at the moment; The memory module includes M memory items , , which are used to record various prototype feature patterns of the normal behavior data of pigs; For each query mapping , a memory item is read by weighted averaging the memory items with corresponding weights to obtain a feature : ​ Among them, represents the weight of the memory item , represents the th memory item; The update formula of the memory item is: Among them, represents an update operation, represents the L2 norm, represents the matching probability value of the reconstruction, represents the query index set of the memory module; The two-stream convolutional autoencoder network is only trained with images of normal pig behaviors; S4: Use the K-means clustering algorithm to cluster the fused feature vectors, and input the results into a binary classifier for training to obtain a trained classifier; S5: In each frame of the image, obtain the classification scores of all target screenshots in the current frame of the image through the classifier, and combine all the classification scores to form an abnormal prediction map of the current frame of the image; S6: Perform Gaussian filtering time series smoothing on the abnormal prediction map of the current frame of the image, and record the highest classification score as the abnormal score of the current frame of the image; S7: Determine whether the abnormal score of the current frame of the image is a positive number; If so, there is no abnormal behavior of the pigs in the current frame of the image; If not, there is an abnormal behavior of the pigs in the current frame of the image.

2. A method for detecting abnormal behaviors of pigs according to claim 1, characterized in that, the channel attention module includes compression, excitation, and scaling operations; wherein, The compression operation is: use global average pooling to compress the dimension H*W*C of the original feature layer into 1*1*C; The excitation operation is: use two fully connected layers to fuse the feature map information of each feature channel, and then use the Sigmoid function to normalize the weights; The scaling operation is: map the weights output after the excitation operation to a set of weights of the feature channels, and then multiply and weight them with the features of the original feature map to achieve feature recalibration of the original features in the channel dimension.

3. A method for detecting abnormal behaviors of pigs according to claim 1, characterized in that, the improvement of the Yolov5n network model also includes adding a 64-fold downsampling detection layer to make the scale of the output feature map 20×20.

4. A method for detecting abnormal behaviors of pigs according to claim 1, characterized in that, In step S5, a target screenshot is selected from the current frame image. The feature vectors of the selected target screenshot are extracted through step S3 and clustered into k clusters, and then the clustering results are respectively input into k classifiers to obtain k classification scores. The highest classification score is selected as the anomaly score of the selected target screenshot. Repeat this step until the anomaly classification scores of all target screenshots in the current frame image are obtained.

5. A method for detecting abnormal behaviors of pigs according to claim 4, characterized in that, The classifier is a binary classifier, and the definition of the i th binary classifier is as follows: Among them, represents the weight vector, represents the bias value, x represents the sample input to the binary classifier, x which can be classified as a normal sample or an abnormal sample, represents the -th element of the sample, represents x the dimension of.

6. A method for detecting abnormal behaviors of pigs according to claim 5, characterized in that, Train through the following steps k binary classifiers: A1: Select images of normal behaviors of pigs from the pig life videos as training images; A2: Use an improved Yolov5n network model to perform object detection and cropping on the training images, and obtain target screenshots of each pig in the training images respectively; A3: Convert the target screenshots into grayscale images, and obtain corresponding frame difference images by subtracting the pixel values from the adjacent frame images of the training images; A4: Use the grayscale frame images and grayscale frame difference images obtained in step A3 as the inputs of the appearance sub-network and the motion sub-network in the pig abnormal behavior convolutional auto-encoder network centered on the object respectively, and extract the appearance feature vectors and motion feature vectors of each pig in the target screenshots through the appearance sub-network and the motion sub-network; The auto-encoder network includes an appearance sub-network and a motion sub-network. The appearance sub-network is used to extract appearance feature vectors from the target screenshots, and the motion sub-network is used to extract motion feature vectors from the frame difference images; A5: Fuse the appearance feature vectors and motion feature vectors to obtain the fused feature vectors of the training images; A6: Perform k-means clustering on the fused feature vectors to obtain the clustering result clusters i , i = 1, 2,...., k ; Input the clustering result into k binary classifiers to obtain k trained binary classifiers.

7. A method for detecting abnormal behaviors of pigs according to claim 6, characterized in that, When updating a memory item, if the score after weighting the frame image is greater than a preset threshold, the frame image is regarded as an abnormal frame, and the abnormal frame is not used to update the memory item; The weighted score is calculated by the following formula : in, represents the weight of the feature, express t A feature in the neighborhood, Indicates t The characteristics of the moment, Represents the spatial index.

8. A method for detecting abnormal behaviors of pigs according to claim 6, characterized in that, The loss function of the autoencoder is as follows: wherein, is the reconstruction error, is the feature compactness loss function, is the feature separation loss function, , are hyperparameters.

9. A method for detecting abnormal behaviors of pigs according to claim 8, characterized in that, The reconstruction error is: The feature compactness loss function is: The feature separation loss function is: Among them, represents the total time, represents the time index, represents the index of the query mapping, represents the total number of query mappings, represents t a certain feature within the neighborhood, represents the t feature at the represents the query mapping of the nearest memory item, is the index of the nearest item for the query mapping , represents the m weight of the th memory item, represents the memory item index, represents the total number of memory items, represents the second - nearest memory item of the query mapping.

Citation Information

Patent Citations

  • Modeling method for prediction model of dissolved gas concentration in transformer oil

    CN112734028A

  • Abnormal behavior detection method based on appearance and action feature dual prediction

    CN113762007A