Animal abnormal behavior detection method based on skeleton sequence
Through spatiotemporal feature modeling and deep learning methods based on skeleton sequences, combined with self-supervised comparative learning, the problem of insufficient accuracy and robustness of animal behavior detection in complex environments is solved, and efficient abnormal behavior recognition is achieved.
Patent Information
- Application Number
- CN202510485443.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-17
- Publication Date
- 2025-07-18
AI Technical Summary
Existing animal behavior detection methods have poor accuracy and robustness in complex environments, making it difficult to effectively distinguish between normal and abnormal behaviors, especially lacking effective mining when processing spatiotemporal information.
Using a skeleton sequence-based method, the spatial and temporal characteristics of animal behavior are extracted through adaptive adjacency matrix and graph convolution network, and combined with deep learning and self-supervised comparison learning, the detection of animal abnormal behavior is carried out.
It realizes efficient and accurate animal abnormal behavior detection in complex environments, improves the robustness and generalization ability of detection, and can adapt to behavioral analysis of different scenarios and animal species.
Smart Images

Figure CN120340069A_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to animal behavior detection, and in particular to an animal abnormal behavior detection method based on skeleton sequence. Background Art
[0002] With the rapid development of artificial intelligence technology, animal behavior analysis has become a research hotspot in the fields of biology, agriculture, environmental protection, etc. Traditional animal behavior analysis methods usually rely on manual observation or simple sensor monitoring, which has problems such as strong subjectivity, low efficiency, and high cost. In recent years, animal behavior recognition technology based on video analysis has gradually become the mainstream of research. Through methods such as computer vision and deep learning, it can realize automatic and real-time monitoring of animal behavior. However, these methods still face some challenges, especially in dealing with abnormal behavior detection in complex environments.
[0003] Abnormal behavior of animals usually manifests as movements or postures that are significantly different from normal behavior patterns, which may be caused by pathological changes, environmental stress, malnutrition and other factors. Early detection of abnormal behavior is of great significance for disease warning, animal welfare monitoring and production efficiency improvement. However, most existing abnormal behavior detection methods rely on traditional image processing technology or behavior pattern comparison, lack of sufficient mining of spatiotemporal information, resulting in poor detection accuracy and robustness.
[0004] As an effective means of posture recognition, skeleton sequence analysis technology can capture the position information of key nodes of animals at different time points and has strong spatiotemporal expression capabilities. In recent years, skeleton sequence analysis methods based on deep learning have been widely used in human behavior recognition, motion analysis and other fields, but their application in animal behavior recognition is still in the exploratory stage. Traditional animal behavior recognition methods mostly rely on image-level feature extraction, ignoring the detailed information of animal movement and unable to effectively distinguish between normal and abnormal behaviors.
[0005] Therefore, how to combine skeleton sequence technology and deep learning methods to accurately extract the spatiotemporal features of animal behavior, especially the detection of abnormal behavior, has become a key issue that needs to be urgently solved in this field. Summary of the invention
[0006] In view of the above-mentioned deficiencies in the prior art, the present invention provides a method for detecting abnormal animal behavior based on skeleton sequences, which solves the problems of low accuracy and poor robustness in detecting abnormal animal behavior in complex environments.
[0007] In order to achieve the above purpose, the technical solution adopted by the present invention is: a method for detecting abnormal behavior of animals based on skeleton sequences, comprising the following steps: S1, collect animal behavior videos and preprocess them, and build a dataset based on skeleton sequences; S2. Spatiotemporal feature modeling of the skeleton sequence based on the dataset; S3. Detection of abnormal animal behaviors based on the time feature modeling; S4. Determine whether all video frames have been processed. If so, go to S5; otherwise, return to S1; S5. Output the animal behavior classification and the abnormal detection result based on the judgment result.
[0008] Further, the specific content of S1 is as follows: S101. Collect animal behavior videos and perform preprocessing; S102. Detect the preprocessed animal behavior videos, and extract the bounding box and feature information of the animal in each frame; S103. Based on the extracted animal bounding box and feature information, perform skeleton key point detection on the animal in each frame, and extract the key point coordinates of each part of the animal body; S104. Based on the extracted key point coordinates, organize the skeleton key point data of each frame into a skeleton sequence in chronological order, and perform data annotation to construct a dataset including normal and abnormal animal behaviors.
[0009] Still further, the specific content of S2 is as follows: S201. Based on the dataset, use an adaptive adjacency matrix to model the two-dimensional coordinate points of the skeleton sequence, dynamically learn the association weights between joints, and extract spatial features; S202. Use sequence modeling technology to mine the dynamic change features of the time series, capture the change rules of animal behaviors, and extract time features; S203. Fuse the spatial features obtained in S201 and the time features obtained in S202 to complete the spatiotemporal feature modeling.
[0010] Still further, the specific content of S201 is as follows: S2011. Based on the dataset, convert each frame of data in the skeleton sequence into the form of two-dimensional coordinate points, where the two-dimensional coordinate point form represents the position of each joint in space; S2012. Based on the conversion result, construct an initial adaptive adjacency matrix , where the initial adjacency matrix determines the preliminary connection relationship between joints according to the prior structure of the animal skeleton; S2013. Dynamically adjust the initial adaptive adjacency matrix using trainable parameters to obtain an updated adaptive adjacency matrix , and use the updated adaptive adjacency matrix Together with two-dimensional coordinate point data as the input of the graph convolutional network, spatial features are extracted by graph convolutional operations, where an adaptive adjacency matrix is used to dynamically learn the association weights between joints;
[0011] Among them, represents the spatial feature representation obtained after graph convolutional operations, represents the t two-dimensional coordinates of all joints in the
[0012] Furthermore, the specific content of S2011 is as follows: Based on the dataset, let the image width and height of each frame in the skeleton sequence be W and H ; Based on the image width and height, the pixel coordinates of each joint are normalized using the following formula to complete the conversion of the two-dimensional coordinate point form:
[0013]
[0014] Among them, and respectively represent the normalized two-dimensional coordinates, and the range is .
[0015] Furthermore, the specific content of S202 is as follows: S2021. Adopt a sliding window mechanism to analyze the changes in the positions of skeleton joints within each time window and extract the dynamic change features of the time series; S2022. Based on the extracted dynamic change features, capture the temporal change trend of joint movement by calculating the change rate of skeleton joints to extract the temporal change features; S2023. Combine the temporal change trend during joint movement to model the time dependence of the skeleton sequence and complete the extraction of time features.
[0016] Furthermore, the specific content of S2021 is as follows: S20211. Set the sliding window size to T , and divide the skeleton sequence into multiple time windows according to the time step. Among them, each window contains T consecutive skeleton frame data; S20212. For each time window, use the following formula to calculate the position change amount of each joint within the window:
[0017] Among them, represents the change amount of the i -th joint between the current moment and the previous moment, represents the t coordinate of the i -th joint at the moment, t represents the coordinate of the i -th joint at the S20213. Extract the dynamic change features of the time series through the position change amounts of each joint to capture the change rules of animal behaviors.
[0018] Furthermore, the specific content of S2022 is as follows: Calculate the displacement change of the skeleton joints between consecutive time frames based on the extracted dynamic change features; Calculate the speed and acceleration of each joint based on the displacement change; Combine the speed and acceleration features, and use weighted average to model the joint changes to obtain the temporal change trend of joint movement, so as to extract the temporal change features. Among them, the expression of the temporal change trend is as follows:
[0019]
[0020]
[0021]
[0022] Among them, represents the temporal change rate of the overall joints at the t moment, and both represent the weight coefficients used to control the contributions of speed and acceleration to the change rate, N represents the total number of joints, represents the i -th joint at the t moment, represents the i -th joint at the t moment, represents the time interval, represents the i -th joint at the moment, i represents the t -th joint at the moment, i represents the t -th joint at the Indicates the i position of the t -1th joint at time
[0023] Furthermore, the S203 is specifically as follows: S2031. Obtain the spatio-temporal joint feature representation of the skeleton sequence according to the spatial and temporal features; S2032. According to the spatio-temporal joint feature representation, use the graph convolutional network to model the spatio-temporal dependence relationship, and model the spatio-temporal features. The expression for modeling the spatio-temporal features is as follows:
[0024]
[0025] where, represents the final spatio-temporal fusion feature, represents the graph convolution operation, represents the time convolution operation, represents the spatio-temporal feature at each time t moment, represents the weighting factor, which controls the program of fusing the spatial and temporal features, represents the spatial feature of the t th frame of the skeleton sequence, represents the spatial feature of the t th frame of the skeleton sequence.
[0026] Furthermore, the S3 is specifically as follows: S301. Use the spatio-temporal features of the skeleton sequence to classify animal behaviors through the trained classification model, and automatically identify normal animal behaviors and abnormal animal behaviors. The expression for the classification probability vector is as follows:
[0027] where, represents the classification probability vector, and respectively represent the weights and biases of the fully connected layer, represents the final spatio-temporal fusion feature; S302. Adopt a self-supervised contrastive learning strategy, construct positive and negative sample pairs, calculate the similarity between the skeleton spatio-temporal features and the normal animal behavior pattern, and use the triplet loss function to strengthen the recognition of the abnormal animal behavior pattern. The expression for the triplet loss function is as follows:
[0028] where, represents the triplet loss function, represents the Euclidean distance, represents the spatio-temporal features of the anchor sample, represents the features of normal animal behavior, represents the features of abnormal animal behavior, represents a preset marginal threshold; S303. When an abnormal animal behavior is detected, use the output result of the classification model based on spatio-temporal features, mark the behavior as abnormal using a preset decision function, and classify the abnormal animal behavior through a clustering algorithm to complete the detection of abnormal animal behavior. Among them, the expression of the decision function is as follows:
[0029] Among them, represents the decision function, represents the abnormal animal behavior, represents the normal animal behavior, represents the probability of abnormal animal behavior, represents the threshold.
[0030] The beneficial effects of the present invention are: (1) The present invention utilizes the spatio-temporal dynamic features of the key points of the animal skeleton in the video image, combines advanced algorithms such as deep learning and graph convolution, and realizes the automatic real-time detection and classification of normal and abnormal animal behaviors. In addition, the present invention combines technologies such as adaptive adjacency matrix and self-supervised contrast learning, improves and innovates on the basis of traditional skeleton sequence analysis methods, not only improves the model's ability to express complex behavior patterns, but also enhances the robustness and generalization ability of the system in actual application scenarios.
[0031] (2) The present invention proposes an efficient method for detecting abnormal animal behaviors based on skeleton sequences by performing spatio-temporal feature modeling on the skeleton sequences in animal behavior videos and combining abnormal behavior recognition algorithms. This method can automatically detect abnormal animal behaviors in complex environments, has high accuracy and robustness, and provides a new technical means for animal behavior monitoring and health management.
[0032] (3) By combining spatio-temporal feature modeling of skeleton sequences with abnormal behavior detection, the present invention can accurately capture the spatial and temporal dynamic changes in animal behaviors, thereby realizing efficient abnormal behavior recognition. The present invention can deeply analyze animal behaviors and effectively improve the accuracy of behavior classification and abnormal detection.
[0033] (4) The present invention adopts a fusion method of spatio-temporal features and time series features, and conducts comprehensive modeling based on multi-dimensional information of skeleton data, thereby enhancing the recognition ability of complex animal behavior patterns and being able to accurately distinguish normal behaviors from abnormal behaviors. This method not only improves the reliability of abnormal behavior detection but also can effectively cope with the diversity challenges under different animal species and behavior scenarios.
[0034] (5) The present invention conducts dynamic modeling on the skeleton sequence, combines deep learning algorithms, processes each moment in the video frame in real time, and can effectively identify long-term dependence relationships. This technical solution is particularly applicable to real-time monitoring systems, providing technical support for real-time detection and emergency handling of animal behaviors.
[0035] (6) The algorithm design of the present invention can adaptively process animal behaviors under various environmental conditions, has strong generalization ability, can adapt to behavior pattern analysis of different scenarios and different animals, and has good application prospects. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] Figure 1 It is a flowchart of the method of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0037] The following describes the specific embodiments of the present invention to facilitate those skilled in the art to understand the present invention. However, it should be clear that the present invention is not limited to the scope of the specific embodiments. For those ordinary skilled in the art, as long as various changes are within the spirit and scope of the present invention defined and determined by the appended claims, these changes are obvious, and all inventions created using the concept of the present invention are within the scope of protection.
[0038] EXAMPLE As Figure 1 shown, the present invention provides a method for detecting abnormal animal behaviors based on skeleton sequences, and its implementation method is as follows: S1. Collect animal behavior videos and perform preprocessing, and construct a data set based on skeleton sequences. The implementation method is as follows: S101. Collect animal behavior videos and perform preprocessing; S102. Detect the preprocessed animal behavior videos, and extract the bounding boxes and feature information of the animals in each frame; S103. Based on the extracted animal bounding boxes and feature information, perform skeleton key point detection on the animals in each frame, and extract the key point coordinates of each part of the animal body; S104. Based on the extracted key point coordinates, organize the skeleton key point data of each frame into a skeleton sequence in chronological order, and perform data annotation to construct a data set including normal and abnormal animal behaviors.
[0039] In this embodiment, collecting and preprocessing the animal behavior video includes: obtaining the animal behavior video using a video acquisition device, and parameters such as the quality and resolution of video frames should be ensured to be as clear as possible during acquisition. Through preprocessing means such as denoising, image enhancement, and cropping, ensure that the video data is suitable for subsequent processing.
[0040] In this embodiment, extracting the bounding box and feature information of the animal includes: using the YOLOv8 object detection algorithm to detect the animals in the video frames, obtaining the bounding box of the animals and their feature information (such as posture, movement trajectory, etc.) in each frame, and using it for subsequent positioning of skeleton key points.
[0041] In this embodiment, in the skeleton key point detection, the HR-Net pose estimation algorithm is used to detect the skeleton key points of the animals in each frame, and extract the key point coordinates of each part of the animal body for subsequent analysis and processing.
[0042] In this embodiment, during the process of organizing the skeleton sequence, annotating data and establishing a dataset, the skeleton key point data of each frame is organized into a skeleton sequence in chronological order, and data annotation is performed in combination with the actual application requirements to establish a dataset containing normal behavior and abnormal behavior labels.
[0043] S2. Based on the dataset, perform spatio-temporal feature modeling on the skeleton sequence, and its implementation method is as follows: S201. Based on the dataset, use an adaptive adjacency matrix to model the two-dimensional coordinate points of the skeleton sequence, dynamically learn the correlation weights between joints, and extract spatial features. Its implementation method is as follows: S2011. Based on the dataset, convert the data of each frame in the skeleton sequence into the form of two-dimensional coordinate points, where the two-dimensional coordinate point form represents the position of each joint in space. Its implementation method is as follows: Based on the dataset, let the image width and height of each frame in the skeleton sequence be W and H ; Based on the image width and height, use the following formula to normalize the pixel coordinates of each joint to complete the conversion of the two-dimensional coordinate point form; S2012. Based on the conversion result, construct an initial adaptive adjacency matrix , where the initial adjacency matrix determines the preliminary connection relationship between joints according to the prior structure of the animal skeleton; S2013. Dynamically adjust the initial adaptive adjacency matrix using trainable parameters to obtain the updated adaptive adjacency matrix , and use the updated adaptive adjacency matrix Together with two-dimensional coordinate point data as the input of the graph convolutional network, spatial features are extracted using graph convolutional operations, where an adaptive adjacency matrix is used to dynamically learn the correlation weights between joints; S202. Use sequence modeling technology to mine the dynamic change characteristics of the time series, capture the changing patterns of animal behavior, and extract time features. The implementation method is as follows: S2021. Adopt a sliding window mechanism to analyze the changes in the positions of skeleton joints within each time window and extract the dynamic change characteristics of the time series. The implementation method is as follows: S20211. Set the sliding window size to T and divide the skeleton sequence into multiple time windows according to the time step. Among them, each window contains T consecutive skeleton frame data; S20212. For each time window, calculate the amount of change in the positions of each joint within the window; S20213. Extract the dynamic change characteristics of the time series through the amount of change in the positions of each joint to capture the changing patterns of animal behavior; S2022. Based on the extracted dynamic change characteristics, capture the temporal change trend of joint movement by calculating the change rate of skeleton joints to extract temporal change features. The implementation method is as follows: Based on the extracted dynamic change characteristics, calculate the displacement change between consecutive time frames of skeleton joints; Based on the displacement change, calculate the velocity and acceleration of each joint; Combining the velocity and acceleration features, use weighted averaging to model the joint changes to obtain the temporal change trend of joint movement and extract temporal change features; S2023. Combine the temporal change trend during joint movement to model the time dependence of the skeleton sequence and complete the extraction of time features; S203. Fuse the spatial features obtained in S201 and the time features obtained in S202 to complete the modeling of spatio-temporal features. The implementation method is as follows: S2031. Obtain the spatio-temporal joint feature representation of the skeleton sequence according to the spatial features and time features; S2032. According to the spatio-temporal joint feature representation, use the graph convolutional network to model the spatio-temporal dependence relationship and model the spatio-temporal features.
[0044] In this embodiment, during the process of spatio-temporal feature modeling of the skeleton sequence to identify animal behaviors, by combining temporal information and spatial features, a spatio-temporal feature model adapted to animal behavior recognition is constructed. Specifically, since animal behaviors exhibit dynamic changes and the spatial positions of each skeleton key point continuously adjust over time, it is necessary to model through a deep learning model to extract the spatio-temporal features in the skeleton sequence.
[0045] During the spatio-temporal modeling process, first, according to the temporal data of the skeleton key points, the dynamic change features in the time dimension are extracted through a temporal neural network. Then, a graph convolutional network is used to model the spatial structure information to capture the spatial dependence relationships between various parts of the animal. In this way, the model can not only learn the posture of the animal at each time point but also identify the spatio-temporal change patterns of animal behaviors.
[0046] In the specific implementation, the model extracts features for each segment of the skeleton sequence and performs temporal analysis on the skeleton data in different time periods to identify potential animal behaviors. The classification basis of behaviors includes the movement patterns of the limbs, the relative position changes between key points, and the overall posture transformation. By fusing spatio-temporal information, this process can effectively identify complex animal behavior patterns and provide accurate input data for subsequent abnormal behavior detection.
[0047] Finally, the behavior recognition results output by the model can be used for subsequent abnormal behavior detection, thereby achieving precise monitoring and analysis of animal behaviors.
[0048] In this embodiment, the specific method for spatio-temporal feature modeling of the skeleton sequence is as follows: An adaptive adjacency matrix is used to model the two-dimensional coordinate points of the skeleton sequence, dynamically learning the correlation weights between joints to extract more accurate spatial relationships and time dynamic features; sequence modeling techniques are used to deeply explore the time characteristics of the skeleton sequence to capture the behavior change rules; the extracted spatial features and time features are fused to generate a more expressive behavior feature representation.
[0049] In this embodiment, an adaptive adjacency matrix is used to model the two-dimensional coordinate points of the skeleton sequence, dynamically learning the correlation weights between joints to extract more accurate spatial relationships and time dynamic features. The specific method is as follows: Each frame of data in the skeleton sequence is converted into the form of two-dimensional coordinate points to represent the position of each joint in space; an initial adjacency matrix is constructed , and the preliminary connection relationships between joints are determined based on the prior structure of the animal skeleton; the updated adaptive adjacency matrix Together with the two-dimensional coordinate point data, it serves as the input of the graph convolutional network, and extracts spatial features through graph convolutional operations; combines time convolutional operations to capture dynamic changes in the time dimension, combines the spatial features extracted by the graph convolutional network with the time convolutional layer, and extracts the dynamic change features of the time series through convolutional operations, so as to realize the modeling of the spatio-temporal features of the skeleton.
[0050] In order to effectively model the spatio-temporal features of the skeleton sequence, in this embodiment, the skeleton joint positions in each frame are first transformed into two-dimensional coordinate points through normalization processing. Next, the initial connection relationship between each joint is determined according to the prior structure of the animal skeleton to construct an initial adjacency matrix. Through these processing steps, it is possible to provide basic data support for subsequent spatio-temporal feature fusion and abnormal behavior detection. The following steps will further describe how to extract and combine spatial and temporal features through graph convolution and time convolution operations.
[0051] In this embodiment, each frame of data in the skeleton sequence is transformed into the form of two-dimensional coordinate points to represent the position of each joint in space. The specific method is as follows: Set the image width and height of each frame to be W and H ; Normalize the pixel coordinates of each joint, and the calculation method is:
[0052]
[0053] where and respectively represent the normalized two-dimensional coordinates, and the range is .
[0054] In this embodiment, an initial adjacency matrix is constructed, and the initial connection relationship between each joint is determined according to the prior structure of the animal skeleton. Subsequently, by introducing an adaptive learning module, feature extraction is performed on each frame of skeleton data, and the adjacency matrix is dynamically adjusted using trainable parameters to update and obtain an adaptive adjacency matrix . This matrix can learn and optimize the association weights between joints according to the feature distribution of the input data, so as to better reflect the real-time spatial dependence relationship.
[0055] In this embodiment, the updated adaptive adjacency matrix and the two-dimensional coordinate point data are jointly used as the input of the graph convolutional network, and spatial features are extracted through graph convolutional operations. The specific method is as follows: Calculate the spatial features of the t th frame:
[0056] where Represents the spatial feature representation obtained after the graph convolution operation. Denotes the t Two-dimensional coordinates of all joints in the
[0057] In this embodiment, a temporal convolution operation is combined to capture the dynamic changes in the temporal dimension. The spatial features extracted by the graph convolutional network are combined with the temporal convolutional layer, and the dynamic change features of the time series are extracted through the convolution operation, so as to realize the modeling of the spatio-temporal features of the skeleton. The specific method is as follows:
[0058] Among them, Represents the dynamic change features of the time series. Denotes the activation function. Represents the temporal convolution kernel. Represents the spatial features of the previous frame. Denotes the bias term.
[0059] In this embodiment, after the spatio-temporal features of the skeleton sequence are extracted, sequence modeling techniques are used to deeply mine the temporal characteristics of the skeleton sequence and further deeply model the temporal features. By analyzing the change law of the skeleton sequence in the temporal dimension, this step can reveal the dynamic characteristics of the behavior and capture the temporal trend of joint movement. Based on the extraction results of the spatio-temporal features, through the sliding window mechanism and the calculation of the change rate, the temporal sequence changes of joint actions are accurately captured, and the corresponding time-dependent model is constructed, so as to effectively enhance the recognition ability of abnormal behavior patterns.
[0060] Based on this, sequence modeling techniques are used to deeply mine the temporal characteristics of the skeleton sequence and capture the behavior change law. The specific method is as follows: The sliding window mechanism is adopted to analyze the change of the skeleton joint position in each time window, and the dynamic features of the time series are extracted; by calculating the change rate of the skeleton joint, the temporal change trend of joint movement is captured, and the temporal features are extracted; combined with the feature change trend within the time window, a time-dependent model is constructed to further enhance the recognition accuracy of abnormal behavior patterns.
[0061] After the spatio-temporal feature modeling of the skeleton sequence is completed, the sliding window mechanism is adopted to analyze the change of the skeleton joint position in each time window and extract the dynamic features of the time series. This step further focuses on the change of the skeleton joint in the temporal dimension. For this purpose, by setting the window size 𝑇 and dividing the skeleton sequence into multiple time windows according to the time step, each window contains 𝑇 consecutive skeleton frame data. In each time window, the displacement change amount of each joint is calculated, so as to lay a foundation for subsequent temporal feature extraction and further capture the temporal dependence of the skeleton sequence.
[0062] In this embodiment, the window size is set to \(T\), and the skeleton sequence is segmented into multiple time windows according to the time step, where each window contains \(T\) consecutive skeleton frame data.
[0063] For each time window, calculate the position changes of each joint within the window, using the following formula:
[0064] Where, represents the change amount of the \( i -th joint between the current moment and the previous moment, represents the coordinate of the \( t -th joint at the \( i -th moment, represents the coordinate of the \( t -1-th moment of the \( i -th joint.
[0065] By calculating the change amounts of each joint, extract the dynamic features of the time series and use them as the input features of the time window to capture the time dependence of the skeleton sequence.
[0066] In this embodiment, first calculate the displacement change of the skeleton joints between consecutive time frames, using the following formula:
[0067] By calculating the velocity and acceleration of each joint, further extract the time series change features. The calculation formulas for velocity and acceleration are:
[0068]
[0069] Combining the velocity and acceleration features, use the weighted average method to model the joint changes and strengthen the recognition ability of abnormal behavior patterns. The calculation formula is:
[0070] Where, represents the time series change rate of the overall joints at the \( t -th moment, and both represent the weight coefficients used to control the contributions of velocity and acceleration to the change rate, N represents the total number of joints, represents the velocity of the \( i -th joint at the \( t -th moment, represents the acceleration of the \( i -th joint at the \( t -th moment, represents the time interval It represents the velocity change of the i th joint between two adjacent time steps. It represents the i th joint's t displacement at time It represents the i th joint's t position at time It represents the i th joint's t position at time - 1.
[0071] In this embodiment, after extracting the spatial and temporal features of the skeleton sequence, by fusing the spatial features with the temporal features, the spatio - temporal expression ability of the skeleton sequence is further enhanced. By splicing or weighted - fusing these two types of features, a more representative spatio - temporal joint feature representation is obtained, laying a foundation for subsequent spatio - temporal dependency modeling. Then, a graph convolutional network is used to model the spatio - temporal dependency relationship of these fused features, and finally the spatio - temporal features are input into the abnormal behavior detection module to accurately judge whether the animal's behavior is abnormal.
[0072] Based on this, the extracted spatial features and temporal features are fused to generate a more expressive behavioral feature representation. The specific method is as follows: By splicing or weighted - fusing the spatial features and temporal features, a spatio - temporal joint feature representation of the skeleton sequence is obtained; on the basis of spatio - temporal feature fusion, a graph convolutional network is further used to model the spatio - temporal dependency relationship; the spatio - temporal features are input into the abnormal behavior detection module, and whether the animal's behavior is abnormal is further determined through a classification or regression model. This model makes a decision on abnormal behavior based on the extracted spatio - temporal features.
[0073] In this embodiment, after the spatio - temporal feature extraction of the skeleton sequence is completed, by further fusing the spatial features with the temporal features, the comprehensive expression ability of the model for animal behavior is enhanced. This process effectively combines spatial and temporal information through weighted fusion, providing richer feature support for subsequent behavioral pattern determination. Through this deep fusion of spatio - temporal features, the model can more accurately capture the changing rules of animal behavior in complex behavioral scenarios, thereby improving the accuracy and reliability of abnormal behavior detection.
[0074] In this embodiment, let the spatial feature of the th frame of the skeleton sequence be , and the temporal feature be , and they can be fused through the following formula:
[0075] In this embodiment, let the spatio - temporal feature at each time t be , capture spatial structure dependencies through graph convolution operations and capture temporal dependencies using temporal convolutions. The specific formula is as follows:
[0076] Among them, represents the final spatio-temporal fusion feature, represents the graph convolution operation, represents the temporal convolution operation, represents at each moment t the spatio-temporal feature, represents the weighting factor, controlling the process of fusing spatial features and temporal features, represents the t spatial feature of the th frame of the skeleton sequence, t represents the temporal feature of the
[0077] S3. Modeling based on temporal features to detect abnormal animal behaviors. The implementation method is as follows: S301. Through the trained classification model, classify animal behaviors using the spatio-temporal features of the skeleton sequence, and automatically identify normal animal behaviors and abnormal animal behaviors; S302. Adopt a self-supervised contrastive learning strategy to construct positive and negative sample pairs, calculate the similarity between the spatio-temporal features of the skeleton and the normal animal behavior pattern, and use the triplet loss function to strengthen the recognition of abnormal animal behavior patterns; S303. When an abnormal animal behavior is detected, based on the output result of the classification model based on spatio-temporal features, use a preset decision function to mark the behavior as abnormal, and classify the abnormal animal behavior through a clustering algorithm to complete the detection of abnormal animal behaviors.
[0078] In this embodiment, after the extraction and fusion of spatio-temporal features are completed, these features will be used to detect abnormal behaviors. Through the trained classification model, first classify the behaviors of the spatio-temporal features of the skeleton sequence. Then, the model calculates the similarity between the current behavior and the normal behavior pattern, and identifies significantly different behavior patterns as abnormal behaviors. Finally, when an abnormal behavior is detected, combined with the classification output of spatio-temporal features, based on the set decision rules, mark and classify the behavior, and further refine the types of abnormal behaviors through a clustering algorithm.
[0079] In this embodiment, the specific method for abnormal behavior detection is as follows: Through a trained classification model, the spatio-temporal features of the skeleton sequence are used for behavior classification to automatically identify normal and abnormal behaviors; A self-supervised contrastive learning strategy is adopted. By constructing positive and negative sample pairs, the similarity between the spatio-temporal features of the skeleton and the normal behavior pattern is calculated, and the triplet loss is used to further strengthen the discriminative ability of the model for abnormal behavior patterns; When an abnormal behavior is detected, based on the output result of the classification model of spatio-temporal features, a preset decision rule is used to mark the behavior as abnormal, and the clustering algorithm is used to further refine the classification of abnormal behavior types.
[0080] In this embodiment, through a trained classification model, the spatio-temporal features of the skeleton sequence are used for behavior classification to automatically identify normal and abnormal behaviors of animals. The specific method is as follows: Using the extracted final spatio-temporal features As the input, behavior classification is performed through a fully connected layer and the Softmax function. The formula is as follows:
[0081] Where, Represents the classification probability vector, And Represent the weight and bias of the fully connected layer respectively, Represents the final spatio-temporal fusion feature. This step can automatically divide behaviors into two major categories: normal and abnormal, laying a foundation for subsequent detection.
[0082] In this embodiment, the training method of the classification model specifically includes the following steps: On the basis of the standard spatio-temporal graph convolutional network, a learnable spatio-temporal attention module is added. Its feature extraction process is defined as:
[0083] Where, Represents the input feature matrix of the l +1 layer, that is, the result after the input feature of the l layer passes through the graph convolution operation and the non-linear activation function processing, Represents the spatio-temporal attention weight (key frames are selected according to this weight when constructing positive sample pairs in S302), Represents the dynamic adjacency matrix, Represents the graph convolution operation, t Represents the time dimension index, T Represents the total number of frames of the input skeleton sequence (i.e., the time step), v Represents the space dimension index, V Represents the total number of skeleton joints.
[0084] Introduce joint optimization of weighted cross-entropy loss and feature compactness loss:
[0085] Among them, represents the joint optimization objective function, represents the total number of behavior categories, represents the true class label of the sample, represents the probability that the model predicts the sample belongs to class , represents class the number of training samples of, represents the th spatio-temporal fusion feature of the sample, is the feature center of class ; represents the L2 norm; In the pre-training stage, the input is the skeleton sequence of the animal behavior video (including normal / abnormal annotations), and the output is the highly discriminative features (the input features of the S303 clustering algorithm); In the fine-tuning stage, the underlying parameters of the spatio-temporal graph convolutional network are frozen, and the parameters of the spatio-temporal attention module are updated (sharing the parameter update range with the contrast learning in S302).
[0086] In this embodiment, a self-supervised contrast learning strategy is adopted. By constructing positive and negative sample pairs, the similarity between the spatio-temporal features of the skeleton and the normal behavior pattern is calculated, and the triplet loss is used to further strengthen the discriminative ability of the model for abnormal behavior patterns. The specific method is as follows: The model is optimized using the triplet loss function. By constructing positive sample pairs (normal behaviors) and negative sample pairs (abnormal behaviors or behavior deviations), the distance relationship between spatio-temporal features is learned. The loss function is defined as:
[0087] Among them, represents the triplet loss function, represents the Euclidean distance, represents the spatio-temporal features of the anchor sample, represents the features of the same class as the anchor sample (normal animal behavior), represents the features of a different class from the anchor sample (abnormal animal behavior), represents the preset margin threshold.
[0088] This self-supervised contrast learning mechanism further enhances the discriminative ability of the model for abnormal behavior patterns by minimizing the triplet loss.
[0089] In this embodiment, when an abnormal behavior is detected, based on the output result of the classification model with spatio-temporal features, a preset decision rule is used to mark the behavior as abnormal, and a clustering algorithm is used to further refine the classification of the abnormal behavior types. The specific method is as follows: When the classification model detects an abnormal behavior, it is marked using a decision rule. The following decision function can be defined:
[0090] Wherein, represents the decision function, represents the abnormal behavior, represents the normal behavior, represents the probability of abnormal behavior, represents the threshold value.
[0091] In addition, to further refine the types of abnormal behaviors, K-means clustering is used to group all samples marked as abnormal. The goal of K-means clustering is to minimize:
[0092] Wherein, represents the set of samples within the i th cluster, represents the i th cluster center. In this way, the abnormal behaviors can be further subdivided into various specific types (such as falling, excessive struggling, abnormal trembling, or attacking, etc.), which is convenient for accurate early warning and subsequent processing.
[0093] S4. Determine whether all video frames have been processed. If so, enter S5; otherwise, return to S1. In this embodiment, it is checked whether the current video frame is the last frame of the video. If the current frame is the last frame, enter S5. If the current frame is not the last frame, return to step S1 to continue processing the next frame.
[0094] S5. Based on the judgment result, output the animal behavior classification and the abnormal detection result.
[0095] In this embodiment, the output behavior classification results include normal behavior categories and abnormal behavior categories. Among them, the normal behavior categories include, but are not limited to, common animal behavior patterns such as standing, walking, running, eating, etc. These behavior patterns usually have stable spatio-temporal characteristics, reflecting the normal activities of animals in specific situations. By training the model, the spatio-temporal laws of these behaviors can be recognized and accurately classified. The abnormal behavior categories include abnormal actions that go beyond the normal behavior patterns, such as falling, excessive struggling, abnormal trembling, or aggressive behaviors, etc. These behaviors usually accompany large spatial changes and temporal fluctuations and rarely occur under normal circumstances. The detection of abnormal behavior categories relies on comparing the differences from the normal behavior patterns, and can effectively identify the deviation of behaviors and give timely feedback alerts.
[0096] Those of ordinary skill in the art will realize that the embodiments described herein are for helping the reader understand the principles of the present invention, and it should be understood that the protection scope of the present invention is not limited to such specific statements and embodiments. Those of ordinary skill in the art can make various other specific deformations and combinations that do not depart from the essence of the present invention according to these technical revelations disclosed in the present invention, and these deformations and combinations are still within the protection scope of the present invention.
Claims
1. An animal abnormal behavior detection method based on skeleton sequences, characterized in that It includes the following steps: S1. Collect animal behavior videos and perform preprocessing, and construct a dataset based on the skeleton sequences; S2. Based on the dataset, perform spatio-temporal feature modeling on the skeleton sequences; S3. Based on the modeling of the time features, detect abnormal animal behaviors; S4. Judge whether all video frames have been processed. If so, enter S5; otherwise, return to S1; S5. Based on the judgment result, output the animal behavior classification and the abnormal detection result.
2. The method for detecting abnormal animal behavior based on the skeleton sequence according to claim 1, wherein, The specific content of S1 is as follows: S101. Collect animal behavior videos and perform preprocessing; S102. Detect the preprocessed animal behavior videos, and extract the bounding boxes and feature information of the animals in each frame; S103. Based on the extracted animal bounding boxes and feature information, perform skeleton key point detection on the animals in each frame, and extract the key point coordinates of each part of the animal body; S104. Based on the extracted key point coordinates, organize the skeleton key point data of each frame into a skeleton sequence in chronological order, and perform data annotation to construct a dataset including normal and abnormal animal behaviors.
3. The method for detecting abnormal animal behavior based on a skeleton sequence according to claim 1, wherein, The specific content of S2 is as follows: S201. Based on the dataset, use an adaptive adjacency matrix to model the two-dimensional coordinate points of the skeleton sequence, dynamically learn the association weights between joints, and extract spatial features; S202. Use sequence modeling techniques to mine the dynamic change features of the time series, capture the change rules of animal behaviors, and extract time features; S203. Fuse the spatial features obtained in S201 and the time features obtained in S202 to complete the modeling of spatio-temporal features.
4. The method for detecting abnormal animal behavior based on a skeleton sequence according to claim 3, wherein The specific content of S201 is as follows: S2011. Based on the dataset, convert the data of each frame in the skeleton sequence into the form of two-dimensional coordinate points, where the two-dimensional coordinate point form represents the position of each joint in space; S2012. Construct an initial adaptive adjacency matrix based on the conversion result , where the initial adjacency matrix determines the preliminary connection relationship between each joint according to the prior structure of the animal skeleton; S2013. Dynamically adjust the initial adaptive adjacency matrix using trainable parameters , and obtain the updated adaptive adjacency matrix . Then, use the updated adaptive adjacency matrix and the two-dimensional coordinate point data as the input to the graph convolutional network, and use graph convolutional operations to extract spatial features. Among them, use the adaptive adjacency matrix to dynamically learn the correlation weights between each joint; Among them, represents the spatial feature representation obtained after the graph convolution operation, represents the t two-dimensional coordinates of all joints in the th frame.
5. The method for detecting abnormal animal behavior based on a skeleton sequence according to claim 4, characterized in that The specific content of S2011 is as follows: Based on the dataset, let the image width and height of each frame in the skeleton sequence be W and H ; Based on the width and height of the image, the pixel coordinates of each joint are normalized using the following formula to complete the conversion in the form of two-dimensional coordinate points: Among them, and respectively represent the normalized two-dimensional coordinates, with the range of .
6. The method for detecting abnormal animal behavior based on a skeleton sequence according to claim 3, wherein The specific content of S202 is as follows: S2021. Adopt a sliding window mechanism to analyze the change of skeleton joint positions within each time window, and extract the dynamic change features of the time series; S2022. Based on the extracted dynamic change features, calculate the displacement change of the skeleton joints between consecutive time frames; S2023. Combine the temporal change trend during joint movement to model the time dependence of the skeleton sequence, and complete the extraction of time features.
7. The method for detecting abnormal animal behavior based on the skeleton sequence according to claim 6, wherein The specific content of S2021 is as follows: S20211. Set the sliding window size to T , and divide the skeleton sequence into multiple time windows according to the time step, where each window contains T consecutive skeleton frame data; S20212. For each time window, use the following formula to calculate the position change amount of each joint within the window: Among them, represents the change amount of the i th joint between the current moment and the previous moment, represents the coordinate of the t th moment of the i th joint, represents the coordinate of the t th joint at the i th moment at -1; S20213. Extract the dynamic change features of the time series through the position change amounts of each joint to capture the change rules of animal behaviors.
8. The method for detecting abnormal animal behavior based on a skeleton sequence according to claim 6, characterized in that The specific content of S2022 is as follows: Based on the extracted dynamic change features, calculate the displacement change of the skeleton joints between consecutive time frames; Based on the displacement change, calculate the speed and acceleration of each joint; Combine the speed and acceleration features, and use weighted average to model the joint change to obtain the temporal change trend of joint movement, so as to extract the temporal change features, where the expression of the temporal change trend is as follows: Among them, represents the temporal change rate of the overall joint at t ; and both represent weight coefficients, which are used to control the contributions of speed and acceleration to the change rate. N represents the total number of joints, represents the i th joint's speed at t ; represents the i th joint's acceleration at t ; represents the time interval, represents the i th joint's speed change amount between two adjacent time steps, represents the i th joint's displacement at t ; represents the i th joint's position at t ; represents the i th joint's position at t - 1.
9. The method for detecting abnormal animal behavior based on the skeleton sequence according to claim 3, wherein The specific content of S203 is as follows: S2031. Obtain the spatio-temporal joint feature representation of the skeleton sequence according to the spatial features and temporal features; S2032. According to the spatio-temporal joint feature representation, use a graph convolutional network to model the spatio-temporal dependence relationship and model the spatio-temporal features. The expression for modeling the spatio-temporal features is as follows: Among them, represents the final spatio-temporal fusion feature, represents the graph convolution operation, represents the temporal convolution operation, represents each moment t of the spatio-temporal feature, represents the weighting factor that controls the process of fusing spatial features and temporal features, represents the spatial feature of the t th frame of the skeleton sequence, represents the temporal feature of the t th frame of the skeleton sequence.
10. The method for detecting abnormal animal behavior based on the skeleton sequence according to claim 1, wherein The specific content of S3 is as follows: S301. Through the trained classification model, that is, the behavior classifier constructed based on the spatio-temporal graph convolutional network, classify animal behaviors using the spatio-temporal features of the skeleton sequence, and automatically identify normal animal behaviors and abnormal animal behaviors. The expression for the classification probability vector is as follows: Among them, represents the classification probability vector, and represent the weights and biases of the fully connected layer respectively, represents the final spatio-temporal fusion feature; S302. Adopt a self-supervised contrastive learning strategy to construct positive and negative sample pairs, calculate the similarity between the spatio-temporal features of the skeleton and the normal animal behavior pattern, and use the triplet loss function to strengthen the recognition of abnormal animal behavior patterns. The expression for the triplet loss function is as follows: Among them, represents the triplet loss function, represents the Euclidean distance, represents the spatio-temporal features of the anchor sample, represents the features of normal animal behavior, represents the features of abnormal animal behavior, represents the preset margin threshold; S303. When an abnormal animal behavior is detected, use the output result of the classification model based on spatio-temporal features, adopt a preset decision function to mark the behavior as abnormal, and classify the abnormal animal behavior through a clustering algorithm to complete the detection of abnormal animal behavior. The expression for the decision function is as follows: Among them, represents the decision function, represents the abnormal behavior of animals, represents the normal behavior of animals, represents the probability of abnormal behavior of animals, represents the threshold value.
Citation Information
Cited By
Space-time sequence data processing method and system for pet abnormal behavior recognition
CN121071757A
Pet abnormal behavior recognition spatio-temporal sequence data processing method and system
CN121071757B
Home health monitoring method and system based on Internet of Things, electronic equipment and storage medium
CN121278668A