Behavior analysis and monitoring method, device and equipment based on pet camera

By extracting motion vector data in the coding domain and combining a streamlined capsule network and a two-way cross-attention mechanism, the problem of insufficient computational burden and three-dimensional understanding of traditional pet monitoring methods is solved, and multi-dimensional accurate description of pet behavior and abnormal behavior recognition are achieved.

CN120219441AInactive Publication Date: 2025-06-27SHENZHEN ANKED SHITONG ELECTRONICS CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510411996.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-02
Publication Date
2025-06-27
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Traditional pet monitoring methods rely on the full decoding processing of RGB video data, resulting in huge storage and computing burdens, which is difficult to meet the real-time monitoring needs of limited computing resources equipment. In addition, existing pet behavior recognition technology lacks a three-dimensional understanding of pet posture and cannot effectively identify abnormal behaviors.

Method used

By directly extracting motion vector data in the encoding domain, avoiding the computational overhead of the decoding process, using residual-based motion vector purification technology and a dual-threshold segmentation algorithm, a streamlined capsule network and a two-way cross-attention mechanism are designed to realize the three-dimensional positioning and behavioral analysis of pet bone key points.

Benefits of technology

A multi-dimensional accurate description of pet behavior is achieved, which eliminates noise caused by camera jitter and lighting changes, improves the system's deployment adaptability in resource-constrained environments, and can identify abnormal behaviors and perform severity grading.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120219441A_ABST
    Figure CN120219441A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of cameras, and discloses a behavior analysis and monitoring method, device and equipment based on a pet camera, and the method comprises the steps: carrying out the coding domain processing of pet activity video data collected by the pet camera, obtaining a first motion vector field tensor, carrying out the residual analysis and dual-threshold segmentation processing, and obtaining a second motion vector field tensor; obtaining a second motion vector field tensor; inputting the second motion vector field tensor into a target capsule network for processing to obtain a behavior capsule activation vector; performing bidirectional cross attention mechanism fusion on the depth features and behavior capsule activation vectors to obtain a three-dimensional coordinate set of pet skeleton key points; the skeleton features are calculated according to the three-dimensional coordinate set of the key points of the pet skeleton, anomaly analysis is carried out according to the skeleton features, a health risk assessment report is generated, non-pet behavior related noise caused by factors such as camera shake and illumination change is effectively eliminated, and multi-dimensional accurate description of pet behaviors is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of cameras, and in particular, to a method, device, and equipment for behavior analysis and monitoring based on a pet camera. Background Art

[0002] Traditional pet monitoring methods mainly rely on the full decoding process of RGB video data, which not only brings huge storage and computing burdens, but also is difficult to meet the real-time monitoring requirements of devices with limited computing resources such as home pet cameras. Especially in scenarios that require long-term continuous monitoring, the storage and processing of a large amount of redundant video data become the main bottleneck restricting system performance.

[0003] Most existing pet behavior recognition technologies directly process the decoded image data using complex deep learning models. Although the recognition accuracy is relatively high, the number of model parameters is large and the computational complexity is high, making it difficult to be efficiently deployed on edge devices with limited computing resources. On the other hand, these methods usually only focus on two-dimensional image information and lack a three-dimensional understanding of pet postures, resulting in obvious limitations in pet behavior recognition, especially in abnormal behavior early warning, and unable to provide a comprehensive and accurate analysis basis for pet health monitoring. Summary of the Invention

[0004] The present invention provides a method, device, and equipment for behavior analysis and monitoring based on a pet camera. The present invention effectively eliminates non-pet behavior-related noises caused by factors such as camera jitter and light changes, and realizes multi-dimensional precise description of pet behaviors.

[0005] In a first aspect, the present invention provides a method for behavior analysis and monitoring based on a pet camera, and the method for behavior analysis and monitoring based on a pet camera includes:

[0006] Perform encoding domain processing on the pet activity video data collected by the pet camera to obtain a first motion vector field tensor, and perform residual analysis and double-threshold segmentation processing on the first motion vector field tensor to obtain a second motion vector field tensor;

[0007] Input the second motion vector field tensor into a target capsule network for processing to obtain a behavior capsule activation vector;

[0008] Extract the depth features of the pet activity video data, and perform bidirectional cross-attention mechanism fusion on the depth features and the behavior capsule activation vector to obtain a three-dimensional coordinate set of pet bone key points;

[0009] Calculate skeleton features based on the three-dimensional coordinate set of the pet bone key points, and perform abnormal analysis based on the skeleton features to generate a health risk assessment report.

[0010] In a second aspect, the present invention provides a behavior analysis and monitoring device based on a pet camera. The behavior analysis and monitoring device based on a pet camera includes:

[0011] A processing module, configured to perform encoding domain processing on pet activity video data collected by the pet camera to obtain a first motion vector field tensor, and perform residual analysis and double-threshold segmentation processing on the first motion vector field tensor to obtain a second motion vector field tensor;

[0012] An activation module, configured to input the second motion vector field tensor into a target capsule network for processing to obtain a behavior capsule activation vector;

[0013] A fusion module, configured to extract depth features of the pet activity video data, and perform two-way cross-attention mechanism fusion on the depth features and the behavior capsule activation vector to obtain a three-dimensional coordinate set of pet bone key points;

[0014] A generation module, configured to calculate skeleton features according to the three-dimensional coordinate set of the pet bone key points, and perform anomaly analysis according to the skeleton features to generate a health risk assessment report.

[0015] In a third aspect of the present invention, there is provided a behavior analysis and monitoring device based on a pet camera, including: a memory and at least one processor, wherein instructions are stored in the memory; the at least one processor invokes the instructions in the memory so that the behavior analysis and monitoring device based on a pet camera executes the above-mentioned behavior analysis and monitoring method based on a pet camera.

[0016] In the technical solution provided by the present invention, by directly extracting motion vector data from the video coding domain instead of processing the fully decoded RGB image frames, the computational overhead of the decoding process is avoided, making the system more suitable for deployment on resource-constrained home pet cameras. The residual-based motion vector purification technology is adopted. Through the dual-threshold segmentation algorithm and temporal domain stabilization processing, the non-pet-behavior-related noises caused by factors such as camera jitter and illumination changes are effectively eliminated. A streamlined capsule network dedicated to pet behavior recognition is designed. By reducing the number of capsules, decreasing the dimension of the pose matrix, and introducing a disentangled capsule routing mechanism, the effective capture ability of pet behavior characteristics is maintained. The depth image information and behavior characteristics are fused. Through the bidirectional cross-attention mechanism and the improved state space model, the accurate three-dimensional positioning of the pet's skeletal key points is achieved. A three-level behavior classification system including a basic behavior classifier, a complex behavior classifier, and a state duration analyzer is constructed. By analyzing the spatial relationship matrix and temporal change characteristics, a complete first behavior pattern code including basic behavior codes, complex behavior codes, and state transition codes is generated, realizing the multi-dimensional accurate description of pet behavior. The behavior autoencoder based on the state space model, by considering both the reconstruction error and the prediction error simultaneously, establishes a more accurate abnormal score calculation mechanism, which can identify abnormal behaviors and classify their severity levels.

[0017] Other features and advantages of the present invention will be described in the following specification, and, in part, will be obvious from the specification, or will be understood by implementing the present invention. The objectives and other advantages of the present invention are achieved and obtained by the structures specifically pointed out in the specification, claims, and drawings.

[0018] To make the above objectives, features, and advantages of the present invention more obvious and understandable, the following specific preferred embodiments are given, and in conjunction with the accompanying drawings, the detailed description is as follows. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 It is a schematic diagram of an embodiment of the behavior analysis and monitoring method based on a pet camera in an embodiment of the present invention;

[0020] Figure 2 It is a schematic diagram of an embodiment of the behavior analysis and monitoring device based on a pet camera in an embodiment of the present invention;

[0021] Figure 3 It is a schematic diagram of an embodiment of the behavior analysis and monitoring device based on a pet camera in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0022] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0023] As used in the embodiments of the present invention, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes other unlisted steps or units, or optionally further includes other steps or units inherent to these processes, methods, products, or devices.

[0024] To facilitate the understanding of this embodiment, first, a behavior analysis and monitoring method based on a pet camera disclosed in the embodiments of the present invention will be introduced in detail. As Figure 1 shown, this method includes the following steps:

[0025] 101. Process the pet activity video data collected by the pet camera in the coding domain to obtain a first motion vector field tensor, and perform residual analysis and double-threshold segmentation processing on the first motion vector field tensor to obtain a second motion vector field tensor;

[0026] It can be understood that the execution subject of the present invention can be a behavior analysis and monitoring device based on a pet camera, or a terminal or a server. Specifically, it is not limited here. In the embodiments of the present invention, the server is taken as an example of the execution subject for illustration.

[0027] Specifically, pet activity video data is collected in real time through a pet monitoring camera deployed in an indoor environment, and the video data is input into a video encoding module. This module uses standard compression encoding methods such as H.264 or H.265 to perform compression encoding on consecutive video frames and outputs the corresponding video encoding bitstream. Without performing a complete image decoding operation, the bitstream is intercepted directly at the encoding end, and the motion vector data in the encoding domain is extracted. These motion vectors are used for inter-frame prediction in video compression and carry rich spatio-temporal motion information. To accurately extract these vectors from the bitstream, the syntax elements in the video encoding stream are parsed frame by frame to identify the information structure at the macroblock level, and the core parameters including the macroblock position coordinates (x, y), the horizontal and vertical motion vector components (mvx, mvy), and the reference frame index are extracted. Since the I-frame (key frame) does not contain motion vector information in the encoding structure, to complement the spatio-temporal motion characteristics corresponding to these missing regions, the motion vector data of the corresponding macroblocks in adjacent P-frames and B-frames is used, and interpolation estimation is performed by the method of weighted average to form the motion vector completion result in the continuous time dimension. When interpolating, the weight decay coefficient is set according to the inter-frame time interval to ensure the rationality of the interpolation result in terms of time consistency. After completing the interpolation and completion, all the vector data at the macroblock level is organized according to the frame order and normalized, that is, the motion vector values under different resolutions and different frame rates are uniformly projected into the [-1, 1] interval through normalization mapping. The normalization step can effectively eliminate the motion vector scale deviation caused by device differences or different video encoding parameters, ensuring the robustness of subsequent algorithms under different input conditions. After the above processing, a structured first motion vector field tensor is constructed, and its tensor structure is in four-dimensional form T×H×W×2, representing the time length, the height and width of the motion vector map, and the horizontal and vertical direction components of the two-dimensional vector respectively. To eliminate noise interference and improve the signal quality related to pet movement, a residual analysis mechanism is introduced to perform inter-frame difference operation on the motion vector field in the time dimension, calculate the motion vector change amplitude of each macroblock between consecutive frames, and construct a residual field accordingly. To distinguish effective motion from background perturbation, two thresholds, a high threshold and a low threshold, are set, and the double-threshold segmentation algorithm is applied to divide the residual values into high-confidence, low-confidence, and undetermined regions. For macroblocks with residual values higher than the high threshold, they are directly marked as pet activity regions, while those lower than the low threshold are marked as background regions. For the uncertain regions in the middle of the thresholds, their connectivity with the high-confidence regions in the spatial neighborhood (such as 8-neighborhood) is combined for judgment. A mask matrix representing the effective motion region is formed, and this mask is applied to the first motion vector field tensor to eliminate the background and interference regions and generate a purified second motion vector field tensor.

[0028] The point-by-point difference calculation is performed on the adjacent frames in the first motion vector field to obtain the residual value of the motion vector in the time dimension and construct a new motion vector residual field. The residual field reflects the change amplitude of the motion state between consecutive frames. The area with a larger value corresponds to the real pet movement behavior, while the area with a smaller value or close to zero mostly belongs to the static part of the background or slight disturbance. According to the two pre-set thresholds, namely the high threshold and the low threshold, the above residual field is subjected to double threshold segmentation processing, so that the residual value is divided into three sub-areas: when the residual value is greater than the high threshold, the macroblock is classified as a high confidence area of ​​pet activity; when the residual value is less than the low threshold, the macroblock is marked as a background area; and the area between the high and low thresholds is temporarily regarded as a pending area. In order to clarify the ownership of the pending area, the spatial connectivity analysis mechanism is introduced to detect whether each pending macroblock has a direct neighborhood connection with the high confidence area in two-dimensional space, especially considering the connection within the eight neighborhoods. When a macroblock to be determined has spatial contact with at least one high-confidence macroblock, the system classifies it as a pet activity area; otherwise, it is determined to belong to the background area. Based on this method, macroblock area classification information is constructed, and a structured binary mask matrix is ​​generated based on this, where the matrix value corresponding to the pet activity area is 1 and the background area is 0. The first motion vector field tensor is multiplied element by element with the above mask matrix to mask the motion information of the background macroblock and obtain a purified motion vector field tensor that preliminarily eliminates the influence of the background area. Considering the occasional errors in instantaneous frames, a multi-frame weighted average strategy is introduced in the time dimension. The mask matrices of consecutive N frames are exponentially weighted to calculate the weighted average value. The weight is set to decay over time to highlight the recent activity characteristics. The obtained average matrix is ​​then converted into a time-domain stable binary mask matrix through threshold judgment. Compared with the single-frame result, this matrix has better temporal consistency and noise resistance. The stable mask matrix is ​​element-wise multiplied with the initial first motion vector field tensor to mask out unstable areas or areas not related to the pet, and finally a second motion vector field tensor with high confidence and low background interference is constructed.

[0029] 102. Input the second motion vector field tensor into the target capsule network for processing to obtain a behavior capsule activation vector;

[0030] Specifically, the continuously collected purified motion vector data is reorganized in chronological order, and the second motion vector field tensors of every 16 frames are combined into a complete data block, thereby constructing an input tensor with spatio-temporal continuity. The input tensor is fed into the initial feature extraction module of the capsule network for processing. Inside this module, there are three interconnected two-dimensional convolutional layers. Each convolutional layer is equipped with a convolutional kernel of a specific size, and after the convolutional operation, an activation function and a batch normalization operation are connected to enhance the network's non-linear expression ability and stabilize the training process. After this process, the input tensor is gradually extracted with basic motion features to form a multi-dimensional feature map. The regions in the basic feature map are divided and transformed into 16 primary capsules, and each primary capsule contains a pose representation describing the motion pattern. To establish more semantically discriminative connections between the primary capsules, a disentangled capsule routing mechanism is adopted, enabling the lower-level capsules to be more selective and explicit when passing information to the higher-level capsules, and avoiding the problem of mixed different behavioral features. Under the action of this mechanism, each primary capsule outputs an activation vector, which reflects the probability of the behavior recognized by the current capsule and retains the directionality and pattern features of the behavior. The activation information of all primary capsules is uniformly fed into the behavior capsule layer for aggregation and classification. In the behavior capsule layer, multiple behavior capsules representing specific pet behaviors are set, and each capsule corresponds to a potential behavior category. After the output information of the primary capsules undergoes transformation processing, it is aggregated into each behavior capsule to form a high-level behavior representation rich in semantics. To make these representations comparable in the numerical range and enhance the system's understanding ability of the behavior intensity, a non-linear compression function is applied to the output of each behavior capsule to scale its vector output to a unified range. This processing method keeps the vector direction unchanged and makes its length directly reflect the confidence of the corresponding behavior. Through the above processing flow, the activation vectors of the behavior capsules are obtained.

[0031] 103. Extract the depth feature tensor of the pet activity video data, and perform a two-way cross-attention mechanism fusion on the depth feature tensor and the behavior capsule activation vector to obtain a three-dimensional coordinate set of the pet bone key points;

[0032] Specifically, the corresponding depth image is synchronously obtained from the video data collected by the pet camera. The depth image records the distance information from each pixel point in the scene to the camera. To eliminate the interference of device errors and environmental noise on the depth value, preprocessing operations are performed on the collected original depth image, including linearly normalizing all valid depth values to a fixed numerical range to improve the processing consistency under different input conditions; at the same time, interpolating and filling the invalid or missing depth regions, and using a filtering algorithm to denoise the overall image, resulting in a clear, continuous, and uniformly scaled preprocessed depth feature map. The preprocessed depth map is input into the depth feature extraction network, which is composed of three convolutional modules in sequence. Each convolutional operation slides a convolutional kernel of a set size to extract the spatial structure information in the image, and cooperates with the activation function and the downsampling mechanism to continuously improve the feature abstraction level. After multi-layer feature extraction, a depth feature tensor containing rich spatial morphological information is output. The depth feature tensor and the behavior capsule activation vector obtained from the capsule network before are used as inputs, and a bidirectional cross-attention mechanism is introduced for fusion. In this mechanism, two paths are constructed to calculate the attention response of the depth feature to the behavior information and the modulation effect of the behavior feature on the depth space respectively, and the high-response regions in the two feature domains are aligned through a bidirectional guiding method. The fusion operation performs weighted interaction according to their respective attention regions, generating a set of unified splicing features containing depth geometric information and behavior semantic information, and this splicing feature has both position perception and behavior perception capabilities. The fused splicing feature is input into a structured state space model for dynamic modeling and information transformation. The state space model mainly undertakes the functions of temporal context integration and feature smooth transformation in this scenario, receives the fused features of consecutive frames, constructs state variables, and outputs stable high-level representations. The output features fuse behavior-driven cues and depth space cues. According to the features output by the state space model, a set of key point heatmaps are generated, and these heatmaps are used to represent the probability distribution of each pet bone point in the image space. Using a coordinate parsing algorithm based on the heatmap, the two-dimensional position coordinates of each key point in the image are extracted. Since these two-dimensional coordinates are still in the planar projection and cannot reflect the depth structure, the depth value at the corresponding position in the original preprocessed depth map is combined to convert each two-dimensional coordinate point into a position in the three-dimensional space. Through this step, a set of three-dimensional coordinate sets of pet bone key points containing complete spatial information are generated, reflecting the body posture and movement trajectory of the pet in the video sequence.

[0033] 104. Calculate the skeleton features based on the three-dimensional coordinate set of the pet bone key points, and perform anomaly analysis based on the skeleton features to generate a health risk assessment report.

[0034] Specifically, the extracted three-dimensional skeletal key point sequence is structurally organized, and the key point coordinates within a continuous time period are combined according to the frame sequence to construct the initial skeleton sequence data. This skeleton sequence is based on the positions of all key points in each frame in three-dimensional space, reflecting the posture change trajectory and action path of the pet over the entire time period. Based on this initial skeleton sequence, a spatial relationship modeling operation between key points is performed. By calculating the Euclidean distance between each key point within each frame, a spatial relationship matrix describing the geometric relationship of the pet's skeleton structure is constructed. This matrix can be regarded as a structural template at each moment, used to depict the current body configuration of the pet and is applicable to identifying typical states such as standing, lying prone, and bending. To reflect the pet's movement trend and dynamic behavior characteristics in the time dimension, the displacement of key points between adjacent frames in the skeleton sequence is calculated to extract the time-varying characteristics of the key points. This characteristic form describes the movement direction and amplitude of each key point on the time axis in the form of frame differences, effectively revealing the continuity and abruptness of the pet's behavior and is suitable for identifying highly dynamic behavior patterns such as jumping, running, and rolling. The spatial relationship matrix is combined with the time-varying characteristics to construct the skeleton features that comprehensively describe the pet's skeleton state. The skeleton features and the behavior capsule activation vectors generated by the capsule network in the early stage are subjected to a bilinear pooling fusion operation. This fusion method combines the two feature tensors element-wise, retains the structural information of the original features, and at the same time enhances the coupling expression ability between the two to generate the fusion features. The fusion features are input into a three-level behavior classifier including a basic behavior classifier, a complex behavior classifier, and a state duration analyzer for behavior recognition. The basic classifier is used to identify conventional behavior categories such as eating, sleeping, standing, and walking. The complex behavior classifier uses time series modeling capabilities to identify complex behaviors with high temporal dependence such as grooming, rolling, and scratching. The state duration analyzer tracks and analyzes the persistence, frequency, and change trend of behaviors in the time dimension to construct the behavior evaluation results. Based on the behavior recognition results at each level, a first behavior pattern code is generated. This pattern code consists of three parts, namely the basic behavior code, the complex behavior code, and the state transition code. The basic behavior code reflects the occurrence probability of each conventional behavior, the complex behavior code represents the activity level of special behaviors, and the state transition code depicts the transition frequency and sequential characteristics between different basic behaviors in matrix form, reflecting the pet's daily rhythm and state switching rules. The behavior pattern code is input into the trained behavior autoencoder for anomaly analysis. This autoencoder has both reconstruction and prediction capabilities, maps the current behavior characteristics to the latent space, and compares them with the historical normal behavior patterns. If the reconstruction error or prediction deviation exceeds the set threshold, the system determines that there is a behavior anomaly.Based on the abnormal analysis results, combined with the behavior intensity, duration, and change trend, automatically generate a health risk assessment report. The report lists the suspicious behavior categories, occurrence times, and severity levels, and puts forward targeted suggestions to help pet owners timely grasp the pet's health status and achieve intelligent and continuous pet health management.

[0035] The first-line pattern code is input into a behavior autoencoder composed of an encoder, a state converter, and a decoder. The structure of this autoencoder is specifically designed to perform efficient encoding, state prediction, and reconstruction learning on behavior pattern data. The encoder part maps the input first-line pattern code, compressing it from the high-dimensional behavior feature space to a low-dimensional latent space representation. This latent space representation retains the key information of pet behavior characteristics and enables the system to model the dynamic changes in behavior in the compressed latent space. The latent space representation and the previous h historical behavior state data are input into the state converter for state prediction. The state converter captures the dynamic laws of behavior in the time dimension and predicts the state value at the next moment. The prediction result reflects the continuity and trend of the current behavior pattern within the normal behavior category. If there is a large deviation between the predicted state value at the next moment and the actual observed value, it indicates a potential abnormal behavior in the system. At the same time, the latent space representation is input into the decoder for reconstruction calculation. Through the reconstruction mechanism, the latent space representation is restored to the second-line pattern code. This reconstruction process aims to minimize the difference between the original behavior pattern code and the reconstructed behavior pattern code to determine whether the current behavior conforms to the historical behavior pattern. To quantify the degree of abnormality of the behavior pattern, a first anomaly score is calculated after completing state prediction and behavior reconstruction calculations. This score is based on two aspects of information: one is the reconstruction error between the second-line pattern code reconstructed at the current moment and the input first-line pattern code. The larger the reconstruction error, the more obvious the difference between the current behavior and the historical behavior; the other is the prediction error between the state value predicted by the state converter at the next moment and the actually observed behavior state. If the prediction result is inconsistent with the actual behavior, it indicates an abnormal behavior pattern. These two errors are combined to calculate the first anomaly score. The second anomaly score is calculated based on the first anomaly score. This score is the result of anomaly quantification after being corrected by a normal behavior baseline model established based on long-term behavior monitoring data. The normal behavior baseline model models the pet behavior data for continuous M days, divides it into normal behavior patterns in different time periods, and sets anomaly score thresholds for each time period. The currently calculated second anomaly score is compared with the threshold in the corresponding time period of the baseline model. If the anomaly score exceeds the threshold range, the system determines that an abnormal behavior has occurred and triggers the abnormal behavior recognition mechanism. Based on the abnormal behavior trigger result, the abnormal behavior is recognized. According to information such as the type, duration, and trend of changes in the behavior pattern of the abnormal behavior, the abnormal behavior is classified and analyzed to identify common abnormal behavior types such as abnormal activity level, abnormal eating behavior, sleep disorders, aggressive behavior, or repetitive behavior.Meanwhile, by combining the duration and degree of abnormality of behavior changes, the abnormal behaviors are classified into severity levels to generate a health risk assessment report. The report includes the specific types and occurrence time periods of abnormal behaviors, and also evaluates the duration and degree of abnormality, and gives specific recommended measures, such as adjusting the pet's activity pattern, adjusting the eating habits, or contacting a veterinarian for a health check.

[0036] In the embodiments of the present invention, by directly extracting motion vector data from the video coding domain instead of processing the fully decoded RGB image frames, the computational overhead of the decoding process is avoided, making the system more suitable for deployment on resource-constrained home pet cameras. The residual-based motion vector purification technology is adopted. Through the double-threshold segmentation algorithm and temporal domain stabilization processing, the non-pet-behavior-related noises caused by factors such as camera jitter and light changes are effectively eliminated. A streamlined capsule network dedicated to pet behavior recognition is designed. By reducing the number of capsules, decreasing the dimension of the pose matrix, and introducing a disentangled capsule routing mechanism, the effective capture ability of pet behavior features is maintained. The depth image information and behavior features are fused. Through the bidirectional cross-attention mechanism and the improved state space model, the accurate three-dimensional positioning of the pet's bone key points is achieved. A three-level behavior classification system including a basic behavior classifier, a complex behavior classifier, and a state duration analyzer is constructed. By analyzing the spatial relationship matrix and temporal change features, a complete first behavior pattern code including a basic behavior code, a complex behavior code, and a state transition code is generated, realizing the multi-dimensional accurate description of pet behaviors. Based on the state space model behavior autoencoder, by considering both the reconstruction error and the prediction error, a more accurate abnormal score calculation mechanism is established, which can identify abnormal behaviors and classify their severity levels.

[0037] In a specific embodiment, the process of executing step 101 may specifically include the following steps:

[0038] Video acquisition of the pet's activities is performed through a pet camera to obtain pet activity video data, and the pet activity video data is encoded to obtain a video coding bitstream;

[0039] The coding information is directly intercepted from the video coding bitstream without performing a full decoding operation to obtain the coding domain motion vector data;

[0040] The syntax elements of the coding domain motion vector data are parsed to extract the macroblock-level motion vector information including the macroblock position coordinates, motion vector components, and reference frame index;

[0041] The I-frame motion vector-free region is identified according to the macroblock-level motion vector information, and the weighted average calculation is performed using the macroblock-level motion vector information of adjacent P / B frames to obtain the complete motion vector data;

[0042] Perform normalization mapping processing on the complete motion vector data to obtain the first motion vector field tensor;

[0043] Perform residual analysis and double-threshold segmentation processing on the first motion vector field tensor to obtain the second motion vector field tensor.

[0044] Specifically, continuous video acquisition of the daily activities of pets is carried out through a camera installed in the pet activity area. The camera is an intelligent monitoring device with night vision, infrared sensing, or motion detection functions, ensuring that pet behaviors can be continuously recorded even in low-light or complex environments. The video data is encoded frame by frame in RGB format or YUV format, and the original video data is input into the encoding module for compression and encoding processing. The encoding module adopts mainstream video compression standards, such as H.264 or H.265, to convert the video data into an encoded bitstream, and uses the inter-frame prediction mechanism to generate motion vectors during the encoding process for efficiently representing the displacement information between different time frames. In this process, instead of performing a complete decoding operation, the motion vector information in the bitstream is directly intercepted in the encoding domain, avoiding the reconstruction of the decoded image frames, thereby reducing the computational overhead and improving the processing efficiency. The intercepted motion vector data in the encoding domain is subjected to syntax element parsing. By parsing the macro-block level data structure in the encoded bitstream, the macro-block position information, motion vector components, and reference frame index are extracted. A macro-block is the smallest encoding unit in the encoding process. The motion vector data of each macro-block reflects the displacement change information of this area in the current frame relative to the reference frame, including the horizontal and vertical components. The macro-block position coordinates (x, y) are used to locate the specific position of the macro-block in the image frame, and the reference frame index indicates which frame the macro-block is based on for motion estimation. Since the I-frame is a key frame and does not contain motion vector information, after identifying the I-frame, the macro-block level motion vector information of adjacent P-frames and B-frames is used for interpolation and completion to ensure the continuity and integrity of the motion vector data. During the interpolation process, a weighted average calculation method is adopted to perform temporal weighting on the motion vector data of frames adjacent in time. The weights decay exponentially according to the inter-frame distance, ensuring that the reference frame closer to the target frame has a greater weight, thereby generating smoother and more accurate motion vector completion data. The complete motion vector data is subjected to normalization mapping processing to eliminate the scale differences caused by different camera configurations, resolutions, or frame rates. The normalization process maps the motion vector data to a fixed numerical interval of [-1, 1], ensuring the scale consistency of the motion vector field tensor under different input conditions. At this time, the first motion vector field tensor is constructed. This tensor has four dimensions, representing the time dimension T, the spatial height H, the spatial width W, and the motion components in the horizontal and vertical directions 2, forming a standard motion vector data structure of T×H×W×2. Residual analysis and double-threshold segmentation processing are performed on the first motion vector field tensor. Residual analysis is to perform a difference operation on the motion vectors of adjacent frames to calculate the motion change amount of each macro-block in the time dimension, forming a motion vector residual field.The system uses the magnitude of the macroblock-level motion change in the residual field as a feature to perform a dual-threshold segmentation process. A high threshold and a low threshold are set, and the macroblock region is divided into three categories: the region above the high threshold is marked as the high-confidence region of pet activity, the region below the low threshold is marked as the background region, and the region between the high threshold and the low threshold is regarded as the region to be determined. A spatial neighborhood connectivity judgment mechanism is introduced for the region to be determined to judge whether there is a spatial association between these regions and the high-confidence region. If it is connected to the high-confidence region within the 8-neighborhood range, the region is classified as the pet activity region; otherwise, it is determined as the background region. Based on the determination result, a binary mask matrix is generated, and the mask matrix is multiplied element-wise with the first motion vector field tensor to obtain the second motion vector field tensor after shielding the background and noise. To improve the temporal stability of the motion vector field, the continuous N-frame mask matrices are weighted and averaged to form a mask matrix that is stable in the time domain, and it is multiplied element-wise with the first motion vector field tensor again to ensure that the final second motion vector field tensor has high stability and accuracy in both the spatial and temporal dimensions.

[0045] In a specific embodiment, the process of performing residual analysis and dual-threshold segmentation on the first motion vector field tensor to obtain the second motion vector field tensor may specifically include the following steps:

[0046] Calculate the difference between adjacent frames of the first motion vector field tensor to obtain the motion vector residual field;

[0047] Perform dual-threshold segmentation on the motion vector residual field using a preset high threshold and low threshold to obtain the high-confidence region of pet activity, the background region, and the region to be determined;

[0048] Based on the neighborhood spatial connectivity relationship between the high-confidence region of pet activity and the region to be determined, determine the region to be determined to obtain the division information of the pet activity region and the background region;

[0049] Generate a binary mask matrix based on the division information of the pet activity region and the background region, and apply the binary mask matrix to the first motion vector field tensor for calculation to obtain the motion vector field with the background region preliminarily shielded;

[0050] Perform weighted average calculation on the motion vector field with the background region preliminarily shielded to obtain the weighted average value, and perform binarization processing on the obtained weighted average value to obtain the mask matrix that is stable in the time domain;

[0051] Perform element-wise multiplication on the mask matrix that is stable in the time domain and the first motion vector field tensor to obtain the second motion vector field tensor.

[0052] Specifically, the difference calculation is performed on adjacent frames in the first motion vector field tensor to capture the change trend of macroblock-level motion vectors between consecutive frames and construct a motion vector residual field. The residual field reflects the intensity of macroblock motion change between the current frame and the previous frame by calculating the difference of motion vector components in the horizontal and vertical directions of two adjacent frames. If the motion change between adjacent frames is large, it means that the pet has moved significantly in this area, while the area with small motion change is usually the background or static area. The motion vector residual field is subjected to double threshold segmentation processing to divide the macroblock area into different confidence areas. According to the preset high threshold and low threshold, the area with large residual value is marked as a high confidence area where the pet is active. These areas are areas with drastic motion changes and have high credibility for behavior recognition; while the area with low residual value is marked as a background area. These areas have no obvious motion changes and correspond to the static background part. For the area with residual value between the high threshold and the low threshold, it is divided into the area to be determined. There is a certain uncertainty in this part of the area, and it is necessary to further determine whether it belongs to the pet activity area. The spatial neighborhood connectivity determination mechanism is introduced to refine the area to be determined. In the process of spatial connectivity determination, it is detected whether there is a direct or indirect spatial connection relationship between the area to be determined and the high confidence area. If there is a spatial connection between the area to be determined and the high confidence area within the eight-neighborhood range, the area to be determined is classified as the pet activity area, otherwise it is determined as the background area. Through the spatial connectivity determination method, the area with blurred boundaries is effectively identified to ensure the recognition accuracy of the pet activity area. After completing the area determination, a binary mask matrix is ​​generated according to the final area division information. The dimension of the matrix is ​​consistent with the original motion vector field tensor. The position corresponding to the pet activity area in the mask matrix is ​​marked as 1, and the position corresponding to the background area is marked as 0. The binary mask matrix is ​​applied to the first motion vector field tensor, and the original motion vector field is shielded by element-by-element multiplication to obtain the motion vector field after the background area interference is initially eliminated. The motion vector field of the preliminary shielded background area is weighted averaged, and the sliding window mechanism in the time dimension is used to weight the sum of the motion vector fields of N consecutive frames. The weight decays over time, so that the weight of the most recent frame is larger to highlight the recent motion trend. The noise interference caused by occasional motion jitter is eliminated through weighted average calculation to obtain a more time-consistent weighted average vector field. The weighted average vector field is binarized, and the area with motion intensity greater than the set threshold is marked as the pet activity area, and the area below the threshold is marked as the background area, generating a time-domain stable mask matrix. The time-domain stable mask matrix is ​​element-by-element multiplied with the original first motion vector field tensor, and the unstable background area interference is eliminated again to obtain the final second motion vector field tensor.

[0053] In a specific embodiment, the process of executing step 102 may specifically include the following steps:

[0054] Organize the second motion vector field tensor into a spatio-temporal data block where each data block contains 16 frames of continuous refined motion vector fields to form an input tensor;

[0055] Input the input tensor into the initial feature extraction module of the target capsule network, and process it through three consecutive two-dimensional convolutional layers in the initial feature extraction module to obtain a basic feature map;

[0056] Convert the basic feature map into the initial input of 16 primary capsules to obtain a set of primary capsules, and use a disentangled capsule routing mechanism to process the set of primary capsules to output primary capsule activation values;

[0057] Input the primary capsule activation values into the behavior capsule layer for a transformation matrix to obtain behavior capsule output features, and perform a squash non-linear function processing on the behavior capsule output features to obtain a behavior capsule activation vector.

[0058] Specifically, the second motion vector field tensor is spatiotemporally organized to form data blocks suitable for the input of the capsule network. Sixteen consecutive frames of purified motion vector field data are stacked in chronological order to form a spatiotemporal data block. Each data block retains the motion vector information of the pet within a 16-frame time window. The spatiotemporal data block structure can effectively capture the motion trajectory of the pet in a short period of time and retain the temporal characteristics of motion changes, providing a complete spatiotemporal feature representation for the subsequent behavior recognition network. The spatiotemporal data block is fed into the initial feature extraction module of the target capsule network. This module consists of three consecutive two-dimensional convolutional layers, and each convolutional layer has different-sized convolutional kernels for extracting motion features at different scales. The first convolutional layer uses a larger convolutional kernel size to capture macroscopic motion features and extract information about large-scale motion patterns. The second convolutional layer further refines the features with a smaller convolutional kernel to enhance local motion details. The third convolutional layer continues to deeply mine the features, compresses the spatial relationships of the motion information, and at the same time preserves the key spatial semantics. After each convolutional layer performs the convolution operation, the ReLU activation function is connected to enhance the non-linear expression ability. At the same time, the batch normalization operation is used to maintain the stability of the feature distribution and avoid the problems of gradient vanishing or gradient explosion. After being processed by three consecutive convolutional layers, a high-dimensional basic feature map is obtained. The basic feature map is converted into the initial input of 16 primary capsules. Each primary capsule corresponds to a motion pattern subspace and represents local motion information through the pose matrix. Different from traditional convolutional networks, primary capsules not only capture the feature intensity but also retain the spatial position and direction information of the features, making the system have stronger spatial invariance and view robustness. The generation process of the primary capsule set maps the feature map to multiple subspaces and maps the features of each subspace to the input matrix of the primary capsule. This matrix contains the pose information and confidence information of the capsule. To improve the information decoupling effect between primary capsules and ensure that the capsule network establishes a stable association between different feature spaces, a disentangled capsule routing mechanism is used for primary capsule processing. Compared with the traditional dynamic routing algorithm, the disentangled capsule routing mechanism reduces the computational complexity while optimizing the information transmission path between capsules, enabling different primary capsules to more effectively focus on different motion patterns and avoiding interference between different motion features. Through the optimized routing coefficient calculation process, the output of the lower-level capsules is weighted and summed with the pose matrix of the higher-level capsules to obtain the activation value of the primary capsule. This activation value contains the existence probability of the motion pattern and also encodes the pose information of the motion features. The activation value of the primary capsule is input into the behavior capsule layer for feature transformation. In the behavior capsule layer, by introducing a specific transformation matrix, the activation value of the primary capsule is mapped into a higher-dimensional behavior feature space. The transformation matrix is dynamically adjusted according to the feature differences of different behavior categories to ensure that the primary capsule features are accurately mapped to various pet behavior patterns.The behavioral capsule layer contains multiple behavioral capsules, and each behavioral capsule corresponds to a specific pet behavior category, such as eating, playing, sleeping, walking, etc. The input features are projected and transformed through a specific pose matrix, enabling each behavioral capsule to focus on specific behavioral semantics. To standardize the norm of the output features of the behavioral capsules and make the output vector of the capsules reflect the confidence of the behavior, a squash non-linear function is applied to the output features of the behavioral capsules. The squash function adjusts by compressing the length of the feature vector, restricting the vector length to the interval [0,1], thus associating the probability of the behavior's existence with the length of the feature vector. After the squash non-linear processing, the activation vector of the behavioral capsule retains the direction information of the behavioral features, and its length directly reflects the probability of a specific behavior occurring.

[0059] In a specific embodiment, the process of performing step 103 may specifically include the following steps:

[0060] Obtain a depth image from the pet activity video data, normalize the depth values of the depth image to obtain a preprocessed depth feature map;

[0061] Input the preprocessed depth feature map into a depth feature extraction network to perform three-layer convolution processing to obtain a depth feature tensor;

[0062] Concatenate the depth feature tensor and the behavioral capsule activation vector through a bidirectional cross-attention module to obtain a concatenated feature;

[0063] Input the concatenated feature into a state space model for processing to obtain a state space model output feature;

[0064] Construct a key point heat map based on the state space model output feature, and determine the two-dimensional coordinates of the key points according to the key point heat map;

[0065] Generate a set of three-dimensional coordinates of the pet bone key points according to the two-dimensional coordinates of the key points and the depth values at the corresponding positions in the preprocessed depth feature map.

[0066] Specifically, depth images are extracted from the video data captured by the pet camera. The depth images record the distance information of each pixel point of the pet relative to the camera in the three-dimensional space, and these depth data help to accurately estimate the bone position and three-dimensional motion trajectory of the pet. During the acquisition of depth images, depth data are generated by using technologies such as binocular stereo vision, structured light, and ToF (time of flight), and are acquired synchronously with the RGB video data. The obtained original depth images have problems such as non-uniformity, noise interference, and scale inconsistency. Therefore, preprocessing operations are performed on the depth images, including steps such as depth value normalization, filling of invalid depth points, and filtering and noise reduction. Normalization maps the original depth data to a fixed [0,1] interval, eliminates the influence of different devices and environments on the depth measurement results, and adjusts the depth information to a unified scale to improve the stability of subsequent depth feature extraction. After normalization, a standardized preprocessed depth feature map is generated. This feature map retains the spatial distribution characteristics of the original depth information and can effectively adapt to the input format of the subsequent network model. The preprocessed depth feature map is input into the depth feature extraction network for feature extraction. This network consists of three two-dimensional convolutional layers. Each layer of convolutional operation combines a convolutional kernel of a specific size to capture spatial feature information at different scales. The first layer of convolution uses a larger convolutional kernel to extract global depth distribution information and capture large-scale spatial structures; the second layer has a smaller convolutional kernel size, which is used to refine the local features of the depth image and enhance the boundary details; the third layer of convolution further refines the features to extract depth feature representations with strong semantic information. After each layer of convolutional operation, a LeakyReLU activation function is equipped to enhance the non-linear expression ability of the network, and a batch normalization mechanism is combined to maintain the stability of the feature distribution. Through these three layers of convolutional processing, a depth feature tensor containing rich spatial features is generated. To effectively fuse depth information and behavior features, a bidirectional cross-attention module is introduced. This module is used to perform feature interaction between the depth feature tensor and the behavior capsule activation vector and generate a fused concatenated feature. The bidirectional cross-attention mechanism calculates the mutual attention weights between the depth feature and the behavior feature, enabling the two to be correlated in the feature dimension. Channel attention calculation is performed on the depth feature tensor, enabling the depth feature to dynamically focus on regions with strong correlation with the behavior feature. At the same time, the behavior capsule activation vector is mapped to the same spatial dimension as the depth feature, and information interaction is achieved through element-wise weighting. In the bidirectional cross-attention module, while the depth feature performs weighted modulation on the behavior feature, the behavior feature also enhances the depth feature, resulting in a concatenated feature tensor that combines spatial depth information and behavior semantic features. The concatenated feature is input into the state space model for processing to capture the dynamic change information in the feature sequence and construct a more temporally consistent feature representation.The main function of the state space model is to establish the dynamic relationship of the feature sequence in the time dimension. This model consists of a state matrix, a projection matrix, and an output matrix. By combining the historical feature states, it models the current spliced features, thereby outputting the feature representation of the state space model. Through state updates of the spliced features in the time dimension, it predicts the potential positions of key points in future frames and generates more stable and behavior-logic-compliant output features. Based on the features output by the state space model, a key point heat map is constructed. The key point heat map is a two-dimensional tensor representing the probability distribution of key points, and each channel corresponds to the spatial probability distribution of a specific key point. Through deconvolution operations, the features output by the state space model are upsampled to the spatial resolution of the original depth image, and multiple heat maps corresponding to the number of key points are generated. The pixel values of each heat map reflect the probability that the corresponding key point appears at that pixel position. After generating the heat maps, through spatial analysis of the heat maps, the two-dimensional coordinates of the key points are calculated using the soft argmax operation. These coordinate points represent the positions of the key parts of the pet on the image plane. To convert the two-dimensional coordinates of the key points into position information in three-dimensional space, the depth values at the corresponding positions in the preprocessed depth feature map are combined to restore the three-dimensional coordinates of each key point. The two-dimensional coordinates are mapped onto the depth feature map to obtain the depth values of the corresponding pixel points, and the two-dimensional coordinates and depth information are combined through the internal parameter matrix of the camera to restore the coordinate positions of the key points in three-dimensional space. A set of three-dimensional coordinate sets of the pet's skeletal key points containing complete spatial information is generated, reflecting the posture and movement trajectory of the pet in different frame time periods.

[0067] In a specific embodiment, the process of executing step 104 may specifically include the following steps:

[0068] Extract spatio-temporal skeleton features from the three-dimensional coordinate set of the pet's skeletal key points to obtain initial skeleton sequence data;

[0069] Perform spatial relationship calculations between key points based on the initial skeleton sequence data to obtain a spatial relationship matrix characterizing the pet's skeletal structure features;

[0070] Perform key point time change feature calculations on adjacent frames in the initial skeleton sequence data to obtain time change features characterizing the pet's motion characteristics;

[0071] Combine the spatial relationship matrix and the time change features to form skeleton features, and perform bilinear pooling fusion with the behavior capsule activation vector to obtain fused features;

[0072] Input the fused features into a three-level behavior classifier including a basic behavior classifier, a complex behavior classifier, and a state duration analyzer for behavior recognition to obtain behavior recognition results at each level;

[0073] Generate a first behavior pattern code based on the behavior recognition results at each level. The first behavior pattern code includes a basic behavior code, a complex behavior code, and a state transition code;

[0074] Input the first behavior pattern code into a behavior autoencoder for anomaly analysis to generate a health risk assessment report.

[0075] Specifically, for the three-dimensional coordinate set of pet bone key points, spatio-temporal skeleton features are extracted. The position data of key points within a continuous time period are arranged in order according to time frames to form initial skeleton sequence data, recording the motion posture information of the pet within a specific time window. The three-dimensional coordinate positions of each key point form spatio-temporal trajectories over time. According to the initial skeleton sequence data, the spatial relationship calculation between key points is performed. Based on the Euclidean distance between key points, a spatial relationship matrix reflecting the pet bone structure features is constructed. All key points in each frame are traversed, the spatial distance between any two key points is calculated, and these distance values are filled into a symmetric matrix. Each element in the matrix reflects the distance magnitude between the corresponding key point pairs, thereby capturing the bone structure changes of the pet in different time frames. At the same time, in order to capture the dynamic change characteristics during the pet's movement, based on the displacement difference of key points between adjacent frames, the time change feature calculation of the initial skeleton sequence data is performed. The difference calculation is performed on the same key point in consecutive frames to obtain the motion displacement feature of the key point in the time dimension, reflecting the motion trend and direction information of the key point in a short period of time. Especially for the dynamic feature changes during rapid movement or action switching, the system can capture the pet's behavior dynamics through time change features, ensuring the sensitivity and recognition accuracy of subsequent behavior analysis for sudden behavior changes. The spatial relationship matrix and time change features are combined to form skeleton features. The skeleton features and the behavior capsule activation vector are fused through bilinear pooling. Bilinear pooling is a non-linear feature interaction mechanism. Through the element-wise product operation of the skeleton features and the behavior capsule activation vector, the information in the two feature spaces is cross-combined to generate a fused feature tensor with high-order feature expression ability. The fused feature is input into a three-level behavior classifier including a basic behavior classifier, a complex behavior classifier, and a state duration analyzer for behavior recognition. The basic behavior classifier quickly classifies the fused feature, identifying basic behavior categories such as eating, sleeping, standing, walking, etc. This classifier adopts a fully connected layer structure, maps the skeleton features using the ReLU activation function, and outputs the classification results of basic behaviors. The basic behavior classification results and higher-dimensional temporal features are input into the complex behavior classifier together. This classifier is based on a long short-term memory network or an attention mechanism to identify complex behavior patterns with high temporal dependence, such as behaviors like grooming, scratching, spinning, running, etc. The complex behavior classifier can model the behavior sequence over a long time period, thereby capturing cross-frame behavior features and improving the recognition accuracy of complex behavior patterns. The state duration analyzer tracks the behavior change trend, models and analyzes the duration, transition frequency, and state changes of different behavior categories. Through the sliding window mechanism, the transition rules of each behavior state within a specified time period are calculated to form the behavior state temporal analysis result.Based on the output results of the basic behavior classifier, complex behavior classifier, and state duration analyzer for behavior recognition at each level, a first behavior pattern code is generated. This pattern code is a structured data representation that includes three parts: the basic behavior code, complex behavior code, and state transition code. The basic behavior code is a vector that reflects the occurrence probability of basic behavior categories. The complex behavior code records the activity level of complex behavior patterns. The state transition code is a matrix used to describe the transition probability between different basic behaviors, reflecting the change law of the pet's behavior state in the time dimension. The first behavior pattern code is input into a behavior autoencoder for anomaly analysis. The behavior autoencoder consists of an encoder, a state converter, and a decoder, which maps the first behavior pattern code to a latent space representation, combines the historical behavior states for state prediction, and then reconstructs the latent space representation to obtain the reconstructed second behavior pattern code. By calculating the reconstruction error between the original pattern code and the reconstructed pattern code, as well as the prediction error between the predicted value of the next moment state and the true state, and combining with the multi-dimensional anomaly score, the anomaly detection result is obtained. Based on the anomaly analysis result, a health risk assessment report is generated. This report includes the type of abnormal behavior, duration, severity of the anomaly, and corresponding recommended measures, and the analysis result is pushed to the pet owner through the user interface to achieve precise monitoring and early warning of the pet's health status.

[0076] In a specific embodiment, the process of performing the step of inputting the first behavior pattern code into the behavior autoencoder for anomaly analysis and generating a health risk assessment report may specifically include the following steps:

[0077] Input the first behavior pattern code into the behavior autoencoder composed of an encoder, a state converter, and a decoder. The encoder performs mapping processing on the first behavior pattern code to obtain a latent space representation;

[0078] Input the latent space representation and the previous h historical states into the state converter for state prediction to obtain the predicted value of the next moment state;

[0079] Input the latent space representation into the decoder for reconstruction calculation to obtain the second behavior pattern code, and calculate the first anomaly score in combination with the predicted value of the next moment state;

[0080] Calculate the second anomaly score based on the reconstruction error between the second behavior pattern code and the first behavior pattern code, and the prediction error between the actual state and the predicted state of the next moment, in combination with the first anomaly score;

[0081] Compare the second anomaly score with the corresponding time period threshold in the normal behavior baseline model established based on the pet behavior data for consecutive M days to determine the abnormal behavior trigger result;

[0082] Based on the triggering results of abnormal behaviors, identify abnormal behaviors and generate a health risk assessment report including the types of abnormal behaviors, duration, severity, and recommended measures.

[0083] Specifically, the first-line pattern code is input into the behavior autoencoder, which consists of three parts: an encoder, a state converter, and a decoder, and has the capabilities of behavior feature modeling and anomaly detection. The first-line pattern code, as input data, contains basic behavior codes, complex behavior codes, and state transition codes. Through structured feature representation, it comprehensively quantifies the behavior states of pets within a specific time period, providing high-dimensional behavior feature inputs for the anomaly detection of the autoencoder. The first-line pattern code is input into the encoder for mapping processing. The encoder performs dimensionality reduction and feature extraction on the input data through a multi-layer fully connected network, mapping the high-dimensional behavior pattern information into a low-dimensional latent space to generate a latent space representation. This latent space representation is a compact behavior feature representation that retains key behavior pattern information and eliminates data redundancy through the dimensionality reduction process, thereby improving the computational efficiency and generalization ability of the model. The encoder introduces non-linear characteristics through activation functions during the mapping process, enabling the latent space representation to better capture complex behavior relationships and temporal features. The generated latent space representation and the previous h historical state data are jointly input into the state converter, which predicts the state at the next moment by constructing the temporal correlation between the latent space representation and the historical state. The state converter adopts a parameterized state update mechanism to infer the evolution trend of future behavior patterns based on the current latent space state and historical state information and generate the predicted value of the state at the next moment. During the prediction process, the state converter enables the system to have the ability to detect abnormal behavior trends in advance by capturing the variation law of pet behavior patterns in the time dimension, ensuring that abnormal behaviors can be effectively predicted before they occur. At the same time, the latent space representation is input into the decoder for reconstruction calculation. The decoder maps the latent space representation back to the same feature dimension as the original behavior pattern code through a multi-layer fully connected network symmetric to the encoder structure to generate the reconstructed second-line pattern code. The goal of the reconstruction process is to minimize the difference between the original behavior pattern code and the reconstructed pattern code, thereby ensuring that the model can accurately reproduce the feature distribution of normal behavior patterns. The reconstruction ability of the decoder is directly related to the learning accuracy of the system for normal behavior features. The smaller the reconstruction error, the better the learning effect of the system on normal behavior patterns. To quantify the abnormal degree of the behavior pattern, the first anomaly score is calculated based on the reconstructed second-line pattern code and the predicted value of the state at the next moment. The first anomaly score combines the reconstruction error of the current moment's behavior pattern and the prediction error of the future behavior state, thereby providing a preliminary judgment of the abnormal degree for anomaly detection. The reconstruction error reflects the deviation size of the system when reproducing the current behavior pattern, while the prediction error represents the prediction accuracy of the system for the future behavior change trend. When both the reconstruction error and the prediction error are large, it indicates that there may be potential abnormalities in the pet's behavior. The second anomaly score is calculated based on the reconstruction error between the second-line pattern code and the first-line pattern code, the prediction error between the actual state and the predicted state at the next moment, and in combination with the first anomaly score.The second anomaly score is a higher-dimensional anomaly quantification result. By weighted summing the anomaly features in different time periods, it captures the trend of long-term behavior pattern changes and generates a more accurate anomaly score. To ensure the accuracy of anomaly detection, the calculated second anomaly score is compared with the corresponding time period threshold in the normal behavior baseline model established based on the pet behavior data of consecutive M days. The normal behavior baseline model is established through statistical analysis of a large amount of historical behavior data, and each time period has a corresponding anomaly score threshold for the behavior pattern. When the second anomaly score exceeds the set threshold, the system determines that there is an abnormal behavior in that time period and triggers the abnormal behavior recognition mechanism. Once the abnormal behavior recognition is triggered, the system enters the anomaly recognition stage. By comparing the abnormal behavior features with the predefined abnormal behavior templates, it identifies the specific types of abnormal behaviors, such as reduced abnormal activity, abnormal eating, sleep disorders, abnormal excretion, aggressive behavior, repetitive behavior, or abnormal vocalization, etc. And based on the duration, change frequency of the behavior anomaly, and the fluctuation trend of the anomaly score, it judges the severity of the abnormal behavior, and quantitatively grades the abnormal behavior by combining these factors, dividing the abnormal behavior into mild anomaly, moderate anomaly, and severe anomaly, and forming an anomaly analysis report. A health risk assessment report is generated according to the abnormal behavior recognition result. This report records the type, occurrence time period, duration, and severity of the abnormal behavior, and provides targeted suggestions, such as suggesting adjusting the pet's activity pattern, improving eating habits, or contacting a veterinarian for a health check. The health risk assessment report is pushed to the pet owner through the user interface to ensure that the pet owner can timely grasp the pet's health status and achieve early warning and effective intervention.

[0084] The above describes the behavior analysis and monitoring method based on a pet camera in the embodiments of the present invention. Next, the behavior analysis and monitoring device based on a pet camera in the embodiments of the present invention will be described. Please refer to Figure 2 , an embodiment of the behavior analysis and monitoring device based on a pet camera in the embodiments of the present invention includes:

[0085] A processing module 201, configured to perform encoding domain processing on the pet activity video data collected by the pet camera to obtain a first motion vector field tensor, and perform residual analysis and double-threshold segmentation processing on the first motion vector field tensor to obtain a second motion vector field tensor;

[0086] An activation module 202, configured to input the second motion vector field tensor into a target capsule network for processing to obtain a behavior capsule activation vector;

[0087] A fusion module 203, configured to extract the depth features of the pet activity video data, and perform two-way cross-attention mechanism fusion on the depth features and the behavior capsule activation vector to obtain a three-dimensional coordinate set of the pet bone key points;

[0088] A generation module 204 is configured to calculate skeleton features based on a set of three-dimensional coordinates of pet bone key points, perform anomaly analysis based on the skeleton features, and generate a health risk assessment report.

[0089] Through the collaborative cooperation of the above-mentioned various components, by directly extracting motion vector data from the video coding domain instead of processing full decoded RGB image frames, the computational overhead of the decoding process is avoided, making the system more suitable for deployment on resource-constrained home pet cameras. The residual-based motion vector purification technology is adopted. Through the double-threshold segmentation algorithm and time-domain stability processing, the noise related to non-pet behaviors caused by factors such as camera jitter and lighting changes is effectively eliminated. A lightweight capsule network dedicated to pet behavior recognition is designed. By reducing the number of capsules, decreasing the dimension of the pose matrix, and introducing a disentangled capsule routing mechanism, the effective capture ability of pet behavior features is maintained. The depth image information and behavior features are fused. Through the bidirectional cross-attention mechanism and the improved state space model, the accurate three-dimensional positioning of pet bone key points is achieved. A three-level behavior classification system including a basic behavior classifier, a complex behavior classifier, and a state persistence analyzer is constructed. By analyzing the spatial relationship matrix and time-varying features, a complete first behavior pattern code including basic behavior codes, complex behavior codes, and state transition codes is generated, realizing the multi-dimensional accurate description of pet behaviors. Based on the state space model behavior autoencoder, by considering both the reconstruction error and the prediction error simultaneously, a more accurate anomaly score calculation mechanism is established, which can identify abnormal behaviors and grade their severity.

[0090] Above Figure 2 The behavior analysis and monitoring device based on a pet camera in the embodiments of the present invention is described in detail from the perspective of modular functional entities. Next, the behavior analysis and monitoring device based on a pet camera in the embodiments of the present invention will be described in detail from the perspective of hardware processing.

[0091] Figure 3FIG. 0 is a schematic structural diagram of a behavior analysis and monitoring device based on a pet camera. The behavior analysis and monitoring device 300 based on the pet camera may vary greatly due to configuration or performance differences, and may include one or more processors (central processing units, CPUs) 310 (for example, one or more processors) and a memory 320, and one or more storage media 330 for storing application programs 333 or data 332 (for example, one or more mass storage device terminals). Among them, the memory 320 and the storage media 330 may be transient storage or persistent storage. The program stored in the storage media 330 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the behavior analysis and monitoring device 300 based on the pet camera. Further, the processor 310 may be configured to communicate with the storage media 330 and execute a series of instruction operations in the storage media 330 on the behavior analysis and monitoring device 300 to implement the steps of the above-mentioned behavior analysis and monitoring method based on the pet camera.

[0092] The behavior analysis and monitoring device 300 based on the pet camera may further include one or more power supplies 340, one or more wired or wireless network interfaces 350, one or more input / output interfaces 360, and / or one or more operating systems 331, such as Windows Serve, Mac OS X, Unix, Linux, FreeBSD, etc. Those skilled in the art can understand that Figure 3 The shown structural diagram of the behavior analysis and monitoring device based on the pet camera does not limit the behavior analysis and monitoring device based on the pet camera provided by the present invention, and may include more or fewer components than shown in the figure, or combine some components, or have different component arrangements.

[0093] Those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working processes of the above-described systems, systems, and units can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0094] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions to enable a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs.

[0095] As described above, the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A behavior analysis and monitoring method based on a pet camera, characterized in that: include: Performing coding domain processing on the pet activity video data collected by the pet camera to obtain a first motion vector field tensor, and performing residual analysis and double threshold segmentation processing on the first motion vector field tensor to obtain a second motion vector field tensor; Inputting the second motion vector field tensor into the target capsule network for processing to obtain a behavior capsule activation vector; Extracting the deep feature tensor of the pet activity video data, and fusing the deep feature tensor with the behavior capsule activation vector using a bidirectional cross-attention mechanism to obtain a three-dimensional coordinate set of pet skeleton key points; Skeleton features are calculated based on the three-dimensional coordinate set of key points of the pet's skeleton, and abnormal analysis is performed based on the skeleton features to generate a health risk assessment report.

2. The behavior analysis and monitoring method based on a pet camera according to claim 1, characterized in that: The coding domain processing is performed on the pet activity video data collected by the pet camera to obtain a first motion vector field tensor, and the residual analysis and double threshold segmentation processing are performed on the first motion vector field tensor to obtain a second motion vector field tensor, including: Capturing pet activity videos through a pet camera to obtain pet activity video data, and encoding the pet activity video data to obtain a video encoding bit stream; Directly intercepting the coding information from the video coding bit stream without performing a complete decoding operation to obtain coding domain motion vector data; Parsing the coded domain motion vector data by syntax elements to extract macroblock-level motion vector information including macroblock position coordinates, motion vector components and reference frame indexes; Identify the I frame without motion vector area according to the macroblock level motion vector information, and use the macroblock level motion vector information of adjacent P / B frames to perform weighted average calculation to obtain complete motion vector data; Performing normalized mapping processing on the complete motion vector data to obtain a first motion vector field tensor; The first motion vector field tensor is subjected to residual analysis and double threshold segmentation processing to obtain a second motion vector field tensor.

3. The behavior analysis and monitoring method based on a pet camera according to claim 2, characterized in that: The performing residual analysis and double threshold segmentation processing on the first motion vector field tensor to obtain a second motion vector field tensor includes: performing difference calculation on adjacent frames of the first motion vector field tensor to obtain a motion vector residual field; Using a preset high threshold and a low threshold to perform double threshold segmentation processing on the motion vector residual field to obtain a pet activity high confidence area, a background area and an area to be determined; According to the neighborhood spatial connectivity relationship between the pet activity high confidence area and the area to be determined, the area to be determined is determined to obtain pet activity area and background area division information; Generate a binary mask matrix based on the pet activity area and background area division information, and apply the binary mask matrix to the first motion vector field tensor for calculation to obtain a motion vector field for preliminarily shielding the background area; Performing weighted average calculation on the motion vector field of the preliminary shielded background area to obtain a weighted average value, and performing binarization processing on the obtained weighted average value to obtain a time-domain stable mask matrix; An element-wise product operation is performed on the time-domain stable mask matrix and the first motion vector field tensor to obtain a second motion vector field tensor.

4. The behavior analysis and monitoring method based on a pet camera according to claim 1, characterized in that: The step of inputting the second motion vector field tensor into a target capsule network for processing to obtain a behavior capsule activation vector includes: Organizing the second motion vector field tensor into spatiotemporal data blocks, each data block containing 16 frames of continuous purified motion vector fields, to form an input tensor; The input tensor is passed into the initial feature extraction module of the target capsule network, and processed by three consecutive two-dimensional convolutional layers in the initial feature extraction module to obtain a basic feature map; The basic feature map is converted into the initial input of 16 main capsules to obtain a main capsule set, and the main capsule set is processed by a disentangled capsule routing mechanism to output a main capsule activation value; The main capsule activation value is input into the behavior capsule layer for conversion matrix to obtain the behavior capsule output feature, and the behavior capsule output feature is processed by squash nonlinear function to obtain the behavior capsule activation vector.

5. The behavior analysis and monitoring method based on a pet camera according to claim 1, characterized in that: The extracting of the deep feature tensor of the pet activity video data and fusing the deep feature tensor with the behavior capsule activation vector by a bidirectional cross-attention mechanism to obtain a three-dimensional coordinate set of pet skeleton key points includes: Acquire a depth image from the pet activity video data, and normalize the depth value of the depth image to obtain a preprocessed depth feature map; Input the preprocessed deep feature map into a deep feature extraction network to perform three-layer convolution processing to obtain a deep feature tensor; Concatenating the deep feature tensor and the behavior capsule activation vector through a bidirectional cross attention module to obtain a concatenated feature; Inputting the splicing features into a state space model for processing to obtain state space model output features; Constructing a key point heat map based on the output features of the state-space model, and determining the two-dimensional coordinates of the key points according to the key point heat map; A three-dimensional coordinate set of pet skeleton key points is generated according to the two-dimensional coordinates of the key points and the depth values ​​of corresponding positions in the preprocessed depth feature map.

6. The behavior analysis and monitoring method based on a pet camera according to claim 1, characterized in that: The calculating of skeleton features according to the three-dimensional coordinate set of the pet's skeletal key points, and performing abnormal analysis according to the skeleton features to generate a health risk assessment report include: Extracting spatiotemporal skeleton features from the three-dimensional coordinate set of the pet's skeleton key points to obtain initial skeleton sequence data; Calculate the spatial relationship between key points based on the initial skeleton sequence data to obtain a spatial relationship matrix that represents the pet's skeletal structure characteristics; Calculating the time-varying features of key points on adjacent frames in the initial skeleton sequence data to obtain time-varying features that characterize the motion characteristics of the pet; Combining the spatial relationship matrix and the time variation feature to form a skeleton feature, and performing bilinear pooling fusion with the behavior capsule activation vector to obtain a fusion feature; Inputting the fusion features into a three-level behavior classifier including a basic behavior classifier, a complex behavior classifier and a state persistence analyzer for behavior recognition, and obtaining behavior recognition results at each level; Generate a first behavior pattern code based on the behavior recognition results of each level, wherein the first behavior pattern code includes a basic behavior code, a complex behavior code and a state transition code; The first behavior pattern code is input into the behavior autoencoder for abnormal analysis to generate a health risk assessment report.

7. The behavior analysis and monitoring method based on a pet camera according to claim 6, characterized in that: The step of inputting the first behavior pattern code into a behavior autoencoder for abnormal analysis to generate a health risk assessment report includes: Inputting the first behavior pattern code into a behavior autoencoder composed of an encoder, a state converter, and a decoder, and mapping the first behavior pattern code through the encoder to obtain a latent space representation; Input the latent space representation and the previous h historical states into the state converter to perform state prediction to obtain a state prediction value at the next moment; Inputting the latent space representation into the decoder for reconstruction calculation to obtain a second behavior pattern code, and calculating a first anomaly score in combination with the next moment state prediction value; A second anomaly score is calculated based on a reconstruction error between the second behavior pattern code and the first behavior pattern code and a prediction error between an actual state at the next moment and a predicted state, combined with the first anomaly score; Compare the second abnormal score with a corresponding time period threshold in a normal behavior baseline model established based on M consecutive days of pet behavior data to determine an abnormal behavior trigger result; Abnormal behavior is identified based on the abnormal behavior triggering result, and a health risk assessment report including the abnormal behavior type, duration, severity and recommended measures is generated.

8. A behavior analysis and monitoring device based on a pet camera, characterized in that: Used to perform the behavior analysis and monitoring method based on a pet camera as described in any one of claims 1 to 7, the behavior analysis and monitoring device based on a pet camera comprises: A processing module, configured to perform coding domain processing on the pet activity video data collected by the pet camera to obtain a first motion vector field tensor, and perform residual analysis and double threshold segmentation processing on the first motion vector field tensor to obtain a second motion vector field tensor; an activation module, configured to input the second motion vector field tensor into a target capsule network for processing to obtain a behavior capsule activation vector; A fusion module is used to extract the depth features of the pet activity video data, and fuse the depth features with the behavior capsule activation vector through a bidirectional cross-attention mechanism to obtain a three-dimensional coordinate set of pet skeleton key points; A generation module is used to calculate skeleton features based on a three-dimensional coordinate set of key points of the pet's skeleton, and to perform abnormal analysis based on the skeleton features to generate a health risk assessment report.

9. A behavior analysis and monitoring device based on a pet camera, characterized in that: The pet camera-based behavior analysis and monitoring device comprises: a memory and at least one processor, wherein instructions are stored in the memory; The at least one processor calls the instructions in the memory so that the pet camera-based behavior analysis and monitoring device executes the pet camera-based behavior analysis and monitoring method as described in any one of claims 1-7.

Citation Information

Cited By

  • Unmanned aerial vehicle RGB-infrared progressive fusion target detection method and device based on state space model

    CN121482655A

  • An unmanned aerial vehicle RGB-infrared progressive fusion target detection method and device based on a state space model

    CN121482655B