A pet dangerous behavior monitoring system based on a visual neural network

Through a pet hazard behavior monitoring system based on visual neural network, a three-dimensional semantic map is generated using image and point cloud data, the key points and gaze direction of pet bones are extracted, the target area of pet movement is predicted, and attention transfer strategies are implemented, which solves the problem of not being able to identify pet behavior intentions in the existing technology, and real-time early warning and intervention of pet dangerous behaviors is achieved.

CN120108006BActive Publication Date: 2025-07-18广州佳可电子科技股份有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510602121.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-12
Publication Date
2025-07-18
Estimated Expiration
2045-05-12

AI Technical Summary

Technical Problem

Existing pet monitoring products cannot understand pet behavior in real time, and cannot identify their action intentions, spatial goals or behavioral trends, resulting in the inability to effectively warn and interfere with pet dangerous behaviors, such as jumping at high places, accidentally eating or bumping into fragile items.

Method used

A pet hazard behavior monitoring system based on visual neural network is adopted to obtain indoor images and point cloud data through the data acquisition module, and a three-dimensional semantic map is generated using the semantic segmentation module. The key points and gaze direction of pet bones are extracted in combination with the posture estimation network, and the pet's bones are predicted through a time series model, and an attention transfer strategy is implemented to intervene.

Benefits of technology

Real-time early warning and intervention on pet dangerous behaviors has been achieved, the accuracy and responsiveness of monitoring have been improved, and the future behavior trends of pets can be predicted and effective guiding measures have been taken to prevent dangerous behaviors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120108006B_ABST
    Figure CN120108006B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of pet behavior monitoring, and particularly to a pet dangerous behavior monitoring system based on a visual neural network. By acquiring image data and spatial point cloud data, using a semantic segmentation module to extract image semantic features and spatial structure features, fusing them to generate a three-dimensional semantic map, and marking dangerous areas such as high platforms, objects that are easily ingested by mistake, and fragile objects. By extracting the sequence of pet bone key points through a pose estimation network, and obtaining the spatial vector of the gaze direction through a convolutional neural network, after constructing a behavior state vector, it is input into a time series model to predict the target area of pet movement, and match it with the three-dimensional semantic map to determine whether there is an intention of dangerous behavior, such as jumping, ingesting by mistake, or approaching a fragile object, and send an alarm and execute an attention transfer strategy, realizing early warning and active intervention for pet dangerous behaviors, and significantly improving the monitoring accuracy and response ability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of pet behavior monitoring, and particularly to a pet dangerous behavior monitoring system based on a visual neural network. Background Art

[0002] The role of pets in the family is becoming increasingly important. However, in the case of the owner being away for a long time or unable to monitor in real time, pets may exhibit a series of dangerous behaviors, which are often difficult to detect, identify, and stop in a timely manner, thus triggering a series of chain consequences.

[0003] Taking common pets such as cats and dogs as an example, their activity space covers the entire indoor environment, and they have a certain jumping ability and exploration desire, and are extremely likely to perform behaviors such as jumping from a height, biting foreign objects, and breaking fragile items. Among them, the behavior of jumping from a height, such as jumping onto a window sill, balcony railing, or the edge of a second-floor structure, may very likely lead to a fall if the action is misjudged or the structure is unstable, causing fractures or internal injuries to the pet; the behavior of accidental ingestion is common in small plastic parts, etc., which may cause serious consequences such as digestive obstruction and poisoning once swallowed; and when running at high speed and chasing and playing, pets are extremely likely to knock down high-value or fragile items such as vases, televisions, stereos, and computer monitors, which may not only cause property losses, but also pose a safety hazard to the pet itself due to scratches from the fragments or to vulnerable groups such as infants and young children.

[0004] Most of the existing pet monitoring products on the market adopt the technical route of general security cameras, and only have the functions of image acquisition and historical video playback, and cannot realize the real-time understanding and intelligent analysis of pet behaviors. A small number of products introduce simple target detection algorithms, which can realize the static judgment of "appearing in the picture" of pets, but cannot identify their action intentions, spatial targets, or behavior trends. For example, when a pet is about to jump onto a balcony railing, approach a fragile object, or try to pick up a foreign object on the ground, the system cannot make an effective response, resulting in the complete absence of the early warning and intervention links. Summary of the Invention

[0005] To solve the above problems, the present invention provides a pet dangerous behavior monitoring system based on a visual neural network.

[0006] To achieve the above object, the technical solution adopted by the present invention is:

[0007] A pet dangerous behavior monitoring system based on a visual neural network, comprising:

[0008] A data acquisition module for acquiring indoor area image data and indoor space point cloud data;

[0009] A semantic segmentation module, which is used to perform semantic recognition of indoor structures and objects based on image data using a visual neural network, and at the same time extract spatial height information, edge structures, and object contour features based on point cloud data, and fuse the image semantic results and point cloud geometric data to generate a three-dimensional semantic map;

[0010] A behavior parameter construction module, which is used to construct a pet behavior dataset based on indoor area image data and indoor space point cloud data. Based on the pet behavior dataset, extract the pet's bone point sequence through a pose estimation network, and obtain the spatial vector of the pet's current gaze direction through a convolutional neural network;

[0011] A behavior recognition module, which is used to predict the pet's moving target area through a time series model based on the bone point sequence and the spatial vector of the gaze direction, and calculate the pet's dangerous intention state based on the matching relationship between the moving target area and the positions of environmental objects in the three-dimensional semantic map. The pet's dangerous intention state includes preparation for high jumps, ingestion of foreign objects, and approaching fragile objects;

[0012] A pet guidance module, which is used to send an alarm to the mobile terminal according to the pet's dangerous intention state and execute an attention transfer strategy.

[0013] Further, the acquisition of indoor area image data and indoor space point cloud data includes the following steps:

[0014] Collect video image frames of the pet's activity area at a fixed frequency through RGB cameras set at several positions indoors to obtain indoor area image data;

[0015] Collect the spatial point cloud data of the indoor area through a lidar.

[0016] Further, the semantic segmentation module is used to perform the following steps:

[0017] Based on the indoor area image data, use a semantic segmentation network to process the image data and extract an image semantic feature map containing object semantic categories and two-dimensional spatial positions;

[0018] Based on the image semantic feature map, combined with the spatially synchronized point cloud data, map the image features to the point cloud space coordinate system through depth projection to obtain a projected semantic feature map;

[0019] Based on the spatial height information, edge structures, and geometric contour features extracted from the projected semantic feature map and the spatial point cloud data, construct a joint feature representation;

[0020] Input the joint feature representation into a three-dimensional semantic fusion network to generate a three-dimensional semantic map, and perform semantic annotation on high platforms, objects prone to ingestion of foreign objects, and fragile objects in the indoor environment in the three-dimensional semantic map.

[0021] Further, the pose estimation network is constructed through the following steps:

[0022] Based on the indoor area image data, a multi-scale convolutional network is used to process the image sequence, and a sequence of image feature maps containing pet contour and structure information is extracted;

[0023] Based on the sequence of image feature maps, a key point detection sub-network is used to perform key point heat map regression processing on predefined key parts in each frame of the image, and a sequence of heat maps representing the key point response probability distribution is output;

[0024] Based on the sequence of heat maps, temporal modeling is used to model the response trajectories of key points between consecutive image frames, and the key point positions are fitted and spatially corrected in combination with the inter-frame time order, and a sequence of pet bone points is output.

[0025] Further, the formula of the key point detection sub-network is as follows:

[0026] ;

[0027] where is the set of output key point heat maps; is the Sigmoid activation function; is the weight of the fully connected regression layer; is the bias term; the feature map extracted by the th multi-scale convolution; is the corresponding residual correction term of the feature map is the weighted coefficient generated by the channel attention mechanism, satisfying ; is the number of multi-scale branches.

[0028] Further, obtaining the spatial vector of the current gaze direction of the pet through the convolutional neural network includes the following steps:

[0029] Construct a gaze direction training data set, which includes a number of pet head images and corresponding gaze direction spatial vectors;

[0030] Using the pet head image as the input, image feature extraction is performed through the convolutional neural network to obtain an image feature vector representing the facial orientation feature;

[0031] Based on the image feature vector, regression prediction of the gaze direction spatial vector is performed through a regression structure. The regression prediction uses the corresponding gaze direction spatial vector in the gaze direction training data set as the supervision target, and the mean square error is used as the loss function to optimize the network parameters;

[0032] Based on the trained convolutional neural network model, during runtime, it receives the image of the pet's current frame head region as input and outputs the spatial vector of the pet's current gaze direction.

[0033] Further, predicting the pet's motion target area through a time series model based on the bone point sequence and the spatial vector of the gaze direction includes the following steps:

[0034] Based on the bone point sequence output by the pose estimation network and the spatial vector of the gaze direction output by the convolutional neural network, a behavior state vector is constructed in each frame image. The behavior state vector includes the combined result of the spatial coordinates of each key point and the unit vector of the gaze direction in the current frame;

[0035] Arrange the behavior state vectors in chronological order to form an input sequence, and input it into the Transformer prediction model. The prediction model includes an encoder sub-module and a decoder sub-module;

[0036] The encoder sub-module uses the multi-head self-attention mechanism to model the key point trajectory changes and gaze direction trends in the input sequence, and outputs the global time-dependent encoding representation;

[0037] The decoder sub-module, based on the global time-dependent encoding representation and the current position embedding information, predicts the pet's spatial motion trend in the next few time steps through the position-aware decoding network, and outputs the corresponding target position prediction sequence;

[0038] Based on the spatial similarity matching between the target position prediction sequence and the positions of the labeled environmental objects in the three-dimensional semantic map, determine the approaching area of the pet's motion path, and output the semantic position label of the approaching area as the prediction result of the motion target area.

[0039] Further, using the multi-head self-attention mechanism to model the key point trajectory changes and gaze direction trends in the input sequence includes the following steps:

[0040] Perform feature embedding mapping on the sequence of the behavior state vectors to generate corresponding query vector sequence, key vector sequence, and value vector sequence respectively;

[0041] Based on the query vector sequence and the key vector sequence, use the multi-head self-attention mechanism to calculate the attention distribution weights between each time step in the behavior state vector sequence;

[0042] Based on the attention distribution weights and the value vector sequence, calculate the temporal behavior feature representation of each attention head to obtain the multi-head temporal dependence feature set;

[0043] Concatenate and fuse the multi-head temporal dependency feature set, and convert it into a global temporal dependency encoding representation through an output mapping layer.

[0044] Furthermore, the formula of the position-aware decoding network is as follows:

[0045] ;

[0046] where is the target position prediction vector for the th frame; is the current time step; is the offset step for predicting into the future; s is the historical time step index; is the current prediction step the attention weight for the historical s-th frame; is the global temporal dependency encoding of the s-th frame output by the encoder; is the position embedding vector of the s-th frame; is the decoding linear mapping weight matrix; the embedding state vector of the current frame; is the projection matrix from the current frame state to the position vector.

[0047] Furthermore, the attention transfer strategy includes voice broadcast, pet snack delivery, and pet toy movement.

[0048] The beneficial effects of the present invention are as follows: The present invention obtains indoor area image data and indoor space point cloud data through the setting of a data acquisition module, and uses a semantic segmentation module to perform visual neural network processing on the image data respectively, extract semantic feature maps, and extract spatial height, edge structure, and object contour information from the point cloud data. Furthermore, by fusing the image semantic results and the point cloud geometric data, a three-dimensional semantic map is constructed, and target areas with dangerous attributes are explicitly marked in the map, including high platforms, easily misingested foreign objects, and fragile objects, etc. Through the collaborative fusion of images and point clouds, the system can not only obtain the two-dimensional position of the pet, but also understand its activity environment in the three-dimensional space, forming a unified understanding of the physical structure and semantic attributes. The bone key point sequence of the pet is extracted from the collected image sequence through a pose estimation network, and the spatial vector of the current gaze direction of the pet is obtained through a convolutional neural network. On this basis, a behavior state vector for prediction is constructed by a behavior parameter construction module. Further, the behavior recognition module inputs the bone key point sequence and the spatial vector of the gaze direction into a time series model, and predicts the movement target area of the pet in the next stage by learning its temporal evolution trend, and matches the spatial position of this area with the dangerous area in the three-dimensional semantic map to determine whether the pet has dangerous behavior intentions, such as jump preparation, ingestion tendency, approaching the fragile area, etc. This processing chain realizes a continuous reasoning process from behavior characteristics to behavior trends to risk states, which is significantly superior to the single-frame judgment ability of the existing system for the target position in static frames. Finally, the system sets a pet guidance module, sends an alarm message to the mobile terminal according to the recognized dangerous intention state, and can execute an attention transfer strategy to assist the user in guiding the pet to avoid high-risk behaviors. In summary, by fusing image and point cloud data to establish a three-dimensional semantic map, modeling the pet's behavior state by combining bone key points and gaze directions, and using a Transformer model for motion trend prediction, a pet dangerous behavior monitoring system with the ability of in-depth structure understanding, intention perception, and trend judgment is formed, effectively preventing the occurrence of pet dangerous behaviors. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] Figure 1 FIG. is a schematic structural diagram of a pet dangerous behavior monitoring system based on a visual neural network in the present invention.

[0050] Figure 2 FIG. is a diagram of predicting the pet's movement target area through a time series model based on the bone key point sequence and the spatial vector of the gaze direction in the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0051] Please refer to Figure 1 - Figure 2 as shown, the present invention relates to a pet dangerous behavior monitoring system based on a visual neural network, including:

[0052] A data acquisition module for obtaining indoor area image data and indoor space point cloud data;

[0053] A semantic segmentation module for performing semantic recognition of indoor structures and objects based on the image data using a visual neural network, and at the same time extracting spatial height information, edge structures, and object contour features based on the point cloud data, and fusing the image semantic results and the point cloud geometric data to generate a three-dimensional semantic map;

[0054] A behavior parameter construction module for constructing a pet behavior dataset based on the indoor area image data and the indoor space point cloud data, extracting the pet bone point sequence through a pose estimation network based on the pet behavior dataset, and obtaining the spatial vector of the current gaze direction of the pet through a convolutional neural network;

[0055] A behavior recognition module for predicting the pet's moving target area through a time series model based on the bone point sequence and the spatial vector of the gaze direction, and calculating the pet's dangerous intention state based on the matching relationship between the moving target area and the positions of environmental objects in the three-dimensional semantic map, where the pet's dangerous intention state includes high jump preparation, foreign body ingestion, and approaching fragile objects;

[0056] A pet guidance module for sending an alarm to the mobile terminal according to the pet's dangerous intention state and implementing an attention transfer strategy.

[0057] Specifically, the data acquisition module includes several RGB cameras and lidar devices arranged at different indoor locations. The RGB cameras collect image frame data of the pet's activity area at a fixed frequency to provide high-resolution texture, semantic, and pose information; the lidar collects spatial point cloud data of the corresponding indoor area to reflect the three-dimensional structure, terrain height changes, and object contours. Through the parallel acquisition of this heterogeneous data, it provides an input basis for the subsequent modules for image-space joint representation. In the semantic segmentation module, the image data is first input into a vision neural network based on an encoder-decoder structure for processing. The network uses a multi-scale feature extraction channel combined with an upsampling structure to output a two-dimensional semantic feature map, where each pixel in the map corresponds to a class label and semantic confidence. At the same time, the point cloud data undergoes processing means such as edge extraction and curvature estimation to extract spatial height distribution, edge morphology, and three-dimensional contour information. The system projects the image semantic feature map into the point cloud coordinate system through depth projection and combines it with the geometric features of the point cloud to form a semantic structure in three-dimensional space. This joint representation is further input into a three-dimensional semantic fusion network to achieve deep fusion of image semantics and spatial geometry, and finally generate a three-dimensional semantic map. The spatial positions and their semantic attributes of high platforms (such as the edge of the second floor, window sills), foreign objects that can be picked up on the ground, and fragile objects such as glass bottles are explicitly marked in this map. This map not only overcomes the limitation that a pure image recognition system cannot understand spatial height but also makes up for the problem that the point cloud system lacks semantic discrimination ability, realizing complementary enhancement of image perception and three-dimensional understanding. Thus, it has a higher accuracy and stronger generalization ability for identifying dangerous areas in complex indoor layouts, demonstrating the multiplicative effect of "semantic + geometric" fusion recognition in risk modeling. The behavior parameter construction module constructs a pet behavior dataset based on the collected image sequences and point cloud sequences. The image sequences are input into a pose estimation network, which constructs a heatmap sequence through a combination of a convolutional network and a key point detection sub-network, and further extracts the sequence of skeletal key points in consecutive frames of the pet through inter-frame trajectory modeling. To capture the pet's behavior intention, the system further processes the pet's head image through a convolutional neural network to regress and predict the spatial vector corresponding to the pet's gaze direction in the current frame. The system jointly constructs a behavior parameter vector of the pet in the current state from the key point sequence and the gaze direction, which serves as the input for subsequent temporal modeling and predictive analysis. The behavior recognition module models and predicts based on the key point sequence and the gaze direction spatial vector through a time series model (such as a Transformer based on a self-attention structure). This module not only models the spatial state of the current frame but also performs long-term modeling on the change trends of key points and the change patterns of gaze directions, and outputs the moving target areas within several future time steps. Subsequently, the system performs a spatial similarity match between the prediction result and the spatial danger annotation positions in the three-dimensional semantic map.If the match is successful and the target area belongs to a preset dangerous area, such as the edge of a high platform, fragile objects or foreign objects on the ground, the system will output the corresponding dangerous behavior intention state, such as "jump preparation", "approach fragile objects" or "tendency to eat by mistake". Compared with the existing method of classification and judgment based on the current frame, this module introduces an explicit time modeling and spatial semantic fusion mechanism, realizes the trend prediction of the future behavior of the pet, and significantly improves the advance and accuracy of the danger warning. After detecting the dangerous behavior intention state, the pet guidance module triggers the corresponding control response according to the current state and synchronously sends the warning information to the mobile terminal to prompt the user. The guidance strategy includes implementing attention transfer means (such as sound playback or calling other strategy components) to intervene in the pet's behavior path. The overall system realizes a closed-loop control process from "perception - modeling - prediction - guidance". The system in the embodiment introduces a deep network structure in both the image recognition and point cloud modeling perception dimensions, and synergistically fuses the two through a three-dimensional semantic map, enabling the system to have the ability of unified expression of space and semantics. In terms of behavior modeling, the system forms a pet behavior state vector by combining pose estimation and gaze direction, and uses a time series model to predict the future behavior trend, significantly enhancing the understanding ability of the dynamic process of behavior. Compared with the existing technology that only relies on single-frame analysis or image target detection, this embodiment realizes the improvement of perception accuracy and response timeliness through the synergistic effect of the algorithms between modules, and shows higher practicality and stability in the actual pet behavior monitoring scenario.

[0058] Further, the obtaining of the indoor area image data and the indoor space point cloud data includes the following steps:

[0059] Collect video image frames of the pet activity area at a fixed frequency by setting RGB cameras at several positions indoors to obtain indoor area image data;

[0060] Collect the spatial point cloud data of the indoor area by lidar.

[0061] In a specific deployment, RGB cameras are set at different viewing positions indoors, covering typical high-frequency pet activity areas including the living room, bedroom, balcony, and stairway corner. Each camera collects images at a preset fixed frequency and records the video stream at a frame rate of 30 frames per second within the system operation cycle. The image resolution is set at 1920×1080 to ensure a high-fidelity presentation of the pet's face, limb contours, and surrounding objects. The captured image frames enter the image data buffer synchronously numbered by timestamps and serve as the input source for subsequent image semantic recognition, pose estimation, and gaze direction modeling. The system selects a 360-degree scanning indoor lidar, which is installed in the carrier device equipped with the camera, ensuring that all visible object three-dimensional point cloud data in the space can be completely obtained during its working cycle. The lidar working frequency is time-aligned with the image acquisition frequency to achieve the pairing and binding of inter-frame space-image data. The collected point cloud data is mainly used to reflect the true height, boundary structure, and geometric features of objects in three-dimensional space, such as the spatial distribution of furniture, the height difference of the stair platform, and the spatial relationship between the pet's location and various dangerous objects. Through the above image and point cloud synchronous acquisition method, not only the visual information acquisition of the pet's dynamic behavior is realized, but also the ability to structurally understand the spatial environment is significantly improved. Especially the reflection ability of the point cloud data on height, boundary, and spatial relationship, while compensating for the lack of depth dimension in RGB images, provides an accurate and quantifiable basis for the subsequent construction of three-dimensional semantic maps, spatial matching, and behavior reasoning.

[0062] Further, the semantic segmentation module is used to perform the following steps:

[0063] Based on the indoor area image data, use a semantic segmentation network to process the image data and extract an image semantic feature map containing object semantic categories and two-dimensional spatial positions;

[0064] Based on the image semantic feature map, combined with the time-synchronized spatial point cloud data, project the image features into the point cloud space coordinate system through depth projection to obtain a projected semantic feature map;

[0065] Based on the spatial height information, edge structure, and geometric contour features extracted from the projected semantic feature map and the spatial point cloud data, construct a joint feature representation;

[0066] Input the joint feature representation into a three-dimensional semantic fusion network to generate a three-dimensional semantic map, and perform semantic annotation on high platforms, easily misingested foreign objects, and fragile objects in the indoor environment in the three-dimensional semantic map.

[0067] In some embodiments, first, the indoor area image data is input into a semantic segmentation network based on an encoding-decoding structure. The main body of the network adopts a multi-scale receptive field structure. The front-end encoder extracts feature maps at different scales in the image, and the back-end decoder upsamples the high-semantic layer information and restores the spatial resolution, thereby outputting the semantic class label and class confidence map corresponding to each pixel, forming a two-dimensional image semantic feature map. This feature map not only provides semantic information such as pets, furniture, the ground, walls, object contours, etc., but also retains the two-dimensional planar position relationship. Subsequently, the above image semantic feature map is aligned and mapped with the spatially synchronized point cloud data. The mapping process uses a depth projection mechanism: through the internal and external parameters of the camera, the image plane coordinates and the point cloud three-dimensional coordinates are projected and matched, so that each point cloud point can obtain its semantic label in the image, generating a projected semantic feature map. This map not only contains the position and depth information of the points, but also embeds the semantic high-dimensional expression learned in the image network, realizing the cross-modal transfer from two-dimensional planar semantics to three-dimensional spatial structure. To further enhance the understanding of the spatial structure and the parsing of semantic boundaries, the system extracts geometric features such as spatial height gradient, local curvature, and edge intensity from the original point cloud and fuses them with the projected semantic feature map to construct a joint feature representation. This joint feature vector contains multiple dimensions such as semantic label probability distribution, local spatial geometric structure, and position embedding at each point, and has the ability to model the semantic-structure boundary consistency in complex spaces, significantly superior to the feature expressions constructed by traditional point cloud clustering or pure image convolution methods. After completing the feature construction, the system inputs the joint feature representation into a three-dimensional semantic fusion network for training and inference. This network can adopt voxel convolution with a residual structure or the PointNet++ framework to achieve the joint optimization of cross-modal semantics and geometric features while maintaining the continuity of spatial distribution. The finally output three-dimensional semantic map can not only label the three-dimensional positions of various semantic entities, such as sofas, coffee tables, pet beds, etc., but more importantly, it can explicitly identify and label the spatial targets with potential risks in the map, including high platforms (such as balcony railings, stair edges), scattered objects on the ground (such as pills, small toys), and low-lying fragile items (such as glass ornaments, ceramic vases), etc.

[0068] Further, the pose estimation network is constructed through the following steps:

[0069] Based on the indoor area image data, a multi-scale convolutional network is used to process the image sequence, and a sequence of image feature maps containing pet contour and structure information is extracted;

[0070] Based on the sequence of image feature maps, a key point detection sub-network is used to perform key point heat map regression processing on predefined key parts in each frame of the image, and a sequence of heat maps representing the key point response probability distribution is output;

[0071] Based on the heatmap sequence, temporal modeling is used to model the response trajectory of key points between consecutive image frames. Combining the temporal order between frames, trajectory fitting and spatial correction are performed on the positions of key points to output a sequence of pet bone points.

[0072] In this embodiment, the pose estimation network is used to accurately extract the sequence of skeletal key points of a pet from consecutive image frames. Its core lies in integrating three types of algorithm modules: multi-scale perception, key point heatmap regression, and time series modeling. By parsing the subtle changes and continuous trends in the pet's movement process in a structured manner, it significantly enhances the capture accuracy of dynamic behaviors. The innovation of this module is not only reflected in the integrity of the perception chain, but more importantly in the enhanced effect formed in spatio-temporal collaborative modeling, that is, considering both spatial image features and time series evolution features simultaneously, thereby obtaining much better skeleton stability and prediction robustness than traditional intra-frame estimation schemes. Specifically, the system first inputs the indoor area image data into a feature extraction network based on a multi-scale convolutional structure. This network parallelly extracts visual features at different scales and different texture scales in the image through convolutional kernels with multiple receptive field configurations in the front layer, such as the pet's contour edges, limb connection points, facial regions, etc.; in the middle layer, batch normalization and activation functions are used to enhance the stability of feature expression; in the back layer, a feature compression method is adopted to output a sequence of image feature maps of a unified size. This feature map not only retains the spatial layout information but also integrates multi-scale structural semantics, providing rich context support for subsequent key point detection. Next, the sequence of image feature maps is fed into the key point detection sub-network. This sub-network is based on a fully convolutional regression structure, maps the high-dimensional features of each frame of the image into a predefined number of key point channels, and models the position probability of each key point in the form of a Gaussian heatmap. The output result is a sequence of heatmaps, and each heatmap represents the position distribution where the corresponding key point may appear in that frame. The heatmap uses a two-dimensional Gaussian center as the regression target, and the optimization objective function is usually the mean squared error loss, which reflects the difference between the predicted heatmap and the annotated heatmap. This method effectively overcomes the problem of unstable positioning of traditional direct coordinate regression in local occlusion and deformation scenarios. However, it is difficult for key point detection based on a single frame of image to handle the continuous pose deformation and behavior switching of a pet during movement. Therefore, this embodiment further introduces a time series modeling mechanism to establish the trajectory continuity constraint between frames. The system inputs the response position sequence of each key point in multiple time frames into the time modeling module, uses a lightweight time convolutional network structure to model the dynamic response path of the key point, and mines the potential laws of inter-frame behavior evolution. At the same time, combining the inter-frame time step and the physical movement consistency rule, spatio-temporal fitting and position correction are performed on the continuous key point trajectory, and a globally consistent sequence of pet skeletal key points is output, significantly improving the robustness to local occlusion, non-linear deformation, and noise drift under complex dynamics. For example, during the process when the system monitors a cat jumping from the living room sofa to the edge of the dining table, since the cat's front limb is blocked by the tail before taking off, it appears blurred in the local image. Traditional static key point estimation methods may result in key point loss or jumps.In this embodiment, through the temporal modeling and heatmap fusion correction of the previous few frames, the system can stably output the positions of the key points and track their continuous trajectories during the jumping process, so as to completely capture the skeletal action sequence from the pre-jump to the landing point, providing accurate support for subsequent dangerous behavior prediction. Compared with the existing pose detection methods based on single-frame estimation or direct coordinate regression, this embodiment uses a three-layer algorithm linkage of "spatial feature extraction + key point heatmap modeling + temporal trajectory fusion" in the network structure, which not only improves the accuracy of key point positioning, but also realizes the in-depth modeling of temporal consistency and motion dynamics, and has stronger behavior understanding and forward prediction capabilities, which is an important technical support for realizing the core function of the present invention.

[0073] Further, the formula of the key point detection sub-network is as follows:

[0074] ;

[0075] Wherein, is the set of output key point heatmaps; is the Sigmoid activation function; is the weight of the fully connected regression layer; is the bias term; The feature map extracted by the th multi-scale convolution; is the corresponding residual correction term of the feature map is the weighted coefficient generated by the channel attention mechanism, satisfying ; is the number of multi-scale branches.

[0076] It should be noted that, first, the system inputs the image features into multiple convolutional branches with the same structure but different receptive field sizes. Each branch independently extracts the local or global structural features of the image at that scale, respectively adapting to the spatial scale features of different key points of the pet. For example, smaller convolutional kernels extract fine-grained features such as the face and ears with rich details, while larger convolutional kernels are suitable for capturing overall pose features such as the torso and limb contours. Then, the original feature maps extracted by each branch will also undergo a round of residual correction operations. This operation calculates the difference between the original feature map and its approximate structure obtained through downsampling and then upsampling to generate a residual feature map representing the structural error or detail compensation. This difference reflects the deviation between the local region and its smooth reconstruction, which is used to enhance the response of the significant regions near the key points and suppress the interference of background or texture noise. Especially in cases of occlusion, uneven lighting, etc., it can significantly enhance the localization robustness. Subsequently, the system takes all the multi-scale feature maps corrected by the residuals as inputs and feeds them into the channel attention mechanism for weight learning and fusion. In this mechanism, the network evaluates the response capabilities of different scales for the current task and assigns a weight value to each feature map, and the sum of all weights is one. The assignment of the weight values reflects the dynamic adaptation ability of the feature at this scale to the contribution degree of the overall key point prediction, enabling the network to automatically adjust the attention degree to a specific scale under different environments, angles, or pet behavior states. Finally, the multi-scale feature maps after weighted fusion are fed into the heatmap generation module. After being processed by the fully connected mapping layer and the activation function, a set of two-dimensional heatmaps are generated, and each heatmap corresponds to the position probability distribution of a key point. In the heatmap, the higher the pixel value, the more likely the position is the central region of the target key point. This output result provides a structural input with high spatial accuracy, continuous response distribution, and strong occlusion robustness for subsequent temporal modeling and behavior prediction. Compared with the traditional method that only relies on single-scale convolution to output the key point coordinates, this network structure has significant advantages in terms of structural expression, response stability, and scale adaptability. Through multi-scale perception, residual enhancement, and attention weighting, the system can accurately distinguish the positions of key points. Even in complex backgrounds, occlusion interference, or cases where key points are close to each other, it can still maintain a stable and continuous skeletal response output, significantly improving the accuracy of subsequent behavior modeling and motion prediction.

[0077] Further, the obtaining of the spatial vector of the current gaze direction of the pet through the convolutional neural network includes the following steps:

[0078] Construct a gaze direction training data set, and the data set includes a number of pet head images and corresponding gaze direction spatial vectors;

[0079] Taking the pet head image as the input, perform image feature extraction through the convolutional neural network to obtain an image feature vector representing the facial orientation feature;

[0080] Based on the image feature vector, a regression structure is used to perform regression prediction on the gaze direction space vector. The regression prediction uses the corresponding gaze direction space vector in the gaze direction training dataset as the supervision target, and the mean square error is used as the loss function to optimize the network parameters;

[0081] Based on the trained convolutional neural network model, during operation, the current frame head region image of the pet is received as input, and the current gaze direction space vector of the pet is output.

[0082] In some embodiments, in this embodiment, in order to predict the next movement intention of the pet in advance, the system constructs a gaze direction prediction model based on a convolutional neural network, which aims to extract the orientation features in the pet's head image and map it into a gaze vector in three-dimensional space, thereby constituting a key parameter in the behavior trend modeling. This module not only has the ability to express image features efficiently, but also converts the two-dimensional facial orientation information into a measurable spatial target prediction factor through a supervised regression training mechanism, breaking through the single-dimensional limitation of "judging behavior only by posture" in the prior art, and realizing the expansion of the spatial intention perception dimension. First, in the model training stage, a structured gaze direction training data set is constructed. The data set consists of a large number of pet head images and their corresponding three-dimensional gaze direction vectors, where the gaze direction vector is obtained by manual annotation or sensor-assisted acquisition, and is usually expressed as a unit space vector to indicate the main gaze direction of the pet at that time. In terms of sample coverage, the data set contains images of multiple species, multiple postures, and different lighting and occlusion conditions to enhance the model's adaptability to complex environments. In the training stage, each head image is input into a convolutional neural network designed based on a residual structure for feature extraction. The front layer of the network mainly extracts local facial structural features, such as the edge contours and texture features of the eyes, nose bridge, and ear roots; the middle layer uses a medium-scale convolutional perceptron to integrate the overall orientation pattern of the entire face area; the back layer obtains a set of low-dimensional, distinguishable feature vectors through global average pooling and feature compression to characterize the facial orientation features corresponding to the image. Then, the image feature vector is sent to the regression structure module, and the continuous spatial prediction of the three-dimensional gaze direction vector is realized through a multi-layer perceptron. The prediction process does not output discrete category labels, but fits a three-dimensional vector to represent the direction unit vector of the pet's gaze line. During training, the real gaze direction in the training set is used as the supervision target, and the mean square error is used as the loss function for optimization to ensure that the spatial angle between the predicted vector and the target direction is as close as possible, thereby enhancing the prediction stability and accuracy. After the system completes training, the model is deployed on the inference end. In the actual operation process, the system captures the head image area of the pet's current frame from the RGB video frame as input, calls the trained convolutional neural network model in real time, and outputs the pet's gaze direction vector in the current frame. This vector will be input into the subsequent time series modeling module together with the skeleton key point sequence output by the posture estimation module to infer the moving target area and spatial approach intention.

[0083] Furthermore, the predicting of the pet movement target area through a time series model based on the skeleton point sequence and the spatial vector of the gaze direction includes the following steps:

[0084] Based on the sequence of skeleton points output by the pose estimation network and the spatial vector of the gaze direction output by the convolutional neural network, a behavior state vector is constructed in each frame of the image. The behavior state vector includes the combined result of the spatial coordinates of each key point and the unit vector of the gaze direction in the current frame;

[0085] Arrange the behavior state vectors in chronological order to form an input sequence, and input it into the Transformer prediction model. The prediction model includes an encoder sub-module and a decoder sub-module;

[0086] The encoder sub-module uses the multi-head self-attention mechanism to model the key point trajectory changes and gaze direction trends in the input sequence, and outputs the global time-dependent encoding representation;

[0087] The decoder sub-module, based on the global time-dependent encoding representation and the current position embedding information, predicts the spatial movement trend of the pet in the next few time steps through the position-aware decoding network, and outputs the corresponding target position prediction sequence;

[0088] Based on the spatial similarity matching between the target position prediction sequence and the positions of the environment objects annotated in the three-dimensional semantic map, determine the approaching area of the pet's movement path, and output the semantic position label of the approaching area as the prediction result of the movement target area.

[0089] In some embodiments, first, based on the sequence of skeleton points output by the front-end pose estimation network and the spatial vectors output by the gaze direction extraction module, a set of behavior state vectors in a unified format is constructed in each frame of the image. This state vector is jointly composed of the three-dimensional spatial coordinates of each skeleton key point in the current frame and the unit vector of the gaze direction of the corresponding frame, constituting a complete "structure-orientation" description unit. This data structure not only reflects the static structure distribution of the current pose of the pet in space but also introduces the dynamic prior of the gaze intention, enabling the model to capture two information dimensions, namely "the basis for action occurrence" and "the trend of target attention", during the training process. Subsequently, the system arranges this series of state vectors in chronological order to form a behavior sequence with a fixed window length and inputs it into a time series prediction model constructed based on the Transformer structure. Different from traditional recurrent neural networks, the Transformer effectively models the global dependencies between different time steps in the input sequence by introducing encoder and decoder sub-modules, as well as the multi-head self-attention mechanism, avoiding the problems of gradient decay and long-term dependencies in temporal modeling, and is particularly suitable for modeling the interaction between the key point trajectories and gaze trends under the continuous behavior of pets. In the encoder module, the multi-head self-attention mechanism realizes the temporal co-modeling of "pose change - gaze offset" by calculating the feature relationships between any two time frames. For example, when the pet moves forward continuously for several frames while gazing at a certain area, the attention mechanism can automatically identify the "trend weight" of this behavior in the sequence and embed it as a high-weight focus point into the global time-dependent expression. This encoded representation not only contains the short-term features of local motion but also reflects the evolution law of the behavior trend in the time dimension, providing a well-structured temporal latent representation for subsequent prediction. In the decoder module, the model realizes the prediction of the pet's movement trajectory in the next several time steps by introducing the position embedding mechanism and the spatial perception decoding structure. Different from directly outputting a single target position, this decoder combines the time position offset and the current skeleton orientation to infer the possible direction and amplitude of the position change of the pet in the next step, and then generates a continuous target position prediction sequence. Each prediction point is located in three-dimensional space, constituting a multi-step fitting result of the future movement path. Finally, the system performs spatial position matching between the prediction sequence and the three-dimensional semantic map output by the semantic segmentation module. Specifically, the spatial coordinates in the prediction path are used to calculate the similarity with the semantic target positions such as high platforms, fragile objects, and ground foreign objects in the three-dimensional semantic map to determine whether the pet has a tendency to approach these dangerous areas. If the matching distance is lower than the preset threshold, the end point of this path is marked as a potential dangerous area, and its corresponding semantic label (such as "balcony railing", "glass vase", "ground foreign object", etc.) is output as the prediction result of the movement target area.Taking the actual situation as an example, if the system continuously observes that the pet keeps the direction of its head gaze on a certain high platform area, and at the same time the skeletal structure shows features such as the body leaning back and the hind legs shaking to accumulate strength, and the position gradually approaches the target area, the Transformer model will learn the "jump preparation" behavior pattern and direct its next predicted path to the high area. After matching, it is concluded that the position is a dangerous platform marked in the three-dimensional semantic map, thereby judging that the pet's behavior has the risk intention of "high jump preparation". This embodiment is different from the existing solution of judging danger through static coordinates in terms of method layout. It jointly models the structural action, orientation trend and time relationship, and introduces a spatial alignment mechanism based on semantic structure. The algorithm uses the Transformer's ability to perform time series modeling, so that the system no longer relies on manually defined behavior rules or static threshold judgments, but has deep semantic reasoning and continuous target prediction capabilities to achieve early warning of dangerous pet behaviors.

[0090] Furthermore, the multi-head self-attention mechanism is used to model the key point trajectory changes and gaze direction trends in the input sequence, including the following steps:

[0091] Performing feature embedding mapping on the sequence of behavior state vectors to generate corresponding query vector sequence, key vector sequence and value vector sequence respectively;

[0092] Based on the query vector sequence and the key vector sequence, a multi-head self-attention mechanism is used to calculate the attention distribution weights between each time step in the behavior state vector sequence;

[0093] Based on the attention distribution weight and value vector sequence, the temporal behavior feature representation of each attention head is calculated to obtain a multi-head temporal dependency feature set;

[0094] The multi-head temporal dependency feature sets are concatenated and fused, and converted into a global temporal dependency encoding representation through an output mapping layer.

[0095] In some embodiments, first, the input behavioral state sequence consists of state vectors at several moments. Each vector contains the coordinates of the skeletal key points and the gaze direction vector of the current frame, which are used to characterize the immediate state of the pet in terms of spatial structure and intended orientation. To adapt to the computational requirements of the multi-head attention mechanism, the system constructs three mutually independent but dimensionally consistent vector sequences through linear mapping, corresponding to the query vector, the key vector, and the value vector respectively. This linear transformation can be regarded as a feature reconstruction process, which projects the original input into subspaces that each attention head focuses on, ensuring that subsequent calculations are distinguishable in different semantic dimensions. Subsequently, the system calculates the correlation score between the query and the key within each attention head respectively, that is, measures the similarity relationship between the current time step and other time steps through the dot product method. This process constructs a set of time-based attention distribution matrices, which are used to weight-encode the importance of each position in the value vector sequence. In the pet behavior recognition scenario, this kind of attention mechanism can effectively capture cross-time-frame pattern associations such as "continuously gazing at a target and gradually approaching" and "instantly jumping after continuous posture adjustment", significantly enhancing the perception ability of time-dependent structures. After that, the system applies the attention weights to the value vector sequence to generate the context feature representations corresponding to each attention head. Multiple attention heads model the time-series data from different dimensions, and have the ability to co-model short-term action changes (such as sudden head turning and turning around) and medium- and long-term behavior trends (such as continuously approaching a target object). They can integrate the skeletal displacement trend and the gaze direction drift dynamics into a set of hierarchical semantic time feature vectors. Finally, the system concatenates the context representations of all attention heads and completes the unified dimension projection through the output mapping layer, outputting the global time-dependent encoding representation. As the core output of the Transformer encoder, this representation not only retains the joint dynamic structure of actions and gaze changes in the input sequence, but also has strong expressiveness and scalability, and can provide high-quality semantic embedding support for the subsequent decoding and prediction module. This module breaks through the performance bottleneck of traditional RNN or single-head attention mechanisms in long sequence modeling in terms of algorithm architecture, and improves the model's ability to model complex pet behavior trends by explicitly constructing a multi-dimensional coupling expression of time-structure-intention. Especially in the initial stage of continuous state changes, the trend of behavior turning can be perceived through time correlation modeling.

[0096] Furthermore, the formula of the position-aware decoding network is as follows:

[0097] ;

[0098] Wherein, is the predicted target position vector of the th frame; is the current time step; is the offset step number predicted into the future; s is the historical time step index; For the current prediction step Attention weights for the s-th frame in history; Global temporal dependence encoding for the s-th frame output by the encoder; Position embedding vector for the s-th frame; Decoding linear mapping weight matrix; Embedding state vector of the current frame; Projection matrix from the current frame state to the position vector.

[0099] Specifically, when predicting a future time step, the system first models the dynamic correlation between the current prediction step and historical moments. By calculating the similarity between the current position embedding state vector and the states at each historical time step, the attention distribution weights of this prediction step for all historical moments are obtained. The above weights are used to weightedly fuse the temporal dependence features output by the encoder to form a dynamic representation of the structural behavior. At the same time, to enhance the semantic perception in spatial positions, the system incorporates the spatial embedding vectors at each historical moment into the weighting process, so that the decoder output not only reflects the temporal behavior change trend but also the approaching direction in spatial semantics. The fused result is then input into the position projection network in the decoder and mapped to the three-dimensional space coordinate domain through a linear transformation, and the output is the spatial target position vector of the current prediction time step. The system iteratively processes each prediction time step in this way to finally form a complete prediction path sequence. It should be noted that to achieve time-space joint modeling, the system introduces the current position embedding mechanism in the process of predicting future behavior trends, which is used to transform the physical space position where the pet is located in historical time steps into a vector form with semantic representation ability, that is, the position embedding vector. Specifically, in the process of constructing the three-dimensional semantic map, the positions of the pet bone center points in each frame of image in the three-dimensional coordinate system have been structurally encoded. This encoding not only contains coordinate values but also additional semantic labels or category information according to the semantic regions they are in (such as near the balcony, high platform, near the vase or table corner, etc.). To introduce this discrete spatial position information into the decoding network, the system vectorizes and embeds the three-dimensional coordinates and their semantic categories. This vectorization process can adopt a combination of two methods: one is to use a coordinate position encoding function, such as after normalizing the spatial coordinates, using sine and cosine periodic functions for embedding to obtain a position encoding with time series modeling ability; the other is to use a trainable embedding matrix for look-up table mapping of the semantic label information, so that similar spatial regions have close vector representations. After combining these two methods, they are mapped to the same feature dimension as the encoder output through a linear transformation layer to form a unified format of position embedding vector.

[0100] Furthermore, the attention transfer strategy includes voice broadcasting, pet snack dispensing, and pet toy movement.

[0101] In this embodiment, in order to achieve effective intervention and behavior guidance for pet dangerous behaviors, the system designs a multimodal attention diversion strategy, which is used to interrupt its behavior path and change its attention focus through active guidance when a high-risk intention state (such as jumping preparation, ingestion tendency or approaching fragile objects) is detected, so as to prevent the occurrence of dangerous behaviors. This strategy includes three specific forms: sound broadcasting, pet snack delivery and pet toy movement, which can be triggered independently or jointly called according to the strategy priority. When the system infers that the pet is about to perform a potentially dangerous action, such as jumping to the window edge, trying to pick up foreign objects on the ground, etc. through the behavior recognition module and the position perception decoding network, its behavior state is judged to have a high-risk tendency. On this basis, the system first matches the corresponding attention diversion strategy according to the category of the behavior intention and the target area attribute of the predicted path. For example, for the intention of jumping near the window edge, the system prefers the sound source diversion scheme to interfere with its concentration; if it is an ingestion behavior, it is more inclined to change its movement direction through an inducement guidance method. The sound broadcasting module is completed by intelligent audio equipment deployed in various places in the room, which has the ability of controllable direction and adjustable timbre. In actual applications, the system will automatically select the most appropriate speaker to play the specified voice command or a prompt tone of a specific frequency based on the pet's current position and predicted path, and guide the pet to divert its attention or terminate the current behavior through the change of the spatial sound field. For example, when the pet is accumulating strength to jump to a high area, the system can play the owner's call behind it, forcing it to turn around and pause its action. For the snack delivery method, the system pre-arranges intelligent feeding devices at several locations indoors, containing an appropriate amount of snacks that the pet likes. When the predicted path approaches a dangerous area, the system can select a device in a relatively safe direction to initiate the delivery action, and the pet can be visually captured, thereby guiding it to divert its movement path in the direction of the snack, reducing the possibility of it continuing to approach the high-risk area. This method performs well in medium and short-distance guidance and is suitable for interrupting behaviors such as low-speed approach to foreign objects. In addition, in the toy movement guidance strategy, the system links electric pet toy modules installed on the ground or in the corners, such as balls or mobile track devices. Based on the inference results of the pet's current position, gaze direction, and skeletal posture, the system can initiate a toy start command to move the pet to a non-dangerous area at a certain speed to construct a new attention stimulus target and attract the pet to change its current path. Different from the existing solutions based on passive alarms or video notifications, the attention diversion mechanism proposed by this system not only realizes the prediction and identification of risky behaviors, but also completes targeted behavior induction based on the prediction results, effectively preventing dangerous behaviors of pets.

[0102] The above embodiments are only descriptions of the preferred embodiments of the present invention, and do not limit the scope of the present invention. Without departing from the design spirit of the present invention, various deformations and improvements made by those of ordinary skill in the art to the technical solutions of the present invention shall fall within the protection scope determined by the claims of the present invention.

Claims

1. A pet dangerous behavior monitoring system based on a visual neural network, characterized in that, Including: A data acquisition module, which is used to obtain indoor area image data and indoor space point cloud data; A semantic segmentation module, which is used to perform semantic recognition of indoor structures and objects based on the image data using a visual neural network. At the same time, based on the point cloud data, it extracts spatial height information, edge structures, and object contour features, and fuses the image semantic results and point cloud geometric data to generate a three-dimensional semantic map; A behavior parameter construction module, which is used to construct a pet behavior dataset based on the indoor area image data and indoor space point cloud data. Based on the pet behavior dataset, it extracts the pet bone point sequence through a pose estimation network, and obtains the spatial vector of the pet's current gaze direction through a convolutional neural network; A behavior recognition module, which is used to predict the pet's movement target area through a time series model based on the bone point sequence and the spatial vector of the gaze direction, and calculate the pet's dangerous intention state based on the matching relationship between the movement target area and the positions of environmental objects in the three-dimensional semantic map. The pet's dangerous intention state includes high jump preparation, foreign body ingestion, and approaching fragile objects; A pet guidance module, which is used to send an alarm to the mobile terminal according to the pet's dangerous intention state and execute an attention transfer strategy; The semantic segmentation module is used to perform the following steps: Based on the indoor area image data, use a semantic segmentation network to process the image data and extract an image semantic feature map containing object semantic categories and two-dimensional spatial positions; Based on the image semantic feature map, combined with the spatially synchronized point cloud data, map the image features to the point cloud space coordinate system through depth projection to obtain a projected semantic feature map; Based on the spatial height information, edge structures, and geometric contour features extracted from the projected semantic feature map and the spatial point cloud data, construct a joint feature representation; Input the joint feature representation into a three-dimensional semantic fusion network to generate a three-dimensional semantic map, and perform semantic annotation on high platforms, foreign objects prone to ingestion, and fragile objects in the indoor environment in the three-dimensional semantic map; The pose estimation network is constructed through the following steps: Based on the indoor area image data, use a multi-scale convolutional network to process the image sequence and extract a sequence of image feature maps containing pet contour and structure information; Based on the sequence of image feature maps, use a key point detection sub-network to perform key point heat map regression processing on predefined key parts in each frame of the image, and output a sequence of heat maps representing the probability distribution of key point responses; Based on the sequence of heat maps, use temporal modeling to model the response trajectories of key points between consecutive image frames, and perform trajectory fitting and spatial correction on the positions of key points in combination with the inter-frame time order to output the pet bone point sequence.

2. The pet dangerous behavior monitoring system based on a visual neural network according to claim 1, characterized in that The obtaining of the indoor area image data and indoor space point cloud data includes the following steps: Collect video image frames of the pet activity area at a fixed frequency through RGB cameras set at several positions indoors to obtain indoor area image data; Collect spatial point cloud data of the indoor area through a lidar.

3. The pet dangerous behavior monitoring system based on a visual neural network according to claim 1, characterized in that, The formula of the key point detection sub-network is as follows: ; Among them, is the set of heatmaps of key points of the output; is the Sigmoid activation function; is the weight of the fully connected regression layer; is the bias term; The feature map extracted by the -th multi-scale convolution; is the corresponding residual correction term of the feature map is the weighted coefficient generated by the channel attention mechanism, satisfying ; is the number of multi-scale branches.

4. The pet dangerous behavior monitoring system based on a visual neural network according to claim 1, characterized in that, The obtaining of the spatial vector of the pet's current gaze direction through a convolutional neural network includes the following steps: Construct a gaze direction training dataset, where the dataset includes a number of pet head images and corresponding gaze direction space vectors; Using the pet head image as input, image features are extracted through a convolutional neural network to obtain an image feature vector representing the facial orientation feature; Based on the image feature vector, the gaze direction space vector is regressively predicted through a regression structure. The regression prediction uses the corresponding gaze direction space vector in the gaze direction training dataset as the supervision target, and the mean square error is used as the loss function to optimize the network parameters; Based on the trained convolutional neural network model, when running, it receives the current frame head region image of the pet as input and outputs the current gaze direction space vector of the pet.

5. The pet dangerous behavior monitoring system based on a visual neural network according to claim 4, wherein, The predicting the pet's motion target area through a time series model based on the skeletal point sequence and the spatial vector of the gaze direction includes the following steps: Based on the skeletal point sequence output by the pose estimation network and the gaze direction space vector output by the convolutional neural network, a behavior state vector is constructed in each frame image. The behavior state vector includes the combined result of the spatial coordinates of each key point and the unit vector of the gaze direction in the current frame; Arrange the behavior state vectors in chronological order to form an input sequence, and input it into the Transformer prediction model. The prediction model includes an encoder sub-module and a decoder sub-module; The encoder sub-module uses the multi-head self-attention mechanism to model the key point trajectory changes and gaze direction trends in the input sequence, and outputs a global time-dependent coding representation; Based on the global time-dependent coding representation and the current position embedding information, the decoder sub-module predicts the spatial motion trend of the pet in the next few time steps through a position-aware decoding network, and outputs a corresponding target position prediction sequence; Based on the spatial similarity matching between the target position prediction sequence and the positions of the environment objects annotated in the three-dimensional semantic map, the approaching area of the pet's motion path is determined, and the semantic position label of the approaching area is output as the prediction result of the motion target area.

6. The pet dangerous behavior monitoring system based on a visual neural network according to claim 5, characterized in that, The modeling of the key point trajectory changes and gaze direction trends in the input sequence using the multi-head self-attention mechanism includes the following steps: Perform feature embedding mapping on the sequence of the behavior state vectors to generate corresponding query vector sequences, key vector sequences, and value vector sequences respectively; Based on the query vector sequence and the key vector sequence, use the multi-head self-attention mechanism to calculate the attention distribution weights between each time step in the behavior state vector sequence; Based on the attention distribution weights and the value vector sequence, calculate the temporal behavior feature representation of each attention head to obtain a multi-head temporal dependence feature set; Concatenate and fuse the multi-head temporal dependence feature set, and convert it into a global time-dependent coding representation through an output mapping layer.

7. The pet dangerous behavior monitoring system based on a visual neural network according to claim 6, characterized in that, The formula of the position-aware decoding network is as follows: ; Among them, is the target position prediction vector of the th frame; is the current time step; is the offset step number predicted into the future; s is the historical time step index; is the current prediction step is the attention weight for the historical s-th frame; is the global time dependence encoding of the s-th frame output by the encoder; is the position embedding vector of the s-th frame; is the decoding linear mapping weight matrix; is the embedding state vector of the current frame; is the projection matrix from the current frame state to the position vector.

8. A pet dangerous behavior monitoring system based on a visual neural network according to claim 1, characterized in that, The attention transfer strategies include voice broadcasting, pet snack dispensing, and pet toy movement.

Citation Information

Patent Citations

  • Pet behavior analysis system based on AI vision

    CN117894078A

  • Systems and methods for deep localization and segmentation with a 3D semantic map

    US20200364554A1