Pet dangerous behavior monitoring system based on optic neural network

Through a pet hazard behavior monitoring system based on visual neural network, a three-dimensional semantic map is constructed using image and point cloud data to identify the pet's movement target area and dangerous intention status, solving the problem that the existing technology cannot understand and intelligently analyze pet behavior in real time, and achieving efficient prevention and intervention in pet dangerous behavior.

CN120108006AActive Publication Date: 2025-06-06广州佳可电子科技股份有限公司

Patent Information

Application Number
CN202510602121.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-12
Publication Date
2025-06-06
Estimated Expiration
2045-05-12

AI Technical Summary

Technical Problem

Existing pet monitoring products cannot understand and intelligently analyze pet behavior in real time, and cannot identify pet's action intentions, spatial goals or behavior trends, resulting in the inability to detect and prevent dangerous behaviors in a timely manner.

Method used

A pet hazard behavior monitoring system based on visual neural network is adopted to obtain indoor image data and point cloud data through the data acquisition module, and a semantic segmentation module is used to perform image semantic recognition and point cloud geometric data extraction, a three-dimensional semantic map is constructed, and a behavioral parameter building module and behavior recognition module are used to predict the pet's movement target area and dangerous intention state.

Benefits of technology

Real-time understanding and intelligent analysis of pet dangerous behaviors is achieved, and can predict the pet's sports target area and dangerous intention status, improving the ability to prevent and intervene pet dangerous behaviors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120108006A_ABST
    Figure CN120108006A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of pet behavior monitoring, in particular to a pet dangerous behavior monitoring system based on an optic neural network, which comprises the following steps: acquiring image data and spatial point cloud data, extracting image semantic features and spatial structure features by using a semantic segmentation module, and fusing to generate a three-dimensional semantic map; and dangerous areas such as high platforms, foreign matters easy to eat by mistake, fragile objects and the like are marked. A pet skeleton key point sequence is extracted through a posture estimation network, a gazing direction space vector is obtained through a convolutional neural network, a behavior state vector is constructed and then input into a time sequence model, a pet movement target area is predicted, and the pet movement target area is matched with a three-dimensional semantic map to judge whether a dangerous behavior intention exists or not, such as jumping, mistakenly eating or approaching a fragile object. According to the invention, early warning and active intervention of dangerous behaviors of the pet are realized, and the monitoring accuracy and the response capability are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of pet behavior monitoring, and in particular to a pet dangerous behavior monitoring system based on visual neural network. Background Art

[0002] Pets play an increasingly important role in the family. However, when their owners are away for a long time or cannot monitor them in real time, pets may exhibit a series of dangerous behaviors that are often difficult to detect, identify and stop in time, thus triggering a series of chain consequences.

[0003] Take common pets such as cats and dogs as an example. Their activity space covers the entire indoor environment. They have a certain jumping ability and desire to explore, and are very likely to jump from high places, chew foreign objects, and break fragile items. Among them, high-altitude jumping behaviors such as jumping onto windowsills, balcony guardrails, or the edge of a second-floor structure may cause falls if the action is wrong or the structure is unstable, causing fractures or internal injuries to the pet; accidental ingestion is common in small plastic parts, which may cause digestive obstruction, poisoning and other serious consequences once swallowed; and when running, chasing and playing at high speed, pets are very likely to knock over high-value or fragile items such as vases, televisions, stereos, and computer monitors. In addition to causing property losses, the fragments may also scratch the pets themselves or pose a safety hazard to vulnerable groups such as infants and young children.

[0004] Most pet monitoring products on the market currently use the technology of general security cameras, which only have image acquisition and historical video playback functions, and cannot achieve real-time understanding and intelligent analysis of pet behavior. A small number of products have introduced simple target detection algorithms, which can achieve static judgment of the pet's "appearance in the picture", but cannot identify its action intentions, spatial targets or behavioral trends. For example, when a pet is about to jump onto the balcony railing, approach fragile objects, or try to pick up foreign objects on the ground, the system cannot respond effectively, resulting in a complete lack of early warning and intervention links. Summary of the invention

[0005] In order to solve the above problems, the present invention provides a pet dangerous behavior monitoring system based on visual neural network.

[0006] To achieve the above object, the technical solution adopted by the present invention is:

[0007] A pet dangerous behavior monitoring system based on visual neural network, comprising:

[0008] Data acquisition module, used to obtain indoor area image data and indoor space point cloud data;

[0009] The semantic segmentation module is used to perform semantic recognition of indoor structures and objects based on image data using a visual neural network. It also extracts spatial height information, edge structures, and object contour features based on point cloud data, and fuses image semantic results with point cloud geometry data to generate a three-dimensional semantic map.

[0010] The behavior parameter construction module is used to construct a pet behavior dataset based on indoor area image data and indoor space point cloud data. Based on the pet behavior dataset, the pet skeleton point sequence is extracted through the posture estimation network, and the spatial vector of the pet's current gaze direction is obtained through the convolutional neural network.

[0011] A behavior recognition module, for predicting the pet's target movement area through a time series model based on the skeleton point sequence and the spatial vector of the gaze direction, and calculating the pet's dangerous intention state based on the matching relationship between the target movement area and the position of the environmental object in the three-dimensional semantic map, wherein the pet's dangerous intention state includes preparation for high-altitude jumping, swallowing foreign objects by mistake, and approaching fragile objects;

[0012] The pet guidance module is used to send alarms to the mobile terminal according to the pet's dangerous intention status and implement attention diversion strategies.

[0013] Furthermore, the acquisition of indoor area image data and indoor space point cloud data comprises the following steps:

[0014] By setting up RGB cameras at several positions indoors to collect video image frames of the pet activity area at a fixed frequency, indoor area image data is obtained;

[0015] The spatial point cloud data of the indoor area is collected by LiDAR.

[0016] Furthermore, the semantic segmentation module is used to perform the following steps:

[0017] Based on the indoor area image data, a semantic segmentation network is used to process the image data and extract the image semantic feature map containing the object semantic category and two-dimensional spatial position;

[0018] Based on the image semantic feature map, combined with the time-synchronized spatial point cloud data, the image features are mapped to the point cloud space coordinate system through deep projection to obtain a projected semantic feature map;

[0019] Constructing a joint feature representation based on the spatial height information, edge structure and geometric contour features extracted from the projected semantic feature map and the spatial point cloud data;

[0020] The joint feature representation is input into a three-dimensional semantic fusion network to generate a three-dimensional semantic map, and semantic annotations are performed on high platforms, easily ingested foreign objects, and fragile objects in the indoor environment in the three-dimensional semantic map.

[0021] Furthermore, the posture estimation network is constructed by the following steps:

[0022] Based on the indoor area image data, a multi-scale convolutional network is used to process the image sequence and extract the image feature map sequence containing the pet's contour and structure information;

[0023] Based on the image feature map sequence, a key point detection subnetwork is used to perform key point heat map regression processing on predefined key parts in each frame of the image, and a heat map sequence representing the probability distribution of key point responses is output;

[0024] Based on the heat map sequence, time series modeling is used to model the response trajectory of key points between consecutive image frames, and the trajectory fitting and spatial correction of the key point positions are performed in combination with the time sequence between frames to output the pet bone point sequence.

[0025] Furthermore, the formula of the key point detection subnetwork is as follows:

[0026] ;

[0027] in, It is the set of key point heat maps output; is the Sigmoid activation function; is the weight of the fully connected regression layer; is the bias term; No. Feature map extracted by multi-scale convolution; The feature map The corresponding residual correction term of ; The weight coefficient generated by the channel attention mechanism satisfies ; is the number of multi-scale branches.

[0028] Furthermore, the method of obtaining the spatial vector of the pet's current gaze direction through a convolutional neural network includes the following steps:

[0029] Constructing a gaze direction training data set, the data set comprising a plurality of pet head images and corresponding gaze direction space vectors;

[0030] Taking the pet head image as input, the convolutional neural network is used to extract image features and obtain the image feature vector representing the facial orientation features;

[0031] Based on the image feature vector, a regression structure is used to perform regression prediction on the gaze direction space vector, wherein the regression prediction takes the gaze direction space vector corresponding to the gaze direction training data set as a supervision target, and uses the mean square error as a loss function to optimize the network parameters;

[0032] Based on the trained convolutional neural network model, the pet's head area image in the current frame is received as input during runtime, and the pet's current gaze direction spatial vector is output.

[0033] Furthermore, the predicting of the pet movement target area through a time series model based on the skeleton point sequence and the spatial vector of the gaze direction includes the following steps:

[0034] Based on the skeleton point sequence output by the posture estimation network and the gaze direction space vector output by the convolutional neural network, a behavior state vector is constructed in each frame image, wherein the behavior state vector includes the combination result of the spatial coordinates of each key point in the current frame and the gaze direction unit vector;

[0035] Arrange the behavior state vectors in chronological order to form an input sequence, and input it into a Transformer prediction model, wherein the prediction model includes an encoder submodule and a decoder submodule;

[0036] The encoder submodule adopts a multi-head self-attention mechanism to model the trajectory changes and gaze direction trends of key points in the input sequence, and outputs a global time-dependent encoding representation;

[0037] The decoder submodule predicts the spatial motion trend of the pet in a few future time steps through a position-aware decoding network based on the global time-dependent coding representation and the current position embedding information, and outputs the corresponding target position prediction sequence;

[0038] Based on the spatial similarity matching between the target position prediction sequence and the environmental object positions marked in the three-dimensional semantic map, the approach area of ​​the pet's movement path is determined, and the semantic position label of the approach area is output as the movement target area prediction result.

[0039] Furthermore, the multi-head self-attention mechanism is used to model the key point trajectory changes and gaze direction trends in the input sequence, including the following steps:

[0040] Performing feature embedding mapping on the sequence of behavior state vectors to generate corresponding query vector sequence, key vector sequence and value vector sequence respectively;

[0041] Based on the query vector sequence and the key vector sequence, a multi-head self-attention mechanism is used to calculate the attention distribution weights between each time step in the behavior state vector sequence;

[0042] Based on the attention distribution weight and value vector sequence, the temporal behavior feature representation of each attention head is calculated to obtain a multi-head temporal dependency feature set;

[0043] The multi-head temporal dependency feature sets are concatenated and fused, and converted into a global temporal dependency encoding representation through an output mapping layer.

[0044] Furthermore, the formula of the position-aware decoding network is as follows:

[0045] ;

[0046] in, For the The target position prediction vector of the frame; is the current time step; is the number of steps predicted into the future; s is the historical time step index; The current prediction step The attention weight for the s-th frame in history; The global temporal dependency encoding of the sth frame output by the encoder; is the position embedding vector of the sth frame; is the weight matrix of the decoding linear map; The embedded state vector of the current frame; It is the projection matrix from the current frame state to the position vector.

[0047] Furthermore, the attention diversion strategy includes sound broadcasting, pet snack delivery, and pet toy movement.

[0048] The beneficial effects of the present invention are as follows: the present invention acquires indoor area image data and indoor space point cloud data by setting a data acquisition module, and uses a semantic segmentation module to perform visual neural network processing on the image data, extract semantic feature maps, and extract spatial height, edge structure and object contour information from the point cloud data. Furthermore, by fusing the image semantic results with the point cloud geometric data, a three-dimensional semantic map is constructed, and target areas with dangerous attributes are explicitly marked in the map, including high platforms, foreign objects that are easily swallowed by mistake, and fragile objects. Through the coordinated fusion of images and point clouds, the system can not only obtain the two-dimensional position of the pet, but also understand its activity environment in three-dimensional space, forming a unified understanding of physical structure and semantic attributes. The pet's skeletal key point sequence is extracted from the acquired image sequence through the posture estimation network, and the spatial vector of the pet's current gaze direction is obtained through the convolutional neural network. On this basis, the behavior parameter construction module constructs a behavior state vector for prediction. Furthermore, the behavior recognition module inputs the skeleton key point sequence and the gaze direction space vector into the time series model, predicts the next stage of the pet's movement target area by learning its time series evolution trend, and matches the area with the dangerous area in the three-dimensional semantic map for spatial position, and determines whether the pet has dangerous behavior intentions, such as jumping preparation, mistaken ingestion tendency, approaching fragile areas, etc. This processing chain realizes the continuous reasoning process from behavioral characteristics to behavioral trends and then to risk status, which is significantly better than the existing system's single-frame judgment ability of the target position in the static frame. Finally, the system sets up a pet guidance module, which sends warning information to the mobile terminal according to the identified dangerous intention state, and can execute attention diversion strategies to assist users in guiding pets away from high-risk behaviors. In summary, by establishing a three-dimensional semantic map through the fusion of image and point cloud data, combining skeleton key points and gaze direction to model the pet behavior state, and using the Transformer model for motion trend prediction, a pet dangerous behavior monitoring system with deep structural understanding, intention perception and trend judgment capabilities is formed, which effectively prevents the occurrence of pet dangerous behaviors. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] Figure 1 It is a structural schematic diagram of a pet dangerous behavior monitoring system based on a visual neural network in the present invention.

[0050] Figure 2 The present invention predicts the pet movement target area map through a time series model based on the skeleton point sequence and the spatial vector of the gaze direction. DETAILED DESCRIPTION

[0051] See also Figure 1-Figure 2 As shown, the present invention relates to a pet dangerous behavior monitoring system based on a visual neural network, comprising:

[0052] Data acquisition module, used to obtain indoor area image data and indoor space point cloud data;

[0053] The semantic segmentation module is used to perform semantic recognition of indoor structures and objects based on image data using a visual neural network. It also extracts spatial height information, edge structures, and object contour features based on point cloud data, and fuses image semantic results with point cloud geometry data to generate a three-dimensional semantic map.

[0054] The behavior parameter construction module is used to construct a pet behavior dataset based on indoor area image data and indoor space point cloud data. Based on the pet behavior dataset, the pet skeleton point sequence is extracted through the posture estimation network, and the spatial vector of the pet's current gaze direction is obtained through the convolutional neural network.

[0055] A behavior recognition module, for predicting the pet's target movement area through a time series model based on the skeleton point sequence and the spatial vector of the gaze direction, and calculating the pet's dangerous intention state based on the matching relationship between the target movement area and the position of the environmental object in the three-dimensional semantic map, wherein the pet's dangerous intention state includes preparation for high-altitude jumping, swallowing foreign objects by mistake, and approaching fragile objects;

[0056] The pet guidance module is used to send alarms to the mobile terminal according to the pet's dangerous intention status and implement attention diversion strategies.

[0057] Specifically, the data acquisition module includes several RGB cameras and laser radar devices arranged at different locations indoors. The RGB camera collects image frame data of the pet's activity area at a fixed frequency to provide high-resolution texture, semantics and posture information; the laser radar collects spatial point cloud data of the corresponding area indoors to reflect the three-dimensional structure, terrain height changes and object contours. Through the parallel acquisition of this heterogeneous data, the input basis of image-space joint expression is provided for subsequent modules. In the semantic segmentation module, the image data is first input into a visual neural network based on an encoding-decoding structure for processing. The network uses a multi-scale feature extraction channel combined with an upsampling structure to output a two-dimensional semantic feature map. Each pixel in the map corresponds to a category label and semantic confidence. At the same time, the point cloud data is processed by edge extraction, curvature estimation and other processing methods to extract spatial height distribution, edge morphology and three-dimensional contour information. The system maps the image semantic feature map to the point cloud coordinate system through depth projection, and combines it with the geometric features of the point cloud to form a semantic structure in three-dimensional space. The joint representation is further input into the 3D semantic fusion network to achieve deep fusion of image semantics and spatial geometry, and finally generate a 3D semantic map, which explicitly annotates the spatial positions and semantic attributes of high platforms (such as the edge of the second floor, window sills), pick-up foreign objects on the ground, glass bottles and other fragile objects. This map not only overcomes the limitation that the simple image recognition system cannot understand the spatial height, but also makes up for the problem that the point cloud system lacks the ability to distinguish semantics, and realizes the complementary enhancement of image perception and 3D understanding, so as to have higher accuracy and stronger generalization recognition ability for dangerous areas in complex indoor layouts, reflecting the multiplicative effect of "semantic + geometry" fusion recognition on risk modeling. The behavior parameter construction module constructs a pet behavior dataset based on the collected image sequence and point cloud sequence. The image sequence is input into the posture estimation network, which constructs a heat map sequence by combining the convolutional network with the key point detection sub-network, and further extracts the skeleton key point sequence in the pet's continuous frames by inter-frame trajectory modeling. In order to capture the pet's behavioral intention, the system further processes the pet's head image through a convolutional neural network, and regresses and predicts the spatial vector corresponding to the pet's gaze direction in the current frame. The system uses the key point sequence and gaze direction to form the pet's behavior parameter vector in the current state as the input for subsequent temporal modeling and predictive analysis. The behavior recognition module models and predicts based on the key point sequence and gaze direction spatial vector through a time series model (such as a Transformer based on a self-attention structure). This module not only models the spatial state of the current frame, but also performs long-term modeling of the key point change trend and gaze direction change pattern, and outputs the moving target area in several future time steps. Subsequently, the system matches the predicted result with the spatial hazard annotation position in the three-dimensional semantic map for spatial similarity.If the match is successful and the target area belongs to the preset dangerous area, such as the edge of a high platform, fragile objects or foreign objects on the ground, the system will output the corresponding dangerous behavior intention state, such as "preparation to jump", "approaching fragile objects" or "ingestion tendency". Compared with the existing method based on the current frame classification judgment, this module introduces explicit time modeling and spatial semantic fusion mechanism, which realizes the trend prediction of the pet's future behavior and significantly improves the advance and accuracy of danger warning. When the dangerous behavior intention state is detected, the pet guidance module triggers the corresponding control response according to the current state, and sends the alarm information to the mobile terminal synchronously to prompt the user. The guidance strategy includes executing attention diversion means (such as sound playback or other strategy component calls) to intervene in the pet's behavior path. The overall system realizes a closed-loop control process from "perception-modeling-prediction-guidance". The system in the embodiment introduces a deep network structure in both image recognition and point cloud modeling, and synergistically integrates the two through a three-dimensional semantic map, so that the system has the ability to express space and semantics in a unified manner. In terms of behavior modeling, the system forms a pet behavior state vector by combining posture estimation and gaze direction, and uses a time series model to predict future behavior trends, significantly enhancing the ability to understand the dynamic process of behavior. Compared with the existing technology that only relies on single-frame analysis or image target detection, this embodiment achieves improved perception accuracy and response timeliness through the synergy of algorithms between modules, showing higher practicality and stability in actual pet behavior monitoring scenarios.

[0058] Furthermore, the acquisition of indoor area image data and indoor space point cloud data comprises the following steps:

[0059] By setting up RGB cameras at several positions indoors to collect video image frames of the pet activity area at a fixed frequency, indoor area image data is obtained;

[0060] The spatial point cloud data of the indoor area is collected by LiDAR.

[0061] In the specific deployment, RGB cameras are set at different viewing angles indoors, covering typical pet high-frequency activity areas including living rooms, bedrooms, balconies, and stair corners. Each camera collects images at a preset fixed frequency and records video streams at a frame rate of 30 frames per second during the system operation cycle. The image resolution is set at 1920×1080 to ensure high-fidelity presentation of the pet's face, limb contours, and surrounding objects. The collected image frames are synchronously entered into the image data cache according to the timestamp number as the input source for subsequent image semantic recognition, posture estimation, and gaze direction modeling. The system uses a 360-degree scanning indoor laser radar, which is deployed in the carrier device equipped with the camera to ensure that it can fully obtain the three-dimensional point cloud data of all visible objects in the space during the working cycle. The laser radar operating frequency is time-aligned with the image acquisition frequency to achieve pairing and binding of space-image data between frames. The collected point cloud data is mainly used to reflect the real height, boundary structure, and geometric features of objects in three-dimensional space, such as the spatial distribution of furniture, the height difference of stair platforms, and the spatial relationship between the pet's location and various dangerous objects. Through the above-mentioned simultaneous acquisition of images and point clouds, not only can the visual information of pets' dynamic behaviors be collected, but also the ability to understand the spatial environment in a structured manner is significantly improved. In particular, the point cloud data's ability to reflect height, boundaries, and spatial relationships not only compensates for the lack of depth dimension in RGB images, but also provides an accurate and quantifiable basis for the subsequent construction of three-dimensional semantic maps, spatial matching, and behavioral reasoning.

[0062] Furthermore, the semantic segmentation module is used to perform the following steps:

[0063] Based on the indoor area image data, a semantic segmentation network is used to process the image data and extract the image semantic feature map containing the object semantic category and two-dimensional spatial position;

[0064] Based on the image semantic feature map, combined with the time-synchronized spatial point cloud data, the image features are mapped to the point cloud space coordinate system through deep projection to obtain a projected semantic feature map;

[0065] Constructing a joint feature representation based on the spatial height information, edge structure and geometric contour features extracted from the projected semantic feature map and the spatial point cloud data;

[0066] The joint feature representation is input into a three-dimensional semantic fusion network to generate a three-dimensional semantic map, and semantic annotations are performed on high platforms, easily ingested foreign objects, and fragile objects in the indoor environment in the three-dimensional semantic map.

[0067] In some embodiments, the indoor area image data is first input into a semantic segmentation network based on an encoding-decoding structure. The main body of the network adopts a multi-scale receptive field structure. The front-end encoder extracts feature maps at different scales in the image, and the back-end decoder upsamples the high semantic layer information and restores the spatial resolution, thereby outputting the semantic category label and category confidence map corresponding to each pixel to form a two-dimensional image semantic feature map. This feature map not only provides semantic information such as pets, furniture, ground, walls, and object contours, but also retains the two-dimensional plane position relationship. Subsequently, the above-mentioned image semantic feature map is aligned and mapped with the spatial point cloud data acquired synchronously in time. The mapping process adopts a deep projection mechanism: through the internal and external parameters of the camera, the image plane coordinates are projected and matched with the three-dimensional coordinates of the point cloud, so that each point cloud point can obtain its semantic label in the image and generate a projected semantic feature map. This map not only contains the position and depth information of the point, but also embeds the semantic high-dimensional expression learned in the image network, realizing cross-modal transmission from two-dimensional plane semantics to three-dimensional spatial structure. In order to further enhance the understanding of spatial structure and semantic boundary resolution, the system extracts geometric features such as spatial height gradient, local curvature, edge strength, etc. from the original point cloud, and splices and fuses them with the projected semantic feature map to construct a joint feature representation. The joint feature vector contains multiple dimensions such as semantic label probability distribution, local spatial geometric structure, position embedding, etc. at each point. It has the ability to model the consistency of semantic-structural boundaries in complex spaces, which is significantly better than the traditional feature expression constructed based on point cloud clustering or pure image convolution. After completing the feature construction, the system inputs the joint feature representation into the three-dimensional semantic fusion network for training and reasoning. The network can use voxel convolution with residual structure or PointNet++ framework, which achieves the joint optimization of cross-modal semantic and geometric features while maintaining the continuity of spatial distribution. The final output three-dimensional semantic map can not only mark the three-dimensional positions of various semantic entities, such as sofas, coffee tables, pet beds, etc., but more importantly, it can explicitly identify and mark spatial targets with potential risks in the map, including high platforms (such as balcony fences, stair edges), scattered objects on the ground (such as pills, small toys) and low-lying fragile items (such as glass ornaments, ceramic vases).

[0068] Furthermore, the posture estimation network is constructed by the following steps:

[0069] Based on the indoor area image data, a multi-scale convolutional network is used to process the image sequence and extract the image feature map sequence containing the pet's contour and structure information;

[0070] Based on the image feature map sequence, a key point detection subnetwork is used to perform key point heat map regression processing on predefined key parts in each frame of the image, and a heat map sequence representing the probability distribution of key point responses is output;

[0071] Based on the heat map sequence, time series modeling is used to model the response trajectory of key points between consecutive image frames, and the trajectory fitting and spatial correction of the key point positions are performed in combination with the time sequence between frames to output the pet bone point sequence.

[0072] In this embodiment, the posture estimation network is used to accurately extract the pet's skeleton key point sequence from continuous image frames. Its core lies in the integration of three types of algorithm modules: multi-scale perception, key point heat map regression, and time series modeling. The small changes and continuity trends in the pet's movements are analyzed in a structured way, and the accuracy of capturing dynamic behaviors is significantly enhanced. The innovation of this module is not only reflected in the integrity of the perception chain, but also in the enhanced effect formed in the spatiotemporal collaborative modeling, that is, considering the spatial image features and the time series evolution features at the same time, so as to obtain skeleton stability and prediction robustness far superior to the traditional intra-frame estimation scheme. Specifically, the system first inputs the indoor area image data into the feature extraction network based on the multi-scale convolution structure. The network extracts visual features of different sizes and texture scales in the image in parallel through multiple convolution kernels configured with receptive fields in the front layer, such as pet contour edges, limb connection points, facial areas, etc.; the middle layer enhances the stability of feature expression through batch normalization and activation functions; the back layer uses feature compression to output a uniform size image feature map sequence. This feature map not only retains the spatial layout information, but also integrates the multi-scale structural semantics, providing rich contextual support for subsequent key point detection. Next, the image feature map sequence is sent to the key point detection subnetwork. Based on the fully convolutional regression structure, this subnetwork maps the high-dimensional features of each frame image to multiple predefined key point channels, and models the position probability of each key point by means of Gaussian heat map. The output result is a set of heat map sequences, each of which represents the possible position distribution of the corresponding key point in the frame. The heat map uses the two-dimensional Gaussian center as the regression target, and the optimization objective function is usually the mean square error loss, which reflects the difference between the predicted heat map and the annotated heat map. This method effectively overcomes the problem of unstable positioning of traditional direct coordinate regression in local occlusion and deformation scenes. However, key point detection based only on a single frame image is difficult to handle the continuous posture deformation and behavior switching of the pet during movement. Therefore, this embodiment further introduces a temporal modeling mechanism to establish trajectory continuity constraints across frames. The system inputs the response position sequence of each key point in multiple time frames into the temporal modeling module, and uses a lightweight temporal convolutional network structure to model the dynamic response path of the key points, and explore the potential laws of inter-frame behavior evolution. At the same time, combined with the inter-frame time step and physical motion consistency rules, the continuous key point trajectories are spatiotemporally fitted and positionally corrected, and a globally consistent sequence of pet skeleton key points is output, which significantly improves the robustness to local occlusion, nonlinear deformation and noise drift under complex dynamics. For example, when the system monitors a cat jumping from the sofa in the living room to the edge of the dining table, the cat's forelimbs are occluded by its tail before jumping, so it appears blurred in the local image. Traditional static key point estimation methods may result in key point loss or jumps.In this embodiment, through the time series modeling and heat map fusion correction of the first few frames, the system can output the key point position in a steady state and track its continuous trajectory during the jump, so as to completely capture the skeletal action sequence from pre-jump to landing point, and provide accurate support for subsequent dangerous behavior prediction. Compared with the existing posture detection method based on single-frame estimation or direct coordinate regression, this embodiment uses the three-layer algorithm linkage of "spatial feature extraction + key point heat map modeling + time series trajectory fusion" in the network structure, which not only improves the accuracy of key point positioning, but also realizes the deep modeling of time series consistency and motion dynamics, and has stronger behavior understanding and forward-looking prediction capabilities, which is an important technical support for realizing the core functions of the present invention.

[0073] Furthermore, the formula of the key point detection subnetwork is as follows:

[0074] ;

[0075] in, It is the set of key point heat maps output; is the Sigmoid activation function; is the weight of the fully connected regression layer; is the bias term; No. Feature map extracted by multi-scale convolution; The feature map The corresponding residual correction term of ; The weight coefficient generated by the channel attention mechanism satisfies ; is the number of multi-scale branches.

[0076] It should be noted that, first, the system inputs the image features into multiple convolution branches with the same structure but different receptive field sizes. Each branch independently extracts the local or global structural features of the image at that scale, and adapts to the spatial scale features of different key points of the pet. For example, smaller convolution kernels extract fine-grained features such as faces and ears with rich details, while larger convolution kernels are suitable for capturing overall posture features such as the trunk and limb contours. Next, the original feature map extracted by each branch will also undergo a round of residual correction operation. This operation generates a residual feature map representing structural error or detail compensation by performing a difference calculation between the original feature map and its approximate structure obtained by downsampling and upsampling. This difference reflects the deviation between the local area and its smooth reconstruction, which is used to strengthen the response of the significant area near the key point and suppress the interference of background or texture noise, especially in the case of occlusion and uneven illumination, which can significantly enhance the positioning robustness. Subsequently, the system takes all multi-scale feature maps that have undergone residual correction as input and sends them to the channel attention mechanism for weight learning and fusion. In this mechanism, the network evaluates the responsiveness of different scales according to the current task, assigns a weight value to each feature map, and the sum of all weights is one. The allocation of weight values ​​reflects the dynamic adaptability of the scale feature to the overall key point prediction, so that the network can automatically adjust the degree of attention to a specific scale in different environments, angles or pet behavior states. Finally, the weighted fused multi-scale feature map is sent to the heat map generation module, and after the fully connected mapping layer and activation function processing, a set of two-dimensional heat maps is generated, each of which corresponds to the position probability distribution of a key point. In the heat map, the higher the pixel value, the more likely the position is to be the center area of ​​the target key point. The output result provides a structural input with high spatial accuracy, continuous response distribution, and strong occlusion robustness for subsequent time series modeling and behavior prediction. Compared with the traditional method of relying only on single-scale convolution to output key point coordinates, this network structure has significant advantages in structural expression, response stability and scale adaptability. Through multi-scale perception, residual enhancement and attention weighting, the system can accurately distinguish the positions of key points and maintain stable and continuous skeletal response output even in complex backgrounds, occlusion interference or when key points are close to each other, significantly improving the accuracy of subsequent behavior modeling and motion prediction.

[0077] Furthermore, the method of obtaining the spatial vector of the pet's current gaze direction through a convolutional neural network includes the following steps:

[0078] Constructing a gaze direction training data set, the data set comprising a plurality of pet head images and corresponding gaze direction space vectors;

[0079] Taking the pet head image as input, the convolutional neural network is used to extract image features and obtain the image feature vector representing the facial orientation features;

[0080] Based on the image feature vector, a regression structure is used to perform regression prediction on the gaze direction space vector, wherein the regression prediction takes the gaze direction space vector corresponding to the gaze direction training data set as a supervision target, and uses the mean square error as a loss function to optimize the network parameters;

[0081] Based on the trained convolutional neural network model, the pet's head area image in the current frame is received as input during runtime, and the pet's current gaze direction spatial vector is output.

[0082] In some embodiments, in this embodiment, in order to predict the next movement intention of the pet in advance, the system constructs a gaze direction prediction model based on a convolutional neural network, which aims to extract the orientation features in the pet's head image and map it into a gaze vector in three-dimensional space, thereby constituting a key parameter in the behavior trend modeling. This module not only has the ability to express image features efficiently, but also converts the two-dimensional facial orientation information into a measurable spatial target prediction factor through a supervised regression training mechanism, breaking through the single-dimensional limitation of "judging behavior only by posture" in the prior art, and realizing the expansion of the spatial intention perception dimension. First, in the model training stage, a structured gaze direction training data set is constructed. The data set consists of a large number of pet head images and their corresponding three-dimensional gaze direction vectors, where the gaze direction vector is obtained by manual annotation or sensor-assisted acquisition, and is usually expressed as a unit space vector to indicate the main gaze direction of the pet at that time. In terms of sample coverage, the data set contains images of multiple species, multiple postures, and different lighting and occlusion conditions to enhance the model's adaptability to complex environments. In the training stage, each head image is input into a convolutional neural network designed based on a residual structure for feature extraction. The front layer of the network mainly extracts local facial structural features, such as the edge contours and texture features of the eyes, nose bridge, and ear roots; the middle layer uses a medium-scale convolutional perceptron to integrate the overall orientation pattern of the entire face area; the back layer obtains a set of low-dimensional, distinguishable feature vectors through global average pooling and feature compression to characterize the facial orientation features corresponding to the image. Then, the image feature vector is sent to the regression structure module, and the continuous spatial prediction of the three-dimensional gaze direction vector is realized through a multi-layer perceptron. The prediction process does not output discrete category labels, but fits a three-dimensional vector to represent the direction unit vector of the pet's gaze line. During training, the real gaze direction in the training set is used as the supervision target, and the mean square error is used as the loss function for optimization to ensure that the spatial angle between the predicted vector and the target direction is as close as possible, thereby enhancing the prediction stability and accuracy. After the system completes training, the model is deployed on the inference end. In the actual operation process, the system captures the head image area of ​​the pet's current frame from the RGB video frame as input, calls the trained convolutional neural network model in real time, and outputs the pet's gaze direction vector in the current frame. This vector will be input into the subsequent time series modeling module together with the skeleton key point sequence output by the posture estimation module to infer the moving target area and spatial approach intention.

[0083] Furthermore, the predicting of the pet movement target area through a time series model based on the skeleton point sequence and the spatial vector of the gaze direction includes the following steps:

[0084] Based on the skeleton point sequence output by the posture estimation network and the gaze direction space vector output by the convolutional neural network, a behavior state vector is constructed in each frame image, wherein the behavior state vector includes the combination result of the spatial coordinates of each key point in the current frame and the gaze direction unit vector;

[0085] Arrange the behavior state vectors in chronological order to form an input sequence, and input it into a Transformer prediction model, wherein the prediction model includes an encoder submodule and a decoder submodule;

[0086] The encoder submodule adopts a multi-head self-attention mechanism to model the trajectory changes and gaze direction trends of key points in the input sequence, and outputs a global time-dependent encoding representation;

[0087] The decoder submodule predicts the spatial motion trend of the pet in a few future time steps through a position-aware decoding network based on the global time-dependent coding representation and the current position embedding information, and outputs the corresponding target position prediction sequence;

[0088] Based on the spatial similarity matching between the target position prediction sequence and the environmental object positions marked in the three-dimensional semantic map, the approach area of ​​the pet's movement path is determined, and the semantic position label of the approach area is output as the movement target area prediction result.

[0089] In some embodiments, first, based on the skeleton point sequence output by the front-end posture estimation network and the spatial vector output by the gaze direction extraction module, a set of behavior state vectors in a unified format is constructed in each frame of the image. The state vector is composed of the three-dimensional spatial coordinates of each skeleton key point in the current frame and the gaze direction unit vector of the corresponding frame, forming a complete "structure-orientation" description unit. This data structure not only reflects the static structural distribution of the pet's current posture in space, but also introduces the dynamic prior of gaze intention, so that the model can capture the two information dimensions of "action basis" and "target attention trend" during the training process. Subsequently, the system arranges this series of state vectors in chronological order to form a behavior sequence of fixed window length and inputs it into the time series prediction model built based on the Transformer structure. Unlike traditional recurrent neural networks, Transformer effectively models the global dependency between different time steps in the input sequence by introducing encoder and decoder submodules, as well as a multi-head self-attention mechanism, avoiding the gradient attenuation and long-term dependency problems in time series modeling, and is particularly suitable for modeling the interaction between key point trajectories and gaze trends under the continuous behavior of pets. In the encoder module, the multi-head self-attention mechanism realizes the temporal collaborative modeling of "posture change-gaze offset" by calculating the feature relationship between any two time frames. For example, when the pet moves forward for several consecutive frames while looking at a certain area, the attention mechanism can automatically identify the "trend weight" of this behavior in the sequence and embed it as a high-weight focus point in the global time-dependent expression. This encoding representation not only contains the short-term characteristics of local motion, but also reflects the evolution law of behavioral trends in the time dimension, providing a well-structured temporal latent representation for subsequent predictions. In the decoder module, the model predicts the pet's motion trajectory in the next few time steps by introducing a position embedding mechanism and a spatial perception decoding structure. Different from directly outputting a single target position, this decoder combines the temporal position offset with the current skeleton orientation to infer the direction and magnitude of the position change that may occur in the next step of the pet, and then generates a continuous target position prediction sequence. Each prediction point is located in three-dimensional space, constituting a multi-step fitting result of the future motion path. Finally, the system matches the predicted sequence with the three-dimensional semantic map output by the semantic segmentation module for spatial position. Specifically, the spatial coordinates in the predicted path are similar to the semantic target positions such as high platforms, fragile objects, and foreign objects on the ground in the three-dimensional semantic map to determine whether the pet has a tendency to approach these dangerous areas. If the matching distance is lower than the preset threshold, the end of the path is marked as a potential dangerous area, and its corresponding semantic label (such as "balcony guardrail", "glass vase", "foreign object on the ground", etc.) is output as the prediction result of the moving target area.Taking the actual situation as an example, if the system continuously observes that the pet keeps the direction of its head gaze on a certain high platform area, and at the same time the skeletal structure shows features such as the body leaning back and the hind legs shaking to accumulate strength, and the position gradually approaches the target area, the Transformer model will learn the "jump preparation" behavior pattern and direct its next predicted path to the high area. After matching, it is concluded that the position is a dangerous platform marked in the three-dimensional semantic map, thereby judging that the pet's behavior has the risk intention of "high jump preparation". This embodiment is different from the existing solution of judging danger through static coordinates in terms of method layout. It jointly models the structural action, orientation trend and time relationship, and introduces a spatial alignment mechanism based on semantic structure. The algorithm uses the Transformer's ability to perform time series modeling, so that the system no longer relies on manually defined behavior rules or static threshold judgments, but has deep semantic reasoning and continuous target prediction capabilities to achieve early warning of dangerous pet behaviors.

[0090] Furthermore, the multi-head self-attention mechanism is used to model the key point trajectory changes and gaze direction trends in the input sequence, including the following steps:

[0091] Performing feature embedding mapping on the sequence of behavior state vectors to generate corresponding query vector sequence, key vector sequence and value vector sequence respectively;

[0092] Based on the query vector sequence and the key vector sequence, a multi-head self-attention mechanism is used to calculate the attention distribution weights between each time step in the behavior state vector sequence;

[0093] Based on the attention distribution weight and value vector sequence, the temporal behavior feature representation of each attention head is calculated to obtain a multi-head temporal dependency feature set;

[0094] The multi-head temporal dependency feature sets are concatenated and fused, and converted into a global temporal dependency encoding representation through an output mapping layer.

[0095] In some embodiments, first, the input behavior state sequence is composed of state vectors at several moments, each vector containing the coordinates of the skeleton key points and the gaze direction vector of the current frame, which is used to characterize the instantaneous state of the pet in terms of spatial structure and intention orientation. In order to adapt to the computational requirements of the multi-head attention mechanism, the system constructs three independent but dimensional vector sequences through linear mapping, corresponding to the query vector, key vector and value vector respectively. This linear transformation can be regarded as a feature reconstruction process, which is used to project the original input to the subspaces that multiple attention heads focus on, ensuring that subsequent calculations are distinguishable in different semantic dimensions. Subsequently, the system calculates the correlation score between the query and the key in each attention head, that is, the similarity relationship between the current time step and other time steps is measured by dot product. This process constructs a set of time-based attention distribution matrices for weighted encoding of the importance of each position in the value vector sequence. In the pet behavior recognition scenario, this attention mechanism can effectively capture pattern associations across time frames such as "continuously looking at a target and gradually approaching" and "jumping instantly after continuous posture adjustment", significantly enhancing the perception of time-dependent structures. Afterwards, the system applies the attention weights to the value vector sequence to generate the context feature representations corresponding to each attention head. Multiple attention heads model time series data from different dimensions, and have the ability to collaboratively model short-term action changes (such as sudden head tilt and turn) and medium- and long-term behavioral trends (such as continuous approach to the target object), and can integrate the skeletal displacement trend and gaze direction drift dynamics into a set of time feature vectors with hierarchical semantics. Finally, the system concatenates the context representations of all attention heads, completes the unified dimension projection through the output mapping layer, and outputs the global time-dependent encoding representation. As the core output of the Transformer encoder, this representation not only retains the joint dynamic structure of action and gaze changes in the input sequence, but also has strong expression ability and scalability, which can provide high-quality semantic embedding support for subsequent decoding prediction modules. In terms of algorithm architecture, this module breaks through the performance bottleneck of traditional RNN or single-head attention mechanism in long sequence modeling, and improves the model's ability to model complex behavioral trends of pets by explicitly constructing a multi-dimensional coupled expression of time-structure-intention. Especially in the early stage of continuous state change, the behavioral turning trend can be perceived through time correlation modeling.

[0096] Furthermore, the formula of the position-aware decoding network is as follows:

[0097] ;

[0098] in, For the The target position prediction vector of the frame; is the current time step; is the number of steps predicted into the future; s is the historical time step index; The current prediction step The attention weight for the s-th frame in history; The global temporal dependency encoding of the sth frame output by the encoder; is the position embedding vector of the sth frame; is the weight matrix of the decoding linear map; The embedded state vector of the current frame; It is the projection matrix from the current frame state to the position vector.

[0099] Specifically, when predicting a future time step, the system first models the dynamic correlation between the current prediction step and the historical moment. By calculating the similarity between the current position embedding state vector and the state of each historical time step, the attention distribution weight of the prediction step for all historical moments is obtained. The above weights are used to weight the time-dependent features of the encoder output to form a dynamic representation of structural behavior. At the same time, in order to enhance the semantic perception of spatial position, the system incorporates the spatial embedding vectors of each historical moment into the weighting process, so that the decoder output not only reflects the trend of behavioral changes in time, but also reflects the direction of spatial semantic convergence. The fusion result is then input into the position projection network in the decoder, mapped to the three-dimensional space coordinate domain through linear transformation, and output as the spatial target position vector of the current prediction time step. The system iterates each prediction time step in this way, and finally forms a complete prediction path sequence. It should be noted that in order to realize time-space joint modeling, the system introduces the current position embedding mechanism in the process of predicting future behavior trends, which is used to convert the physical spatial position of the pet in the historical time step into a vector form with semantic representation capabilities, that is, the position embedding vector. Specifically, in the process of constructing a three-dimensional semantic map, the system has structuredly encoded the position of the pet's bone center point in the three-dimensional coordinate system in each frame of the image. The code not only contains the coordinate value, but also has semantic labels or category information based on the semantic area where it is located (for example, near a balcony, a high platform, near a vase or a table corner, etc.). In order to introduce this discrete spatial position information into the decoding network, the system vectorized and embedded the three-dimensional coordinates and their semantic categories. The vectorization process can be combined in two ways: one is to use a coordinate position encoding function, such as embedding the spatial coordinates using sine and cosine periodic functions after normalizing them, so as to obtain a position encoding with time series modeling capabilities; the other is to use a trainable embedding matrix to perform a lookup table mapping on the semantic label information, so that similar spatial regions have close vector representations. After these two methods are combined, they are mapped to the same feature dimension as the encoder output through a linear transformation layer to form a position embedding vector in a unified format.

[0100] Furthermore, the attention diversion strategy includes sound broadcasting, pet snack delivery, and pet toy movement.

[0101] In this embodiment, in order to achieve effective intervention and behavior guidance for pet dangerous behaviors, the system designs a multimodal attention diversion strategy, which is used to interrupt its behavior path and change its attention focus through active guidance when a high-risk intention state (such as jumping preparation, ingestion tendency or approaching fragile objects) is detected, so as to prevent the occurrence of dangerous behaviors. This strategy includes three specific forms: sound broadcasting, pet snack delivery and pet toy movement, which can be triggered independently or jointly called according to the strategy priority. When the system infers that the pet is about to perform a potentially dangerous action, such as jumping to the window edge, trying to pick up foreign objects on the ground, etc. through the behavior recognition module and the position perception decoding network, its behavior state is judged to have a high-risk tendency. On this basis, the system first matches the corresponding attention diversion strategy according to the category of the behavior intention and the target area attribute of the predicted path. For example, for the intention of jumping near the window edge, the system prefers the sound source diversion scheme to interfere with its concentration; if it is an ingestion behavior, it is more inclined to change its movement direction through an inducement guidance method. The sound broadcasting module is completed by intelligent audio equipment deployed in various places in the room, which has the ability of controllable direction and adjustable timbre. In actual applications, the system will automatically select the most appropriate speaker to play the specified voice command or a prompt tone of a specific frequency based on the pet's current position and predicted path, and guide the pet to divert its attention or terminate the current behavior through the change of the spatial sound field. For example, when the pet is accumulating strength to jump to a high area, the system can play the owner's call behind it, forcing it to turn around and pause its action. For the snack delivery method, the system pre-arranges intelligent feeding devices at several locations indoors, containing an appropriate amount of snacks that the pet likes. When the predicted path approaches a dangerous area, the system can select a device in a relatively safe direction to initiate the delivery action, and the pet can be visually captured, thereby guiding it to divert its movement path in the direction of the snack, reducing the possibility of it continuing to approach the high-risk area. This method performs well in medium and short-distance guidance and is suitable for interrupting behaviors such as low-speed approach to foreign objects. In addition, in the toy movement guidance strategy, the system links electric pet toy modules installed on the ground or in the corners, such as balls or mobile track devices. Based on the inference results of the pet's current position, gaze direction, and skeletal posture, the system can initiate a toy start command to move the pet to a non-dangerous area at a certain speed to construct a new attention stimulus target and attract the pet to change its current path. Different from the existing solutions based on passive alarms or video notifications, the attention diversion mechanism proposed by this system not only realizes the prediction and identification of risky behaviors, but also completes targeted behavior induction based on the prediction results, effectively preventing dangerous behaviors of pets.

[0102] The above implementation modes are merely descriptions of the preferred implementation modes of the present invention, and are not intended to limit the scope of the present invention. Without departing from the design spirit of the present invention, various modifications and improvements made to the technical solutions of the present invention by ordinary engineering and technical personnel in the field shall fall within the protection scope determined by the claims of the present invention.

Claims

1. A pet dangerous behavior monitoring system based on visual neural network, characterized in that: include: Data acquisition module, used to obtain indoor area image data and indoor space point cloud data; The semantic segmentation module is used to perform semantic recognition of indoor structures and objects based on image data using a visual neural network. It also extracts spatial height information, edge structures, and object contour features based on point cloud data, and fuses image semantic results with point cloud geometry data to generate a three-dimensional semantic map. The behavior parameter construction module is used to construct a pet behavior dataset based on indoor area image data and indoor space point cloud data. Based on the pet behavior dataset, the pet skeleton point sequence is extracted through the posture estimation network, and the spatial vector of the pet's current gaze direction is obtained through the convolutional neural network. A behavior recognition module, for predicting the pet's target movement area through a time series model based on the skeleton point sequence and the spatial vector of the gaze direction, and calculating the pet's dangerous intention state based on the matching relationship between the target movement area and the position of the environmental object in the three-dimensional semantic map, wherein the pet's dangerous intention state includes preparation for high-altitude jumping, swallowing foreign objects by mistake, and approaching fragile objects; The pet guidance module is used to send alarms to the mobile terminal according to the pet's dangerous intention status and implement attention diversion strategies.

2. A pet dangerous behavior monitoring system based on visual neural network according to claim 1, characterized in that: The acquisition of indoor area image data and indoor space point cloud data comprises the following steps: By setting up RGB cameras at several positions indoors to collect video image frames of the pet activity area at a fixed frequency, indoor area image data is obtained; The spatial point cloud data of the indoor area is collected by LiDAR.

3. A pet dangerous behavior monitoring system based on visual neural network according to claim 1, characterized in that: The semantic segmentation module is used to perform the following steps: Based on the indoor area image data, a semantic segmentation network is used to process the image data and extract the image semantic feature map containing the object semantic category and two-dimensional spatial position; Based on the image semantic feature map, combined with the time-synchronized spatial point cloud data, the image features are mapped to the point cloud space coordinate system through deep projection to obtain a projected semantic feature map; Constructing a joint feature representation based on the spatial height information, edge structure and geometric contour features extracted from the projected semantic feature map and the spatial point cloud data; The joint feature representation is input into a three-dimensional semantic fusion network to generate a three-dimensional semantic map, and semantic annotations are performed on high platforms, easily ingested foreign objects, and fragile objects in the indoor environment in the three-dimensional semantic map.

4. A pet dangerous behavior monitoring system based on visual neural network according to claim 1, characterized in that: The posture estimation network is constructed by the following steps: Based on the indoor area image data, a multi-scale convolutional network is used to process the image sequence and extract the image feature map sequence containing the pet's contour and structure information; Based on the image feature map sequence, a key point detection subnetwork is used to perform key point heat map regression processing on predefined key parts in each frame of the image, and a heat map sequence representing the probability distribution of key point responses is output; Based on the heat map sequence, time series modeling is used to model the response trajectory of key points between consecutive image frames, and the trajectory fitting and spatial correction of the key point positions are performed in combination with the time sequence between frames to output the pet bone point sequence.

5. A pet dangerous behavior monitoring system based on visual neural network according to claim 4, characterized in that: The formula of the key point detection subnetwork is as follows: ; in, It is the output key point heat map set; is the Sigmoid activation function; is the weight of the fully connected regression layer; is the bias term; No. Feature map extracted by multi-scale convolution; The feature map The corresponding residual correction term of ; The weight coefficient generated by the channel attention mechanism satisfies ; is the number of multi-scale branches.

6. A pet dangerous behavior monitoring system based on visual neural network according to claim 1, characterized in that: The method of obtaining the spatial vector of the pet's current gaze direction through a convolutional neural network comprises the following steps: Constructing a gaze direction training data set, the data set comprising a plurality of pet head images and corresponding gaze direction space vectors; Taking the pet head image as input, the convolutional neural network is used to extract image features and obtain the image feature vector representing the facial orientation features; Based on the image feature vector, a regression structure is used to perform regression prediction on the gaze direction space vector, wherein the regression prediction takes the gaze direction space vector corresponding to the gaze direction training data set as a supervision target, and uses the mean square error as a loss function to optimize the network parameters; Based on the trained convolutional neural network model, the pet's head area image in the current frame is received as input during runtime, and the pet's current gaze direction spatial vector is output.

7. A pet dangerous behavior monitoring system based on visual neural network according to claim 6, characterized in that: The method of predicting the pet movement target area through a time series model based on the skeleton point sequence and the spatial vector of the gaze direction comprises the following steps: Based on the skeleton point sequence output by the posture estimation network and the gaze direction space vector output by the convolutional neural network, a behavior state vector is constructed in each frame image, wherein the behavior state vector includes the combination result of the spatial coordinates of each key point in the current frame and the gaze direction unit vector; Arrange the behavior state vectors in chronological order to form an input sequence, and input it into a Transformer prediction model, wherein the prediction model includes an encoder submodule and a decoder submodule; The encoder submodule adopts a multi-head self-attention mechanism to model the trajectory changes and gaze direction trends of key points in the input sequence, and outputs a global time-dependent encoding representation; The decoder submodule predicts the spatial motion trend of the pet in a few future time steps through a position-aware decoding network based on the global time-dependent coding representation and the current position embedding information, and outputs the corresponding target position prediction sequence; Based on the spatial similarity matching between the target position prediction sequence and the environmental object positions marked in the three-dimensional semantic map, the approach area of ​​the pet's movement path is determined, and the semantic position label of the approach area is output as the movement target area prediction result.

8. A pet dangerous behavior monitoring system based on visual neural network according to claim 7, characterized in that: The method of using a multi-head self-attention mechanism to model the trajectory changes and gaze direction trends of key points in the input sequence includes the following steps: Performing feature embedding mapping on the sequence of behavior state vectors to generate corresponding query vector sequence, key vector sequence and value vector sequence respectively; Based on the query vector sequence and the key vector sequence, a multi-head self-attention mechanism is used to calculate the attention distribution weights between each time step in the behavior state vector sequence; Based on the attention distribution weight and value vector sequence, the temporal behavior feature representation of each attention head is calculated to obtain a multi-head temporal dependency feature set; The multi-head temporal dependency feature sets are concatenated and fused, and converted into a global temporal dependency encoding representation through an output mapping layer.

9. A pet dangerous behavior monitoring system based on visual neural network according to claim 8, characterized in that: The formula of the position-aware decoding network is as follows: ; in, For the The target position prediction vector of the frame; is the current time step; is the number of steps predicted into the future; s is the historical time step index; The current prediction step The attention weight for the s-th frame in history; The global temporal dependency encoding of the sth frame output by the encoder; is the position embedding vector of the sth frame; is the weight matrix of the decoding linear map; The embedded state vector of the current frame; It is the projection matrix from the current frame state to the position vector.

10. The pet dangerous behavior monitoring system based on visual neural network according to claim 1, characterized in that: The attention diversion strategies include sound broadcasting, pet treat delivery, and pet toy movement.

Citation Information

Patent Citations

  • Pet behavior analysis system based on AI vision

    CN117894078A

  • Systems and methods for deep localization and segmentation with a 3D semantic map

    US20200364554A1

Cited By

  • Automatic track and field action recognition and capture method based on image data

    CN120673479A

  • Athletic action automatic recognition and capture method based on image data

    CN120673479B

  • Storage material monitoring method and system for guiding attention based on multi-view space

    CN121121629A

  • Warehouse material monitoring method and system based on multi-view spatial attention guidance

    CN121121629B

  • Bemisia tabaci feeding trend analysis method and system based on artificial intelligence

    CN121982709A