Remote snack feeding method, device and equipment based on pet camera

Through the pet camera, images are collected and combined with multi-level prediction networks and behavior recognition algorithms, the problems of low detection accuracy and poor environmental adaptability in the existing system are solved, stable and accurate analysis of pet behavior status and accurate judgment of feeding needs are achieved, and the accuracy and efficiency of snack delivery are improved.

CN120108037AInactive Publication Date: 2025-06-06SHENZHEN ANKED SHITONG ELECTRONICS CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510237119.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-01
Publication Date
2025-06-06
Estimated Expiration
Not applicable · inactive patent

Smart Images

  • Figure CN120108037A_ABST
    Figure CN120108037A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of pet cameras, and discloses a remote snack feeding method, device and equipment based on a pet camera, and the method comprises the steps: collecting a target pet image through the pet camera, and carrying out the spatial positioning and feature recognition enhancement of the target pet image, and obtaining pet spatial position information and pet behavior features; according to the pet spatial position information and the pet behavior characteristics, executing activity-based behavior identification to obtain pet feeding demand confidence; inputting the pet spatial position information, the pet behavior characteristics and the feeding demand confidence into a multi-level prediction network for processing to obtain a pet detection frame, a behavior category and a feeding judgment result; and executing pet behavior tracking and feeding analysis based on the pet detection frame, the behavior category and the feeding judgment result, and outputting a snack feeding control signal. According to the invention, stable and accurate analysis of the behavior state of the pet is realized, and accurate judgment of the feeding demand of the pet is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of pet cameras, and in particular to a remote snack delivery method, device and equipment based on a pet camera. Background Art

[0002] Most existing automatic feeders adopt a fixed-time delivery mode and are unable to intelligently deliver food based on the actual behavior and needs of pets, resulting in low delivery efficiency and failure to meet the personalized needs of pets.

[0003] Existing pet camera snack delivery systems generally have technical defects such as low detection accuracy, slow response speed, and poor environmental adaptability. Especially in complex home environments, factors such as pet occlusion, light changes, and rapid pet movement often cause the delivery system to be unable to accurately identify pet behavior and needs, resulting in problems such as missed deliveries, misdeliveries, and inaccurate delivery locations. At the same time, due to the lack of in-depth understanding and precise analysis of pet behavior, it is difficult to determine the pet's true feeding needs. The determination of feeding timing often relies on preset fixed rules, lacking flexibility and individualized adaptability. Summary of the invention

[0004] The present invention provides a remote snack delivery method, device and equipment based on a pet camera, which realizes stable and accurate analysis of the pet's behavioral state and accurate judgment of the pet's feeding needs.

[0005] In a first aspect, the present invention provides a remote snack delivery method based on a pet camera, the remote snack delivery method based on a pet camera comprising:

[0006] The target pet image is captured by a pet camera, and the target pet image is spatially positioned and feature-recognized and enhanced to obtain the pet spatial position information and pet behavior characteristics;

[0007] Performing activity-based behavior recognition according to the pet spatial position information and the pet behavior characteristics to obtain a pet feeding demand confidence level;

[0008] Input the pet spatial position information, the pet behavior characteristics and the feeding demand confidence into a multi-level prediction network for processing to obtain a pet detection frame, a behavior category and a feeding judgment result;

[0009] Based on the pet detection frame, the behavior category and the feeding judgment result, pet behavior tracking and feeding analysis are performed, and a snack delivery control signal is output.

[0010] In a second aspect, the present invention provides a remote snack delivery device based on a pet camera, the remote snack delivery device based on a pet camera comprising:

[0011] A collection module, used to collect a target pet image through a pet camera, and perform spatial positioning and feature recognition enhancement on the target pet image to obtain the pet's spatial position information and pet behavior characteristics;

[0012] A behavior recognition module, used to perform activity-based behavior recognition according to the pet's spatial position information and the pet's behavior characteristics, and obtain a pet's feeding demand confidence level;

[0013] A processing module, used to input the pet spatial position information, the pet behavior characteristics and the feeding demand confidence into a multi-level prediction network for processing to obtain a pet detection frame, a behavior category and a feeding judgment result;

[0014] A feeding analysis module is used to perform pet behavior tracking and feeding analysis based on the pet detection frame, the behavior category and the feeding judgment result, and output a snack delivery control signal.

[0015] The third aspect of the present invention provides a computer device, comprising: a memory and at least one processor, wherein the memory stores instructions; the at least one processor calls the instructions in the memory so that the computer device executes the above-mentioned remote snack delivery method based on a pet camera.

[0016] In the technical solution provided by the present invention, by improving the multi-scale cascade GhostNet feature extraction network, setting up multi-scale convolution modules with three different sizes of convolution kernels, and combining channel attention and spatial attention calculations, the feature extraction capability of pet targets is significantly enhanced. The spatial positioning and feature recognition enhancement module realizes centimeter-level precise positioning of the pet's spatial position through the comprehensive application of spatial attention, feature pyramid transformation and position-sensitive convolutional network, and extracts pet behavior characteristics with global context information through channel interaction enhancement algorithm and non-local relationship modeling algorithm, which greatly improves the accuracy of pet behavior recognition in complex home scenes. The activity-based behavior recognition algorithm does not require a predefined activity threshold. It calculates pet activity evaluation indicators, constructs activity time series curves and dynamic threshold judgments, and combines time window filtering technology to achieve accurate judgment of pet feeding needs. The multi-level prediction network conducts a comprehensive analysis of the pet's detection frame, behavior category and feeding needs through feature splicing, multi-scale feature conversion and parallel calculation. It combines time series smoothing filters and Bayesian update mechanisms to achieve stable and accurate analysis of the pet's behavioral state. Through deep appearance description, behavioral state transfer model and feature compensation technology, it significantly improves the pet tracking performance in low-light environments. The self-adjusting feeding analysis algorithm achieves high-precision snack delivery through trajectory extrapolation calculation, feeding possibility scoring, reinforcement learning model and feasibility verification, combined with precise control of mechanical execution instruction sequences, significantly improving the intelligence level and user experience of remote pet care. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying creative work.

[0018] Figure 1 Schematic diagram of the steps of a remote snack delivery method based on a pet camera in an embodiment of the present invention;

[0019] Figure 2 It is a structural schematic diagram of a remote snack delivery device based on a pet camera in an embodiment of the present invention;

[0020] Figure 3 It is a schematic block diagram of the structure of a computer device in an embodiment of the present invention. DETAILED DESCRIPTION

[0021] Embodiments of the present invention provide a remote snack delivery method, device and equipment based on a pet camera. The terms "first", "second", "third", "fourth", etc. (if any) in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "including" or "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0022] For ease of understanding, the specific process of the embodiment of the present invention is described below. Figure 1 In one embodiment of the present invention, a remote snack delivery method based on a pet camera includes:

[0023] Step S1, collecting a target pet image through a pet camera, and performing spatial positioning and feature recognition enhancement on the target pet image to obtain the pet's spatial position information and pet behavior characteristics;

[0024] It is understandable that the execution subject of the present invention may be a remote snack delivery device based on a pet camera, or may be a terminal or a server, which is not specifically limited here. The embodiment of the present invention is described by taking a server as the execution subject as an example.

[0025] Specifically, the target pet is captured in real time through a pet camera to obtain the original image data. The original image is adaptively adjusted in brightness, and an image enhancement method based on histogram equalization is used to maintain good contrast and detail information in the image brightness under different environmental conditions, so as to obtain the preprocessed pet image. Multi-scale feature extraction is performed on the preprocessed pet image. A multi-scale convolution module containing three different sizes of convolution kernels is constructed, and convolution operations are performed on the image using 1×1, 3×3, and 5×5 convolution kernels respectively. Among them, the 1×1 convolution kernel is used for linear transformation of information between channels, so that the features between different channels are fused and reduced in dimension, while reducing the amount of calculation; the 3×3 convolution kernel extracts local detail information and captures the local texture features of the pet's body, such as the outline of the ears, paws or tail; and the 5×5 convolution kernel focuses on a wider range of contextual information, which helps to extract global structural information, such as the overall shape of the pet and its relative position in the image. The three convolution kernels of different sizes work together to enable the network to extract features from different receptive field scales, enhance the ability to depict pet targets, and output pet feature maps of different scales. Channel attention and spatial attention calculations are performed on pet feature maps of different scales. Channel attention calculation uses global average pooling or global maximum pooling to obtain the importance information of each channel, and calculates feature weights through a fully connected network to generate a channel weight map. The role of channel attention is to strengthen important semantic information and suppress irrelevant features, so that the model pays more attention to features that are valuable for pet behavior recognition. Spatial attention calculation performs maximum pooling and average pooling operations on feature maps of different scales, and after splicing them in the channel dimension, a small convolution layer is used to calculate the pixel-level importance distribution to generate a spatial attention weight map. Through the dual weighted processing of channel attention and spatial attention, the pet target area is highlighted, while background interference is reduced, so that the feature extraction process is more focused on key parts, such as the pet's eyes, mouth, ears and other areas. The weighted pet feature map is fused at multiple levels to fully retain the information expression at different levels. A bottom-up cascade connection method is adopted, that is, feature maps of different scales are cascaded in order from shallow to deep. Shallow features contain rich low-level visual information such as edges and textures, while deep features focus more on global semantic information, such as the overall shape and behavior category of pets. Through cascade connection, these different levels of information are fused to form a multi-level feature representation of pets. Based on the multi-level feature representation, spatial positioning and feature recognition enhancement are performed to obtain the spatial position information and behavioral characteristics of pets. Spatial positioning is achieved by using target detection or key point detection methods, including target detection networks based on FasterR-CNN, YOLO or Transformer architectures, and the position coordinates of the pet in the image are obtained by regressing the detection box.At the same time, in terms of feature recognition enhancement, deep neural networks are used to analyze the pet's posture and dynamic features, and extract the pet's behavioral change characteristics at different times.

[0026] The multi-level feature representation is input into the spatial attention module for adaptive spatial weight calculation. The global pooling mechanism is combined with the self-attention mechanism to weight the feature responses of different regions, so that the network pays more attention to the spatial position of the pet and reduces the influence of background interference. The feature map is subjected to maximum pooling and average pooling operations in the spatial dimension and input into a lightweight convolutional neural network to calculate the importance distribution of each position point and generate a spatial attention weight map. The weight map is multiplied point by point with the original multi-level feature representation to obtain a spatial weighted feature map. The spatial weighted feature map is subjected to feature transformation to obtain a multi-scale spatial feature representation. The pyramid feature transformation method is adopted to combine convolution kernels of different step sizes with pooling operations of different scales so that the same feature map can simultaneously express local detail information and global structural information. For example, by using 1×1, 3×3 and 5×5 convolution kernels to perform convolution transformation on the spatial weighted feature map, feature representations under different receptive fields are obtained, so that the model can take into account both local texture features and overall spatial relationships during the learning process. At the same time, in order to optimize the feature distribution, feature normalization strategies such as Batch Normalization are adopted to ensure that multi-scale spatial features can maintain numerical stability at different resolutions and avoid the problem of gradient vanishing or gradient exploding. The multi-scale spatial feature representation is input into the position-sensitive convolutional network for spatial coordinate encoding to obtain the pet's spatial position information. The position-sensitive convolutional network adopts a method based on the coordinate perception mechanism, that is, additional position information encoding is introduced on the basis of conventional convolution operations, so that each feature point carries the information of its relative spatial position. A coordinate encoding matrix is ​​attached to each channel of the feature map. The matrix consists of normalized information of horizontal and vertical coordinates and is nonlinearly transformed through a lightweight fully connected network, so that each feature point contains local visual features and carries its absolute position in the image, thereby improving the accuracy of spatial positioning. After obtaining the pet's spatial position information, channel interaction enhancement is performed on the multi-scale spatial feature representation to obtain cross-channel fused behavioral representation features. Channel attention mechanisms, such as the Squeeze-and-Excitation (SE) module, are used to enable key information to be shared between different channels. For example, by globally pooling each channel and calculating the dependencies between channels, the weights of each channel are adjusted so that the network pays more attention to features related to pet behavior, such as motion trajectory, posture changes, and local actions. The cross-channel fused behavior representation features are cascaded with the pet's spatial position information to form a position-aware pet behavior feature map. The behavior representation features are fused with the spatial position information in the channel dimension by channel-by-channel splicing, and feature compression is performed through 1×1 convolution to reduce redundant information while maintaining a close association between spatial information and behavior features.Through this step, the network combines location information and behavioral features during the learning process to accurately determine the behavior category of the pet in a specific area, such as standing, running, rolling, or stillness. Non-local relationship modeling is performed on the location-aware pet behavior feature map to obtain higher-level pet behavior features. Non-local relationship modeling uses global feature similarity calculation to make the network focus on the information of the local area and combine long-distance feature relationships to improve the understanding of complex behavior patterns. A non-local attention mechanism is used, that is, the similarity of each pixel in the entire feature map is calculated with other pixels, and features are weighted according to the similarity to capture long-range dependencies.

[0027] Step S2: performing activity-based behavior recognition according to the pet's spatial position information and pet's behavior characteristics to obtain the pet's feeding demand confidence;

[0028] Specifically, a series of activity-related indicators are calculated based on the pet's spatial position information and behavioral characteristics to evaluate the pet's current movement state and behavior changes. The activity level of the pet is comprehensively evaluated by analyzing the pet's movement frequency, movement amplitude, behavior change rate, and the frequency of specific behaviors. Among them, the movement frequency is obtained by calculating the number of times the pet appears in the camera's field of view or the number of times its position changes per unit time, while the movement amplitude is obtained by averaging the displacement vector between consecutive frames to measure the size of the pet's movement range. The behavior change rate reflects the frequency of switching between different behaviors of the pet within a certain period of time, such as the number of times from standing to walking, from walking to running, and multiple posture changes in a short period of time. At the same time, the frequency of occurrence of specific behaviors is counted, such as the frequency of occurrence of food-related behaviors such as wagging the tail, turning in circles, and looking up at the camera. These indicators together constitute the evaluation system of pet activity. The continuously collected pet activity evaluation indicators are subjected to time series sampling and sliding window processing to construct a time series curve of pet activity. The role of the sliding window is to smooth short-term fluctuations while ensuring temporal continuity, so as to more stably extract trend information of pet behavior changes. By performing weighted averaging or exponential smoothing on the activity data in the sliding window, the impact of sudden outliers on the overall evaluation is reduced, making the time series curve more stable. Perform statistical distribution analysis based on the pet activity time series curve to calculate the dynamic threshold of pet activity. The calculation of the dynamic threshold is based on the mean, standard deviation of historical data and the activity change trend in a specific time period. The upper and lower limits of pet activity are determined by Gaussian distribution fitting or quantile analysis to distinguish normal activity from abnormally active or abnormally depressed states. The dynamic threshold is updated over time to adapt to the changes in the activity patterns of pets in different time periods and environments, ensuring that the judgment of feeding needs is more personalized and accurate. Make rule judgments on pet activity evaluation indicators and pet activity dynamic thresholds. Match the activity evaluation indicators with the dynamic thresholds according to rules. If the current activity exceeds a certain threshold or the frequency of certain specific behaviors reaches the preset standard, the initial feeding demand score is calculated. The initial feeding demand score is filtered by time window to obtain a stable feeding demand score. The time window filter uses a sliding mean filter or an exponentially weighted average filter to reduce misjudgments caused by instantaneous fluctuations, making the feeding demand score more stable and reflecting the long-term behavior trend of the pet. A comprehensive score calculation is performed based on the stable feeding demand score, the time interval from the last feeding, and the daily feeding pattern of the pet to obtain the confidence level of the pet's feeding demand.

[0029] Step S3, inputting the pet's spatial position information, pet's behavior characteristics and feeding demand confidence into a multi-level prediction network for processing to obtain a pet detection frame, behavior category and feeding judgment result;

[0030] Specifically, the pet spatial location information, pet behavioral characteristics and feeding demand confidence are feature spliced ​​and normalized to construct a combined feature representation. The feature normalization method is used to keep the features from different data sources at a uniform scale to ensure that the neural network will not experience uneven gradient updates due to differences in feature ranges during training. Feature splicing uses channel expansion, that is, spatial location information, behavioral characteristics and feeding demand confidence are used as additional feature channels and fused with the target pet image to form an input data structure containing multi-dimensional information. The combined feature representation and the target pet image are input to the upsampling and downsampling modules for multi-scale feature conversion, so that the network can extract key features at different resolutions. In this process, the original image is downsampled, and deep features are extracted through multi-layer convolution and pooling operations, so that high-level features capture global semantic information, such as the overall behavior pattern of the pet, while shallow features retain more local detail information, such as the fine morphology of the ears, tail and paws. At the same time, the deep features are gradually restored through the upsampling module to ensure that the low-resolution features are fused with the high-resolution features to form a feature pyramid structure, so that the network can take into account both small-scale targets (such as local actions of pets) and large-scale targets (such as overall behavior). The prediction branch networks are connected at feature levels of different resolutions and parallel calculations are performed to obtain the pet detection box coordinates, behavior category probability distribution, and feeding judgment scores at each level. Each feature level performs target detection independently, that is, the pet's bounding box coordinates are predicted by regression, and the probability distribution of behavior categories is calculated using the classification network. And through an additional scoring network, the confidence of the feeding demand is predicted to determine whether the pet's current state meets the feeding conditions. In the fusion stage, the prediction results of each level are weighted and fused to improve the stability and accuracy of the overall prediction. In the process of fusing the prediction results, the attention weighting mechanism is used to weight the pet detection box coordinates, behavior category probability distribution, and feeding judgment scores at different levels. Combined with contextual information, through feature fusion strategies such as adaptive weighted fusion or feature splicing fusion, the information at different levels complements each other to ensure the stability of the final prediction results. After completing the weighted fusion, the pet detection frames in the prediction results are filtered for redundant frames to ensure that there will be no multiple duplicate frames in the final detection results. The non-maximum suppression method is used to remove detection frames with high overlap, that is, for the same target, if the IoU value of multiple detection frames exceeds a certain threshold, only the detection frame with the highest confidence is retained, and the remaining redundant frames are removed. This effectively avoids the problem of repeated detection, improves the accuracy of the detection frame, and ensures that subsequent behavior analysis is based on a unique and optimal target position. The target detection frame position, behavior category probability distribution, and feeding judgment score are input into the temporal smoothing filter and Bayesian update module to reduce the impact of burst noise in a short period of time, so that the system can respond more smoothly to the behavior changes of the pet.The Bayesian update module is used to dynamically optimize the feeding judgment, that is, based on the statistical characteristics of historical data, the calculation method of the confidence of feeding demand is continuously adjusted to improve the accuracy of the prediction. For example, if the system detects that the activity of the pet has continued to increase over the past period of time, and the previous feeding decision results show that the pet does need food in similar situations, the Bayesian update module dynamically adjusts the current feeding judgment score to improve the confidence of feeding, thereby optimizing the decision logic.

[0031] Step S4: perform pet behavior tracking and feeding analysis based on the pet detection frame, behavior category and feeding judgment result, and output a snack delivery control signal.

[0032] Specifically, the spatial position information provided by the detection frame is cascaded with the multi-level feature representation of the pet, so that the model uses local detail information (such as changes in parts such as ears and tails) and overall semantic information (such as body contours and posture states) to generate a deep appearance description of the pet target. A behavior state transition model is constructed based on the pet behavior category to predict the future movement trajectory and state changes of the pet. The behavior state transition model uses a Markov decision process or a hidden Markov model to learn the transition probability of the pet in different states through historical behavior sequences. The output of the behavior state transition model is input into the Kalman filter for motion prediction. The Kalman filter predicts the position and behavior state of the pet in the short future based on the current observation data and historical motion trends, reduces the tracking error caused by short-term occlusion or camera shaking, and ensures that the pet's trajectory prediction is more accurate and continuous. After completing the motion prediction, the pet's deep appearance description, position and behavior state prediction values, and feeding judgment results are comprehensively used to calculate the multidimensional matching cost matrix to measure the target association metric between different time steps. The matching cost matrix is ​​used to calculate the degree of match between each detected target in the current frame and the previous frame. The calculation method is based on metrics such as Euclidean distance, Mahalanobis distance or cosine similarity. After calculating the target association metric, target association optimization is performed to obtain the temporal correspondence of pet targets. The target association optimization uses the Hungarian algorithm or cascade matching method to maintain the consistency of target identity during multi-target tracking. Through the optimal cost matching method, the target in the current frame is uniquely associated with the target in the previous frame to ensure that the same pet maintains a stable identity throughout the tracking process. If a new target appears in a frame, a new identity tag will be assigned to it. If a target is not detected within a certain period of time, the target loss processing mechanism will be triggered to prevent the occurrence of misidentification. At the same time, in order to adapt to complex ambient lighting conditions, especially tracking tasks in low-light environments, brightness enhancement and contrast adjustment processing are performed on the real-time input image collected by the camera to improve the image recognizability. Brightness enhancement enhances details in dark areas through adaptive histogram equalization or gamma correction, while contrast adjustment uses Laplace enhancement or color transformation technology to improve the clarity of target edges. Feature compensation is performed on the depth appearance description of the pet target to adapt to the visual feature deviation caused by lighting changes. Through illumination invariant feature mapping or color normalization methods, it is ensured that the tracking algorithm can still stably identify and track pet targets even in low light conditions. Based on the optimized time series correspondence and enhanced tracking features, pet target state updates and trajectory management are performed to construct pet behavior tracking data. Each target is assigned a unique identity and its historical trajectory information is stored, including spatial position, movement direction, behavior category, etc.For example, if a pet moves within the camera's field of view, its trajectory path is recorded, and its behavior pattern is judged based on its movement trend, such as whether it is approaching the feeding device or showing a posture of expecting feeding. Through the trajectory management module, the behavior changes of the pet within a certain period of time are analyzed, such as calculating its activity trend, behavior pattern change rate, etc. Pet feeding analysis is performed based on the pet behavior tracking data, and a snack delivery control signal is output. The feeding analysis combines the pet's behavior sequence, trajectory information, and feeding demand confidence to determine whether a feeding operation is currently required. This process uses a rule-based approach combined with a machine learning model for decision optimization. When the feeding decision is determined, a snack delivery control signal is generated, and the feeding action is completed by remotely controlling the feeding device, and the pet's feeding log is updated at the same time.

[0033] Trajectory extrapolation is performed based on pet behavior tracking data to predict the pet's position when the snacks arrive. Since there is a certain time delay from the time the snacks are placed to the time they arrive on the ground, time series modeling methods, such as motion prediction models based on Kalman filtering or long short-term memory networks, are used to analyze the pet's movement trend and extrapolate its future movement trajectory. The pet's position data for the last few time steps is used to calculate its velocity vector and acceleration change, and combined with factors such as ground friction and pet walking habits, the expected movement path within a period of time after the snacks are placed is estimated to obtain the predicted position coordinates when the snacks arrive. After obtaining the predicted position, the feeding possibility score is calculated based on the pet behavior tracking data to evaluate the success probability of feeding. The pet's behavior sequence, position stability parameters, environmental suitability parameters, and historical feeding success rate are comprehensively analyzed. Among them, the behavior sequence reflects the pet's behavior pattern over the past period of time, such as whether it continues to pay attention to the feeding device, whether it shows behavior of expecting food (such as raising its head, sitting and waiting, or moving in circles), and these factors are automatically extracted through a deep learning behavior recognition network. At the same time, the position stability parameter measures the stability of the pet in the camera's field of view. For example, if the pet moves frequently in a short period of time, it means that its position is unstable, resulting in the failure of snack delivery. If the pet stays in a fixed area for a long time, the feeding success rate is higher. Consider environmental suitability parameters, such as whether the camera's field of view is unobstructed, whether there are obstacles around the pet, whether the snacks may be disturbed by external factors (such as furniture or other pets), etc. These environmental factors are analyzed through depth map detection or object recognition technology. Through historical feeding success rate data, the actual feeding success rate of the pet under different conditions is counted, and then the feeding strategy is optimized. These parameters are weighted and fused to obtain the basic score of feeding decision, which measures the rationality and feasibility of feeding in the current situation. The basic score of feeding decision is input into the reinforcement learning model to adjust the feeding parameters to optimize the feeding angle, strength and timing. The reinforcement learning model continuously learns the best strategy based on past feeding experience, where the state space includes the pet's position, speed, environmental factors, etc., and the action space includes the feeding angle, strength and release time. By training the feeding strategy network based on deep reinforcement learning, the system continuously optimizes the feeding parameters in the trial and error process to maximize the success probability of feeding. For example, if the model finds that a pet needs to increase the force of feeding at a long distance, and adjust the feeding angle at a close distance to avoid the snacks bouncing away, the system will automatically adjust these parameters in subsequent feedings to increase the probability of the pet successfully receiving the food. At the same time, the reinforcement learning model dynamically adjusts the timing of feeding. For example, when the pet is stationary or moving slowly, feeding is prioritized, and when the pet is moving at high speed, feeding is postponed to improve the accuracy of feeding. After optimizing the feeding parameters, the feasibility of the feeding angle, force and timing parameters is verified, and the target feeding execution plan is generated in combination with the pet's predicted position coordinates.For example, if the feeding angle calculated by the system exceeds the mechanical limit of the device, or the feeding force is not enough to deliver the snacks to the target area, the parameters are adjusted to ensure the feasibility of the feeding plan. Combined with the predicted position of trajectory extrapolation, feeding correction is performed to obtain the target feeding plan. According to the target feeding execution plan, the mechanical execution instruction sequence is calculated, and the mechanical execution instruction sequence includes the rotation angle of the drive motor, the ejection force and the release time point. For example, if a mechanical ejection feeding device is used, the specific angle of motor rotation is set to control the ejection direction of the snacks, and the tension of the spring or motor is set to adjust the flight distance of the snacks. According to the timing parameters optimized by reinforcement learning, the time point of releasing the snacks is accurately controlled to ensure that the snacks fall into the pet's accessible range according to the predetermined trajectory. If the feeding device adopts a track delivery method, the parameters such as the sliding distance of the track and the exit switch time are set to ensure that the snacks are accurately delivered to the target area. The mechanical execution instruction sequence is encapsulated by communication protocol and data integrity check to ensure that the instructions are not lost or erroneous during transmission. During the encapsulation process, the instructions are encoded to adapt to the control interface of the feeding device. At the same time, a cyclic redundancy check is used for integrity detection to ensure that the feeding command is not interfered with during the transmission process. When all checks are passed, the snack delivery control signal is output to drive the feeding device to execute snack delivery.

[0034] In the embodiment of the present invention, by improving the multi-scale cascade GhostNet feature extraction network, setting up a multi-scale convolution module with three different sizes of convolution kernels, and combining channel attention and spatial attention calculations, the feature extraction capability of the pet target is significantly enhanced. The spatial positioning and feature recognition enhancement module realizes centimeter-level precise positioning of the pet's spatial position through the comprehensive application of spatial attention, feature pyramid transformation and position-sensitive convolutional network, and extracts pet behavior characteristics with global context information through the channel interaction enhancement algorithm and non-local relationship modeling algorithm, which greatly improves the accuracy of pet behavior recognition in complex home scenes. The activity-based behavior recognition algorithm does not require a predefined activity threshold. It calculates pet activity evaluation indicators, constructs activity time series curves and dynamic threshold judgments, and combines time window filtering technology to achieve accurate judgment of pet feeding needs. The multi-level prediction network conducts a comprehensive analysis of the pet's detection frame, behavior category and feeding needs through feature splicing, multi-scale feature conversion and parallel calculation. It combines time series smoothing filters and Bayesian update mechanisms to achieve stable and accurate analysis of the pet's behavioral state. Through deep appearance description, behavioral state transfer model and feature compensation technology, it significantly improves the pet tracking performance in low-light environments. The self-adjusting feeding analysis algorithm achieves high-precision snack delivery through trajectory extrapolation calculation, feeding possibility scoring, reinforcement learning model and feasibility verification, combined with precise control of mechanical execution instruction sequences, significantly improving the intelligence level and user experience of remote pet care.

[0035] In a specific embodiment, the process of executing step S1 may specifically include the following steps:

[0036] Adaptively adjust the brightness of the original image captured by the pet camera to obtain a pre-processed pet image;

[0037] The preprocessed pet images are input into a multi-scale convolution module with three different convolution kernel sizes of 1×1, 3×3, and 5×5 for feature extraction to obtain pet feature maps of different scales;

[0038] Perform channel attention calculation and spatial attention calculation on pet feature maps of different scales to obtain weighted pet feature maps, and then perform cascade connection processing on the weighted pet feature maps in order from shallow to deep layers to obtain a multi-level feature representation of the pet;

[0039] The multi-level feature representation is spatially positioned and feature recognized to enhance the pet's spatial location information and pet's behavior characteristics.

[0040] Specifically, the original image captured by the camera is adaptively adjusted in brightness to improve the image quality. A brightness enhancement method based on histogram equalization is adopted. Assuming that the pixel value distribution of the original image is , its histogram distribution function is:

[0041]

[0042] in, and are the width and height of the image, respectively. represents the pixel value, is the unit impulse function. The cumulative distribution function is calculated, and the brightness of the image is adjusted based on the normalization transformation so that its pixel values ​​are evenly distributed, thereby improving the visibility of details in a dark environment. A brightness enhancement method based on gamma correction is adopted, and its core transformation is:

[0043]

[0044] in, For the adjusted image, is the normalization coefficient, is the brightness adjustment parameter, to increase brightness, or To reduce the brightness of the overexposed area. Through the adaptive brightness adjustment method, a more balanced pre-processed pet image is obtained. The pre-processed pet image is input into the , 3×3 and 5×5 convolution kernels to extract pet feature information of different scales. Among them, the 1×1 convolution kernel is used for linear transformation, adjusting channel information and reducing computational complexity. The calculation process is as follows:

[0045]

[0046] in, is the feature map after 1×1 convolution, is the number of channels of the input image, is the weight of the 1×1 convolution kernel. The 3×3 convolution kernel is used to capture local detail features, and the formula is as follows:

[0047]

[0048] in, is the weight of the 3×3 convolution kernel, which is used to extract local edge features. The 5×5 convolution kernel is used to capture a wider range of context information. The calculation formula is as follows:

[0049]

[0050] Feature maps of different scales are stitched together to form a multi-scale feature representation, so that the model can focus on both local texture and global morphology. Channel attention and spatial attention calculations are performed on pet feature maps of different scales to enhance effective features and suppress irrelevant background information. In terms of channel attention calculation, the Squeeze-and-Excitation (SE) module is used, that is, the features of each channel are globally averaged and pooled to calculate the global channel features:

[0051]

[0052] in, For Channel The average activation value of , and then generate the channel attention weight through the fully connected layer and nonlinear activation function :

[0053]

[0054] in, and is the weight matrix of the channel attention network, is the ReLU activation function, is the sigmoid normalization function, and the obtained attention weight Used to adjust the channel distribution of feature maps:

[0055]

[0056] The spatial attention calculation uses a combination of maximum pooling and average pooling to first calculate the global spatial feature map:

[0057]

[0058] in, is the weight of the spatial attention network, which performs spatial weighting on the original feature map:

[0059]

[0060] By combining channel attention and spatial attention, a weighted pet feature map is obtained. The weighted feature map is cascaded from shallow to deep to construct a multi-level feature representation of the pet. Shallow features mainly contain low-level visual information such as edges and textures, while deep features contain richer high-level semantic information, such as the pet's shape, posture and behavior category. Finally, these features are integrated through cascade operations, so that the model pays attention to local details and overall shape at the same time, and obtains a multi-level feature representation of the pet. Spatial positioning and feature recognition enhancement are performed on the multi-level feature representation to obtain the spatial position information and behavioral characteristics of the pet. Spatial positioning uses target detection methods, such as YOLO or Faster R-CNN, by regressing the target bounding box coordinates:

[0061]

[0062] in, is the center coordinate of the pet, and are the width and height of the pet detection frame respectively. Feature recognition enhancement calculates the pet's behavior category probability through the classification network:

[0063]

[0064] in, Indicates that the pet belongs to the behavioral category The probability of pet behavior recognition is finally optimized by combining this information with trajectory tracking methods (such as Kalman filtering).

[0065] In this embodiment, a target pet image is captured by a pet camera, and features are extracted from the target pet image. Before obtaining a multi-level feature representation of the pet, a step of enhancing the target pet image is also included. The enhancing the target pet image includes: inputting the original image captured by the pet camera into an adaptive image quality assessment module, and obtaining an image quality score and an enhancement requirement mapping by calculating brightness distribution, contrast, noise level and blurriness quantification indicators; executing a selective bilateral filtering algorithm based on the image quality score and the enhancement requirement mapping, and performing differential noise suppression processing on the pet area and the background area respectively, to obtain a denoised image that retains pet details; performing an adaptive histogram equalization operation on the denoised image, and dynamically adjusting the enhancement parameters according to the brightness distribution of the local area of ​​the image, to obtain an enhanced image after illumination correction; inputting the enhanced image after illumination correction into an adversarial residual enhancement network, and performing feature fusion The residual learning mechanism is used to improve the image resolution and details to obtain a super-resolution enhanced image; an artifact detection and removal algorithm is performed based on the super-resolution enhanced image, and the stripes, color shift and distortion artifacts caused by the camera hardware are identified and suppressed through frequency domain analysis to obtain an artifact corrected image; an edge-preserving sharpening process is performed on the artifact corrected image, and the pet contour and texture details are enhanced through an adaptive sharpening factor to obtain a detail enhanced image; the detail enhanced image is input into the time domain stability enhancement module, and the similarity between consecutive multi-frame images is analyzed, and the time domain filtering algorithm is performed to smooth the fluctuations in the time dimension to obtain a time consistency enhanced image; a quality assessment and optimization adjustment based on a convolutional neural network is performed on the time consistency enhanced image, and the final enhancement parameter fine-tuning is guided by the perceptual quality loss function to obtain a target pet image with optimized visual quality, and the target pet image with optimized visual quality is used for subsequent feature extraction and behavior recognition processing.

[0066] In a specific embodiment, the execution step performs spatial positioning and feature recognition enhancement on the multi-level feature representation to obtain the pet spatial position information and pet behavior characteristics, which may specifically include the following steps:

[0067] The multi-level feature representation is input into the spatial attention module for adaptive spatial weight calculation to obtain a spatial weighted feature map, and feature transformation is performed on the spatial weighted feature map to obtain a multi-scale spatial feature representation;

[0068] The multi-scale spatial feature representation is input into the position-sensitive convolutional network for spatial coordinate encoding to obtain the pet's spatial position information, and channel interaction enhancement is performed on the multi-scale spatial feature representation to obtain cross-channel fused behavior representation features;

[0069] The cross-channel fused behavior representation features are cascaded with the pet's spatial position information to obtain a position-aware pet behavior feature map, and non-local relationship modeling is performed on the position-aware pet behavior feature map to obtain the pet behavior features.

[0070] Specifically, the multi-level feature representation is input into the spatial attention module to calculate the adaptive spatial weights and obtain the spatial weighted feature map. Spatial attention enables the network to pay more attention to the pet body area and suppress irrelevant background information by calculating the importance of different positions in the feature map. Assume that the multi-level feature representation is in represents the spatial coordinates of the feature map, Represents the channel index, then the global average pooling and maximum pooling are calculated for it:

[0071]

[0072]

[0073] Will and To perform channel dimension splicing, enter a Convolutional Layer , and the spatial attention weight is calculated through the sigmoid activation function :

[0074]

[0075] in, is the convolution kernel parameter of the spatial attention network, is the sigmoid function, and we get As a spatial weighting factor, the original feature map is weighted:

[0076]

[0077] Get the spatial weighted feature map , the feature map highlights the pet target area while reducing background interference. Perform feature transformation on the spatial weighted feature map to obtain a multi-scale spatial feature representation. The multi-scale pyramid feature method is adopted, that is, through pooling operations of different scales combined with convolution kernels with variable steps, the feature map can simultaneously express local details and global information at different resolutions. The average pooling operations of 2×2, 4×4 and 8×8 are used to calculate the features at different scales:

[0078]

[0079]

[0080]

[0081] The feature maps of different scales are then restored to the same resolution through upsampling operations, and feature fusion is performed to form a multi-scale spatial feature representation:

[0082]

[0083] in, is a learnable fusion weight. The multi-scale spatial feature representation is input into the position-sensitive convolutional network for spatial coordinate encoding to obtain the precise spatial position information of the pet. Based on the conventional convolution calculation, the position-sensitive convolutional network introduces additional position information encoding, so that each pixel not only contains local feature information, but also reflects its absolute position in the entire image. Construct two position encoding matrices and ,in and , respectively represent the normalized coordinates of the pixel in the horizontal and vertical directions. These position information are concatenated to the original feature map and input into a 3×3 position-sensitive convolution layer :

[0084]

[0085] The final output This is the encoded spatial feature, which contains the precise location coordinate information of the pet. At the same time, channel interaction enhancement is performed on the multi-scale spatial feature representation to obtain cross-channel fused behavioral features. The key to channel interaction enhancement is to utilize the interdependence between different channels to improve the network's sensitivity to behavioral patterns. The Squeeze-and-Excitation (SE) module is used to perform global average pooling on each channel and calculate the activation value of each channel:

[0086]

[0087] Generate channel weights through nonlinear transformation through the fully connected layer :

[0088]

[0089] in, and is the weight of the channel attention network, is the ReLU activation function, is the sigmoid normalization function, and the final calculated Used to adjust channel characteristics:

[0090]

[0091] The cross-channel fused behavior representation features are obtained. The cross-channel fused behavior representation features are cascaded with the pet's spatial position information to form a location-aware pet behavior feature map. and Splicing is done in the channel dimension and through Convolution performs channel compression to reduce redundant information while ensuring close association between behavioral features and spatial information:

[0092]

[0093] Non-local relationship modeling is performed on the location-aware pet behavior feature map, allowing the network to capture long-range feature dependencies. A non-local attention mechanism is used to calculate the similarity of global features and weight features based on the similarity:

[0094]

[0095] in, and They are mapping functions used to calculate feature similarity, and non-local features are expressed as:

[0096]

[0097] As a pet behavior feature, it reflects the pet's dynamic behavior pattern and provides high-precision input for feeding decisions.

[0098] Among them, before performing activity-based behavior recognition based on the pet's spatial position information and the pet's behavior characteristics, it also includes a pet behavior characteristic adaptive optimization step based on the autonomous model sequence selection, and the pet behavior characteristic adaptive optimization step includes: performing multi-time window segmentation processing on the pet behavior characteristics to obtain pet behavior sequence samples of different durations; constructing a Hankel matrix based on the pet behavior sequence samples, and performing a singular value decomposition operation to obtain a system eigenvalue distribution; applying a threshold decay analysis algorithm to the system eigenvalue distribution to determine the optimal dimension of the pet behavior state space and obtain the behavior model complexity parameter; performing a subspace projection operation on the pet behavior characteristics based on the behavior model complexity parameter , obtain the core behavior features after dimensionality reduction; construct a cross-Gramian matrix based on the core behavior features, quantitatively characterize the interactive relationship between the pet behavior states, and obtain the behavior state correlation index; input the core behavior features into the state estimator for state space reconstruction to obtain the pet behavior state estimation result; calculate the state energy error based on the pet behavior state estimation result and the actual observed pet behavior features, and obtain the behavior model accuracy evaluation index; perform adaptive optimization operations on the core behavior features according to the behavior model accuracy evaluation index, eliminate redundant features and enhance key features, and obtain the optimized pet behavior features, which provide more accurate feature input for subsequent activity-based behavior recognition.

[0099] In a specific embodiment, the process of executing step S2 may specifically include the following steps:

[0100] The movement frequency, movement amplitude, behavior change rate and frequency of specific behaviors are calculated based on the pet's spatial location information and pet behavior characteristics to obtain the pet activity evaluation index;

[0101] Perform time series sampling and sliding window processing on the continuously collected pet activity evaluation indicators to obtain a pet activity time series curve, and perform statistical distribution analysis based on the pet activity time series curve to obtain a pet activity dynamic threshold;

[0102] Make rule judgments on pet activity evaluation indicators and pet activity dynamic thresholds to obtain an initial feeding requirement score, and perform time window filtering on the initial feeding requirement score to obtain a stable feeding requirement score;

[0103] A comprehensive score calculation is performed based on the stable feeding demand score, the time interval from the last feeding, and the pet's daily feeding pattern to obtain the pet's feeding demand confidence.

[0104] Specifically, the pet's spatial location information and behavioral characteristics are used to calculate its movement frequency, movement amplitude, behavior change rate, and the frequency of specific behaviors, and to construct an activity evaluation index for the pet. Collect pet's location , then the movement frequency is defined as the number of displacements of the pet per unit time:

[0105]

[0106] in, Observation time for pets The effective number of moves within a time step is the displacement distance of the pet in two adjacent time steps. Exceeding the set inactivity threshold The displacement distance is calculated as follows:

[0107]

[0108] like , it is considered as one move and accumulated to The range of motion is a measure of the pet's overall range of motion and is calculated as:

[0109]

[0110] in, Represents the average distance the pet moves and is used to assess the intensity of exercise. Reflects the switching frequency between different behaviors of the pet within a certain period of time, and is defined as follows:

[0111]

[0112] in, is the number of behavior category switches, assuming For pets in time , then when At the same time, the frequency of specific behaviors is calculated, such as the pet's sitting, jumping, moving around the food device, etc. In the time window The number of occurrences in , then its frequency of occurrence is expressed as:

[0113]

[0114] Combining the above four indicators, construct a pet activity evaluation vector:

[0115]

[0116] The activity evaluation indicators are sampled in time series and processed with sliding windows to construct a time series curve of pet activity. The sampled activity data is , the sliding window average method is used for smoothing to make the activity curve more stable:

[0117]

[0118] in, is the length of the sliding window. Based on the activity time series curve, the dynamic threshold of the pet's activity is calculated using the statistical distribution analysis method. Assuming that the pet's activity data follows a normal distribution, its dynamic threshold is expressed as:

[0119]

[0120] in, and are the mean and standard deviation of the activity data, is the adjustment factor, and takes 1.5-2 to cover the normal data range of 85%-95%. The activity evaluation index and the dynamic threshold are judged by rules to calculate the initial feeding demand score. And the frequency of specific behaviors Exceeding the set threshold When , set the initial feeding requirement score:

[0121]

[0122] in, and is the weight parameter, is an indicator function, which takes 1 if the condition is met, otherwise takes 0. In order to avoid misjudgment of the feeding demand score due to short-term fluctuations, the time window filtering method is used to calculate the stable feeding demand score:

[0123]

[0124] in, is the smoothing coefficient, The feeding demand score at the last moment makes the feeding decision more stable. The final pet feeding demand confidence is calculated by comprehensively considering the stable feeding demand score, the time interval from the last feeding, and the pet's daily feeding pattern. Define the time interval from the last feeding And normalize it:

[0125]

[0126] in, The maximum time interval is set. Consider the daily feeding pattern of the pet and use long-term statistical data to calculate the feeding probability within a specific time period. , the comprehensive score is calculated as follows:

[0127]

[0128] in, is the weight parameter, and the final calculated Confidence in feeding needs as a pet.

[0129] In a specific embodiment, the process of executing step S3 may specifically include the following steps:

[0130] Perform feature concatenation and standardization on the pet's spatial location information, pet's behavioral characteristics, and feeding demand confidence to obtain a combined feature representation;

[0131] The combined feature representation is used to perform multi-scale feature transformation with the target pet image input upsampling and downsampling modules to obtain feature prediction levels of different resolutions;

[0132] Each layer of the feature prediction hierarchy is connected to the prediction branch network for parallel calculation to obtain the pet detection frame coordinates, behavior category probability distribution and feeding judgment score at each level;

[0133] Based on the semantic information depth of each level of features, weighted fusion calculation is performed on the pet detection frame coordinates, behavior category probability distribution and feeding judgment score at each level to obtain the fusion prediction result;

[0134] Perform redundant frame filtering on the pet detection frame in the fusion prediction result to obtain the target detection frame position;

[0135] The target detection frame position, behavior category probability distribution and feeding judgment score are input into the time series smoothing filter and Bayesian update module for dynamic optimization to obtain the pet detection frame, behavior category and feeding judgment results.

[0136] Specifically, the pet's spatial location information, behavioral characteristics, and feeding demand confidence are concatenated and standardized to obtain a combined feature representation. Assume that the pet's spatial location information is represented as ,in is the center coordinate of the pet, is the width and height of the detection box; the behavioral characteristics of the pet are ,in represents the probability of a certain category of behavior characteristics; the confidence level of feeding demand is , then the combined feature representation is defined as:

[0137]

[0138] Since the numerical ranges of different features vary greatly, they are standardized to ensure that different features have similar scales during calculation. The standardization adopts the mean-standard deviation normalization method, that is,

[0139]

[0140] in, and are the mean and standard deviation of the combined features in the training data, respectively, so that the normalized data mean is close to 0 and the standard deviation is close to 1, which improves the training stability of the neural network. The standardized combined feature representation and the target pet image are input into the upsampling and downsampling modules to perform multi-scale feature conversion and obtain feature prediction levels of different resolutions. The input image and the combined feature representation are downsampled, and deep semantic features are extracted through multiple convolutional layers and pooling operations. Assume that the input feature map is , the convolutional layer parameters are , then The feature extraction process of the layer is expressed as:

[0141]

[0142] Among them, * represents the convolution operation, represents the activation function (such as ReLU), Indicates At the same time, in order to improve the detection accuracy, the upsampling module is used to restore the high-resolution features to obtain more detailed information. Upsampling uses bilinear interpolation or deconvolution operations:

[0143]

[0144] By combining up and down sampling, a pyramid feature hierarchy is constructed, so that the model can learn information of different scales at the same time. Each layer of the feature prediction hierarchy is connected to the prediction branch network for parallel calculation to obtain the pet detection frame coordinates, behavior category probability distribution and feeding judgment score of each layer. The characteristics of the layer are , then the prediction of the detection box coordinates is calculated through the regression network:

[0145]

[0146] in, is the learned weight matrix. Similarly, the probability distribution of the behavior category is calculated by softmax:

[0147]

[0148] The feeding judgment score is calculated by sigmoid:

[0149]

[0150] Each level independently predicts the detection frame, behavior category probability, and feeding score. Based on the semantic information depth of the features at each level, the pet detection frame coordinates, behavior category probability distribution, and feeding judgment score at each level are weighted and fused to obtain a comprehensive prediction result. Suppose the feature importance of each level is , then the weighted fusion calculation formula is as follows:

[0151]

[0152]

[0153]

[0154] in, Determined by the attention mechanism or the confidence-based adaptive weighting method to ensure that high-level features provide more global detection results, while low-level features provide more refined spatial information. Redundant frame filtering is performed on the pet detection frame in the fusion prediction result to remove the repeatedly detected targets and obtain the target detection frame position. Redundant frame filtering uses a non-maximum suppression algorithm. For all predicted frames, they are first sorted by confidence, and then starting from the high-confidence frame, the intersection over union (IoU) with all other frames is calculated. If the IoU of a frame exceeds the set threshold, , then remove the box. The calculation formula for non-maximum suppression is as follows:

[0155]

[0156]

[0157] Through non-maximum suppression filtering, the target detection frames with the highest confidence and without overlap are retained. The target detection frame position, behavior category probability distribution, and feeding judgment score are input into the time series smoothing filter and Bayesian update module for dynamic optimization. The time series smoothing filter uses a first-order exponential weighted average filter:

[0158]

[0159]

[0160]

[0161] in, is the smoothing coefficient to reduce detection jitter. At the same time, the Bayesian update module dynamically adjusts the feeding confidence based on past detection results. For example, if the pet's behavior pattern after the system has fed it several times in the past shows that it does have a need for food, the feeding confidence is updated based on the prior probability:

[0162]

[0163] in, represents past observation data, Indicates the current feeding confidence The data observed below The probability of is the prior probability of feeding confidence, is the total probability of the data. In this way, the feeding strategy can be adjusted more intelligently.

[0164] In a specific embodiment, the process of executing step S4 may specifically include the following steps:

[0165] The pet detection frame is fused with the pet's multi-level feature representation to construct an enhanced appearance feature, and a deep appearance description of the pet target is obtained;

[0166] A pet behavior state transition model is constructed based on the behavior category and input into the Kalman filter for motion prediction to obtain the pet's position and behavior state prediction values;

[0167] Based on the deep appearance description of the pet target, the predicted value of the pet's position and behavior state, and the feeding judgment result, a multi-dimensional matching cost matrix is ​​calculated to obtain the target association metric;

[0168] Perform target association optimization on the target association metric to obtain the temporal correspondence between pet targets;

[0169] Perform brightness enhancement and contrast adjustment on the real-time input image captured by the pet camera, and perform feature compensation on the depth appearance description of the pet target to obtain enhanced tracking features in low-light environments;

[0170] Based on the time sequence correspondence and enhanced tracking features, the pet target state is updated and the trajectory is managed to obtain pet behavior tracking data including pet identity, location trajectory and behavior sequence;

[0171] Perform pet feeding analysis based on pet behavior tracking data and output snack delivery control signals.

[0172] Specifically, the pet detection frame is fused with the multi-level feature representation to construct an enhanced appearance feature and obtain a deep appearance description of the pet target. Assume that the pet detection frame is represented as ,in is the center coordinate of the detection box, is the width and height of the detection box; the multi-level feature representation is ,in Represents feature channel information at different levels. Through feature fusion operations, enhanced apparent features are obtained :

[0173]

[0174] in, is a weight matrix used to adjust the contribution ratio of different information. The enhanced appearance feature embeds the spatial information of the target in the visual feature to improve the detection stability. After completing the deep appearance description, in order to predict the pet behavior, a state transition model is constructed based on the behavior category, and the Kalman filter is combined for motion prediction. Assume that the pet behavior category set is ,in Representing the different behavioral states of the pet, the state transition probability matrix is ​​constructed :

[0175]

[0176] in, Indicates the pet's behavior status Transfer to The probability is calculated by using past observation data. At the same time, the Kalman filter is used to predict the pet's motion trajectory. Assuming that the pet's motion state vector is ,in is the speed in the horizontal and vertical directions, then the state update formula is:

[0177]

[0178] in, is the state transfer matrix, is the process noise, which is assumed to be zero-mean Gaussian noise. The observation model is:

[0179]

[0180] in, is the observed pet location, is the observation matrix, is the observation noise. Through the prediction and update steps of the Kalman filter, the future position of the pet is estimated Its behavior status . Based on the deep appearance description of the pet target, the predicted value of the pet's position and behavior state, and the feeding judgment result, a multi-dimensional matching cost matrix is ​​calculated to perform target association. Suppose the pet target set at the previous moment is , the newly detected target set is , then the matching cost matrix Each element of It is composed of Euclidean distance, apparent feature similarity and behavioral similarity:

[0181]

[0182] in, is the weighting coefficient, Represents cosine similarity, and the closer it is to 1, the more similar the apparent features are. The cost matrix is ​​optimally matched through the Hungarian algorithm to obtain the temporal correspondence of the pet target. After completing the target association, the real-time input image captured by the camera is brightness enhanced and contrast adjusted to improve the tracking effect in low-light environments. Brightness enhancement is achieved through histogram equalization or adaptive gamma correction:

[0183]

[0184] in, is the original image, is the brightness adjustment parameter, is the normalization coefficient. In low-light environments, feature compensation is performed on the depth appearance description of pet targets to reduce the impact of lighting changes. Feature compensation uses color normalization method:

[0185]

[0186] Obtain enhanced tracking features in low-light environments. Update the pet's status and manage its trajectory information based on the target temporal correspondence and enhanced tracking features. is the pet's trajectory set, then the update formula is:

[0187]

[0188] The current spatial position and behavioral status Append to the trajectory collection. If the pet target does not match a new detection result within a certain period of time, the old target is removed using the forgetting strategy to ensure the timeliness of the tracking data. Perform pet feeding analysis based on the pet behavior tracking data to determine whether to trigger feeding. Assume that the feeding judgment function is , the pet's activity, exercise status and feeding demand confidence are comprehensively considered:

[0189]

[0190] in, For your pet's activity level, Represents the pet's behavior categories related to eating in the past observations. Feeding demand confidence is the weighting coefficient. Exceeding the set threshold , the system outputs the snack delivery control signal:

[0191]

[0192] Thereby driving the feeding device to complete the snack delivery.

[0193] In a specific embodiment, the execution step performs pet feeding analysis based on pet behavior tracking data, and the process of outputting a snack delivery control signal may specifically include the following steps:

[0194] Based on the pet's behavior tracking data, the trajectory is extrapolated to obtain the pet's predicted location coordinates when the snacks arrive;

[0195] The feeding possibility score is calculated based on the behavior sequence, position stability parameter, environmental suitability parameter and historical feeding success rate in the pet behavior tracking data to obtain the basic score for feeding decision;

[0196] The basic score of feeding decision is input into the reinforcement learning model to adjust the feeding parameters, and the optimized feeding angle, strength and timing parameters are obtained;

[0197] Perform feasibility verification on the optimized feeding angle, force and timing parameters, and obtain the target feeding execution plan based on the predicted position coordinates of the pet when the snack arrives;

[0198] Calculate the mechanical execution instruction sequence according to the target feeding execution plan, the mechanical execution instruction sequence includes the driving motor rotation angle, ejection force and release time point;

[0199] The machine execution instruction sequence is encapsulated in communication protocol and data integrity is checked, and the snack delivery control signal is output.

[0200] Specifically, trajectory extrapolation calculation is performed based on the pet behavior tracking data to predict the location coordinates of the pet when the snack arrives. Assume that the current location of the pet is , the velocity vector is , the acceleration is , then in the time interval After that, the predicted position of the pet is calculated by the uniformly accelerated motion equation:

[0201]

[0202]

[0203] in, , Indicates the calculation method of horizontal velocity and acceleration, vertical The same is true for the direction. By extrapolating the trajectory, we can get the predicted position of the snack landing point after the snack is delivered to the pet. The feeding possibility score is calculated based on the behavior sequence, position stability parameter, environmental suitability parameter, and historical feeding success rate in the pet behavior tracking data to determine whether the current moment is suitable for feeding. Assume that the behavior sequence Representing pets in the past Behavior categories of frames, position stability parameters Reflects the degree of fluctuation of the pet's spatial position. The calculation formula is as follows:

[0204]

[0205] in, It's the past The average position of the frame, The smaller it is, the more stable the pet's position is. Environmental suitability parameters Evaluate the obstacles around the pet detected by the camera, analyze the scene complexity through the deep learning model, and set Suitable for feeding Indicates that there are obstacles and it is not suitable for feeding. Historical feeding success rate The probability of a pet successfully receiving food in different positions and behavior patterns is calculated through long-term data analysis. Taking all the above factors into consideration, the feeding probability score is calculated as follows:

[0206]

[0207] in, is the weight parameter, Represents the degree of match between the current behavior sequence and feeding-related behaviors. The basic score of feeding decision is input into the reinforcement learning model to optimize the feeding angle, strength and timing parameters. As input, take action ,in For the feeding angle, For feeding intensity, Free up time for snacks. The reinforcement learning goal is to maximize the reward function for the pet to successfully receive food:

[0208]

[0209] in, Indicates the historical probability of the pet successfully receiving food. Represents the deviation between the snack drop point and the predicted position. Update the policy function using the policy gradient method:

[0210]

[0211] in, represents the expected reward, is the learning rate. After the reinforcement learning iteration, the optimized feeding angle, strength and timing parameters are obtained. The feasibility of the optimized feeding parameters is verified, and the target feeding execution plan is calculated based on the predicted position of the pet. Feeding Angle The physical constraints of the feeding device need to be met, and the maximum feeding angle is set , then it is required Feeding intensity It is necessary to ensure that the snack lands close to the predicted position of the pet and satisfies the parabolic motion equation:

[0212]

[0213]

[0214] in, is the initial velocity, is the acceleration due to gravity, Flight time for snacks, For snack quality. Adjust Make Approach . According to the target feeding execution plan, the mechanical execution instruction sequence is calculated, and the instructions include the rotation angle of the drive motor , Ejection Strength and release time The motor angle control is achieved through a stepper motor, and the step angle is:

[0215]

[0216] in, is the total number of steps of the stepper motor. By controlling the spring preload or motor speed:

[0217]

[0218] in, is the device conversion factor. Release time It is necessary to ensure that snacks are delivered when the pet enters the feeding area, and the timing prediction is used to adjust:

[0219]

[0220] The communication protocol is encapsulated and the data integrity is checked for the mechanical execution instruction sequence to ensure that the feeding device accurately executes the delivery instruction. The communication protocol uses serial communication (UART, I2C) or wireless communication (Wi-Fi, Bluetooth), and the data integrity check uses CRC check:

[0221]

[0222] in, For data to be sent, is the calculated checksum. The feeding instruction encapsulation format is as follows:

[0223]

[0224] Once the device receives the command and passes the verification, snack delivery can be executed.

[0225] The above describes the remote snack delivery method based on the pet camera in the embodiment of the present invention. The following describes the remote snack delivery device based on the pet camera in the embodiment of the present invention. Figure 2 In one embodiment of the present invention, a remote snack delivery device based on a pet camera includes:

[0226] The acquisition module is used to acquire the target pet image through the pet camera, and to perform spatial positioning and feature recognition enhancement on the target pet image to obtain the pet spatial position information and pet behavior characteristics;

[0227] A behavior recognition module is used to perform activity-based behavior recognition based on the pet's spatial location information and pet behavior characteristics to obtain the pet's feeding demand confidence;

[0228] A processing module is used to input the pet's spatial position information, pet's behavioral characteristics and feeding demand confidence into a multi-level prediction network for processing to obtain a pet detection frame, behavior category and feeding judgment result;

[0229] The feeding analysis module is used to perform pet behavior tracking and feeding analysis based on the pet detection frame, behavior category and feeding judgment results, and output snack delivery control signals.

[0230] Through the coordinated cooperation of the above components, by improving the multi-scale cascade GhostNet feature extraction network, setting up multi-scale convolution modules with three different sizes of convolution kernels, and combining channel attention and spatial attention calculations, the feature extraction ability of pet targets is significantly enhanced. The spatial positioning and feature recognition enhancement module achieves centimeter-level precise positioning of the pet's spatial position through the comprehensive application of spatial attention, feature pyramid transformation and position-sensitive convolutional network, and extracts pet behavior characteristics with global context information through channel interaction enhancement algorithm and non-local relationship modeling algorithm, which greatly improves the accuracy of pet behavior recognition in complex home scenes. The activity-based behavior recognition algorithm does not require a predefined activity threshold. It calculates pet activity evaluation indicators, constructs activity time series curves and dynamic threshold judgments, and combines time window filtering technology to achieve accurate judgment of pet feeding needs. The multi-level prediction network conducts a comprehensive analysis of the pet's detection frame, behavior category and feeding needs through feature splicing, multi-scale feature conversion and parallel calculation. It combines time series smoothing filters and Bayesian update mechanisms to achieve stable and accurate analysis of the pet's behavioral state. Through deep appearance description, behavioral state transfer model and feature compensation technology, it significantly improves the pet tracking performance in low-light environments. The self-adjusting feeding analysis algorithm achieves high-precision snack delivery through trajectory extrapolation calculation, feeding possibility scoring, reinforcement learning model and feasibility verification, combined with precise control of mechanical execution instruction sequences, significantly improving the intelligence level and user experience of remote pet care.

[0231] Reference Figure 3 In an embodiment of the present invention, a computer device is also provided. The computer device may be a server, and its internal structure may be as follows: Figure 3 As shown. The computer device includes a processor, a memory, a display screen, an input device, a network interface and a database connected through a system bus. Among them, the processor designed by the computer is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store the corresponding data in this embodiment. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, the above method is implemented.

[0232] Those skilled in the art will understand that Figure 3 The structure shown in the figure is merely a block diagram of a portion of the structure related to the solution of the present invention, and does not constitute a limitation on the computer device to which the solution of the present invention is applied.

[0233] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described systems, systems and units can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0234] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art or the whole or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk and other media that can store program code.

[0235] As described above, the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features thereof may be replaced by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A remote snack delivery method based on a pet camera, characterized in that: The method comprises: The target pet image is captured by a pet camera, and the target pet image is spatially positioned and feature-recognized and enhanced to obtain the pet spatial position information and pet behavior characteristics; Performing activity-based behavior recognition according to the pet spatial position information and the pet behavior characteristics to obtain a pet feeding demand confidence level; Input the pet spatial position information, the pet behavior characteristics and the feeding demand confidence into a multi-level prediction network for processing to obtain a pet detection frame, a behavior category and a feeding judgment result; Based on the pet detection frame, the behavior category and the feeding judgment result, pet behavior tracking and feeding analysis are performed, and a snack delivery control signal is output.

2. The remote snack delivery method based on a pet camera according to claim 1 is characterized in that: The target pet image is collected by the pet camera, and the target pet image is spatially positioned and feature-recognized and enhanced to obtain the pet spatial position information and pet behavior characteristics, including: Adaptively adjust the brightness of the original image captured by the pet camera to obtain a pre-processed pet image; Inputting the preprocessed pet image into a multi-scale convolution module having three different-sized convolution kernels of 1×1, 3×3, and 5×5 for feature extraction to obtain pet feature maps of different scales; Perform channel attention calculation and spatial attention calculation on the pet feature maps of different scales to obtain weighted pet feature maps, and perform cascade connection processing on the weighted pet feature maps in order from shallow to deep layers to obtain a multi-level feature representation of the pet; The multi-level feature representation is spatially positioned and feature recognition enhanced to obtain pet spatial position information and pet behavior characteristics.

3. The remote snack delivery method based on a pet camera according to claim 2 is characterized in that: The performing spatial positioning and feature recognition enhancement on the multi-level feature representation to obtain pet spatial position information and pet behavior characteristics includes: Inputting the multi-level feature representation into a spatial attention module to perform adaptive spatial weight calculation to obtain a spatial weighted feature map, and performing feature transformation on the spatial weighted feature map to obtain a multi-scale spatial feature representation; Inputting the multi-scale spatial feature representation into a position-sensitive convolutional network for spatial coordinate encoding to obtain the spatial position information of the pet, and performing channel interaction enhancement on the multi-scale spatial feature representation to obtain cross-channel fused behavior representation features; A feature cascade operation is performed on the cross-channel fused behavior representation features and the pet spatial position information to obtain a position-aware pet behavior feature map, and non-local relationship modeling is performed on the position-aware pet behavior feature map to obtain pet behavior features.

4. The remote snack delivery method based on a pet camera according to claim 1, characterized in that: The performing of activity-based behavior recognition according to the pet spatial position information and the pet behavior characteristics to obtain the pet feeding demand confidence level includes: Calculate the movement frequency, movement amplitude, behavior change rate and specific behavior occurrence frequency according to the pet spatial position information and the pet behavior characteristics to obtain a pet activity evaluation index; Performing time series sampling and sliding window processing on the continuously collected pet activity evaluation indicators to obtain a pet activity time series curve, and performing statistical distribution analysis based on the pet activity time series curve to obtain a pet activity dynamic threshold; Performing rule judgment on the pet activity evaluation index and the pet activity dynamic threshold to obtain an initial feeding requirement score, and performing time window filtering on the initial feeding requirement score to obtain a stable feeding requirement score; A comprehensive score calculation is performed based on the stable feeding demand score, the time interval from the last feeding, and the daily feeding pattern of the pet to obtain the pet feeding demand confidence.

5. The remote snack delivery method based on a pet camera according to claim 1, characterized in that: The pet spatial position information, the pet behavior characteristics and the feeding demand confidence are input into a multi-level prediction network for processing to obtain a pet detection frame, a behavior category and a feeding judgment result, including: Performing feature concatenation and standardization processing on the pet spatial position information, the pet behavior characteristics and the feeding demand confidence to obtain a combined feature representation; Performing multi-scale feature conversion on the combined feature representation and the target pet image input upsampling and downsampling modules to obtain feature prediction levels of different resolutions; Each layer of the feature prediction level is connected to the prediction branch network for parallel calculation to obtain the pet detection frame coordinates, behavior category probability distribution and feeding judgment score of each level; Based on the semantic information depth of the features at each level, weighted fusion calculation is performed on the pet detection frame coordinates, behavior category probability distribution, and feeding judgment score at each level to obtain a fusion prediction result; Perform redundant frame filtering on the pet detection frame in the fusion prediction result to obtain the target detection frame position; The target detection frame position, behavior category probability distribution and feeding judgment score are input into the time series smoothing filter and the Bayesian update module for dynamic optimization to obtain the pet detection frame, behavior category and feeding judgment result.

6. The remote snack delivery method based on a pet camera according to claim 2, characterized in that: The pet behavior tracking and feeding analysis are performed based on the pet detection frame, the behavior category and the feeding judgment result, and the snack delivery control signal is output, including: The pet detection frame is fused with the pet multi-level feature representation to construct an enhanced appearance feature, so as to obtain a deep appearance description of the pet target; Constructing a pet behavior state transfer model according to the behavior category, and inputting it into a Kalman filter for motion prediction to obtain a predicted value of the pet's position and behavior state; Calculate a multi-dimensional matching cost matrix based on the depth appearance description of the pet target, the predicted value of the pet position and behavior state, and the feeding judgment result to obtain a target association metric; Performing target association optimization on the target association metric to obtain a temporal correspondence relationship between pet targets; Performing brightness enhancement and contrast adjustment processing on the real-time input image captured by the pet camera, and performing feature compensation on the depth appearance description of the pet target, so as to obtain enhanced tracking features in a low-light environment; Based on the time sequence correspondence and the enhanced tracking feature, pet target state update and trajectory management are performed to obtain pet behavior tracking data including pet identity, location trajectory and behavior sequence; A pet feeding analysis is performed based on the pet behavior tracking data, and a snack delivery control signal is output.

7. The remote snack delivery method based on a pet camera according to claim 1, characterized in that: The performing pet feeding analysis according to the pet behavior tracking data and outputting a snack delivery control signal comprises: Perform trajectory extrapolation calculation based on the pet behavior tracking data to obtain the predicted position coordinates of the pet when the snacks arrive; Calculating a feeding possibility score based on the behavior sequence, position stability parameter, environmental suitability parameter, and historical feeding success rate in the pet behavior tracking data to obtain a feeding decision basic score; Inputting the feeding decision basic score into the reinforcement learning model to adjust the feeding parameters to obtain optimized feeding angle, strength and timing parameters; Perform feasibility verification on the optimized feeding angle, force and timing parameters, and obtain a target feeding execution plan in combination with the predicted position coordinates of the pet when the snack arrives; Calculating a mechanical execution instruction sequence according to the target feeding execution plan, wherein the mechanical execution instruction sequence includes a driving motor rotation angle, a ejection force, and a release time point; The machine executes the instruction sequence and performs communication protocol encapsulation and data integrity verification, and outputs a snack delivery control signal.

8. A remote snack delivery device based on a pet camera, characterized in that: Used to execute the remote snack delivery method based on a pet camera according to any one of claims 1 to 7, the remote snack delivery device based on a pet camera comprises: A collection module, used to collect a target pet image through a pet camera, and perform spatial positioning and feature recognition enhancement on the target pet image to obtain the pet's spatial position information and pet behavior characteristics; A behavior recognition module, used to perform activity-based behavior recognition according to the pet's spatial position information and the pet's behavior characteristics, and obtain a pet's feeding demand confidence level; A processing module, used to input the pet spatial position information, the pet behavior characteristics and the feeding demand confidence into a multi-level prediction network for processing to obtain a pet detection frame, a behavior category and a feeding judgment result; A feeding analysis module is used to perform pet behavior tracking and feeding analysis based on the pet detection frame, the behavior category and the feeding judgment result, and output a snack delivery control signal.

9. A computer device, characterized in that: It includes a memory and a processor, the memory stores a computer program that can be run on the processor, and is characterized in that when the processor executes the computer program, it implements the remote snack delivery method based on a pet camera as described in any one of claims 1 to 7.

Citation Information

Cited By

  • Pet behavior training system based on AI vision

    CN120982435A