Pet Perspective-Based Adaptive Camera Tracking Method and System

By collecting pet motion data and using machine learning algorithms to identify behaviors, combining background difference method and target thing segmentation technology for target recognition and tracking, dynamically adjusting camera parameters to achieve image stability, solving the problem of focusing on target recognition and picture stability in pet video processing, improving video quality and user experience.

CN119967289BActive Publication Date: 2025-06-24SINGULARITY INTERNET OF THINGS CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510443434.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-10
Publication Date
2025-06-24
Estimated Expiration
2045-04-10

AI Technical Summary

Technical Problem

The prior art is difficult to effectively identify and track pet's attention targets in video processing, resulting in video images rotating, translating and tilting, affecting clarity and stability.

Method used

By collecting pet motion data, using machine learning algorithms to identify pet behaviors, and combining background difference method and target thing segmentation technology, a target thing recognition model is built for tracking. At the same time, an adaptive viewing angle optimization algorithm is built to dynamically adjust the camera viewing angle, focal length and frame rate to achieve image stability and dynamic adjustment.

Benefits of technology

It realizes accurate identification and tracking of pet attention targets, corrects the rotation, translation and tilt of the video screen, improves the clarity and stability of the video, and enhances the diversity and fun of the shooting effects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119967289B_ABST
    Figure CN119967289B_ABST
Patent Text Reader

Abstract

This application relates to the technical field of pet monitoring, specifically to an adaptive camera tracking method and system based on the pet's perspective. The method includes: collecting the pet's motion data, using machine learning algorithms to identify pet behaviors, and setting the camera working state according to the behaviors; constructing a target object recognition model to track the target objects that the pet is interested in and exclude dynamic interferences in the environment; inserting the pet's favorite objects into the target object recognition model and determining whether the pet enters the focused mode; constructing an adaptive perspective optimization algorithm to automatically adjust the camera parameters according to the pet's behaviors and the target objects of interest; performing adaptive image stabilization and dynamic adjustment on the camera video data, and dynamically adjusting the picture cropping area; shooting the pet's activity process from a third-party perspective by an unmanned aerial vehicle. This application provides an intelligent and personalized pet monitoring solution, which can comprehensively record the pet's real behaviors and provide data support for pet health management and human-pet interaction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of pet monitoring, specifically to an adaptive camera tracking method and system based on the pet's perspective. Background Art

[0002] Currently, intelligent devices for pets mainly focus on the monitoring of motion data. For example, motion data of pets is collected through sensors such as accelerometers and gyroscopes, and then their motion states are analyzed. However, these devices are usually limited to simple direction detection or activity recording, lacking in-depth analysis of pet behavior patterns and being unable to effectively identify specific pet behaviors.

[0003] The following difficulties exist in the prior art in video processing: When a pet is moving, the rapid swinging and jittering of its head will cause problems such as rotation, translation, and tilt in the video image, affecting the clarity and stability of the video. Pet behavior is closely related to the objects they are interested in, but the prior art lacks the ability to combine motion data and video data for target object recognition and tracking, and cannot accurately determine the objects that pets are interested in.

[0004] In view of this, this application proposes an adaptive camera tracking method based on the pet's perspective. Summary of the Invention

[0005] To achieve the above object, this application provides an adaptive camera tracking method based on the pet's perspective, and the specific technical solutions are as follows:

[0006] Collect the motion data of the pet, including acceleration data, angular velocity data, and heart rate data when the pet is moving, and construct the collected motion data into a data set;

[0007] Based on the motion data of the pet, use machine learning algorithms to identify pet behavior and set the camera working state according to the pet behavior;

[0008] Construct a target object recognition model to identify the target objects that the pet is interested in, and combine the background difference method with the target object segmentation technology to track the target objects;

[0009] Insert the pet's favorite objects into the target object recognition model. When the pet's favorite objects are recognized, determine whether the pet enters the focused mode;

[0010] Construct an adaptive perspective optimization algorithm to automatically adjust the camera perspective, focal length, and frame rate according to the recognized pet behavior and the target objects that the pet is interested in in the activity scene;

[0011] Perform adaptive image stabilization and dynamic adjustment on the video data captured by the camera, correct the rotation, translation, and tilt of the video image according to the pet behavior, and dynamically adjust the picture cropping area;

[0012] Build a third - party perspective tracking of the pet's movement process by shooting the pet's activity process with a drone.

[0013] Preferably, collect the pet's acceleration data, angular velocity data, and real - time heart rate data;

[0014] Define the pet movement data set as a time - series data set. Each data sample includes pet movement data and a behavior label, and each data sample contains multiple time steps.

[0015] Preferably, divide the time - series data set into a training set, a validation set, and a test set; normalize the data in the training set, validation set, and test set, and scale the data;

[0016] Build an LSTM model for pet behavior recognition; the input of the LSTM model is the normalized training set data, and the output is the corresponding behavior sequence;

[0017] The architecture of the LSTM model includes an input layer, an LSTM layer, a fully - connected layer, and an output layer. The input layer is used to input the normalized training set data into the LSTM model; the LSTM layer is used to perform feature extraction and temporal modeling on the input sequence using multiple LSTM layers; the fully - connected layer is used to map the hidden state to obtain the probability of the behavior category for each time step; the output layer is used to output the category prediction of the behavior.

[0018] Use the training set to train the LSTM model and optimize the model parameters to minimize the prediction loss;

[0019] Evaluate the performance of the trained LSTM model on the test set and calculate the evaluation metrics; integrate the trained LSTM model into the camera system to predict pet behavior in real - time, and adjust the camera working state according to the pet behavior recognition result and the target object recognition result.

[0020] Preferably, build a target object recognition model based on the YOLO model. The input data is the pet - perspective video frame data, and the target objects are detected in each frame of the image; the output is a set of bounding boxes of the target objects and a set of target object categories;

[0021] Calculate the pixel difference between the current frame and the background reference frame to generate a difference mask for extracting the dynamic region; where the background reference frame is generated from the static scene frame at the differential time;

[0022] Use a target object segmentation model based on the MaskR - CNN model to generate a target object segmentation mask in the video frame to separate the target object region from the background;

[0023] Combine the bounding box of the target object and the target object segmentation mask to filter out the target object with the largest overlap with the difference mask, which is used as the current focus target object of the pet.

[0024] Preferably, use the DeepSORT tracking algorithm to track the pet's focus target object across frames to generate the target object trajectory;

[0025] Filter the detected set of target object categories to exclude dynamic interference target objects that have nothing to do with the pet's behavior.

[0026] Preferably, build a pet preference object database. By analyzing the pet's daily activity videos and image data, identify the object categories that the pet focuses on and interacts with as the pet preference object categories, and import the pet preference object categories into the target object recognition model;

[0027] Create a concentration analysis strategy. When a pet preference object is detected in a video frame, trigger the concentration analysis process:

[0028] Calculate the pet's head orientation through head pose calculation technology;

[0029] Track the positional relationship between the pet's head orientation and the preference object in multiple consecutive video frames, and count the duration that the pet's head orientation falls within the range of the preference object;

[0030] Calculate the comprehensive concentration score of the pet during the period of gazing at the preference object, set a concentration score threshold, and determine whether the pet enters the concentration mode according to the comprehensive concentration score.

[0031] Preferably, when the pet enters the concentration mode, centered on the pet's focus preference object, dynamically calculate the size of the cropping area according to the size of the preference object and the distance to the camera, and enlarge the cropped area image;

[0032] During the cropping process, continuously track the position of the pet preference object and dynamically adjust the center of the cropping area to keep the preference object always in the center of the screen;

[0033] At the same time, continuously track and analyze the change of the pet's concentration, and set the lower and upper limits of the continuous concentration duration; if the continuous concentration duration of the pet is less than the lower limit of the concentration duration, smoothly restore the normal view; if the continuous concentration duration is greater than the upper limit of the concentration duration, expand the cropping area;

[0034] If the pet does not enter the concentration mode, optimize the image content through an adaptive perspective optimization algorithm.

[0035] Preferably, according to the identified pet behavior and the obtained pet perspective video frame data; build an adaptive perspective optimization algorithm, and dynamically adjust the camera shooting parameters according to the pet behavior and the pet perspective video frame data, including the camera view range, focal length, and frame rate;

[0036] Build an adaptive parameter adjustment strategy: Set the adaptive adjustment range of the camera shooting parameters according to the pet's behavior; Build an adaptive perspective optimization algorithm based on the pet's motion data and behavior to dynamically adjust the camera's perspective range;

[0037] Lock the perspective of the target object, update the camera's perspective direction in combination with the position of the target object; The focal length is dynamically adjusted according to the distance of the target object; The frame rate is dynamically adjusted according to the pet's behavior and angular velocity; Combine the perspective range, focal length, and frame rate to output the camera shooting parameters.

[0038] Preferably, build an adaptive image stabilization algorithm. The image stabilization corrects the rotation, translation, and tilt of the video frame through electronic image stabilization technology in combination with motion data;

[0039] Calculate the pet's motion trajectory. According to the angular velocity data of the gyroscope, calculate the angular offset at the time step of the time step;

[0040] According to the acceleration data, calculate the translation offset of the video frame; Build an image transformation matrix based on the angular offset and translation offset for correcting the rotation and translation of the image; Apply the image transformation matrix to the video frame to obtain the stabilized image;

[0041] Calculate the jitter amplitude of the video frame according to the standard deviation of the acceleration and the standard deviation of the angular velocity; Dynamically adjust the size of the cropping area according to the jitter amplitude, and the cropping range shrinks as the jitter amplitude increases; Calculate the cropping center according to the current position of the target object of the pet and the position of the camera; Apply cropping to the stabilized image to output the finally adjusted image and output the video frame after data stabilization.

[0042] Preferably, deploy a drone in the pet's activity area. Using the pet as the shooting fiducial point, obtain the pet's position information in real time through the visual sensor, and autonomously adjust the position and angle of the drone in combination with the pet behavior recognition result. Taking the pet as the shooting center, obtain the pet activity video from a third-party perspective.

[0043] The pet perspective adaptive camera tracking system is used to implement the pet perspective adaptive camera tracking method described above, including: a data acquisition module, a pet behavior recognition module, a target recognition and tracking module, a focus mode module, a perspective optimization module, a picture adjustment module, and a third-party perspective shooting module;

[0044] The data acquisition module is used to collect the pet's motion data, including the acceleration data, angular velocity data, and heart rate data when the pet is moving, and construct the collected motion data into a data set;

[0045] The pet behavior recognition module recognizes the pet behavior based on the pet's motion data using a machine learning algorithm and sets the camera working state according to the pet's behavior;

[0046] The target recognition and tracking module constructs a target object recognition model, identifies the target object of the pet's attention, and tracks the target object by combining the background difference method and the target object segmentation technology;

[0047] The focus mode module inserts the pet's favorite object into the target object recognition model, and when the pet's favorite object is recognized, determines whether the pet enters the focus mode;

[0048] The viewing angle optimization module is used to construct an adaptive viewing angle optimization algorithm to automatically adjust the camera viewing angle, focal length and frame rate according to the recognized pet behavior and the target objects of the pet's attention in the activity scene;

[0049] The picture adjustment module performs adaptive image stabilization and dynamic adjustment on the video data captured by the camera, corrects the rotation, translation and tilt of the video picture according to the pet's behavior, and dynamically adjusts the picture cropping area;

[0050] The third-party perspective shooting module is used to build third-party perspective tracking of the pet's movement process and shoot the pet's activity process through a drone.

[0051] Beneficial effects of this application: This application provides high-quality data support for pet behavior analysis and machine learning model training by providing comprehensive motion and physiological data.

[0052] This application intelligently adjusts the camera's working status according to pet behavior, achieves targeted shooting, and improves the relevance and practicality of video content.

[0053] This application can accurately identify the target objects that the pet is paying attention to, eliminate environmental interference, ensure the main body of the picture is clear, and improve the accuracy of the captured picture.

[0054] This application describes the pet's concentration state in a quantitative way, and dynamically adjusts the camera shooting parameters accordingly to obtain a more intelligent and personalized pet perspective tracking effect.

[0055] This application dynamically adjusts the camera's viewing angle, focal length, and frame rate to adapt to pet behavior and scene changes, ensuring the rationality and quality of the captured image composition.

[0056] This application improves the stability and visual effects of the video image by correcting the rotation, translation and tilt of the image and dynamically cropping the image area.

[0057] This application uses drones to record the pet's movement process in real time, providing a global perspective and enhancing the diversity and fun of the shooting effects. BRIEF DESCRIPTION OF THE DRAWINGS

[0058] Figure 1 Flow chart of the pet - perspective - based adaptive camera tracking method provided for this application;

[0059] Figure 2 Pet behavior recognition and camera working state adjustment diagram of the pet - perspective - based adaptive camera tracking method provided for this application;

[0060] Figure 3 Pet focus mode determination and image cropping diagram of the pet - perspective - based adaptive camera tracking method provided for this application;

[0061] Figure 4 Schematic diagram of the field - of - view determination strategy of the pet - perspective - based adaptive camera tracking method provided for this application;

[0062] Figure 5 Structural diagram of the pet - perspective - based adaptive camera tracking system provided for this application. Detailed implementation manners

[0063] To better understand this application, more detailed descriptions of various aspects of this application will be made with reference to the accompanying drawings. It should be understood that these detailed descriptions are only descriptions of the exemplary embodiments of this application and do not limit the scope of this application in any way. Throughout the specification, the same reference numerals refer to the same elements. The expression “and / or” includes any and all combinations of one or more of the associated listed items.

[0064] In the drawings, for ease of illustration, the sizes, dimensions, and shapes of the elements have been slightly adjusted. The drawings are only examples and are not drawn to an exact scale. As used herein, terms such as “substantially,” “approximately,” and similar terms are used as terms of approximation and not as terms of degree, and are intended to account for the inherent deviations in measured or calculated values that would be recognized by a person of ordinary skill in the art. Additionally, in this application, the order in which the steps are described does not necessarily represent the order in which these processes occur in actual operation, unless otherwise clearly specified or can be deduced from the context.

[0065] It should also be understood that expressions such as "including", "comprising", "having", "containing" and / or "comprising of" in this specification are open rather than closed expressions, which mean that the stated features, elements and / or components exist, but do not exclude the existence of one or more other features, elements, components and / or their combinations. In addition, when an expression such as "at least one of..." appears after a list of listed features, it modifies the entire list of features, rather than just individual elements in the list. In addition, when describing the embodiments of the present application, the use of "may" means "one or more embodiments of the present application". And the term "exemplary" is intended to refer to an example or illustration.

[0066] Unless otherwise defined, all terms used herein (including engineering terms and technical terms) have the same meaning as commonly understood by those of ordinary skill in the art to which this application belongs. It should also be understood that unless there is a clear statement in this application, words defined in common dictionaries should be interpreted as having a meaning consistent with their meaning in the context of the related art, and should not be interpreted in an idealized or overly formal sense.

[0067] It should be noted that, without conflict, the embodiments in this application and the features in the embodiments can be combined with each other. The present application will be described in detail below with reference to the drawings and in combination with the embodiments.

[0068] Embodiment 1

[0069] Refer to Figures 1 to 4 , which is the first embodiment of the present application, and provides a pet perspective adaptive camera tracking method.

[0070] Step 1: Collect the motion data of the pet, including the acceleration data, angular velocity data, and heart rate data of the pet during motion, and construct the collected motion data into a data set.

[0071] Through the acceleration sensor, collect the acceleration values of the pet on the coordinate axes , with the unit of , expressed as ; through the gyroscope sensor, collect the angular velocity values of the pet on the coordinate axes , with the unit of , expressed as ; through the heart rate monitor, collect the real-time heart rate value of the pet, with the unit of (times / minute), expressed as .

[0072] Exemplarily, a method for collecting pet acceleration: By selecting a three-axis accelerometer, such as ADXL345, with a measurement range of ±16g, a resolution of 13 bits, and a sampling rate of 100Hz; fixing the acceleration sensor on a pet collar or harness so that its x-axis, y-axis, and z-axis are aligned with the forward direction, left-right direction, and up-down direction of the pet respectively.

[0073] Exemplarily, a method for collecting pet angular velocity: By selecting a three-axis gyroscope, such as ITG-3200, with a measurement range of ±2000° / s, a resolution of 16 bits, and a sampling rate of 100Hz; integrating the gyroscope and the acceleration sensor in the same sensor module and ensuring that their coordinate systems are consistent.

[0074] Exemplarily, a method for detecting pet heart rate: Selecting a photoplethysmogram sensor, such as MAX30100, with a sampling rate of 100Hz; wearing the heart rate monitor on the pet's chest or the inner side of the ear and calculating the heart rate by measuring the absorption difference between red light and infrared light in the blood.

[0075] Define the pet motion dataset as a time series dataset , consisting of data samples, each data sample including pet motion data and a behavior label, and each data sample containing time steps, , where represents the motion data of the Nth data sample, represents the behavior label of the Nth data sample.

[0076] For the data sample , define:

[0077] ;

[0078] where represents the data of the th sample at time step , represents the behavior label of the th sample at time step , , , is the total number of behavior labels.

[0079] contains the pet's motion data at time step :

[0080] ;

[0081] where respectively represent the th sample at time step Acceleration data on the x-axis, y-axis, and z-axis at time respectively represent the th sample at time step Angular velocity data on the x-axis, y-axis, and z-axis at time represents the th sample at time step Heart rate data.

[0082] The definition of time step is: ;

[0083] where is the start time of the data sample, is the time step size, is the time step number index, is the total number of time steps.

[0084] Through the above mathematical formula, pet movement data, behavior labels, and time are constructed into a time series dataset .

[0085] The time series dataset constructed in Step 1 can be conveniently applied to various time series analysis and prediction tasks, such as pet behavior recognition and anomaly detection, etc.; at the same time, this mathematical representation also helps to understand the structure and characteristics of the dataset, providing a clear idea for the design and implementation of subsequent pet behavior recognition algorithms.

[0086] Step 2: Based on the pet's movement data, use machine learning algorithms to identify pet behaviors and set the camera working state according to the pet's behaviors.

[0087] Divide the time series dataset into a training set , a validation set and a test set , where the training set is used for model training, the validation set is used for model hyperparameter selection and early stopping strategy, and the test set is used for evaluating model performance.

[0088] Normalize the data of the training set, validation set, and test set, and scale the data to the range [0, 1]; this helps to accelerate the model convergence speed and improve numerical stability; the normalization formula is: ; where is the normalized input data, and are the minimum and maximum values of the input data respectively.

[0089] Build an LSTM model for pet behavior recognition; the input of the LSTM model is the normalized data , the output is the corresponding behavior sequence , where represents the th sample's normalized data at time step , is the th sample's behavior class probability at time step .

[0090] The architecture of the LSTM model includes an input layer, an LSTM layer, a fully connected layer, and an output layer. The input layer includes: inputting the data sequence into the LSTM model; the LSTM layer includes: using multiple LSTM layers to perform feature extraction and temporal modeling on the input data sequence; the fully connected layer includes: mapping the hidden state to obtain the behavior class probability at each time step; the output layer includes: outputting the class prediction of the behavior.

[0091] The output dimension of each LSTM layer is , representing the dimension of the hidden state;

[0092] The calculation formula of the LSTM layer is:

[0093]

[0094]

[0095]

[0096]

[0097]

[0098]

[0099] Among them, is the input gate, is the forget gate, is the output gate, is the current input, and are the weight matrices of the input gate, forget gate, output gate, and candidate cell state respectively, and are the bias terms of the input gate, forget gate, output gate, and candidate cell state respectively, is the candidate memory cell state, is the memory cell state, represents the cell state at the previous time step, is the hidden state, is the Sigmoid activation function, is the hyperbolic tangent activation function, is the element-wise multiplication.

[0100] The calculation formula for the fully connected layer is:

[0101]

[0102] where, is the probability of the behavior category of the th sample at time step , is the hidden state of the th sample at time step ; is the weight matrix of the fully connected layer, is the bias term; is the activation function that converts the output into a probability distribution.

[0103] Use the training set to train the LSTM model and optimize the model parameters to minimize the prediction loss; use the cross-entropy loss as the loss function of the LSTM model, and the calculation formula is:

[0104]

[0105] where, are the model parameters, is the pet behavior of the th sample at time step , is the probability of the behavior category of the th sample predicted by the model at time step .

[0106] During the training process, use optimization algorithms (such as Adam, RMSprop, etc.) to update the model parameters to minimize the loss function.

[0107] By monitoring the performance on the validation set , adjust the model hyperparameters (such as the hidden state dimension, number of layers, learning rate, etc.), and adopt an early stopping strategy to prevent overfitting.

[0108] Evaluate the performance of the trained LSTM model on the test set , and calculate evaluation metrics such as accuracy, precision, recall, and F1 score; a confusion matrix can be used to visualize the classification performance of the model on different behaviors.

[0109] Integrate the trained LSTM model into the camera system to predict pet behaviors in real time, and adjust the working state of the camera according to the pet behavior recognition results and the target object recognition results.

[0110] In Step 2, through steps such as data preparation, model construction, training optimization, and evaluation, a behavior recognition model with excellent performance can be obtained; applying this model to the camera system can achieve adaptive camera tracking shooting from the pet's perspective, improving the quality of video data and the user experience.

[0111] Step 3: Construct a target object recognition model to recognize the target objects that the pet is interested in, and combine the background subtraction method with the target object segmentation technology to track the target objects.

[0112] The target object recognition model aims to extract the target objects that the pet is interested in (such as toys, food, other animals, or humans, etc.) from the video frame data from the pet's perspective and achieve target object tracking by combining background subtraction and target object segmentation technology.

[0113] Construct a target object recognition model based on the YOLO model, and the input data is the video frame data from the pet's perspective , representing a single-frame image captured by the camera at time step t; detect the target objects in each frame of the image; the input of the target object recognition model is , and the output is the set of bounding boxes of the target objects and the set of target object categories :

[0114]

[0115] Among them, is the bounding box of the target object at time step t , representing the top-left coordinates and the width and height ; is the category of the target object at time step t (such as toy, food, human, etc.); is the number of detected target objects, is the target object index, .

[0116] Calculate the pixel difference between the current frame and the background reference frame to generate a difference mask for extracting the dynamic region; among them, ; among them, the background reference frame can be generated from the static scene frames at time step t.

[0117] Use a target object segmentation model based on the MaskR-CNN model to generate a target object segmentation mask in the current frame , separate the target object area from the background: ; where is the binary segmentation mask of the target object at time step t ; the difference mask generates a binary dynamic region map after threshold segmentation , where .

[0118] Combine the set of bounding boxes of the target object and the target object segmentation mask , and filter out the target object with the largest overlap with the difference mask as the current attention target object of the pet : ; where represents the intersection over union of the bounding box of the target object and the binary dynamic region map.

[0119] Use the DeepSORT tracking algorithm to track the pet attention target object across frames and generate the target object trajectory : ; where represents the 3D position of the target object at time step .

[0120] Filter the detected set of target object categories to exclude dynamic interference target objects (such as other pets or passersby) unrelated to the pet behavior .

[0121] Exemplarily, the exclusion rules are as follows: If indicates that the pet is recognized as in an interactive state, preferentially retain target objects with categories of toys or humans; if indicates that the pet is recognized as in an exploratory state, preferentially retain target objects with categories of other animals or food.

[0122] Output the bounding box and category of the pet's current frame attention target object; the segmentation mask of the pet attention target object.

[0123] According to the recognized pet behavior and the acquired pet perspective video frame data , represents a single-frame image at time step ; construct an adaptive perspective optimization algorithm according to the pet behavior and the pet perspective video frame data , dynamically adjust the camera shooting parameters, including the camera viewing angle range, focal length, and frame rate; the viewing angle range (Field of View, FOV) is used to adjust the shooting range of the camera; the focal length (Focal Length, ) is used to adjust the clarity of the target object; the frame rate (Frame Rate, FPS) is used to adjust the sampling frequency of the video.

[0124] Build an adaptive parameter adjustment strategy: According to the pet's behavior , set the adaptive adjustment range of the camera shooting parameters.

[0125] Exemplarily, assume , indicating that the pet is recognized as in a stationary state, narrow the viewing angle range, focus on the current observation area of the pet; increase the focal length , highlight the area of interest; reduce the frame rate to , reduce power consumption;

[0126] Assume , indicating that the pet is recognized as in an interactive state, expand the viewing angle range, capture the movement path of the pet; keep the default focal length to avoid frequent adjustment; increase the frame rate to , reduce motion blur;

[0127] Assume , indicating that the pet is recognized as in a focused state, adjust the viewing angle range to lock the target object (such as a toy, a human) with which the pet is interacting; dynamically adjust the focal length to ensure the clarity of the target object; set the frame rate to , balance clarity and power consumption;

[0128] Assume 3, indicating that the pet is recognized as in an exploratory state, quickly switch the viewing angle direction, and give priority to capturing the area of the target object explored by the pet; dynamically adjust the focal length and automatically zoom according to the distance of the target object; set the frame rate to , ensure the smoothness of the picture.

[0129] Step 4: Insert the pet's favorite object into the target object recognition model. When the pet's favorite object is recognized, determine whether the pet enters the focused mode.

[0130] Build a pet's favorite object database. By analyzing the pet's daily activity video and image data, identify the object categories that the pet pays attention to and interacts with, such as pet toys, food, etc., as the pet's favorite object categories, establish pet's favorite object category labels, and import the pet's favorite object categories into the target object recognition model.

[0131] Create a focus analysis strategy. When a pet's favorite object is detected in a video frame, trigger the focus analysis process.

[0132] Through the head pose estimation technology, the head orientation of the pet is calculated in real time;

[0133] Specifically, in each frame of video image, the pet's head region is detected and located, head feature points are extracted, and a three-dimensional head pose model is constructed; Let the unit vector of the pet's head orientation be , and the unit vector of the object position be , then calculate the included angle of the vectors: ; where represents the dot product of the unit vector of the pet's head orientation and the unit vector of the object position, and respectively represent the magnitudes of the unit vector of the pet's head orientation and the unit vector of the object position; If the included angle is less than the preset threshold (such as 30°, i.e., ), then it is preliminarily determined that the pet is staring at the object.

[0134] In consecutive multiple frames of video, track the positional relationship between the pet's head orientation and the favorite object, and count the duration during which the pet's head orientation falls within the range of the favorite object; Let the included angle between the pet's head orientation vector and the object position vector in the th frame of video be , and the frame rate be , then the duration during which the pet stares at the object within consecutive frames is: ; where is an indicator function that takes the value of 1 when the condition is met and 0 otherwise; If the staring duration exceeds the preset threshold (such as 3 seconds), then further verify the pet's degree of concentration on the object.

[0135] Calculate the comprehensive concentration score of the pet during the period of staring at the favorite object, taking into account both the staring duration and the included angle size. The formula for the comprehensive concentration score is:

[0136]

[0137] where is the average value of the included angles between the pet's head orientation vector and the object position vector in each video frame during the staring period, and are weight coefficients, satisfying ; The longer the staring duration and the smaller the average included angle, the higher the comprehensive concentration score.

[0138] Set the concentration score threshold and (such as 0.8 and 0.2); if , it is determined that the pet enters the focused state and triggers the cropping function of the focused mode; if , it is determined that the pet does not show obvious focus and the conventional cropping method is maintained; if , the observation time is extended, and whether to switch the focused mode is determined according to the field of view determination strategy.

[0139] The field of view determination strategy specifically includes:

[0140] When the preferred object is within the pet's field of view, the first timing is triggered and the first timing is recorded as . If the preferred object gradually moves from the edge of the field of view to the center of the field of view, after the preferred object enters the center of the field of view, the second timing is triggered and the second timing is recorded as . If the time interval between the two times is less than or equal to the interval setting threshold , that is , and the duration of the preferred object in the center of the field of view is greater than the duration setting threshold, it is determined as the focused mode. At the same time, the interval setting threshold and the duration setting threshold are calibrated according to the average value of the pet's historical data. The advantage of the field of view determination strategy is simple calculation and high accuracy. Refer to Figure 4 , which is the schematic diagram of the field of view determination strategy provided by this application.

[0141] In the focused mode, centered on the preferred object that the pet is concerned about, according to the size of the preferred object and the distance to the camera, the size of the cropping area is dynamically calculated, and the image of the cropping area is smoothly enlarged; specifically, it includes: Let the pixel size of the preferred object in the image be , the distance from the preferred object to the camera be , the total size of the image be , then the size of the cropping area can be calculated by the following formula:

[0142]

[0143] Among them, is the reference distance (such as 1 meter), and are the cropping coefficients, which control the magnification factor of the cropping area relative to the size of the preferred object; the cropping area should satisfy and . When necessary, the value ranges of and can be appropriately restricted; during the cropping process, continuously track the position of the pet's preferred object and dynamically adjust the center of the cropping area to keep the preferred object always in the center of the screen.

[0144] At the same time, continuously track and analyze the change of the pet's focus degree, and set the lower limit of the continuous focus duration (such as 5 seconds) and the upper limit (such as 30 seconds); if the duration of the pet's concentration is below, then smoothly restore the normal perspective; if the duration of the pet's concentration is above, then appropriately expand the cropping area to avoid visual fatigue caused by long-term high focus; and set the minimum interval duration of the focus mode to avoid overly frequent perspective switching from affecting the video coherence.

[0145] In this step, by constructing an analysis of the pet's concentration and an adaptive cropping method in the focus mode, the focus state of the pet is characterized in a quantitative manner, and the camera shooting parameters are dynamically adjusted accordingly to obtain a more intelligent and personalized pet perspective tracking effect.

[0146] Step 5: Construct an adaptive perspective optimization algorithm to automatically adjust the camera perspective, focal length, and frame rate according to the recognized pet behaviors and the pet's attention target objects in the activity scene;

[0147] According to the pet movement data and pet behaviors , construct an adaptive perspective optimization algorithm to adjust the camera's perspective range.

[0148] Define the camera perspective range as and , is the horizontal perspective, is the vertical perspective, and dynamically adjust the camera perspective range according to the movement data and pet behaviors:

[0149]

[0150] Among them, and are the initial camera perspective ranges, and are the adaptive adjustment amounts for the pet behavior : if , that is, the pet is in a stationary state, then , ; if , that is, the pet is in a moving state, then , .

[0151] Lock the perspective of the target object, and combine the position of the target object to update the camera's perspective direction :

[0152]

[0153] Among them, represents the camera position, Indicates calculating the Euclidean distance, Indicates the arctangent function.

[0154] Focal length According to the distance of the target object Dynamically adjust:

[0155]

[0156] Wherein, is the position of the target object, is the maximum recognizable distance of the target object, is the initial focal length, is the focal length scaling factor.

[0157] Frame rate According to the pet's behavior and the angular velocity Dynamically adjust:

[0158] Wherein, and are the maximum frame rate and the minimum frame rate respectively, is the maximum value of the angular velocity.

[0159] Combining the viewing angle range, focal length and frame rate, output the camera shooting parameters:

[0160] Wherein, represents the camera shooting parameters at the time step of.

[0161] This step combines the pet's behavior and motion data, dynamically adjusts the camera's viewing angle range, focal length and frame rate, and realizes adaptive viewing angle optimization; the introduction of the control formula ensures the precise control of viewing angle change, focal length adjustment and frame rate optimization, providing technical support for shooting high-quality videos.

[0162] Step 6: Perform adaptive image stabilization and dynamic adjustment on the camera shooting video data, correct the rotation, translation and tilt of the video picture according to the pet's behavior, and dynamically adjust the picture cropping area.

[0163] Construct an adaptive image stabilization algorithm, and the image stabilization is achieved through electronic image stabilization technology, combined with motion data to correct the rotation, translation and tilt of the video picture.

[0164] Calculate the motion trajectory of the pet, and calculate the angular offset at the time step according to the angular velocity data of the gyroscope .

[0165] When the angular offset is in continuous form, it is calculated using the following formula:

[0166] where is the angular velocity data;

[0167] When the angular offset is in discrete form, it is calculated using the following formula:

[0168]

[0169] Based on the acceleration data , the translation offset at time step is calculated.

[0170] When the translation offset is in continuous form, it is calculated using the following formula:

[0171]

[0172] When the translation offset is in discrete form, the translation amount is calculated using discrete integration:

[0173]

[0174] where is the acceleration data;

[0175] represents the translation offset in the three-axis directions.

[0176] Construct an image transformation matrix based on the angular offset and the translation offset for correcting the rotation and translation of the image; the image transformation matrix is:

[0177]

[0178] where, and are the rotation matrices around the axis, respectively, defined as:

[0179]

[0180]

[0181] where is the sine value of the rotation angle around the axis, is the rotation angle around the axis, For around Axis rotation angle.

[0182] Further time Video frame Apply image transformation matrix , and get the stabilized image ;

[0183] According to the standard deviation of the acceleration and the standard deviation of the angular velocity , calculate the jitter amplitude at time step t , according to the jitter amplitude ;

[0184]

[0185] in, ; and are the components of acceleration and angular velocity, respectively, and is its average value;

[0186] Dynamically resize the crop area , the cropping range decreases as the jitter amplitude increases, ensuring the stability of the video image;

[0187] in, and is the width and height of the original image, is the cropping amplitude adjustment coefficient, and is the width and height of the cropping area;

[0188] Based on the pet's current target location and camera position , calculate the size of the clipping area ; ;

[0189] The image after stabilization Apply cropping to the output, output the final adjusted image, output data stabilized video frame : , Represents a crop operation.

[0190] Step 5: Through real-time analysis of pet motion data and combined with electronic image stabilization technology, the rotation, translation and tilt of the video image are dynamically corrected, the cropping area and image quality are optimized, and the stability and clarity of the video are ensured.

[0191] Step 6: Construct a third - party perspective tracking of the pet's movement process by using a drone to film the pet's activity process.

[0192] Calculate the real - time shooting center position. Taking the current position of the pet as the fiducial point and combining with the position of the target object, calculate the shooting center position of the drone:

[0193]

[0194] Among them, represents the time step when the shooting center position of the drone; represents the time step when the position of the pet; represents the shooting offset, which is used to adjust the relative position between the shooting center and the pet.

[0195] Adjust the flight altitude of the drone. According to the pet's motion state, dynamically adjust the flight altitude of the drone:

[0196]

[0197] Among them, represents the time step when the flight altitude of the drone; represents the initial flight altitude of the drone; represents the altitude adjustment coefficient; represents the time step when the magnitude of the pet's speed (obtained by integrating the acceleration).

[0198] Calculate the viewing direction of the drone. According to the positions of the pet and the target object, adjust the viewing direction of the drone camera in real - time:

[0199]

[0200] Among them, represents the time step when the pitch angle of the drone camera; represents the time step when the position of the target object; represents the time step when the position of the pet; represents the time step when the flight altitude of the drone.

[0201] In this step, by taking the pet as the fiducial point and dynamically adjusting the flight altitude, viewing direction and focus in combination with the position of the target object, it is ensured that the pet and the target object are always in the core area of the picture; this process can adapt to the pet's motion state in real - time, improve the accuracy and stability of shooting, and at the same time optimize the picture parameters to adapt to multi - scene environments and generate high - quality video content.

[0202] Example 2

[0203] Reference Figure 5 , which is the second embodiment of this application, provides an adaptive camera tracking system based on the perspective of pets.

[0204] The system includes: a data acquisition module, a pet behavior recognition module, a target recognition and tracking module, a focus mode module, a perspective optimization module, a picture adjustment module, and a third-party perspective shooting module.

[0205] The data acquisition module is used to collect the motion data of the pet, including the acceleration data, angular velocity data, and heart rate data when the pet is moving, and construct the collected motion data into a data set.

[0206] The pet behavior recognition module, based on the motion data of the pet, uses machine learning algorithms to recognize pet behaviors and sets the camera working state according to the pet behaviors.

[0207] The target recognition and tracking module constructs a target object recognition model to recognize the pet's attention target object, and combines the background difference method and the target object segmentation technology to track the target object.

[0208] The focus mode module inserts the pet's favorite objects into the target object recognition model, and when the pet's favorite objects are recognized, it determines whether the pet enters the focus mode.

[0209] The perspective optimization module is used to construct an adaptive perspective optimization algorithm, and automatically adjusts the camera perspective, focal length, and frame rate according to the recognized pet behaviors and the pet's attention target object in the activity scene.

[0210] The picture adjustment module performs adaptive image stabilization and dynamic adjustment on the video data captured by the camera, corrects the rotation, translation, and tilt of the video picture according to the pet behaviors, and dynamically adjusts the picture cropping area.

[0211] The third-party perspective shooting module is used to construct a third-party perspective tracking during the pet's movement, and shoots the pet's activity process through a drone.

[0212] The above order of the steps for the method is only for illustration, and the steps of the method of this application are not limited to the above specifically described order, unless otherwise specifically stated.

[0213] In addition, in some embodiments, the present application can also be implemented as a program recorded in a recording medium, and these programs include machine-readable instructions for implementing the method according to the present application. Therefore, the present application also covers a recording medium storing a program for executing the method according to the present application.

[0214] In addition, the parts of the above technical solutions provided in the embodiments of the present application that are consistent with the corresponding technical solutions in the prior art in terms of implementation principles are not described in detail to avoid excessive elaboration.

[0215] As described above in the specific embodiments, the purpose, technical solutions, and beneficial effects of the present application have been further described in detail. It should be understood that the above are only specific embodiments of the present application and are not used to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application shall be included within the protection scope of the present application.

Claims

1. A pet-viewing adaptive camera tracking method, characterized in that: include: Collecting the pet's motion data, including acceleration data, angular velocity data, and heart rate data when the pet is in motion, and constructing the collected motion data into a data set; Based on the pet's motion data, the machine learning algorithm is used to identify the pet's behavior and set the camera's working state according to the pet's behavior; Build a target object recognition model to identify the target objects that the pet is paying attention to, and track the target objects by combining background difference method and target object segmentation technology; Insert the pet's favorite object into the target object recognition model, and when the pet's favorite object is recognized, determine whether the pet enters the focus mode; When the pet enters the focus mode, the pet's favorite object is the center, and the size of the cropping area is dynamically calculated according to the size of the favorite object and the distance to the camera, and the cropping area image is enlarged; During the cropping process, the position of the pet's favorite object is continuously tracked, and the center of the cropping area is dynamically adjusted to keep the favorite object always in the center of the picture; At the same time, it continuously tracks and analyzes the changes in the pet's concentration, and sets the lower and upper limits of the continuous concentration duration; If the pet's concentration duration is less than the lower limit of the concentration duration, the normal viewing angle will be restored smoothly; If the duration of concentration is longer than the upper limit of concentration duration, the cropping area will be expanded; If the pet does not enter the focus mode, the image content is optimized through the adaptive perspective optimization algorithm; Build an adaptive perspective optimization algorithm to automatically adjust the camera perspective, focal length, and frame rate based on the identified pet behavior and the pet's focus in the activity scene; Adaptive image stabilization and dynamic adjustment of the video data captured by the camera, correct the rotation, translation and tilt of the video screen according to the pet's behavior, and dynamically adjust the screen cropping area; Build third-party perspective tracking of the pet's movement process and use drones to capture the pet's activities.

2. The pet-viewing-based adaptive camera tracking method according to claim 1, characterized in that: Collect the acceleration data of the pet, collect the angular velocity data of the pet, and collect the real-time heart rate data of the pet; The pet motion dataset is defined as a time series dataset. Each data sample includes pet motion data and behavior labels, and each data sample contains multiple time steps.

3. The pet-viewing-based adaptive camera tracking method according to claim 2, characterized in that: Divide the time series data set into training set, validation set and test set; normalize the data of the training set, validation set and test set, and scale the data; Construct an LSTM model for pet behavior recognition; the input of the LSTM model is the normalized training set data, and the output is the corresponding behavior sequence; The architecture of the LSTM model includes an input layer, an LSTM layer, a fully connected layer, and an output layer. The input layer is used to input the normalized training set data into the LSTM model; the LSTM layer is used to extract features and perform time series modeling on the input sequence using multiple LSTM layers; the fully connected layer is used to map the hidden state to obtain the behavior category probability at each time step; the output layer is used to output the behavior category prediction; Use the training set to train the LSTM model and optimize the model parameters to minimize the prediction loss; Evaluate the performance of the trained LSTM model on the test set and calculate the evaluation indicators; integrate the trained LSTM model into the camera system to predict pet behavior in real time, and adjust the camera working status based on the pet behavior recognition results and target object recognition results.

4. The pet-viewing-based adaptive camera tracking method according to claim 3, characterized in that: Build a target object recognition model based on the YOLO model. The input data is the pet's perspective video frame data. The target object is detected in each frame image. The output is the target object's bounding box set and target object category set. Calculate the pixel difference between the current frame and the background reference frame to generate a difference mask for extracting dynamic areas; The background reference frame is generated by differential time static scene frames; Use the target object segmentation model based on the Mask R-CNN model to generate the target object segmentation mask in the video frame to separate the target object area from the background; Combine the bounding box of the target object and the segmentation mask of the target object, and filter out the target object with the largest overlap with the differential mask as the pet's current target object of attention.

5. The pet-viewing-based adaptive camera tracking method according to claim 4, characterized in that: Use the DeepSORT tracking algorithm to track the target object that the pet is paying attention to across frames and generate the trajectory of the target object; The detected target object category set is filtered to exclude dynamic interference target objects that are not related to the pet's behavior.

6. The pet-viewing-based adaptive camera tracking method according to claim 5, characterized in that: Build a database of pets' favorite objects. By analyzing the videos and images of pets' daily activities, identify the categories of objects that pets pay attention to and interact with, and use them as pets' favorite object categories. Import the pets' favorite object categories into the target object recognition model. Create a concentration analysis strategy. When the pet's favorite object is detected in the video frame, the concentration analysis process is triggered: Calculate the pet's head orientation through head posture calculation technology; Track the positional relationship between the pet's head orientation and the favorite object in multiple consecutive video frames, and count the duration of the pet's head orientation falling within the range of the favorite object; The comprehensive concentration score of the pet during the period of looking at the favorite object is calculated, and the concentration score threshold is set to determine whether the pet enters the concentration mode based on the comprehensive concentration score.

7. The pet-viewing-based adaptive camera tracking method according to claim 6, characterized in that: Based on the identified pet behavior and the obtained pet-perspective video frame data; construct an adaptive perspective optimization algorithm to dynamically adjust the camera shooting parameters, including the camera perspective range, focal length and frame rate, based on the pet behavior and the pet-perspective video frame data; Build an adaptive parameter adjustment strategy: set the adaptive adjustment range of the camera shooting parameters according to the pet's behavior; build an adaptive viewing angle optimization algorithm based on the pet's motion data and behavior to dynamically adjust the camera's viewing angle range; Lock the viewing angle of the target object, and update the viewing direction of the camera based on the position of the target object; The focal length is dynamically adjusted according to the distance of the target object; the frame rate is dynamically adjusted according to the pet's behavior and angular velocity; the camera shooting parameters are output based on the viewing angle range, focal length and frame rate.

8. The pet-viewing-based adaptive camera tracking method according to claim 7, characterized in that: Build an adaptive image stabilization algorithm. Image stabilization uses electronic image stabilization technology and combines motion data to correct the rotation, translation and tilt of the video image. Calculate the pet's motion trajectory and calculate the time step based on the angular velocity data of the gyroscope Angular deviation when According to the acceleration data, the translation offset of the video frame is calculated; an image transformation matrix based on the angle offset and the translation offset is constructed to correct the rotation and translation of the image; Apply the image transformation matrix to the video frame to obtain a stabilized image; The jitter amplitude of the video frame is calculated based on the standard deviation of acceleration and the standard deviation of angular velocity; the size of the cropping area is dynamically adjusted based on the jitter amplitude, and the cropping range is reduced as the jitter amplitude increases; The cropping center is calculated based on the pet's current target object position and the camera position; cropping is applied to the stabilized image, and the final adjusted image is output, and the video frame after data stabilization is output.

9. The pet-viewing-based adaptive camera tracking method according to claim 8, characterized in that: Deploy drones in pet activity areas, use pets as shooting beacons, obtain pet location information in real time through visual sensors, and autonomously adjust the drone position and angle based on pet behavior recognition results, use pets as the shooting center, and obtain pet activity images from a third-party perspective.

10. A pet-perspective-based adaptive camera tracking system, which is used to implement the pet-perspective-based adaptive camera tracking method according to any one of claims 1 to 9, characterized in that: include: Data collection module, pet behavior recognition module, target recognition and tracking module, focus mode module, perspective optimization module, picture adjustment module and third-party perspective shooting module; The data acquisition module is used to collect the pet's motion data, including acceleration data, angular velocity data, and heart rate data when the pet is in motion, and construct the collected motion data into a data set; The pet behavior recognition module recognizes the pet behavior based on the pet's motion data using a machine learning algorithm and sets the camera working state according to the pet's behavior; The target recognition and tracking module constructs a target object recognition model, identifies the target object of the pet's attention, and tracks the target object by combining the background difference method and the target object segmentation technology; The focus mode module inserts the pet's favorite object into the target object recognition model, and when the pet's favorite object is recognized, determines whether the pet enters the focus mode; The viewing angle optimization module is used to construct an adaptive viewing angle optimization algorithm to automatically adjust the camera viewing angle, focal length and frame rate according to the recognized pet behavior and the target objects of the pet's attention in the activity scene; The picture adjustment module performs adaptive image stabilization and dynamic adjustment on the video data captured by the camera, corrects the rotation, translation and tilt of the video picture according to the pet's behavior, and dynamically adjusts the picture cropping area; The third-party perspective shooting module is used to build third-party perspective tracking of the pet's movement process and shoot the pet's activity process through a drone.

Citation Information

Patent Citations

  • Pet tracking method and device, XR equipment and medium

    CN117788765A

  • Remote communication alarm smart home system

    CN118865627A