Self-adaptive camera tracking method and system based on pet visual angle

By collecting pet motion data and using machine learning algorithms to identify pet behavior, combining background difference method and target thing segmentation technology, adaptively adjusting camera parameters, the problems of pet focus on target recognition and video stability in the existing technology are solved, and high-quality pet viewing video shooting is achieved.

CN119967289AActive Publication Date: 2025-05-09SINGULARITY INTERNET OF THINGS CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510443434.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-10
Publication Date
2025-05-09
Estimated Expiration
2045-04-10

AI Technical Summary

Technical Problem

The prior art is difficult to effectively identify and track pet's attention targets in video processing, resulting in video images rotating, translating and tilting, affecting clarity and stability.

Method used

By collecting pet motion data, using machine learning algorithms to identify pet behavior, building a target thing recognition model, combining background difference method and target thing segmentation technology, adaptively adjusting the camera viewing angle, focal length and frame rate to achieve accurate identification and tracking of pets' targets.

Benefits of technology

It realizes accurate identification and tracking of pet attention targets, corrects the rotation, translation and tilt of the video screen, improves the stability and clarity of the video, and enhances the diversity and fun of the shooting effects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119967289A_ABST
    Figure CN119967289A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of pet monitoring, in particular to an adaptive camera tracking method and system based on a pet visual angle. The method comprises the following steps: collecting motion data of a pet, identifying pet behaviors by using a machine learning algorithm, and setting a camera working state according to the behaviors; a target object identification model is constructed, target objects concerned by pets are tracked, and dynamic interference in the environment is eliminated; inserting pet favorite objects into the target object identification model, and judging whether the pet enters a concentration mode or not; constructing an adaptive visual angle optimization algorithm, and automatically adjusting camera parameters according to pet behaviors and concerned targets; carrying out adaptive image stabilization and dynamic adjustment on camera video data, and dynamically adjusting a picture cutting area; the pet activity process is shot through the unmanned aerial vehicle from a third-party perspective. The invention provides an intelligent and personalized pet monitoring scheme, real behaviors of the pet can be comprehensively recorded, and data support is provided for pet health management and human-pet interaction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of pet monitoring technology, and specifically to a pet-viewing-based adaptive camera tracking method and system. Background Art

[0002] At present, smart devices for pets mainly focus on motion data monitoring, such as collecting pet motion data through sensors such as accelerometers and gyroscopes, and then analyzing their motion status. However, these devices are usually limited to simple direction detection or activity recording, lack in-depth analysis of pet behavior patterns, and cannot effectively identify the specific behavior of pets.

[0003] The existing technology has the following difficulties in video processing: When a pet is moving, the rapid shaking and shaking of its head will cause the video screen to rotate, pan and tilt, affecting the clarity and stability of the video. The behavior of a pet is closely related to the target it pays attention to, but the existing technology lacks the ability to combine motion data and video data to identify and track the target object, and cannot accurately determine the target of the pet's attention.

[0004] In view of this, the present application proposes an adaptive camera tracking method based on the pet's perspective. Summary of the invention

[0005] To achieve the above purpose, this application provides a pet-perspective adaptive camera tracking method, and the specific technical solution is as follows:

[0006] Collecting the pet's motion data, including acceleration data, angular velocity data, and heart rate data when the pet is in motion, and constructing the collected motion data into a data set;

[0007] Based on the pet's motion data, the machine learning algorithm is used to identify the pet's behavior and set the camera's working state according to the pet's behavior;

[0008] Build a target object recognition model to identify the target objects that the pet is paying attention to, and track the target objects by combining background difference method and target object segmentation technology;

[0009] Insert the pet's favorite object into the target object recognition model, and when the pet's favorite object is recognized, determine whether the pet has entered the focus mode;

[0010] Build an adaptive perspective optimization algorithm to automatically adjust the camera perspective, focal length, and frame rate based on the identified pet behavior and the pet's focus in the activity scene;

[0011] Adaptive image stabilization and dynamic adjustment of the video data captured by the camera, correct the rotation, translation and tilt of the video screen according to the pet's behavior, and dynamically adjust the screen cropping area;

[0012] Build third-party perspective tracking of the pet's movement process and use drones to capture the pet's activities.

[0013] Preferably, the acceleration data of the pet, the angular velocity data of the pet, and the real-time heart rate data of the pet are collected;

[0014] The pet motion dataset is defined as a time series dataset. Each data sample includes pet motion data and behavior labels, and each data sample contains multiple time steps.

[0015] Preferably, the time series data set is divided into a training set, a validation set and a test set; the data of the training set, the validation set and the test set are normalized and scaled;

[0016] Construct an LSTM model for pet behavior recognition; the input of the LSTM model is the normalized training set data, and the output is the corresponding behavior sequence;

[0017] The architecture of the LSTM model includes an input layer, an LSTM layer, a fully connected layer, and an output layer. The input layer is used to input the normalized training set data into the LSTM model; the LSTM layer is used to extract features and perform time series modeling on the input sequence using multiple LSTM layers; the fully connected layer is used to map the hidden state to obtain the behavior category probability at each time step; the output layer is used to output the behavior category prediction;

[0018] Use the training set to train the LSTM model and optimize the model parameters to minimize the prediction loss;

[0019] Evaluate the performance of the trained LSTM model on the test set and calculate the evaluation indicators; integrate the trained LSTM model into the camera system to predict pet behavior in real time, and adjust the camera working status based on the pet behavior recognition results and target object recognition results.

[0020] Preferably, a target object recognition model based on the YOLO model is constructed, the input data is the pet perspective video frame data, and the target object is detected in each frame image; the output is a bounding box set of the target object and a target object category set;

[0021] Calculate the pixel difference between the current frame and the background reference frame to generate a difference mask for extracting dynamic areas; the background reference frame is generated by the static scene frame of the difference time;

[0022] Use the target object segmentation model based on the Mask R-CNN model to generate the target object segmentation mask in the video frame to separate the target object area from the background;

[0023] Combine the bounding box of the target object and the segmentation mask of the target object, and filter out the target object with the largest overlap with the differential mask as the pet's current target object of attention.

[0024] Preferably, the DeepSORT tracking algorithm is used to track the target object that the pet is paying attention to across frames to generate a trajectory of the target object;

[0025] The detected target object category set is filtered to exclude dynamic interference target objects that are not related to the pet's behavior.

[0026] Preferably, a pet favorite object database is constructed, and by analyzing the pet's daily activity video and image data, the object categories that the pet pays attention to and interacts with are identified as the pet's favorite object categories, and the pet's favorite object categories are imported into the target object recognition model;

[0027] Create a concentration analysis strategy. When the pet's favorite object is detected in the video frame, the concentration analysis process is triggered:

[0028] Calculate the pet's head orientation through head posture calculation technology;

[0029] Track the positional relationship between the pet's head orientation and the favorite object in multiple consecutive video frames, and count the duration of the pet's head orientation falling within the range of the favorite object;

[0030] The comprehensive concentration score of the pet during the period of looking at the favorite object is calculated, and the concentration score threshold is set to determine whether the pet enters the concentration mode based on the comprehensive concentration score.

[0031] Preferably, when the pet enters the focus mode, the size of the cropping area is dynamically calculated based on the size of the favorite object and the distance to the camera, with the favorite object as the center, and the cropping area image is enlarged;

[0032] During the cropping process, the position of the pet's favorite object is continuously tracked, and the center of the cropping area is dynamically adjusted to keep the favorite object always in the center of the picture;

[0033] At the same time, the pet's concentration changes are continuously tracked and analyzed, and the lower and upper limits of continuous concentration duration are set; if the pet's concentration duration is less than the lower limit, the normal perspective is smoothly restored; if the concentration duration is greater than the upper limit, the cropping area is expanded;

[0034] If the pet does not enter the focus mode, the image content is optimized through an adaptive perspective optimization algorithm.

[0035] Preferably, based on the identified pet behavior and the acquired pet-perspective video frame data; an adaptive perspective optimization algorithm is constructed to dynamically adjust the camera shooting parameters, including the camera perspective range, focal length and frame rate, based on the pet behavior and the pet-perspective video frame data;

[0036] Build an adaptive parameter adjustment strategy: set the adaptive adjustment range of the camera shooting parameters according to the pet's behavior; build an adaptive viewing angle optimization algorithm based on the pet's motion data and behavior to dynamically adjust the camera's viewing angle range;

[0037] Lock the viewing angle of the target object, and update the viewing direction of the camera based on the position of the target object; dynamically adjust the focal length according to the distance of the target object; dynamically adjust the frame rate according to the pet's behavior and angular velocity; output the camera shooting parameters based on the viewing angle range, focal length and frame rate.

[0038] Preferably, an adaptive image stabilization algorithm is constructed, and the image stabilization corrects the rotation, translation and tilt of the video image by combining the electronic image stabilization technology with the motion data;

[0039] Calculate the pet's motion trajectory and calculate the time step based on the angular velocity data of the gyroscope Angular deviation when

[0040] According to the acceleration data, the translation offset of the video frame is calculated; an image transformation matrix based on the angle offset and the translation offset is constructed to correct the rotation and translation of the image; the image transformation matrix is ​​applied to the video frame to obtain a stabilized image;

[0041] The jitter amplitude of the video frame is calculated based on the standard deviation of acceleration and the standard deviation of angular velocity; the size of the cropping area is dynamically adjusted based on the jitter amplitude, and the cropping range decreases as the jitter amplitude increases; the cropping center is calculated based on the pet's current target object position and camera position; cropping is applied to the stabilized image, and the final adjusted image is output, which is a video frame with stabilized data.

[0042] Preferably, a drone is deployed in the pet activity area, with the pet as the shooting beacon point, and the pet's location information is obtained in real time through visual sensors. In combination with the pet's behavior recognition results, the drone's position and angle are adjusted autonomously, with the pet as the shooting center, to obtain pet activity images from a third-party perspective.

[0043] A pet-perspective adaptive camera tracking system, which is used to implement the pet-perspective adaptive camera tracking method, includes: a data acquisition module, a pet behavior recognition module, a target recognition and tracking module, a focus mode module, a perspective optimization module, a picture adjustment module, and a third-party perspective shooting module;

[0044] The data acquisition module is used to collect the pet's motion data, including acceleration data, angular velocity data, and heart rate data when the pet is in motion, and construct the collected motion data into a data set;

[0045] The pet behavior recognition module recognizes the pet behavior based on the pet's motion data using a machine learning algorithm and sets the camera working state according to the pet's behavior;

[0046] The target recognition and tracking module constructs a target object recognition model, identifies the target object of the pet's attention, and tracks the target object by combining the background difference method and the target object segmentation technology;

[0047] The focus mode module inserts the pet's favorite object into the target object recognition model, and when the pet's favorite object is recognized, determines whether the pet enters the focus mode;

[0048] The viewing angle optimization module is used to construct an adaptive viewing angle optimization algorithm to automatically adjust the camera viewing angle, focal length and frame rate according to the recognized pet behavior and the target objects of the pet's attention in the activity scene;

[0049] The picture adjustment module performs adaptive image stabilization and dynamic adjustment on the video data captured by the camera, corrects the rotation, translation and tilt of the video picture according to the pet's behavior, and dynamically adjusts the picture cropping area;

[0050] The third-party perspective shooting module is used to build third-party perspective tracking of the pet's movement process and shoot the pet's activity process through a drone.

[0051] Beneficial effects of this application: This application provides high-quality data support for pet behavior analysis and machine learning model training by providing comprehensive motion and physiological data.

[0052] This application intelligently adjusts the camera's working status according to pet behavior, achieves targeted shooting, and improves the relevance and practicality of video content.

[0053] This application can accurately identify the target objects that the pet is paying attention to, eliminate environmental interference, ensure the main body of the picture is clear, and improve the accuracy of the captured picture.

[0054] This application describes the pet's concentration state in a quantitative way, and dynamically adjusts the camera shooting parameters accordingly to obtain a more intelligent and personalized pet perspective tracking effect.

[0055] This application dynamically adjusts the camera's viewing angle, focal length, and frame rate to adapt to pet behavior and scene changes, ensuring the rationality and quality of the captured image composition.

[0056] This application improves the stability and visual effects of the video image by correcting the rotation, translation and tilt of the image and dynamically cropping the image area.

[0057] This application uses drones to record the pet's movement process in real time, providing a global perspective and enhancing the diversity and fun of the shooting effects. BRIEF DESCRIPTION OF THE DRAWINGS

[0058] Figure 1 A flow chart of the pet-viewing adaptive camera tracking method provided in this application;

[0059] Figure 2 A diagram of pet behavior recognition and camera working state adjustment based on the pet perspective adaptive camera tracking method provided in this application;

[0060] Figure 3 A pet focus mode determination and image cropping diagram based on the pet perspective adaptive camera tracking method provided in this application;

[0061] Figure 4 A schematic diagram of the field of view determination strategy based on the pet perspective adaptive camera tracking method provided in this application;

[0062] Figure 5 This is a structural diagram of the pet-perspective adaptive camera tracking system provided for this application. DETAILED DESCRIPTION

[0063] In order to better understand the present application, a more detailed description will be made of various aspects of the present application with reference to the accompanying drawings. It should be understood that these detailed descriptions are only descriptions of exemplary embodiments of the present application, and are not intended to limit the scope of the present application in any way. Throughout the specification, the same reference numerals refer to the same elements. The expression "and / or" includes any and all combinations of one or more of the associated listed items.

[0064] In the accompanying drawings, the size, dimensions and shapes of the elements have been slightly adjusted for ease of illustration. The drawings are for illustration only and are not strictly drawn to scale. As used herein, the terms "substantially", "approximately" and similar terms are used as terms of approximation, not as terms of degree, and are intended to illustrate the inherent deviations in measurements or calculations that will be recognized by those of ordinary skill in the art. In addition, in this application, the order in which the steps are described does not necessarily represent the order in which these processes occur in actual operation, unless otherwise specified or can be derived from the context.

[0065] It should also be understood that expressions such as "include", "including", "have", "contain" and / or "comprising" are open rather than closed expressions in this specification, which indicate the presence of the stated features, elements and / or components, but do not exclude the presence of one or more other features, elements, components and / or combinations thereof. In addition, when expressions such as "at least one of..." appear after a list of listed features, they modify the entire list of features rather than just the individual elements in the list. In addition, when describing embodiments of the present application, "may" is used to mean "one or more embodiments of the present application". And, the term "exemplary" is intended to refer to an example or illustration.

[0066] Unless otherwise defined, all words (including engineering terms and scientific and technological terms) used in this article have the same meaning as those commonly understood by ordinary technicians in the field to which this application belongs. It should also be understood that unless there is a clear explanation in this application, words defined in commonly used dictionaries should be interpreted as having the same meaning as their meaning in the context of the relevant technology, and should not be interpreted in an idealized or overly formal sense.

[0067] It should be noted that, in the absence of conflict, the embodiments and features of the embodiments in the present application can be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.

[0068] Example 1

[0069] Reference Figures 1 to 4 , which is the first embodiment of the present application, provides a pet-perspective-based adaptive camera tracking method.

[0070] Step 1: Collect the pet's motion data, including the pet's acceleration data, angular velocity data, and heart rate data during motion, and construct the collected motion data into a data set.

[0071] Through the acceleration sensor, the pet's Acceleration value on the coordinate axis , the unit is , expressed as ; Use the gyroscope sensor to collect the pet's Angular velocity value on the coordinate axis , the unit is , expressed as ; Collect the pet's real-time heart rate value through the heart rate monitor, the unit is (times / minute), expressed as .

[0072] Exemplary pet acceleration acquisition method: by selecting a three-axis accelerometer, such as ADXL345, with a measurement range of ±16g, a resolution of 13 bits, and a sampling rate of 100Hz; fix the accelerometer on the pet collar or harness so that its x-axis, y-axis, and z-axis are aligned with the pet's forward direction, left-right direction, and up-down direction, respectively.

[0073] Exemplary method for collecting pet angular velocity: by selecting a three-axis gyroscope, such as ITG-3200, with a measurement range of ±2000° / s, a resolution of 16 bits, and a sampling rate of 100Hz; the gyroscope and acceleration sensor are integrated into the same sensor module, and their coordinate systems are ensured to be consistent.

[0074] Exemplary pet heart rate detection method: select a photoelectric volumetric pulse wave sensor, such as MAX30100, with a sampling rate of 100Hz; wear the heart rate monitor on the pet's chest or inside the ear, and calculate the heart rate by measuring the difference in absorption of red light and infrared light in the blood.

[0075] Define the pet movement dataset as a time series dataset ,Depend on data samples, each of which includes pet movement data and behavior labels. time steps, ,in, represents the motion data of the Nth data sample, Represents the behavior label of the Nth data sample.

[0076] For the data sample ,definition:

[0077] ;

[0078] in, Indicates samples at time step data, Indicates samples at time step Behavior tags, , , is the total number of behavior labels.

[0079] Include pets at time step Sports data:

[0080] ;

[0081] in, Respectively represent samples at time step The acceleration data on the x-axis, y-axis, and z-axis. Respectively represent samples at time step Angular velocity data on the x-axis, y-axis, and z-axis. Indicates samples at time step heart rate data.

[0082] Time step is defined as: ;

[0083] in, is the starting time of the data sample, is the time step, is the time step index, is the total number of time steps.

[0084] Through the above mathematical formula, pet movement data, behavior labels and time are constructed into a time series data set .

[0085] The time series dataset constructed in step 1 can be easily applied to various time series analysis and prediction tasks, such as pet behavior recognition and anomaly detection. At the same time, this mathematical representation also helps to understand the structure and characteristics of the dataset, providing a clear idea for the design and implementation of the subsequent pet behavior recognition algorithm.

[0086] Step 2: Based on the pet's motion data, use machine learning algorithms to identify the pet's behavior and set the camera's working state according to the pet's behavior.

[0087] The time series dataset Divide into training set , validation set and test set , where the training set is used for model training, the validation set is used for model hyperparameter selection and early stopping strategy, and the test set is used to evaluate model performance.

[0088] The data of the training set, validation set, and test set are normalized and scaled to the range of [0, 1]. This helps to speed up the convergence of the model and improve numerical stability. The normalization formula is: ;in, is the normalized input data, and are the minimum and maximum values ​​of the input data respectively.

[0089] Build an LSTM model for pet behavior recognition; the input of the LSTM model is normalized data , the output is the corresponding behavior sequence ,in, Indicates samples at time step After normalization, For the samples at time step The probability of the behavior category.

[0090] The architecture of the LSTM model includes an input layer, an LSTM layer, a fully connected layer, and an output layer. The input layer includes: Input into the LSTM model; the LSTM layer includes: using multiple LSTM layers to extract features and perform time series modeling on the input data sequence; the fully connected layer includes: mapping the hidden state to obtain the behavior category probability for each time step; the output layer includes: outputting the category prediction of the behavior.

[0091] The output dimension of each LSTM layer is , represents the dimension of the hidden state;

[0092] The calculation formula of the LSTM layer is:

[0093]

[0094]

[0095]

[0096]

[0097]

[0098]

[0099] in, is the input gate, For the Gate of Forgetfulness, is the output gate, is the current input, and are the weight matrices of the input gate, forget gate, output gate, and candidate cell state, respectively. and They are the bias items of input gate, forget gate, output gate and candidate cell state respectively. is the candidate memory cell state, is the state of memory cells, represents the cell state at the previous time step, is the hidden state, is the Sigmoid activation function, is the hyperbolic tangent activation function, is element-wise multiplication.

[0100] The calculation formula of the fully connected layer is:

[0101]

[0102] in, For the samples at time step The probability of the behavior category, It is samples at time step The hidden state of is the fully connected layer weight matrix, is the bias term; is the activation function that converts the output into a probability distribution.

[0103] Using the training set The LSTM model is trained and the model parameters are optimized to minimize the prediction loss. The cross entropy loss is used as the loss function of the LSTM model, and the calculation formula is:

[0104]

[0105] in, are model parameters, For the samples at time step Pet behavior, The model predicts samples at time step The probability of the behavior category.

[0106] During the training process, the model parameters are updated using optimization algorithms (such as Adam, RMSprop, etc.) to minimize the loss function.

[0107] By monitoring the validation set Based on the performance, we adjusted the model hyperparameters (such as hidden state dimension, number of layers, learning rate, etc.) and adopted the early stopping strategy to prevent overfitting.

[0108] In the test set Evaluate the performance of the trained LSTM model and calculate the accuracy, precision, recall, and F1 score evaluation indicators. Use the confusion matrix to visualize the classification performance of the model on different behaviors.

[0109] The trained LSTM model is integrated into the camera system to predict pet behavior in real time and adjust the camera working state based on the pet behavior recognition results and target object recognition results.

[0110] In step 2, a behavior recognition model with excellent performance can be obtained through data preparation, model construction, training optimization and evaluation. Applying this model to the camera system can realize adaptive camera tracking shooting from the pet's perspective, improving the quality of video data and user experience.

[0111] Step 3: Build a target object recognition model to identify the pet's target objects of interest, and track the target objects by combining background difference method and target object segmentation technology.

[0112] The object recognition model is designed to detect video frames from the pet's perspective The target objects that the pet is concerned about (such as toys, food, other animals or humans, etc.) are extracted, and the background difference and target object segmentation technology are combined to achieve target object tracking.

[0113] Build a target object recognition model based on the YOLO model, with the input data being the pet's perspective video frame data , represents the time step A single frame image captured by the camera at the time; the target object is detected in each frame image; the input of the target object recognition model is , output the bounding box set of the target object and the target thing category set :

[0114]

[0115] in, is the target object at time step t The bounding box of and width and height ; is the target object at time step t the category of the object (e.g., toy, food, human, etc.); is the number of detected target objects, Index the target object. .

[0116] Calculate the current frame With background reference frame Pixel difference of , generating a difference mask , used to extract dynamic areas; among them, ; The background reference frame Through the time step Static scene frame generation.

[0117] Use the target object segmentation model based on the Mask R-CNN model in the current frame Generate the target object segmentation mask , separate the target area from the background: ;in is the target object at time step t Binary segmentation mask of ; Difference mask Generate a binary dynamic area map after threshold segmentation ,in .

[0118] Combine the bounding box set of the target object and target object segmentation mask , filter out the difference mask The target object with the largest overlap, which is the pet's current target object of attention : ;in The intersection-over-union ratio of the bounding box of the target object and the binary dynamic area map.

[0119] Use the DeepSORT tracking algorithm to track the target object that the pet is paying attention to across frames and generate the trajectory of the target object : ;in Represents the time step The three-dimensional position of the target object.

[0120] The detected target object category set Screening to exclude pet behaviors Extraneous dynamic distractions (such as other pets or passersby)

[0121] Exemplarily, the exclusion rules are as follows: Indicates that the pet is identified as an interactive state, and the target objects of the category of toys or humans are retained first; if Indicates that the pet is identified as being in an exploratory state, and target objects of the category other animals or food are preferentially retained.

[0122] Output the bounding box of the pet's current frame and categories ; Segmentation mask of the pet’s attention target .

[0123] Based on the identified pet behavior and the obtained pet perspective video frame data , Represents the time step Single frame image at each moment; build an adaptive viewing angle optimization algorithm to and pet perspective video frame data , dynamically adjust the camera shooting parameters, including the camera viewing angle range, focal length and frame rate; the viewing angle range (Field of View, FOV) is used to adjust the camera's shooting range; the focal length (Focal Length, ) is used to adjust the clarity of the target object; the frame rate (FrameRate, FPS) is used to adjust the sampling frequency of the video.

[0124] Building an adaptive parameter adjustment strategy: Based on pet behavior , set the adaptive adjustment range of camera shooting parameters.

[0125] For example, , indicating that the pet is recognized as stationary, the viewing angle is narrowed, and the pet's current observation area is focused; increase the focal length , highlight the area of ​​interest; reduce the frame rate to , reduce power consumption;

[0126] set up , indicating that the pet is recognized as interactive, the viewing angle is expanded, and the pet's movement path is captured; the default focal length is maintained to avoid frequent adjustments; the frame rate is increased to , reduce motion blur;

[0127] set up , indicating that the pet is recognized as being in a focused state, and the viewing angle is adjusted to lock the target object (such as toys, humans) for the pet to interact with; the focal length is dynamically adjusted to ensure that the target object is clear; the frame rate is set to , balance clarity and power consumption;

[0128] set up 3, indicating that the pet is recognized as exploring, the viewing direction is quickly switched, and the target area explored by the pet is captured first; the focal length is adjusted dynamically, and the zoom is automatically adjusted according to the distance of the target object; the frame rate is set to , to ensure picture smoothness.

[0129] Step 4: Insert the pet's favorite object into the target object recognition model. When the pet's favorite object is recognized, determine whether the pet enters the focus mode.

[0130] Build a database of pets' favorite objects. By analyzing the video and image data of pets' daily activities, identify the categories of objects that pets pay attention to and interact with, such as pet toys and food, as the categories of pets' favorite objects. Establish labels for the categories of pets' favorite objects and import them into the target object recognition model.

[0131] Create a concentration analysis strategy to trigger the concentration analysis process when the pet's favorite object is detected in the video frame.

[0132] Through head posture estimation technology, the pet's head orientation is calculated in real time;

[0133] Specifically, in each frame of video image, the pet's head area is detected and located, the head feature points are extracted, and a three-dimensional head posture model is constructed; let the pet's head heading unit vector be , the object position unit vector is , then calculate the angle between the vectors : ;in, represents the dot product of the pet's head orientation unit vector and the object's position unit vector, and Respectively represent the modulus of the pet's head direction unit vector and the object's position unit vector; if the angle Less than the preset threshold (For example, 30°, ), it is preliminarily determined that the pet is looking at the object.

[0134] Track the positional relationship between the pet's head orientation and the favorite object in a continuous multi-frame video, and count the duration of the pet's head orientation falling within the favorite object range; The angle between the pet's head orientation vector and the object position vector in the frame video is , the frame rate is , then in continuous The duration of the pet's gaze on the object within the frame for: ;in, is an indicative function, which takes the value 1 when the condition is met, otherwise it takes the value 0; Exceeding the preset threshold (such as 3 seconds), it further verifies the pet's concentration on the object.

[0135] Calculate the pet's overall concentration score while looking at the preferred object , taking into account the two factors of gaze duration and angle size, the comprehensive concentration score The calculation formula is:

[0136]

[0137] in, is the angle between the pet's head orientation vector and the object position vector in each video frame during the fixation period The average value, and is the weight coefficient, satisfying The longer the fixation duration and the smaller the average angle, the higher the comprehensive concentration score.

[0138] Setting the Focus Score Threshold and (such as 0.8 and 0.2); if , then the pet is judged to have entered the focus state, triggering the cropping function of the focus mode; if , then the pet is judged to be not showing obvious concentration and the normal cutting method is maintained; if , the observation time is extended, and whether to switch to the focus mode is decided according to the visual field judgment strategy.

[0139] The vision determination strategy specifically includes:

[0140] When the favorite object is within the pet's field of vision, the first timing is triggered and recorded as If the favorite object gradually moves from the edge of the visual field to the center of the visual field, when the favorite object enters the center of the visual field, the second timing is triggered and the second timing is recorded as , if the time interval between two times is less than or equal to the interval setting threshold ,Right now , and the time the favorite object is in the center of the field of vision is greater than the duration setting threshold, then it is determined to be in focus mode. At the same time, the interval setting threshold and the duration setting threshold are calibrated according to the average value of the pet's historical data. The advantage of the field of vision determination strategy is that the calculation is simple and the accuracy is high. Figure 4 , which is a schematic diagram of the field of view determination strategy provided in this application.

[0141] In the focus mode, the pet's favorite object is the center, and the size of the cropping area is dynamically calculated according to the size of the favorite object and the distance to the camera, and the cropping area image is smoothly enlarged; specifically, the pixel size of the favorite object in the image is set to , the distance from the favorite object to the camera is , the total image size is , then the cropping area size It can be calculated as follows:

[0142]

[0143] in, is the reference distance (e.g. 1 meter), and is the cropping factor, which controls the magnification of the cropping area relative to the size of the preferred object; the cropping area should meet and , appropriate restrictions may be imposed when necessary and During the cropping process, the position of the pet's favorite object is continuously tracked, and the center of the cropping area is dynamically adjusted to keep the favorite object always in the center of the picture.

[0144] At the same time, continuously track and analyze the changes in the pet's concentration, and set a lower limit for the continuous concentration time. (such as 5 seconds) and upper limit (e.g. 30 seconds); if the pet's concentration lasts for If the pet's concentration lasts for If the above is the case, the cropping area should be appropriately expanded to avoid visual fatigue caused by long-term high focus; and the minimum interval duration of the focus mode should be set to avoid too frequent perspective switching affecting video continuity.

[0145] This step constructs a pet concentration analysis and adaptive cropping method in the concentration mode to quantitatively characterize the pet's concentration state, and dynamically adjusts the camera shooting parameters accordingly to obtain a more intelligent and personalized pet perspective tracking effect.

[0146] Step 5: Build an adaptive perspective optimization algorithm to automatically adjust the camera perspective, focal length, and frame rate based on the identified pet behavior and the pet's focus in the activity scene;

[0147] According to pet movement data and pet behavior , build an adaptive viewing angle optimization algorithm to adjust the viewing angle range of the camera.

[0148] Define the camera viewing angle as and , is the horizontal viewing angle, For vertical viewing angle, the camera viewing angle range is dynamically adjusted according to motion data and pet behavior:

[0149]

[0150] in, and is the initial viewing angle of the camera, and For pet behavior Adaptive adjustment amount: If , that is, the pet is in a stationary state, then , ;like , that is, the pet is in motion, then , .

[0151] Lock the target object's perspective and combine it with the target object's position , update the camera's viewing direction :

[0152]

[0153] in, represents the camera position, Indicates the calculation of Euclidean distance, Represents the inverse tangent function.

[0154] focal length According to the distance of the target Dynamic Adjustment:

[0155]

[0156] in, is the location of the target object, is the maximum distance at which the target object can be identified, is the initial focal length, is the focal length scaling factor.

[0157] Frame rate Based on pet behavior and angular velocity Dynamic Adjustment:

[0158] in, and are the maximum frame rate and the minimum frame rate, is the maximum value of the angular velocity.

[0159] Combine the viewing angle, focal length and frame rate to output the camera shooting parameters:

[0160] in, Represents the time step The camera shooting parameters at that moment.

[0161] This step combines the pet's behavior and movement data to dynamically adjust the camera's viewing angle, focal length, and frame rate to achieve adaptive viewing angle optimization; the introduction of the control formula ensures precise control of viewing angle changes, focal length adjustment, and frame rate optimization, providing technical support for shooting high-quality videos.

[0162] Step 6: Adaptively stabilize and dynamically adjust the video data captured by the camera, correct the rotation, translation and tilt of the video image according to the pet's behavior, and dynamically adjust the image cropping area.

[0163] Build an adaptive image stabilization algorithm. Image stabilization is achieved through electronic image stabilization technology combined with motion data , correct the rotation, translation and tilt of the video image.

[0164] Calculate the pet's motion trajectory based on the angular velocity data of the gyroscope , calculate the time step Angular deviation .

[0165] When the angle offset When it is in continuous form, the following formula is used for calculation:

[0166] in is the angular velocity data;

[0167] When the angle offset When it is in discrete form, the following formula is used for calculation:

[0168]

[0169] Based on acceleration data , calculate the time step The translation offset .

[0170] When the translation offset is in continuous form, the following formula is used for calculation:

[0171]

[0172] When the translation offset is in discrete form, discrete integration is used to calculate the translation:

[0173]

[0174] in is the acceleration data;

[0175] Indicates the translation offset in the three-axis direction.

[0176] Constructing based on angle offset and translation offset The image transformation matrix , used to correct the rotation and translation of the image; image transformation matrix for:

[0177]

[0178] in, and Respectively around The rotation matrix of an axis is defined as:

[0179]

[0180]

[0181] in, It is around Axis rotation angle The sine value of For around Axis rotation angle, For around Axis rotation angle.

[0182] Further time Video frame Apply image transformation matrix , and get the stabilized image ;

[0183] According to the standard deviation of the acceleration and the standard deviation of the angular velocity , calculate the jitter amplitude at time step t , according to the jitter amplitude ;

[0184]

[0185] in, ; and are the components of acceleration and angular velocity, respectively, and is its average value;

[0186] Dynamically resize the crop area , the cropping range decreases as the jitter amplitude increases, ensuring the stability of the video image;

[0187] in, and is the width and height of the original image, is the cropping amplitude adjustment coefficient, and is the width and height of the cropping area;

[0188] Based on the pet's current target location and camera position , calculate the size of the clipping area ; ;

[0189] The image after stabilization Apply cropping to the output, output the final adjusted image, output data stabilized video frame : , Represents a crop operation.

[0190] Step 5: Through real-time analysis of pet motion data and combined with electronic image stabilization technology, the rotation, translation and tilt of the video image are dynamically corrected, the cropping area and image quality are optimized, and the stability and clarity of the video are ensured.

[0191] Step 6: Build a third-party perspective tracking of the pet's movement process and use a drone to shoot the pet's activities.

[0192] Calculate the real-time shooting center position, use the pet's current position as the beacon point, and combine the position of the target object to calculate the drone's shooting center position:

[0193]

[0194] in, Represents the time step The drone takes photos of the center position; Represents the time step The location of the pet at the time; Indicates the shooting offset, which is used to adjust the relative position between the shooting center and the pet.

[0195] Adjust the flight altitude of the drone. Dynamically adjust the flight altitude of the drone according to the pet's movement status:

[0196]

[0197] in, Represents the time step The flight altitude of the drone at that time; Indicates the initial flight altitude of the drone; Indicates the height adjustment factor; Represents the time step The velocity modulus of the pet at time (obtained by integrating the acceleration).

[0198] Calculate the direction of the drone's viewing angle and adjust the drone's camera's viewing angle in real time according to the position of the pet and the target object:

[0199]

[0200] in, Represents the time step The pitch angle of the drone camera; Represents the time step The location of the target object; Represents the time step The location of the pet at the time; Represents the time step The flight altitude of the drone at that time.

[0201] This step uses the pet as a beacon point and dynamically adjusts the flight altitude, viewing direction, and focus based on the location of the target object to ensure that the pet and the target object are always in the core area of ​​the picture. This process can adapt to the pet's movement state in real time, improve the accuracy and stability of shooting, and optimize the picture parameters at the same time, adapt to multiple scene environments, and generate high-quality video content.

[0202] Example 2

[0203] Reference Figure 5 , which is the second embodiment of the present application, provides an adaptive camera tracking system based on pet perspective.

[0204] The system includes: a data acquisition module, a pet behavior recognition module, a target recognition and tracking module, a focus mode module, a perspective optimization module, a picture adjustment module and a third-party perspective shooting module.

[0205] The data acquisition module is used to collect the pet's motion data, including acceleration data, angular velocity data, and heart rate data when the pet is in motion, and construct the collected motion data into a data set.

[0206] The pet behavior recognition module recognizes the pet behavior based on the pet's motion data using a machine learning algorithm and sets the camera working state according to the pet's behavior.

[0207] The target recognition and tracking module constructs a target object recognition model, identifies the target object of the pet's attention, and tracks the target object by combining the background difference method and the target object segmentation technology.

[0208] The focus mode module inserts the pet's favorite object into the target object recognition model, and when the pet's favorite object is recognized, it determines whether the pet enters the focus mode.

[0209] The viewing angle optimization module is used to construct an adaptive viewing angle optimization algorithm to automatically adjust the camera viewing angle, focal length and frame rate according to the recognized pet behavior and the target objects of the pet's attention in the activity scene.

[0210] The picture adjustment module performs adaptive image stabilization and dynamic adjustment on the video data captured by the camera, corrects the rotation, translation and tilt of the video picture according to the pet's behavior, and dynamically adjusts the picture cropping area.

[0211] The third-party perspective shooting module is used to build third-party perspective tracking of the pet's movement process and shoot the pet's activity process through a drone.

[0212] The above sequence of steps for the method is for illustration only, and the steps of the method of the present application are not limited to the sequence specifically described above unless otherwise specifically stated.

[0213] In addition, in some embodiments, the present application can also be implemented as a program recorded in a recording medium, which includes machine-readable instructions for implementing the method according to the present application. Therefore, the present application also covers a recording medium storing a program for executing the method according to the present application.

[0214] In addition, the parts of the above-mentioned technical solutions provided in the embodiments of the present application that are consistent with the implementation principles of the corresponding technical solutions in the prior art are not described in detail to avoid excessive redundancy.

[0215] The specific implementation modes as described above further describe the purpose, technical solutions and beneficial effects of the present application in detail. It should be understood that the above description is only the specific implementation modes of the present application and is not intended to limit the present application. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present application shall be included in the protection scope of the present application.

Claims

1. A pet-viewing adaptive camera tracking method, characterized in that: include: Collecting the pet's motion data, including acceleration data, angular velocity data, and heart rate data when the pet is in motion, and constructing the collected motion data into a data set; Based on the pet's motion data, the machine learning algorithm is used to identify the pet's behavior and set the camera's working state according to the pet's behavior; Build a target object recognition model to identify the target objects that the pet is paying attention to, and track the target objects by combining background difference method and target object segmentation technology; Insert the pet's favorite object into the target object recognition model, and when the pet's favorite object is recognized, determine whether the pet enters the focus mode; Build an adaptive perspective optimization algorithm to automatically adjust the camera perspective, focal length, and frame rate based on the identified pet behavior and the pet's focus in the activity scene; Adaptive image stabilization and dynamic adjustment of the video data captured by the camera, correct the rotation, translation and tilt of the video screen according to the pet's behavior, and dynamically adjust the screen cropping area; Build third-party perspective tracking of the pet's movement process and use drones to capture the pet's activities.

2. The pet-viewing-based adaptive camera tracking method according to claim 1, characterized in that: Collect the acceleration data of the pet, collect the angular velocity data of the pet, and collect the real-time heart rate data of the pet; The pet motion dataset is defined as a time series dataset. Each data sample includes pet motion data and behavior labels, and each data sample contains multiple time steps.

3. The pet-viewing-based adaptive camera tracking method according to claim 2, characterized in that: Divide the time series data set into training set, validation set and test set; normalize the data of the training set, validation set and test set, and scale the data; Construct an LSTM model for pet behavior recognition; the input of the LSTM model is the normalized training set data, and the output is the corresponding behavior sequence; The architecture of the LSTM model includes an input layer, an LSTM layer, a fully connected layer, and an output layer. The input layer is used to input the normalized training set data into the LSTM model; the LSTM layer is used to extract features and perform time series modeling on the input sequence using multiple LSTM layers; the fully connected layer is used to map the hidden state to obtain the behavior category probability at each time step; the output layer is used to output the behavior category prediction; Use the training set to train the LSTM model and optimize the model parameters to minimize the prediction loss; Evaluate the performance of the trained LSTM model on the test set and calculate the evaluation indicators; integrate the trained LSTM model into the camera system to predict pet behavior in real time, and adjust the camera working status based on the pet behavior recognition results and target object recognition results.

4. The pet-viewing-based adaptive camera tracking method according to claim 3, characterized in that: Build a target object recognition model based on the YOLO model. The input data is the pet's perspective video frame data. The target object is detected in each frame image. The output is the target object's bounding box set and target object category set. Calculate the pixel difference between the current frame and the background reference frame to generate a difference mask for extracting dynamic areas; The background reference frame is generated by differential time static scene frames; Use the target object segmentation model based on the Mask R-CNN model to generate the target object segmentation mask in the video frame to separate the target object area from the background; Combine the bounding box of the target object and the segmentation mask of the target object, and filter out the target object with the largest overlap with the differential mask as the pet's current target object of attention.

5. The pet-viewing-based adaptive camera tracking method according to claim 4, characterized in that: Use the DeepSORT tracking algorithm to track the target object that the pet is paying attention to across frames and generate the trajectory of the target object; The detected target object category set is filtered to exclude dynamic interference target objects that are not related to the pet's behavior.

6. The pet-viewing-based adaptive camera tracking method according to claim 5, characterized in that: Build a database of pets' favorite objects. By analyzing the videos and images of pets' daily activities, identify the categories of objects that pets pay attention to and interact with, and use them as pets' favorite object categories. Import the pets' favorite object categories into the target object recognition model. Create a concentration analysis strategy. When the pet's favorite object is detected in the video frame, the concentration analysis process is triggered: Calculate the pet's head orientation through head posture calculation technology; Track the positional relationship between the pet's head orientation and the favorite object in multiple consecutive video frames, and count the duration of the pet's head orientation falling within the range of the favorite object; The comprehensive concentration score of the pet during the period of looking at the favorite object is calculated, and the concentration score threshold is set to determine whether the pet enters the concentration mode based on the comprehensive concentration score.

7. The pet-viewing-based adaptive camera tracking method according to claim 6, characterized in that: When the pet enters the focus mode, the pet's favorite object is the center, and the size of the cropping area is dynamically calculated according to the size of the favorite object and the distance to the camera, and the cropping area image is enlarged; During the cropping process, the position of the pet's favorite object is continuously tracked, and the center of the cropping area is dynamically adjusted to keep the favorite object always in the center of the picture; At the same time, it continuously tracks and analyzes the changes in the pet's concentration, and sets the lower and upper limits of the continuous concentration duration; If the pet's concentration duration is less than the lower limit of the concentration duration, the normal viewing angle will be restored smoothly; If the duration of concentration is longer than the upper limit of concentration duration, the cropping area will be expanded; If the pet does not enter the focus mode, the image content is optimized through an adaptive perspective optimization algorithm.

8. The pet-viewing-based adaptive camera tracking method according to claim 7, characterized in that: Based on the identified pet behavior and the obtained pet-perspective video frame data; construct an adaptive perspective optimization algorithm to dynamically adjust the camera shooting parameters, including the camera perspective range, focal length and frame rate, based on the pet behavior and the pet-perspective video frame data; Build an adaptive parameter adjustment strategy: set the adaptive adjustment range of the camera shooting parameters according to the pet's behavior; build an adaptive viewing angle optimization algorithm based on the pet's motion data and behavior to dynamically adjust the camera's viewing angle range; Lock the viewing angle of the target object, and update the viewing direction of the camera based on the position of the target object; The focal length is dynamically adjusted according to the distance of the target object; the frame rate is dynamically adjusted according to the pet's behavior and angular velocity; the camera shooting parameters are output based on the viewing angle range, focal length and frame rate.

9. The pet-viewing-based adaptive camera tracking method according to claim 8, characterized in that: Build an adaptive image stabilization algorithm. Image stabilization uses electronic image stabilization technology and combines motion data to correct the rotation, translation and tilt of the video image. Calculate the pet's motion trajectory and calculate the time step based on the angular velocity data of the gyroscope Angular deviation when According to the acceleration data, the translation offset of the video frame is calculated; an image transformation matrix based on the angle offset and the translation offset is constructed to correct the rotation and translation of the image; Apply the image transformation matrix to the video frame to obtain a stabilized image; The jitter amplitude of the video frame is calculated based on the standard deviation of acceleration and the standard deviation of angular velocity; the size of the cropping area is dynamically adjusted based on the jitter amplitude, and the cropping range is reduced as the jitter amplitude increases; The cropping center is calculated based on the pet's current target object position and the camera position; cropping is applied to the stabilized image, and the final adjusted image is output, and the video frame after data stabilization is output.

10. The pet-viewing-based adaptive camera tracking method according to claim 9, characterized in that: Deploy drones in pet activity areas, use pets as shooting beacons, obtain pet location information in real time through visual sensors, and autonomously adjust the drone position and angle based on pet behavior recognition results, use pets as the shooting center, and obtain pet activity images from a third-party perspective.

11. A pet-perspective-based adaptive camera tracking system, which is used to implement the pet-perspective-based adaptive camera tracking method according to any one of claims 1 to 10, characterized in that: include: Data collection module, pet behavior recognition module, target recognition and tracking module, focus mode module, perspective optimization module, picture adjustment module and third-party perspective shooting module; The data acquisition module is used to collect the pet's motion data, including acceleration data, angular velocity data, and heart rate data when the pet is in motion, and construct the collected motion data into a data set; The pet behavior recognition module recognizes the pet behavior based on the pet's motion data using a machine learning algorithm and sets the camera working state according to the pet's behavior; The target recognition and tracking module constructs a target object recognition model, identifies the target object of the pet's attention, and tracks the target object by combining the background difference method and the target object segmentation technology; The focus mode module inserts the pet's favorite object into the target object recognition model, and when the pet's favorite object is recognized, determines whether the pet enters the focus mode; The viewing angle optimization module is used to construct an adaptive viewing angle optimization algorithm to automatically adjust the camera viewing angle, focal length and frame rate according to the recognized pet behavior and the target objects of the pet's attention in the activity scene; The picture adjustment module performs adaptive image stabilization and dynamic adjustment on the video data captured by the camera, corrects the rotation, translation and tilt of the video picture according to the pet's behavior, and dynamically adjusts the picture cropping area; The third-party perspective shooting module is used to build third-party perspective tracking of the pet's movement process and shoot the pet's activity process through a drone.

Citation Information

Patent Citations

  • Specific object tracking shooting method and system

    CN116489516A

  • Pet tracking method and device, XR equipment and medium

    CN117788765A

  • Remote communication alarm smart home system

    CN118865627A

  • Automatic pet shooting device and pet behavior identification method

    CN119233072A

  • Method and Machine for Predictive Animal Behavior Analysis

    US20210089945A1