Pet posture and expression tracking method and device, electronic equipment and storage medium
By combining pet motion data and image data, using machine learning and neural network models to fuse pet posture and expression information, the problem of degraded recognition accuracy in traditional systems is solved, and high-precision pet posture and expression tracking is achieved.
Patent Information
- Application Number
- CN202510732865.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-04
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-06-04
AI Technical Summary
Traditional pet monitoring systems rely on a single data source, resulting in a decrease in recognition accuracy and making it difficult to capture pet's delicate expression changes and specific postures.
Combining pet motion data and image data, processing is done through machine learning models and neural network models, integrating pet posture and expression information, and using TensorFlow Lite and YOLO neural networks for data processing and prediction.
It improves the accuracy and stability of pet posture and expression recognition, reduces dependence on a single data source, and can continuously and accurately output pet information in complex environments.
Smart Images

Figure CN120260080A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing technology, and particularly to a method, device, electronic device, and storage medium for tracking pet postures and expressions. Background Art
[0002] Nowadays, the status of pets as family members is increasing day by day, and it is also becoming more and more important for pet owners to pay attention to their behavior, emotions, and health status. Traditional pet monitoring systems often rely on a single monitoring technology. For example, video monitoring using a camera or activity tracking using a wearable motion sensor. Video monitoring is easily affected by environmental light and obstacles, resulting in a decrease in recognition accuracy. Or although the motion sensor can record the pet's activity, it is difficult to capture delicate expression changes and specific postures. Summary of the Invention
[0003] In order to overcome the deficiencies of the prior art, the present invention provides a method, device, electronic device, and storage medium for tracking pet postures and expressions, which realizes the processing of high-precision pet posture recognition and expression tracking.
[0004] The first aspect of this application provides a method for tracking pet postures and expressions, and the method includes: Obtain pet motion data and pet images; Process the pet motion data according to a preset machine learning model to obtain a predicted pet posture; Process the pet image according to a preset neural network model to obtain a predicted pet expression; Perform fusion processing on the predicted pet posture and the predicted pet expression, and output the pet posture and expression.
[0005] In an optional embodiment, the fusion of the predicted pet posture and the predicted pet expression to output the pet posture and expression includes: Perform time and space alignment on the predicted pet posture and the predicted pet expression; Determine the first modal weight of the predicted pet posture and the second modal weight of the predicted pet expression; Perform weighted processing according to the predicted pet posture and the first modal weight, and the predicted pet expression and the second modal weight to obtain a fusion result; Output the pet posture and expression according to the fusion result.
[0006] The method further includes: Obtain the first fitness score of the predicted pet posture and the second fitness score of the predicted pet expression; When it is determined that both the first fitness score and the second fitness score are greater than a preset threshold, output the pet posture and pet expression; When it is determined that the first fitness score is greater than the preset threshold and the second fitness score is less than the preset threshold, output the pet posture; When it is determined that the first fitness score is less than the preset threshold and the second fitness score is greater than the preset threshold, output the pet expression.
[0007] In an alternative embodiment, the method further includes: Step 1: Obtain the sensor raw data corresponding to each motion sensor unit among multiple motion sensor units to obtain a set of sensor raw data, where each sensor raw data in the set of sensor raw data includes annotation information; Step 2: Preprocess each sensor raw data in the set of sensor raw data to obtain a first training sample set; Step 3: Construct an initial machine learning model based on the TensorFlow Lite machine learning architecture; Step 4: Input the first training sample into the initial machine learning model to obtain a first prediction result corresponding to the first training sample, where the first training sample is any training sample in the first training sample set; Step 5: Adjust the initial machine learning model based on the first prediction result and the annotation information corresponding to the first training sample; Step 6: Execute Step 4 and Step 5 based on the adjusted initial machine learning model until a preset iteration termination condition is met; Step 7: Determine the initial machine learning model when the preset iteration termination condition is met as the preset machine learning model.
[0008] In an alternative embodiment, the method further includes: Step 1: Obtain the pet raw images corresponding to each image acquisition unit among multiple image acquisition units to obtain a set of pet raw images, where each pet raw image in the set of pet raw images includes annotation information; Step 2: Preprocess each pet raw image in the set of pet raw images to obtain a second training sample set; Step 3: Construct an initial neural network model based on the YOLO neural network; Step 4: Input the second training sample into the initial neural network model to obtain a second prediction result corresponding to the second training sample, where the second training sample is any training sample in the second training sample set; Step 5. Adjust the initial neural network model based on the second prediction result and the annotation information corresponding to the second training sample; Step 6. Execute Step 4 and Step 5 based on the adjusted initial neural network model until a preset iteration termination condition is met; Step 7. Determine the initial neural network model when the preset iteration termination condition is met as the preset neural network model.
[0009] In an alternative embodiment, the method further includes: Judge whether the pet has abnormal state behaviors according to the pet posture and expression; When it is determined that the pet has abnormal state behaviors, send a warning signal to prompt the pet owner.
[0010] In an alternative embodiment, the method further includes: Determine the pet's movement trajectory, pet posture change and pet emotion state according to the pet movement data and the pet image; Visualize the pet's movement trajectory, the pet's posture change, and the pet's emotion state.
[0011] The second aspect of the present application provides a pet posture and expression tracking device, and the device includes: The third aspect of the present application provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the steps of the pet posture and expression tracking method are implemented.
[0012] The fourth aspect of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the above-mentioned pet posture and expression tracking method are implemented.
[0013] In summary, the pet posture and expression tracking method, device, electronic device, and storage medium provided by the present application have at least one of the following beneficial effects: 1. Obtain pet movement data and pet images, and analyze them in combination with pet movement data and image data, reducing the error caused by a single data source; 2. Pet movement data and pet image data contain information in different dimensions of the pet. By fusing these two types of data, the posture and expression information of the pet can be comprehensively obtained; 3. The fusion method reduces the system's dependence on a single data source through the mutual complementation and verification of multiple data, improves the stability and reliability of the system, and can continuously and accurately output the pet posture and expression information under various complex conditions. Description of the Drawings
[0014] Figure 1 It is a schematic structural diagram of a pet posture and expression tracking system shown in an embodiment of the present application; Figure 2 It is a schematic flowchart of a pet posture and expression tracking method shown in an embodiment of the present application; Figure 3 It is another schematic flowchart of a pet posture and expression tracking method shown in an embodiment of the present application; Figure 4 It is a functional module diagram of a pet posture and expression tracking device shown in an embodiment of the present application; Figure 5 It is a schematic structural diagram of an electronic device shown in an embodiment of the present application. Detailed implementation manners
[0015] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.
[0016] The concept, specific structure and technical effects of the present invention will be clearly and completely described below in conjunction with the embodiments and the accompanying drawings, so as to fully understand the purpose, features and effects of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, other embodiments obtained by those skilled in the art without creative efforts shall fall within the scope of protection of the present invention. In addition, all the connection / connection relationships involved in the patent do not refer only to the direct connection of components, but refer to the more optimal connection structure that can be formed by adding or reducing connection accessories according to the specific implementation situation. Each technical feature in the present invention can be combined interactively without conflicting with each other.
[0017] Referring to Figure 1 As shown, it is a schematic structural diagram of a pet posture and expression tracking system shown in an embodiment of the present application. The pet posture and expression tracking system includes two terminal devices, namely the first terminal device and the second terminal device. Among them, the first terminal device is an MCU-based embedded terminal, which is used to run a lightweight machine learning algorithm, classify and infer the collected motion sensor data after filtering, and then upload the prediction result to the second terminal device through wifi, Bluetooth or cellular network. The second terminal device is a board embedded terminal with NPU computing power, which is used to run a convolutional neural network model to perform expression recognition on the collected camera data, perform multi-modal fusion algorithm operations according to the expression recognition result and in combination with the activity type uploaded by the first terminal device, and display the operation result.
[0018] In some embodiments, the first terminal device may include a motion sensor unit, which integrates a high-precision three-axis accelerometer, gyroscope, and magnetometer, and is worn on a pet collar or vest to collect pet motion sensor data in real time, including multi-axis raw data such as acceleration, angular velocity, and direction. After the first terminal device collects multi-axis raw data in thousands, a new machine learning model can be trained based on the multi-axis raw data using a lightweight TensorFlow Lite machine learning architecture. The trained new machine learning model is used to classify the motion sensor data collected in real time to identify the activity types of the pet (e.g., walking, running, jumping, and resting, etc.). The second terminal may include an image acquisition unit, where the image acquisition unit uses an embedded SOC with NPU computing power and is connected to a high-definition camera with night vision function and wide-angle lens, and is installed at key positions in the pet activity area to capture the raw images of the pet in real time. At the same time, the second terminal device can pre-train a new neural network model that combines pet facial feature and expression recognition based on the YOLO convolutional neural network model that supports image classification, object detection, and instance segmentation tasks, to perform fast and accurate expression recognition on the pet images captured by the high-definition image acquisition unit (e.g., camera). When obtaining the pet posture output by the new machine learning model, the first terminal device can send the pet posture output by the new machine learning model to the second terminal device. The second terminal device combines the pet posture output by the new machine learning model and the pet expression output by the new neural network model to output the pet posture and expression. Or when obtaining the pet expression output by the new neural network model, the second terminal device can send the pet expression output by the new neural network model to the first terminal device. The first terminal device combines the pet posture output by the new machine learning model and the pet expression output by the new neural network model to output the pet posture and expression.
[0019] In an alternative real-time manner, when the motion sensor data (i.e., pet motion data) is obtained through the first terminal device and the pet image is obtained through the second terminal device, the pet motion data and the pet image can be transmitted to the electronic device in an active or passive manner. The electronic device processes the pet motion data based on the pre-trained new machine learning model to obtain a first prediction result; and processes the pet image based on the pre-trained new neural network model to obtain a second prediction result. Then, the electronic device can fuse the first prediction result and the second prediction result to obtain the pet posture and expression corresponding to the pet, and visualize the pet posture and expression.
[0020] Referring to Figure 2 shown, is a schematic flowchart of a method for tracking pet posture and expression shown in an embodiment of the present application. The method for tracking pet posture and expression includes the following steps.
[0021] S21. Obtain pet movement data and pet images.
[0022] Refer to Figure 3 , in some embodiments, the motion sensor unit can be worn on the pet collar or vest, and the pet movement data, including multi-axis raw data such as acceleration, angular velocity, and direction, can be collected in real time through the motion sensor unit. Among them, the motion sensor unit integrates a high-precision three-axis accelerometer, gyroscope, and magnetometer. At the same time, the image acquisition unit is installed at key positions in the pet activity area, and the pet images are collected in real time through the image acquisition unit. Among them, the image acquisition unit integrates an embedded SOC with NPU computing power and is connected to a high-definition camera with night vision function and wide-angle lens. The motion sensor unit collects pet movement data in real time and transmits the pet movement data to the electronic device; similarly, the image acquisition unit collects pet images in real time and transmits the pet images to the electronic device, so that the electronic device can process the obtained pet movement data and pet images in real time.
[0023] Refer to Figure 3 , in some embodiments, when the pet movement data is obtained, since the pet movement data is multi-axis, the electronic device needs to verify the integrity of the pet movement data to ensure that the data of all axes are complete and accurate. Then, the Kalman filter algorithm is used to further process the pet movement data to improve the accuracy and stability of the data. At the same time, when the pet images are obtained, the electronic device can perform data cleaning on the pet images, and after the image data cleaning is completed, noise reduction processing and image enhancement processing are performed.
[0024] S22. Process the pet movement data according to a preset machine learning model to obtain a predicted pet posture.
[0025] Refer to Figure 3 , when the preprocessed pet movement data is obtained, the pet movement data is input into a pre-trained machine learning model (referred to as a preset machine learning model) to output a corresponding prediction result through the preset machine learning model, which is called a predicted pet posture. At the same time, when the preset machine learning model outputs the predicted pet posture, it can also output a fitness score corresponding to the predicted pet posture (referred to as the first fitness score) to evaluate the accuracy of the output result of the predicted machine learning model.
[0026] An electronic device can build an initial machine learning model based on the lightweight TensorFlow Lite machine learning architecture. In the embodiments of this application, the training of the initial machine learning model depends on the edge impulse platform (hereinafter referred to as the Edge platform). Among them, the Edge platform is a machine learning framework specifically designed for MCU platforms with limited computing resources, which can realize data acquisition, model training, model export, etc. Specifically, first install the corresponding environment according to the requirements of the Edge platform, and connect the electronic device to the Edge platform. That is, the electronic device can obtain a large amount of raw sensor data, and the raw sensor data has been manually labeled or expert-labeled with the corresponding pet posture, and model training is carried out. In addition, the Edge platform is also integrated in the electronic device. For the convenience of understanding, the specific training process of the preset machine learning model will be described below.
[0027] Step 1: Obtain the raw sensor data corresponding to each motion sensor unit in multiple motion sensor units to obtain a set of raw sensor data, and each raw sensor data in the set of raw sensor data includes annotation information.
[0028] The electronic device first obtains the raw sensor data corresponding to each sensor unit in multiple motion sensor units, as well as the annotation information corresponding to each raw sensor data, to obtain a set of raw sensor data. For example, the electronic device can obtain a large amount of raw sensor data with standard information from a pet hospital or a cloud public database. The above-mentioned method for obtaining the raw sensor data is only an example, and it can also be obtained by other methods, which is not limited in this application.
[0029] In the embodiments of this application, when the electronic device obtains the set of raw sensor data, it sends it to the Edge platform according to the format required by the Edge platform. When the Edge platform receives the set of raw sensor data, it saves the data through the corresponding page. Before training, the electronic device can pre-determine configuration information such as data format and algorithm selection; during training, users can be allowed to set configurable training parameters.
[0030] Step 2: Preprocess each raw sensor data in the set of raw sensor data to obtain a first training sample set.
[0031] After obtaining the set of raw sensor data, the electronic device can preprocess each piece of raw sensor data in the set of raw sensor data to obtain a first training sample set. Specifically, the electronic device can traverse each piece of raw sensor data in the set of raw sensor data to verify the integrity of the raw sensor data. Then, it can use the Kalman filter algorithm to select state variables according to the pet's motion characteristics, define observation variables based on the raw sensor data, and then define the initial state covariance matrix, process noise covariance, observation noise covariance, state transition matrix, and observation matrix. Finally, it performs iterative processing on each data point and outputs the current optimal estimate, that is, the smoothed raw sensor data.
[0032] Step 3: Build an initial machine learning model based on the TensorFlow Lite machine learning architecture.
[0033] In the embodiment of the present application, the electronic device can build an initial machine learning model based on the lightweight TensorFlow Lite machine learning architecture.
[0034] Step 4: Input the first training sample into the initial machine learning model to obtain a first prediction result corresponding to the first training sample, where the first training sample is any training sample in the first training sample set.
[0035] After building the initial machine learning model and preprocessing the raw sensor data to obtain the first training sample set, a training sample can be arbitrarily selected from the first training sample set as the first training sample, and the first training sample is input into the initial machine learning model to obtain the prediction result corresponding to the first training sample, which is called the first prediction result.
[0036] Step 5: Adjust the initial machine learning model based on the first prediction result and the annotation information corresponding to the first training sample.
[0037] Since each training sample in the first training sample set includes standard information, after the electronic device predicts and identifies the first training sample through the initial machine learning model to obtain the first prediction result, it can adjust the model parameters of the initial machine learning model based on the difference between the first prediction result and the standard information corresponding to the first training sample. That is, the electronic device can determine the difference between the output values of the last layer structure during training of the training sample and adjust the initial machine learning model through this difference.
[0038] Step 6: Execute Step 4 and Step 5 based on the adjusted initial machine learning model until a preset iteration termination condition is met.
[0039] Step 7: Determine the initial machine learning model when the preset iteration termination condition is satisfied as the preset machine learning model.
[0040] During the iterative training process, the electronic device can determine whether the preset iteration termination condition is satisfied. If the preset iteration termination condition is satisfied, stop the iterative training and execute Step 7; if the preset iteration termination condition is not satisfied, repeat Steps 4 and 5 until the preset iteration termination condition is satisfied. Among them, the iteration termination condition can be set such that the number of iterations meets a preset number threshold (for example, 1000 times), or the iteration time meets a preset time threshold (for example, 5 minutes), or the loss function of the initial machine learning model converges, that is, the loss function of the initial machine learning model does not change significantly after multiple iterations.
[0041] It should be noted that in practical applications, the electronic device can also use other conditions as the iteration termination condition, which is not specifically limited here.
[0042] When it is determined that the preset iteration termination condition is satisfied, the initial machine learning model when the preset iteration termination condition is satisfied is determined as the preset machine learning model, thereby completing the training.
[0043] In some embodiments, when it is determined that the preset iteration termination condition is satisfied, the electronic device can also use a pre-prepared validation set to verify the recognition accuracy of the model. When it is determined that the recognition accuracy meets a preset threshold (for example, 80%), the model is allowed to be exported to obtain the preset machine learning model.
[0044] S23. Process the pet image according to the preset neural network model to obtain a predicted pet expression.
[0045] Refer to together Figure 3 , when the pre-processed pet image is obtained, input the pet image into a pre-trained neural network model (referred to as the preset neural network model) so as to output a corresponding prediction result through the preset neural network model, which is called the predicted pet expression. At the same time, when the preset neural network model outputs the predicted pet expression, it can also output a fitness score corresponding to the predicted pet expression (referred to as the second fitness score) to evaluate the accuracy of the output result of the predicted neural network model.
[0046] The electronic device can construct a prediction neural network model based on the YOLO convolutional neural network. In the embodiments of the present application, based on YOLOV8 as the basic model, that is, using YOLOV8 as the initial neural network model to train the pet original image dataset. For the convenience of understanding, the specific training process of the preset machine learning model is described below.
[0047] Step 1: Obtain the pet original images corresponding to each image acquisition unit among multiple image acquisition units to obtain a set of pet original images, where each pet original image in the set of pet original images includes annotation information.
[0048] The electronic device can first obtain the pet original images corresponding to each image acquisition unit among multiple image acquisition units, as well as the standard information corresponding to each pet original image, to obtain a set of pet original images. For example, the electronic device can download the relevant original picture dataset of pet expressions from open source websites such as kaggle and roboflow, and perform data marking through the open source tool labelImg. The above-mentioned method for obtaining pet original images is only an example, and it can also be obtained through other methods, which is not limited in this application.
[0049] After obtaining a sufficient amount of marked pet original images, they can be stored according to the following format and path rules, with a part as the training set and a part as the validation set. The optimal data volume ratio between the training set and the validation set is 8 / 2. To facilitate understanding of the inventive concept of this embodiment, the following provides a tool code example to represent the storage format and storage path planning of pet original images: my_dataset / ├── train / │├── images / # Training set images │└── labels / # Training set labels (YOLO format) └── valid / ├── images / # Training set images └── labels / # Validation set labels (YOLO format) In addition, the electronic device needs to specify the paths of the training set and the validation set, and the categories included in the known set of pet original images, and save the set of pet original images in the format of a yaml configuration file, which is convenient for data import during subsequent training. The following provides a tool code example to represent the yaml configuration file: train:. / data / train / images val:. / data / valid / images nc: 4 # class names names : 0: angry 1: happy 2: relaxed 3: sad Step 2: Preprocess each pet original image in the pet original image set to obtain a second training sample set.
[0050] After the electronic device obtains the pet original image set, it can preprocess each pet original image in the pet original image set to obtain a second training sample set. Specifically, the electronic device can traverse each pet original image in the pet original image set and perform data cleaning on the pet images. For example, filtering out invalid data (including blank / damaged image detection, EXIF information verification, etc.), screening low-quality images (including blur detection, abnormal lighting detection, etc.) to eliminate invalid / low-quality images, standardizing the data format. After the cleaning of the pet original images is completed, noise reduction processing can be performed using filtering algorithms (such as Gaussian filtering, bilateral filtering, etc.) to eliminate image noise, and image enhancement can be performed, such as enhancing the contrast using histogram equalization, and / or adjusting the image brightness using Gamma correction, etc.
[0051] Step 3: Build an initial neural network model based on the YOLO neural network.
[0052] In the embodiment of this application, for model training based on YOLOV8, it depends on python. The electronic device first needs to install the python environment and the Ultralytics library, and verify whether the python environment and the Ultralytics library are installed successfully. When model training is required, import the initial neural network model configuration file, set configurations such as the number of training rounds, and then start training. For example: Yolo detect train data=my_dataset.yaml model=yolov8n.pt epochs=100 imgsz=640 batch=16 Step 4: Input the second training sample into the initial neural network model to obtain a second prediction result corresponding to the second training sample, where the second training sample is any training sample in the second training sample set.
[0053] After constructing the initial neural network model and preprocessing the pet original images to obtain the second training sample set, any training sample can be selected from the second training sample set as the second training sample, and this second training sample is input into the initial neural network model to obtain the prediction result corresponding to the second training sample, which is called the second prediction result. Step 5: Adjust the initial neural network model based on the second prediction result and the annotation information corresponding to the second training sample.
[0054] Since each training sample in the second training sample set includes standard information, after the electronic device predicts and identifies the second training sample through the initial neural network model and obtains the second prediction result, it can adjust the model parameters of the initial neural network model based on the difference between the second prediction result and the standard information corresponding to the second training sample. That is, the electronic device can determine the difference between the output values of the last layer structure during training of the training samples and adjust the initial neural network model through this difference.
[0055] Step 6: Based on the adjusted initial neural network model, execute Step 4 and Step 5 until the preset iteration termination condition is met.
[0056] Step 7: Determine the initial neural network model when the preset iteration termination condition is met as the preset neural network model.
[0057] During the iterative training process, the electronic device can determine whether the preset iteration termination condition is met. If the preset iteration termination condition is met, stop the iterative training and execute Step 7; if the preset iteration termination condition is not met, repeat Step 4 and Step 5 until the preset iteration termination condition is met. Among them, the iteration termination condition can be set such that the number of iterations meets the preset number threshold (for example, 1000 times), or it can be set such that the iteration time meets the preset time threshold (for example, 5 minutes), or it can be set that the loss function of the initial machine learning model converges, that is, the loss function of the initial machine learning model no longer changes significantly after multiple iterations.
[0058] It should be noted that in practical applications, the electronic device can also use other conditions as the iteration termination condition, which is not specifically limited here.
[0059] When it is determined that the preset iteration termination condition is met, the initial neural network model when the preset iteration termination condition is met is determined as the preset neural network model, thus completing the training. In some embodiments, when it is determined that the preset iteration termination condition is met, the electronic device can also use a pre-prepared validation set to verify the recognition accuracy of the model. When it is determined that the recognition accuracy meets the preset threshold (for example, 80%), the model is allowed to be exported to obtain the preset neural network model. To understand the inventive concept of this embodiment, the following provides a tool code example to represent model verification: import os from ultralytics import Y0L0 # Import the yolo module model=Y0L0('best.pt') # Load the trained model weight file model.predict(filename, save_txt=True, save=True, save_conf=True) It should be noted that the execution order of step S22 and step S23 can be swapped. Additionally, step S22 and step S23 can be executed simultaneously.
[0060] S24, perform a fusion process on the predicted pet posture and the predicted pet expression, and output the pet posture and expression.
[0061] After the predicted pet posture of the pet is output through a preset machine learning model and the predicted pet expression of the pet is output through a preset neural network model, the electronic device can verify whether the predicted pet posture and the predicted pet expression are inconsistent, so as to correctly output the pet posture and pet expression of the pet. Specifically, the electronic device can perform a fusion process on the predicted pet posture and the predicted pet expression, and use the output fusion result as the pet posture and expression. In some embodiments, the electronic device has the ability to display the pet posture and pet expression to the pet owner in real time through a mobile application and / or a web platform.
[0062] In an alternative embodiment, the fusing of the predicted pet posture and the predicted pet expression to output the pet posture and expression includes: Perform time and space alignment on the predicted pet posture and the predicted pet expression; Determine the first modal weight of the predicted pet posture and the second modal weight of the predicted pet expression; Perform a weighted process based on the predicted pet posture and the first modal weight, and the predicted pet expression and the second modal weight to obtain a fusion result; Output the pet posture and expression according to the fusion result.
[0063] In some embodiments, since the sampling and operation frequency of pet movement data is relatively high (e.g., the movement sensor is 100Hz), while the data frequency of pet images collected by the camera device is lower (e.g., 30fps), and both pet movement data and pet images include timestamp information, the electronic device can use a sliding window to synchronize the data, so as to achieve the alignment of predicted pet postures and predicted pet expressions in time and space. Specifically, a mapping relationship between pixel coordinates and sensor coordinates is established through camera calibration (such as checkerboard calibration). For example, the pet's head position captured by the camera is projected into the sensor coordinate system. Exemplarily, the size of the sliding window is set to an integer multiple of the camera frame interval (e.g., 100ms, corresponding to 3 frames of sensor data + 1 frame of camera data). After achieving the alignment of predicted pet postures and preset pet expressions in time and space, due to the rapid movement of the pet resulting in blurred images, more reliance on pet movement data is required, while in static scenarios, pet images can better display emotions. Therefore, the algorithm flexibly adjusts the modal weights according to the scene requirements, rather than fusing at a fixed ratio. In the embodiments of the present application, an attention mechanism is used to dynamically combine the features of the two modalities. Specifically, the electronic device is based on a pre-trained lightweight CNN (such as MobileNetV2) to classify the camera scene. For example, a static scene (such as a pet lying down): label = 0; a dynamic scene (such as running, jumping): label = 1. And the attention weights are calculated using the attention mechanism. For example, in a static scene, the camera weight = 0.8, and the sensor weight = 0.2; in a dynamic scene, the camera weight = 0.3, and the sensor weight = 0.7, and the weights are updated once every preset time period (e.g., 100ms). The weight values are smoothed through the Sigmoid function to avoid feature jumps caused by sudden weight changes. Finally, weighted splicing is performed based on the sensor features (i.e., predicted pet postures), camera features (i.e., predicted pet expressions), and the corresponding modal weights. Specifically, , where is the fusion result, is the modal weight corresponding to the predicted pet posture (i.e., the first modal weight), is the predicted pet posture, is the modal weight corresponding to the predicted pet expression (i.e., the second modal weight), is the predicted pet expression.
[0064] In an alternative embodiment, before fusing the predicted pet posture and the predicted pet expression, the electronic device may also verify the accuracy of the predicted pet posture and the predicted pet expression output by the model. Specifically, the electronic device may calculate the kurtosis of the acceleration signal. If the kurtosis > 5, it is determined as hair interference (the kurtosis of normal movement ≈ 3), and it is detected whether the data loss rate within 100 ms continuously exceeds 30%; at the same time, calculate the image entropy of the current frame. If it is lower than the threshold (such as 5.5), it is determined as low light / blur, and it is detected whether the pet face area is complete (by whether the aspect ratio of the face detection box is within the range of [0.8, 1.2]). Then, define two gating variables, which are and . If , , only the predicted pet posture is output; if , , only the predicted pet expression is output; if , , indicating that both the predicted pet posture and the predicted pet expression are reliable, then the fusion process is performed.
[0065] In an alternative embodiment, the method further includes: Obtain a first fitness score of the predicted pet posture and a second fitness score of the predicted pet expression; When it is determined that both the first fitness score and the second fitness score are greater than a preset threshold, output the pet posture and the pet expression; When it is determined that the first fitness score is greater than the preset threshold and the second fitness score is less than the preset threshold, output the pet posture; When it is determined that the first fitness score is less than the preset threshold and the second fitness score is greater than the preset threshold, output the pet expression.
[0066] In some embodiments, while the preset machine learning model outputs the predicted pet posture of the pet, it also outputs the first fitness score corresponding to the predicted pet posture. Similarly, while the prediction neural network model outputs the predicted pet expression of the pet, it also outputs the second fitness score corresponding to the predicted pet expression. Among them, the first fitness score represents the prediction accuracy of the predicted pet posture, and the second fitness score represents the prediction accuracy of the predicted pet expression. The electronic device can preset a fitness threshold (for example, 0.8), and compare the first fitness score and the second fitness score with the preset threshold. When both the first fitness score and the second fitness score are greater than the preset threshold, the combined result of the pet posture and the pet expression is output, such as lying prone + happy; when the first fitness score is greater than the preset threshold and the second fitness score is less than the preset threshold, only the pet posture is output, such as jumping, and the pet expression is discarded; when the first fitness score is less than the preset threshold and the second fitness score is greater than the preset threshold, only the pet expression is output, such as angry, and the pet posture is discarded. Or when both the first fitness score and the second fitness score are less than the preset threshold, no result is output, an error code is returned to the pet owner, and the reprocessing of the next frame of data is awaited. In other embodiments, the electronic device can set different fitness thresholds, including a first threshold and a second threshold, compare the first fitness score with the first threshold, compare the second fitness score with the second threshold. When the first fitness score is greater than the first threshold and the second fitness score is greater than the second threshold, the combined result of the pet posture and the pet expression is output; when the first fitness score is greater than the first threshold and the second fitness score is less than the second threshold, only the pet posture is output; when the first fitness score is less than the first threshold and the second fitness score is greater than the second threshold, only the pet expression is output; when the first fitness score is less than the first threshold and the second fitness score is less than the second threshold, no result is output.
[0067] In an alternative embodiment, the electronic device can pre-acquire known pet postures and known pet expressions, store the known pet postures and known pet expressions in different databases to obtain corresponding data lists, and at the same time determine the corresponding matching degree between the known pet postures and the known pet expressions. A pet posture-expression relationship table is constructed based on the data list and the corresponding matching degree. When the predicted pet posture and the predicted pet expression are obtained, the predicted pet posture and the predicted pet expression are input into the pet posture-expression relationship table to determine the matching degree between them. When it is determined that the matching degree is greater than the preset matching degree threshold, the fusion result of the predicted pet posture and the predicted pet expression is output. When it is determined that the matching degree is less than the preset matching degree threshold, it can be determined that the predicted pet posture and the predicted pet expression are contradictory. For example, if the pet posture is sleeping and the pet expression is excited, only the pet posture or the pet expression is output.
[0068] In an alternative embodiment, the electronic device can extract behavioral characteristics based on the collected pet movement data, including the daily total exercise volume (which can be obtained by integrating the L2 norm of the triaxial acceleration of the sensor), the daily lying duration (which can be obtained by the duration of the stationary state of the sensor), the gait stability (which can be obtained by the standard deviation of the step frequency of the acceleration signal), etc. Train a Gaussian mixture model (GMM) based on the pet movement data for a preset time period (such as seven days), and calculate the mean value of each feature i and the standard deviation of each feature, and define an anomaly threshold. For example, if the exercise volume on a certain day is lower than a certain value, it is marked as abnormal. Then, calculate the pain probability based on the anomaly degree of multiple features. Specifically, where is the pain probability, is the feature weight. For example, the exercise volume weight = 0.6, the lying duration weight = 0.4, and is the actual measured value of the feature i on a certain day. When the pain probability is obtained, the pain probability can be compared with a preset pain threshold (for example, 0.7). When the pain probability is greater than the preset pain threshold, a pain warning is issued to inform the pet owner.
[0069] In an alternative embodiment, the method further includes: Judging whether the pet has abnormal state behaviors according to the pet posture and expression; When it is determined that the pet has abnormal state behaviors, send a warning signal to prompt the pet owner.
[0070] Refer to Figure 3 as well. In some embodiments, the electronic device can pre - construct an abnormal behavior database. When obtaining the pet posture and pet expression, the electronic device can detect in real - time whether the pet has abnormal behavior states based on the pet posture and pet expression in the abnormal behavior database. For example, abnormal gaits, high - frequency tremors, excessive howling, etc., and then immediately send a warning signal to prompt the pet owner. Among them, the abnormal behavior database can include feature templates of the following behaviors (including falls, epileptic seizures, rear - end collisions, etc.), with 500 samples for each type of behavior. For example, continuous stillness after an acceleration mutation (peak value > 5g) represents a fall, periodic high - frequency vibrations (main frequency 10 - 15Hz) represent epileptic seizures, and rapid head rotation (angular velocity > 100° / s) and self - intersecting trajectories represent rear - end collisions. When the fusion result is obtained, calculate the DTW distance (Dynamic Time Warping) between the fusion result and each feature template in the abnormal behavior database. When there is a certain feature template and the DTW distance satisfies If it is less than a preset distance threshold (which can be set through cross-validation), then the abnormal behavior corresponding to the feature template is determined. When it is determined that there is an abnormal behavior, when the electronic device can trigger an alarm, the corresponding information (including the type of abnormal behavior, the occurrence timestamp, the fitness score, etc.) is uploaded to the service platform, and a behavior log is generated.
[0071] In an alternative embodiment, the method further includes: Determine the pet's movement trajectory, pet's posture change, and pet's emotional state based on the pet movement data and the pet image; Visualize the pet's movement trajectory, the pet's posture change, and the pet's emotional state.
[0072] In some embodiments, after obtaining the pet movement data and the pet image, the electronic device can obtain the position information of the pet at different times based on the pet movement data, so as to draw the pet's movement trajectory. For example, record the longitude and latitude coordinates of the pet at regular intervals (such as 1 second) during outdoor activities, and connect these coordinate points in chronological order to form the pet's movement trajectory. Or based on the pet image, track the position of the pet in the image through target tracking algorithms (such as Kalman filtering, optical flow method, etc.). For example, in consecutive image frames, identify the physical characteristics (such as color, shape, etc.) of the pet and track the changes of these characteristics to determine the movement path of the pet in the image, and then map it to the movement trajectory in the actual scenario. At the same time, the electronic device can also measure the acceleration and angular velocity changes of the pet through the pet movement data, and infer the posture change of the pet through analysis. For example, when the accelerometer detects a large acceleration change in the vertical direction of the pet, it may indicate that the pet is jumping or sitting down and other posture changes; in addition, the electronic device can also use a deep learning model (such as convolutional neural network CNN) to estimate the posture of the pet image. For example, by training a model to identify the positions of the key points of the pet's body (such as the head, limbs, tail, etc.), and then comparing the relative position changes of the key points at different times to determine the posture change of the pet, such as changing from a standing posture to a lying posture, etc. And, the electronic device can also determine the pet's emotional state based on the fused output of the pet's expression. After obtaining the pet's movement trajectory, the pet's posture change, and the pet's emotional state, the electronic device can visualize the pet's movement trajectory, the pet's posture change, and the pet's emotional state to prompt the pet owner about the movement, posture, and emotional changes of the pet.
[0073] Refer to together Figure 3, in some embodiments, the electronic device can also design emotional interaction games and communication tools according to pet expressions to enhance the emotional connection between the pet and the owner. The electronic device can preset different emotional interaction strategies according to different pet emotional states, where the emotional interaction strategies can include emotional interaction methods and emotional execution strategies. For example, assuming the pet's emotional state is happy, the corresponding emotional interaction method can be an interactive game (such as chasing light spots, sound rewards, etc.), and the corresponding emotional execution strategy can be to start the laser projector, play cheerful sound effects, etc. Assuming the pet's emotional state is tense or scared, the corresponding emotional interaction mode can be a soothing mode (such as soft music, slow interaction, etc.). The corresponding emotional execution strategy can be to dim the lights, play white noise, etc. Assuming the pet's emotional state is tired, the corresponding emotional interaction method can be low-intensity interaction (such as slowly moving a toy), and the corresponding emotional execution strategy can be to control the automatic cat teaser to swing at a low speed. Assuming the pet's emotional state is angry, the corresponding emotional interaction method is to stop the stimulation and avoid further irritation, and the corresponding emotional execution strategy can be to turn off the interactive device and remind the pet owner to intervene, etc. In addition, the electronic device can also optimize the interaction strategy by using reinforcement learning (PPO algorithm) and dynamically adjust the emotional interaction strategy according to the pet's reaction. In addition, the electronic device can generate corresponding personalized feeding suggestions, exercise plans, health monitoring programs, etc. based on the default data provided by the service platform, including the pet's normal activity status and feeding data, when obtaining the pet's behavior pattern and health status.
[0074] Referring to Figure 4 shown, is a functional block diagram of the pet posture and expression tracking device shown in the embodiments of the present application.
[0075] In some embodiments, the pet posture and expression tracking device 40 may include a plurality of functional modules composed of computer program segments. The computer programs of each program segment of the pet posture and expression tracking device 40 can be stored in the memory of the electronic device and executed by at least one processor to perform the functions of pet posture and expression tracking (see details in Figure 1 description). According to the functions it performs, it can be divided into multiple functional modules. The functional modules may include: an acquisition module 401, a first prediction module 402, a second prediction module 403, a fusion module 404, and a visualization module 405. The module referred to in the present application means a series of computer program segments that can be executed by at least one processor and can complete fixed functions, and are stored in the memory. In this embodiment, the functions of each module will be described in detail in subsequent embodiments.
[0076] The acquisition module 401 is used to acquire pet movement data and pet images.
[0077] The first prediction module 402 is configured to process the pet motion data according to a preset machine learning model to obtain a predicted pet posture.
[0078] The second prediction module 403 is configured to process the pet image according to a preset neural network model to obtain a predicted pet expression.
[0079] The fusion module 404 is configured to perform a fusion process on the predicted pet posture and the predicted pet expression, and output the pet posture and expression.
[0080] The fusion module 404 is further specifically configured to: perform time and space alignment on the predicted pet posture and the predicted pet expression; determine a first modality weight of the predicted pet posture and a second modality weight of the predicted pet expression; perform a weighted process according to the predicted pet posture and the first modality weight, and the predicted pet expression and the second modality weight to obtain a fusion result; output the pet posture and expression according to the fusion result.
[0081] The fusion module 404 is further configured to: obtain a first fitness score of the predicted pet posture and a second fitness score of the predicted pet expression; when it is determined that both the first fitness score and the second fitness score are greater than a preset threshold, output the pet posture and the pet expression; when it is determined that the first fitness score is greater than the preset threshold and the second fitness score is less than the preset threshold, output the pet posture; when it is determined that the first fitness score is less than the preset threshold and the second fitness score is greater than the preset threshold, output the pet expression.
[0082] The first prediction module 402 is further configured to: Step 1, obtain the sensor raw data corresponding to each motion sensor unit in a plurality of motion sensor units to obtain a sensor raw data set, and each sensor raw data in the sensor raw data set includes annotation information; Step 2, preprocess each sensor raw data in the sensor raw data set to obtain a first training sample set; Step 3, construct an initial machine learning model based on the TensorFlow Lite machine learning architecture; Step 4, input the first training sample into the initial machine learning model to obtain a first prediction result corresponding to the first training sample, where the first training sample is any training sample in the first training sample set; Step 5, adjust the initial machine learning model based on the first prediction result and the annotation information corresponding to the first training sample; Step 6, perform Step 4 and Step 5 based on the adjusted initial machine learning model until a preset iteration termination condition is met; Step 7, determine the initial machine learning model that meets the preset iteration termination condition as the preset machine learning model.
[0083] The second prediction module 403 is further configured to: Step 1, obtain the pet original images corresponding to each image acquisition unit in multiple image acquisition units to obtain a pet original image set, where each pet original image in the pet original image set includes annotation information; Step 2, preprocess each pet original image in the pet original image set to obtain a second training sample set; Step 3, construct an initial neural network model based on the YOLO neural network; Step 4, input the second training sample into the initial neural network model to obtain a second prediction result corresponding to the second training sample, where the second training sample is any training sample in the second training sample set; Step 5, adjust the initial neural network model based on the second prediction result and the annotation information corresponding to the second training sample; Step 6, execute Step 4 and Step 5 based on the adjusted initial neural network model until a preset iteration termination condition is met; Step 7, determine the initial neural network model when the preset iteration termination condition is met as the preset neural network model.
[0084] The visualization module 405 is configured to: judge whether the pet has abnormal state behaviors according to the pet posture and expression; when it is determined that the pet has abnormal state behaviors, send a warning signal to prompt the pet owner.
[0085] The visualization module 405 is further configured to: determine the pet movement trajectory, pet posture change and pet emotion state according to the pet movement data and the pet image; visualize the pet movement trajectory, the pet posture change and the pet emotion state.
[0086] It should be understood that the various change modes and specific embodiments in the pet posture and expression tracking method provided in the above embodiments are equally applicable to the pet posture and expression tracking device in this embodiment. Through the foregoing detailed description of the pet posture and expression tracking method, those skilled in the art can clearly know the implementation method of the pet posture and expression tracking device in this embodiment. For the sake of brevity of the specification, it will not be elaborated here.
[0087] Refer to Figure 5 As shown in the schematic structural diagram of the electronic device shown in the embodiments of the present application. In a preferred embodiment of the present application, the electronic device 5 includes a memory 51, at least one processor 52 and at least one communication bus 53.
[0088] Those skilled in the art should understand that Figure 5 the structure of the electronic device shown does not constitute a limitation to the embodiments of the present application. It can be a bus structure or a star structure. The electronic device 5 may further include more or fewer other hardware or software than shown in the figure, or different component arrangements.
[0089] In some embodiments, the electronic device 5 is a device capable of automatically performing numerical calculations and / or information processing according to pre-set or stored instructions. Its hardware includes, but is not limited to, microprocessors, application-specific integrated circuits, programmable gate arrays, digital signal processors, and embedded devices, etc. The electronic device 5 may further include a user device, and the user device includes, but is not limited to, any electronic product that can perform human-computer interaction with the user through means such as a keyboard, mouse, remote control, touchpad, or voice control device. For example, a personal computer, a tablet computer, a smart phone, a digital camera, etc.
[0090] In the above embodiments provided in the present application, it should be understood that the disclosed methods, devices, computer-readable storage media, and electronic devices can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division. In actual implementation, there may be other division methods. For example, multiple components or modules can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, indirect couplings or communication connections of devices or components or modules, and can be electrical, mechanical, or other forms.
[0091] The components described as separate components may or may not be physically separated, and the components shown as components may or may not be physical modules, that is, they may be located in one place, or may be distributed to multiple network modules. Some or all of the components can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0092] In addition, in each embodiment of the present invention, the functional modules can be integrated in a processing module, or each component can exist physically alone, or two or more modules can be integrated in one module. The above integrated modules can be implemented in the form of hardware or in the form of software functional modules.
[0093] When the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs.
[0094] It should be noted that, for the foregoing method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations. However, those skilled in the art should be aware that the present invention is not limited by the described action sequence, because according to the present invention, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily all essential to the present invention.
[0095] In the above embodiments, the descriptions of the various embodiments have their own focuses. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0096] The above is a specific description of the preferred embodiments of the present invention, but the present invention is not limited to the described embodiments. Those skilled in the art can also make various equivalent deformations or substitutions without departing from the spirit of the present invention, and these equivalent deformations or substitutions are all included in the scope defined by the claims of this application.
Claims
1. A method for tracking the posture and expression of a pet, characterized in that, The method includes: Obtaining pet movement data and pet images; Processing the pet movement data according to a preset machine learning model to obtain a predicted pet posture; Processing the pet images according to a preset neural network model to obtain a predicted pet expression; Performing fusion processing on the predicted pet posture and the predicted pet expression, and outputting the pet posture and expression.
2. The pet posture and expression tracking method according to claim 1, characterized in that The performing fusion on the predicted pet posture and the predicted pet expression and outputting the pet posture and expression includes: Performing temporal and spatial alignment on the predicted pet posture and the predicted pet expression; Determining a first modality weight of the predicted pet posture and a second modality weight of the predicted pet expression; Performing weighted processing according to the predicted pet posture and the first modality weight, and the predicted pet expression and the second modality weight to obtain a fusion result; Outputting the pet posture and expression according to the fusion result.
3. The pet posture and expression tracking method according to claim 1, wherein The method further includes: Obtaining a first fitness score of the predicted pet posture and a second fitness score of the predicted pet expression; When it is determined that both the first fitness score and the second fitness score are greater than a preset threshold, outputting the pet posture and the pet expression; When it is determined that the first fitness score is greater than the preset threshold and the second fitness score is less than the preset threshold, outputting the pet posture; When it is determined that the first fitness score is less than the preset threshold and the second fitness score is greater than the preset threshold, outputting the pet expression.
4. The pet posture and expression tracking method according to claim 1, characterized in that The method further includes: Step 1: Obtaining the sensor raw data corresponding to each motion sensor unit in a plurality of motion sensor units to obtain a sensor raw data set, where each sensor raw data in the sensor raw data set includes annotation information; Step 2: Preprocessing each sensor raw data in the sensor raw data set to obtain a first training sample set; Step 3: Constructing an initial machine learning model based on the TensorFlow Lite machine learning architecture; Step 4: Inputting the first training sample into the initial machine learning model to obtain a first prediction result corresponding to the first training sample, where the first training sample is any training sample in the first training sample set; Step 5: Adjusting the initial machine learning model based on the first prediction result and the annotation information corresponding to the first training sample; Step 6: Executing Step 4 and Step 5 based on the adjusted initial machine learning model until a preset iteration termination condition is met; Step 7: Determining the initial machine learning model that meets the preset iteration termination condition as the preset machine learning model.
5. The pet posture and expression tracking method according to claim 1, wherein The method further includes: Step 1: Obtaining the pet raw images corresponding to each image acquisition unit in a plurality of image acquisition units to obtain a pet raw image set, where each pet raw image in the pet raw image set includes annotation information; Step 2: Preprocessing each pet raw image in the pet raw image set to obtain a second training sample set; Step 3: Constructing an initial neural network model based on the YOLO neural network; Step 4: Input the second training sample into the initial neural network model to obtain a second prediction result corresponding to the second training sample, where the second training sample is any training sample in the second training sample set; Step 5: Adjust the initial neural network model based on the second prediction result and the annotation information corresponding to the second training sample; Step 6: Execute Step 4 and Step 5 based on the adjusted initial neural network model until a preset iteration termination condition is met; Step 7: Determine the initial neural network model when the preset iteration termination condition is met as the preset neural network model.
6. The pet posture and expression tracking method according to claim 1, characterized in that The method further includes: Judging whether the pet has abnormal state behaviors according to the pet posture and expression; When it is determined that the pet has abnormal state behaviors, sending a warning signal to prompt the pet owner.
7. The pet posture and expression tracking method according to claim 1, wherein The method further includes: Determining the pet movement trajectory, pet posture change and pet emotion state according to the pet movement data and the pet image; Visualizing the pet movement trajectory, the pet posture change and the pet emotion state.
8. A pet posture and expression tracking device, characterized in that, The device includes: An acquisition module, configured to acquire pet movement data and pet images; A first prediction module, configured to process the pet movement data according to a preset machine learning model to obtain a predicted pet posture; A second prediction module, configured to process the pet image according to a preset neural network model to obtain a predicted pet expression; A fusion module, configured to perform a fusion process on the predicted pet posture and the predicted pet expression and output the pet posture and expression.
9. An electronic device, characterized in that, It includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the steps of the pet posture and expression tracking method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, the steps of the pet posture and expression tracking method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Pet recognition method combining face and voiceprint
CN108734114A
Pet searching method and device based on image recognition
CN111507302A
Pet state evaluation system, pet camera, server, pet state evaluation method, and program
CN115885313A
Emotion recognition method and device based on facial expressions and human body postures
CN119027997A
AI-based pet emotion recognition system
CN119049086A
Cited By
Pet purifier control method and device, pet purifier, medium and product
CN120858895A