Animal three-dimensional attitude estimation method, device and equipment and storage medium
By constructing a model based on key point error in animal three-dimensional pose estimation, the problem of manual labeling error impact is solved, the accuracy of the estimation and anti-interference ability are improved, and the adaptability to different individuals and environments is adapted.
Patent Information
- Application Number
- CN202510181372.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-19
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-02-19
AI Technical Summary
The errors of manual labeling in the prior art seriously affect the accuracy of animal three-dimensional pose estimation, and most methods are fully supervised and rely on a large number of labeled training data.
By building a multi-view shooting platform to collect video data, build training data sets and prediction data sets, and build a three-dimensional pose estimation model of animal based on key point errors. The model includes a multi-view volume 3D pose estimation network and a framework designed for different key point errors, using L1 loss and time smoothness constraints for training, adjusting the weight of the loss to reduce key point errors.
It effectively reduces key point errors, improves the accuracy and anti-interference ability of animal three-dimensional pose estimation, enhances the generalization ability of the model, and adapts to differences between different types and individuals.
Smart Images

Figure CN120108036A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a method, device, equipment and storage medium for estimating three-dimensional posture of an animal. Background Art
[0002] Animal 3D posture estimation is a key technology for studying animal behavior and is widely used in fields such as neuroscience, biology, and animal behavior. Its goal is to recover the posture information of animals in 3D space from 2D video data of animals.
[0003] However, due to the diversity of animal morphology, complex background environment, and the influence of factors such as perspective and lighting, the 3D pose estimation of animals is extremely challenging. In existing technologies, most animal 3D pose estimation methods rely on deep learning technology, especially convolutional neural networks. These methods usually require labeled training data to train the model to predict the key points of animals in video frames. However, manually labeled training samples will have errors caused by manual annotation, which greatly affects the accuracy of these methods. In addition, most of the existing animal 3D pose estimation methods are fully supervised, which means that they rely entirely on labeled training data during the training process. This method will cause the errors of manual annotation to seriously affect the final training results, thereby affecting the final results. Therefore, how to effectively avoid key point errors and improve the performance of animal 3D pose estimation is a technical problem that needs to be solved urgently.
[0004] The above contents are only used to assist in understanding the technical solution of the present invention and do not constitute an admission that the above contents are prior art. Summary of the invention
[0005] The main purpose of the present invention is to provide a method, device, equipment and storage medium for estimating animal three-dimensional posture, aiming to solve the technical problems in the prior art that the errors of manual annotation seriously affect the final training results and the performance of animal three-dimensional posture estimation is poor.
[0006] To achieve the above object, the present invention provides a method for estimating a three-dimensional posture of an animal, the method comprising:
[0007] Collect experimental animal video data based on the built multi-view shooting platform;
[0008] Constructing a training data set and a prediction data set based on the experimental animal video data;
[0009] Performing frame screening processing on the training data set;
[0010] Constructing an animal 3D posture estimation model based on key point errors, wherein the animal 3D posture estimation model includes a multi-view volumetric 3D posture estimation network and a framework designed for different key point errors;
[0011] The training data set after frame screening and the prediction data set are combined into a time block and input into the animal 3D posture estimation model for training, and the loss during training is weighted and adjusted using a framework designed for different key point errors;
[0012] The trained model is used to perform key point recognition and posture estimation on experimental animals.
[0013] Preferably, the multi-view shooting platform comprises a plurality of cameras, and the plurality of cameras are calibrated using a checkerboard calibration method to obtain their intrinsic and extrinsic parameters;
[0014] The multi-view shooting platform constructed to collect experimental animal video data includes:
[0015] The experimental animal video data is collected by setting the cameras of the multi-view shooting platform to synchronously collect the video data.
[0016] Preferably, constructing a training data set and a prediction data set based on the experimental animal video data includes:
[0017] Using a 3D marking software package to mark key points of the experimental animal video data to construct a training data set;
[0018] Construct an unlabeled dataset as a prediction dataset.
[0019] Preferably, the frame screening process of the training data set includes:
[0020] Extracting high-dimensional features from each frame of the training data set using a deep learning pre-training model, and representing the features of all frames as a feature matrix;
[0021] Clustering the feature matrix using a k-means clustering algorithm to obtain K clusters;
[0022] Calculate the distance between each frame of the K clusters and the center of the corresponding cluster, and select the frame with the smallest distance as the key frame;
[0023] According to the selected key frame index, the corresponding frame is extracted from the original video and saved as a representative frame of the video.
[0024] Preferably, the multi-view volumetric 3D pose estimation network is a 3D convolutional neural network, and the framework designed for different key point errors is constructed based on L1 loss and temporal smoothness constraints;
[0025] The method of constructing an animal three-dimensional posture estimation model based on key point errors includes:
[0026] Use projective geometry to construct a metric 3D feature space;
[0027] Inferring landmark locations using shared features across cameras and learned spatial statistics of animal poses via the 3D convolutional neural network;
[0028] The L1 loss is used to perform a standard supervised pose regression loss on labeled frames and establish a temporal constraint expression for the temporal smoothness constraint between keypoint coordinates.
[0029] Preferably, inferring landmark positions using shared features across cameras and learned spatial statistics of animal postures by the 3D convolutional neural network comprises:
[0030] Use a standard 2D U-Net to detect the center of mass of the animal in each view and infer the 3D center of mass of the animal through triangulation;
[0031] A 3D volume frame that can accommodate the entire animal is constructed according to the position relationship of multiple cameras and the 3D center of mass of the animal, and the 3D volume frame is processed by the 3D convolutional neural network to predict the position of 3D landmarks.
[0032] Preferably, the step of using the trained model to perform key point recognition and posture estimation on the experimental animal includes:
[0033] The trained model is used to predict the key points of each frame of the experimental animal, and the relative position change relationship of the key points is used to estimate the animal's posture.
[0034] In addition, to achieve the above-mentioned purpose, the present invention also proposes an animal three-dimensional posture estimation device comprising:
[0035] A collection module is used to collect experimental animal video data based on the built multi-view shooting platform;
[0036] A construction module, used to construct a training data set and a prediction data set based on the experimental animal video data;
[0037] A screening module, used for performing frame screening processing on the training data set;
[0038] The construction module is also used to construct an animal 3D posture estimation model based on key point errors, wherein the animal 3D posture estimation model includes a multi-view volume 3D posture estimation network and a framework designed for different key point errors;
[0039] A training module, used for inputting a time block of a training data set and a prediction data set after frame screening into the animal 3D posture estimation model for training, and adjusting the weight of the loss during training by using a framework designed for different key point errors;
[0040] The recognition module is used to use the trained model to perform key point recognition and posture estimation on experimental animals.
[0041] In addition, to achieve the above-mentioned purpose, the present invention also proposes an animal three-dimensional posture estimation device, on which an animal three-dimensional posture estimation program is stored, and when the animal three-dimensional posture estimation program is executed by a processor, the steps of the animal three-dimensional posture estimation method described above are implemented.
[0042] In addition, to achieve the above-mentioned purpose, the present invention also proposes a storage medium, on which an animal three-dimensional posture estimation program is stored. When the animal three-dimensional posture estimation program is executed by a processor, the steps of the animal three-dimensional posture estimation method described above are implemented.
[0043] Compared with the prior art, the present invention can achieve at least the following beneficial effects:
[0044] (1) Targeted error correction: The method of this embodiment can focus on key point errors. In the three-dimensional posture estimation of animals, different key points have different importance for the accurate description of the posture. By analyzing the key point errors, those key points with inaccurate estimates can be accurately found. For example, for the estimation of the walking posture of quadrupeds, the errors of the key points of the limb joints and the spine have a greater impact on the overall posture. Through error analysis, the weights of the estimated errors of these key parts are adjusted, which can more accurately restore the true three-dimensional posture of the animal and reduce the deviation of the overall posture estimation.
[0045] (2) Adapting to individual differences: Different species of animals, and even different individuals of the same species, have different body structures and movement patterns. This method can adjust the weights based on the errors in key point estimation for each individual animal. For example, dogs and mice have different limb proportions. When errors occur in key point estimation of the legs when estimating the posture of a dog, the weight adjustment strategy required may be different from that required when estimating the same part of the legs for a mouse. Through error analysis, more appropriate posture estimation models can be customized for different individuals, thereby improving the accuracy of posture estimation for various animals.
[0046] (3) Improved anti-interference ability: In actual scenes, animal posture estimation may be interfered by many factors, such as changes in lighting, changes in animal surface features (such as hair length, color changes, etc.), and occlusion of some body parts. By adjusting the weights through key point error analysis, the model can pay more attention to key points that are less interfered with and more accurately estimated, reducing the impact of interference factors on the overall posture estimation. For example, when one side of the animal's body is obscured by a shadow, by giving higher weights to the key points on the other side that are not obscured and have more accurate estimates, the model can overcome the shadow interference to a certain extent and still estimate the animal's three-dimensional posture more accurately.
[0047] (4) Enhanced generalization ability: This method helps the model work better in different data sets and scenarios. If the model has been trained with weight adjustment based on key point error analysis, it can adapt and adjust the posture estimation strategy more quickly when faced with new animal species, new movement patterns, or new environmental conditions. For example, when converting from an animal posture estimation dataset in a laboratory environment to a dataset in a wild environment, the model can readjust the weights based on the key point errors in the new dataset, thereby improving the accuracy of animal 3D posture estimation in complex wild environments and enhancing the generalization ability of the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] Figure 1 It is a structural schematic diagram of an animal three-dimensional posture estimation method in a hardware operating environment involved in an embodiment of the present invention;
[0049] Figure 2 It is a schematic diagram of the flow chart of the first embodiment of the method for estimating the three-dimensional posture of an animal according to the present invention;
[0050] Figure 3 (a) is a schematic diagram of a multi-view shooting platform in an embodiment of the animal three-dimensional posture estimation method of the present invention, Figure 3 (b) is a schematic diagram of the shooting angle of view of the multi-view shooting platform;
[0051] Figure 4 This is a schematic diagram of the key point structure marked in the embodiment of the animal three-dimensional posture estimation method of the present invention;
[0052] Figure 5 An image of a camera view in an embodiment of the method for estimating three-dimensional posture of an animal of the present invention;
[0053] Figure 6 This is a structural block diagram of the first embodiment of the animal three-dimensional posture estimation device of the present invention.
[0054] The realization of the purpose, functional features and advantages of the present invention will be further explained in conjunction with embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION
[0055] It should be understood that the specific embodiments described herein are only used to explain the present invention, and are not used to limit the present invention.
[0056] Reference Figure 1 , Figure 1 The present invention is a schematic diagram of the structure of an animal three-dimensional posture estimation device in the hardware operating environment involved in the embodiment of the present invention.
[0057] like Figure 1 As shown, the animal three-dimensional posture estimation device may include: a processor 1001, such as a central processing unit (CPU), a communication bus 1002, a user interface 1003, a network interface 1004, and a memory 1005. Among them, the communication bus 1002 is used to realize the connection and communication between these components. The user interface 1003 may include a display screen (Display), and the optional user interface 1003 may also include a standard wired interface and a wireless interface. The wired interface of the user interface 1003 may be a USB interface in the present invention. The network interface 1004 may optionally include a standard wired interface and a wireless interface (such as a wireless fidelity (WIreless-FIdelity, WI-FI) interface). The memory 1005 may be a high-speed random access memory (Random Access Memory, RAM) memory, or a stable memory (Non-volatile Memory, NVM), such as a disk memory. The memory 1005 may also be a storage device independent of the aforementioned processor 1001.
[0058] Those skilled in the art will understand that Figure 1 The structure shown in the figure does not constitute a limitation on the animal three-dimensional posture estimation device, and may include more or less components than shown in the figure, or combine certain components, or arrange the components differently.
[0059] like Figure 1 As shown, the memory 1005 as a computer storage medium may include an operating system, a network communication module, a user interface module, and an animal three-dimensional posture estimation program.
[0060] exist Figure 1 In the animal three-dimensional posture estimation device shown, the network interface 1004 is mainly used to connect to the background server and communicate data with the background server; the user interface 1003 is mainly used to connect to the user device; the animal three-dimensional posture estimation device calls the animal three-dimensional posture estimation program stored in the memory 1005 through the processor 1001, and executes the animal three-dimensional posture estimation method provided by the embodiment of the present invention.
[0061] Based on the above hardware structure, an embodiment of the animal three-dimensional posture estimation method of the present invention is proposed.
[0062] Reference Figure 2 , Figure 2 The flowchart of the first embodiment of the method for estimating the three-dimensional posture of an animal of the present invention is provided.
[0063] In a first embodiment, the animal three-dimensional posture estimation method comprises the following steps:
[0064] Step S10: Collecting experimental animal video data based on the constructed multi-view shooting platform.
[0065] In a specific implementation, the executor of this embodiment is the animal three-dimensional posture estimation device, wherein the animal three-dimensional posture estimation device can be an electronic device such as a personal computer or a server, and this embodiment does not impose any restrictions on this.
[0066] Furthermore, in this embodiment, the multi-view shooting platform includes a plurality of cameras, and the plurality of cameras are calibrated using a checkerboard calibration method to obtain their intrinsic and extrinsic parameters;
[0067] The step S10 comprises:
[0068] The experimental animal video data is collected by setting the cameras of the multi-view shooting platform to synchronously collect the video data.
[0069] It should be noted that the number of the multiple cameras may be 6 or other numbers. In this embodiment, 6 cameras are taken as an example for explanation. The multi-view shooting platform includes 6 cameras. Figure 3 As shown in (a) in the figure, the shooting angle is Figure 3 As shown in (b), the six cameras are calibrated using the checkerboard calibration method to obtain their intrinsic and extrinsic parameters. The camera extrinsic parameter matrix is expressed as:
[0070]
[0071] in, and are the three-dimensional rotation matrix and translation matrix relative to the i-th camera in the standard coordinate system, and the internal geometric expression of each camera is:
[0072]
[0073] in and are the focal length of the i-th camera, and is the coordinate of the camera principal point, s iis the skew parameter. Then, the animal video data is collected by setting the camera synchronous acquisition.
[0074] Step S20: constructing a training data set and a prediction data set based on the experimental animal video data.
[0075] It is understandable that the training data set is labeled data, and the prediction data set is an unlabeled data set, wherein the amount of data in the prediction data set is much larger than that in the training data. The labeled key points include 22 key points, with the structure as follows Figure 4 As shown, they are left ear, right ear, nose, top of spine, middle of spine, top of tail, middle of tail, end of tail, left front paw, left front wrist, left front elbow, left front shoulder, right front paw, right front wrist, right front elbow, right front shoulder, left hind paw, left hind ankle, left hind knee, right hind paw, right hind ankle and right hind knee. The labeling method uses the existing 3D labeling software package Label3D, which can display images from all camera views simultaneously, such as Figure 5 As shown, six views can be marked at the same time, and during the marking process, the calibration parameters of the camera can be used to triangulate the labels to obtain three-dimensional marked data. In this embodiment, step S20 includes: using a 3D marking software package to mark key points of the experimental animal video data to construct a training data set; and constructing the unlabeled data set as a prediction data set.
[0076] Step S30: performing frame screening processing on the training data set.
[0077] It should be understood that the method used in the preprocessing of the training data set for the key point annotation occlusion problem is to use frame screening based on deep learning feature embedding, and select key frames as training sets through deep learning methods.
[0078] Furthermore, in this embodiment, step S30 includes:
[0079] Extracting high-dimensional features from each frame of the training data set using a deep learning pre-training model, and representing the features of all frames as a feature matrix;
[0080] Clustering the feature matrix using a k-means clustering algorithm to obtain K clusters;
[0081] Calculate the distance between each frame of the K clusters and the center of the corresponding cluster, and select the frame with the smallest distance as the key frame;
[0082] According to the selected key frame index, the corresponding frame is extracted from the original video and saved as a representative frame of the video.
[0083] In the specific implementation, frame screening based on deep learning feature embedding is a process centered on extracting deep features of video frames and performing cluster analysis. It is divided into the following steps: First, use the deep learning pre-trained model to extract high-dimensional features from each frame. Assuming that the video contains N frames, each frame can obtain a d-dimensional feature vector after being extracted by the model. The features of all frames are represented as a matrix:
[0084]
[0085] in, is the feature vector of the i-th frame. In order to screen out representative frames, the k-means clustering algorithm is used to cluster the feature matrix X. The goal is to divide N frames into K clusters and minimize the sum of the squares of the distances from each frame to its cluster center:
[0086]
[0087] Among them, c j is the center of the jth cluster. After clustering is completed, each cluster center represents the feature mean of the cluster. In order to select the frame that best represents the cluster center from each cluster, the distance from each frame to its cluster center can be calculated, and the frame with the smallest distance can be selected as the key frame. Specifically, for each cluster j, the key frame index k j for:
[0088]
[0089] Among them, S j is the index set of all frames in cluster j. Finally, according to the selected key frame index {k 1 ,k 2 ,…,k K}, extract the corresponding frames from the original video and save them as representative frames of the video. The high-level semantic features of the frames are extracted through the deep learning model, and the frames are grouped and filtered in combination with k-means clustering. It can effectively remove redundant frames and retain important and representative key frames.
[0090] Step S40: constructing an animal three-dimensional posture estimation model based on key point errors, wherein the animal three-dimensional posture estimation model includes a multi-view volumetric three-dimensional posture estimation network and a framework designed for different key point errors.
[0091] It should be noted that the multi-view volumetric 3D pose estimation network is currently a relatively advanced 3D convolutional neural network, which uses projective geometry to construct a metric 3D feature space that is robust to perspective changes, and then uses the 3D convolutional neural network to use shared features across cameras and learned spatial statistics of animal postures to infer landmark positions; further, in order to control the size of the 3D feature space, the volume is concentrated on the 3D center of mass of the inferred animal. The center of mass of the animal in each view is detected by using a standard 2D U-Net, and the 3D center of mass is inferred by triangulation. Furthermore, the mathematical relationship between the camera positions in S1 and the information of the 3D center of mass are used to construct a 3D volume framework that can accommodate the entire animal, and then these volumes are processed by a 3D convolutional neural network to directly predict the 3D landmark positions;
[0092] The training framework designed for different key point errors is constructed based on L1 loss and temporal smoothness constraints;
[0093] The L1 loss is a standard supervised pose regression loss that is applied only to labeled frames. Given the ground truth and predicted 3D keypoint coordinates, the supervised regression loss is defined as follows:
[0094]
[0095] Where L sup represents supervised regression loss, which represents the loss function of supervised learning and is used to measure the difference between the predicted value and the actual value; N represents the number of samples, which represents the number of samples in the data set; Y i : The actual 3D keypoint coordinates of the ith sample. The predicted 3D keypoint coordinates of the i-th sample. In this formula, the supervised regression loss L sup It is calculated by calculating the L1 distance between the actual keypoint coordinates and the predicted keypoint coordinates for each sample. The average loss is obtained by summing the losses of all samples and dividing by the number of samples N, which is used to measure the accuracy of the model in predicting 3D keypoint coordinates.
[0096] The temporal smoothness constraint means that at a high frame rate, the animal's motion speed per frame is very low, and their overall motion trajectory should generally be smooth, rather than abrupt or discontinuous. Here, the dataset is set to contain T frames of data, with a small amount of labeled data represented as y t ,i, t belongs to a frame between 0 and T, where y t represents a marked frame, i represents a key point in the marked frame, Represents the three-dimensional coordinates of key point i, for The temporal smoothness constraint between the key point coordinates establishes the temporal constraint expression:
[0097]
[0098] Where, L T Expressed as a time constraint, Respectively represent the 3D coordinates of key point i at time t-1, t, and t+1, and N is the number of 3D key points. This expression not only takes into account the first-order change of posture (i.e., velocity), but also the second-order change (i.e., acceleration). In this way, the predicted posture sequence is not only smooth in position, but also satisfies the smoothness constraints in velocity and acceleration. The goal is to minimize the second-order change of posture, so that the predicted posture sequence is as smooth as possible in velocity and acceleration.
[0099] The training framework designed for different key point errors is to adjust L sup and L T A framework for training with two loss weights. During the training process, the two loss functions form a total loss function: L = w L1 ·L sup +w smooth ·L T ;
[0100] Where L represents the total loss; w L1 Indicates the loss L sup The weight of smooth Indicates the loss L T Here, w is adjusted dynamically according to the error of each key point during the training process. L1 and w smooth , so as to achieve the following goals: key points with large errors pay more attention to L1 loss to reduce position error; key points with small errors pay more attention to temporal smoothness loss to improve temporal consistency. The weights are dynamically allocated according to the L1 error of each key point. For the position error of key point i: Dynamically adjust weights:
[0101]
[0102] Among them, w L1,i represents the total loss of key point i. β>0 is a balancing factor used to control the minimum weight of the smooth loss. i When w is larger, L1,i approaches 1, and w smooth,i Approaching 0, the model focuses more on L1 loss. i When w is smaller, smooth,iIncrease, the model pays more attention to temporal smoothness. In this embodiment, the multi-view volume 3D pose estimation network is a 3D convolutional neural network, and the framework designed for different key point errors is constructed based on L1 loss and temporal smoothness constraints;
[0103] The method of constructing an animal three-dimensional posture estimation model based on key point errors includes:
[0104] Use projective geometry to construct a metric 3D feature space;
[0105] Inferring landmark locations using shared features across cameras and learned spatial statistics of animal poses via the 3D convolutional neural network;
[0106] The L1 loss is used to perform a standard supervised pose regression loss on labeled frames and establish a temporal constraint expression for the temporal smoothness constraint between keypoint coordinates.
[0107] Further, in this embodiment, the inferring the landmark position by using the shared features across cameras and the learned spatial statistics of the animal posture through the 3D convolutional neural network includes:
[0108] Use a standard 2D U-Net to detect the center of mass of the animal in each view and infer the 3D center of mass of the animal through triangulation;
[0109] A 3D volume frame that can accommodate the entire animal is constructed according to the position relationship of multiple cameras and the 3D center of mass of the animal, and the 3D volume frame is processed by the 3D convolutional neural network to predict the position of 3D landmarks.
[0110] Step S50: The training data set after frame screening and the prediction data set are combined into a time block and input into the animal 3D posture estimation model for training, and the loss during training is weighted and adjusted using a framework designed for different key point errors.
[0111] It should be understood that the time chunks composed of labeled data and unlabeled data mean that each training batch contains one labeled sample and three unlabeled samples extracted from its local neighborhood, while introducing additional temporally continuous unlabeled segments. These segments are only used to calculate the unsupervised temporal loss, helping the model achieve better smoothness and consistency in the temporal dimension. Figure 4 The L sup and L T Two loss and time block inputs are used.
[0112] Step S60: Use the trained model to perform key point recognition and posture estimation on the experimental animal.
[0113] It is understandable that the optimized model is used to predict the key points of each frame, and the relative position change relationship of the key points is used to analyze the behavior of the animal. The three-dimensional key point effect is as follows: Figure 5 As shown, the top three in the figure are pictures taken from three perspectives, and the bottom one shows the predicted three-dimensional key point effect. Figure 5 (1) to (6) show different postures of animals.
[0114] In this embodiment, a new training strategy is provided, which sets different key point training weights according to the annotation errors of key points calculated in advance during the training phase to avoid result errors caused by key point errors. In experiments with freely moving animals, this innovative method optimizes the existing state-of-the-art multi-view volumetric 3D pose estimation performance and further enhances the stability of 3D key point tracking. The beneficial effects that can be achieved are at least as follows:
[0115] (1) Targeted error correction: The method of this embodiment can focus on key point errors. In the three-dimensional posture estimation of animals, different key points have different importance for the accurate description of the posture. By analyzing the key point errors, those key points with inaccurate estimates can be accurately found. For example, for the estimation of the walking posture of quadrupeds, the errors of the key points of the limb joints and the spine have a greater impact on the overall posture. Through error analysis, the weights of the estimated errors of these key parts are adjusted, which can more accurately restore the true three-dimensional posture of the animal and reduce the deviation of the overall posture estimation.
[0116] (2) Adapting to individual differences: Different species of animals, and even different individuals of the same species, have different body structures and movement patterns. This method can adjust the weights based on the errors in key point estimation for each individual animal. For example, dogs and mice have different limb proportions. When errors occur in key point estimation of the legs when estimating the posture of a dog, the weight adjustment strategy required may be different from that required when estimating the same part of the legs for a mouse. Through error analysis, more appropriate posture estimation models can be customized for different individuals, thereby improving the accuracy of posture estimation for various animals.
[0117] (3) Improved anti-interference ability: In actual scenes, animal posture estimation may be interfered by many factors, such as changes in lighting, changes in animal surface features (such as hair length, color changes, etc.), and occlusion of some body parts. By adjusting the weights through key point error analysis, the model can pay more attention to key points that are less interfered with and more accurately estimated, reducing the impact of interference factors on the overall posture estimation. For example, when one side of the animal's body is obscured by a shadow, by giving higher weights to the key points on the other side that are not obscured and have more accurate estimates, the model can overcome the shadow interference to a certain extent and still estimate the animal's three-dimensional posture more accurately.
[0118] (4) Enhanced generalization ability: This method helps the model work better in different data sets and scenarios. If the model has been trained with weight adjustment based on key point error analysis, it can adapt and adjust the posture estimation strategy more quickly when faced with new animal species, new movement patterns, or new environmental conditions. For example, when converting from an animal posture estimation dataset in a laboratory environment to a dataset in a wild environment, the model can readjust the weights based on the key point errors in the new dataset, thereby improving the accuracy of animal 3D posture estimation in complex wild environments and enhancing the generalization ability of the model.
[0119] In addition, an embodiment of the present invention further proposes a storage medium, on which an animal three-dimensional posture estimation program is stored. When the animal three-dimensional posture estimation program is executed by a processor, the steps of the animal three-dimensional posture estimation method described above are implemented.
[0120] In addition, refer to Figure 6 The embodiment of the present invention further provides a device for estimating a three-dimensional posture of an animal, the device comprising:
[0121] A collection module 10, used to collect experimental animal video data based on the constructed multi-view shooting platform;
[0122] A construction module 20, used to construct a training data set and a prediction data set according to the experimental animal video data;
[0123] A screening module 30, used for performing frame screening processing on the training data set;
[0124] The construction module 20 is further used to construct an animal 3D posture estimation model based on key point errors, wherein the animal 3D posture estimation model includes a multi-view volume 3D posture estimation network and a framework designed for different key point errors;
[0125] A training module 40 is used to input the time block composed of the training data set and the prediction data set after frame screening into the animal 3D posture estimation model for training, and to adjust the weight of the loss during training by using a framework designed for different key point errors;
[0126] The recognition module 50 is used to perform key point recognition and posture estimation on the experimental animal using the trained model.
[0127] Other embodiments or specific implementations of the animal three-dimensional posture estimation device of the present invention can refer to the above-mentioned method embodiments, which will not be repeated here.
[0128] It should be noted that, in this article, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or system including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or system. In the absence of further restrictions, an element defined by the sentence "comprises a ..." does not exclude the existence of other identical elements in the process, method, article or system including the element.
[0129] The serial numbers of the embodiments of the present invention are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments. In a unit claim that lists several means, several of these means may be embodied by the same hardware item. The use of the words first, second, and third, etc. does not indicate any order and these words may be interpreted as identifiers.
[0130] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as a disk, an optical disk), and includes a number of instructions for enabling a terminal device (which can be a mobile phone, a computer, or a network device, etc.) to execute the methods described in each embodiment of the present invention.
[0131] The above are only preferred embodiments of the present invention, and are not intended to limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made using the contents of the present invention specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present invention.
Claims
1. A method for estimating a three-dimensional posture of an animal, characterized in that: The animal three-dimensional posture estimation method comprises: Collect experimental animal video data based on the built multi-view shooting platform; Constructing a training data set and a prediction data set based on the experimental animal video data; Performing frame screening processing on the training data set; Constructing an animal 3D posture estimation model based on key point errors, wherein the animal 3D posture estimation model includes a multi-view volumetric 3D posture estimation network and a framework designed for different key point errors; The training data set after frame screening and the prediction data set are combined into a time block and input into the animal 3D posture estimation model for training, and the loss during training is weighted and adjusted using a framework designed for different key point errors; The trained model is used to perform key point recognition and posture estimation on experimental animals.
2. The animal three-dimensional posture estimation method according to claim 1, characterized in that: The multi-view shooting platform includes multiple cameras, and the multiple cameras are calibrated using a chessboard calibration method to obtain their internal and external parameters; The multi-view shooting platform constructed to collect experimental animal video data includes: The experimental animal video data is collected by setting the cameras of the multi-view shooting platform to synchronously collect the video data.
3. The animal three-dimensional posture estimation method according to claim 1, characterized in that: The step of constructing a training data set and a prediction data set based on the experimental animal video data includes: Using a 3D marking software package to mark key points of the experimental animal video data to construct a training data set; Construct an unlabeled dataset as a prediction dataset.
4. The animal three-dimensional posture estimation method according to claim 1, characterized in that: The frame screening process of the training data set includes: Extracting high-dimensional features from each frame of the training data set using a deep learning pre-training model, and representing the features of all frames as a feature matrix; Clustering the feature matrix using a k-means clustering algorithm to obtain K clusters; Calculate the distance between each frame of the K clusters and the center of the corresponding cluster, and select the frame with the smallest distance as the key frame; According to the selected key frame index, the corresponding frame is extracted from the original video and saved as a representative frame of the video.
5. The animal three-dimensional posture estimation method according to claim 2, characterized in that: The multi-view volumetric 3D pose estimation network is a 3D convolutional neural network, and the framework designed for different key point errors is constructed based on L1 loss and temporal smoothness constraints; The method of constructing an animal three-dimensional posture estimation model based on key point errors includes: Use projective geometry to construct a metric 3D feature space; Inferring landmark locations using shared features across cameras and learned spatial statistics of animal poses via the 3D convolutional neural network; The L1 loss is used to perform a standard supervised pose regression loss on labeled frames and establish a temporal constraint expression for the temporal smoothness constraint between keypoint coordinates.
6. The animal three-dimensional posture estimation method according to claim 5, characterized in that: The 3D convolutional neural network uses shared features across cameras and learned spatial statistics of animal postures to infer landmark locations, including: Use a standard 2D U-Net to detect the center of mass of the animal in each view and infer the 3D center of mass of the animal through triangulation; A 3D volume frame that can accommodate the entire animal is constructed according to the position relationship of multiple cameras and the 3D center of mass of the animal, and the 3D volume frame is processed by the 3D convolutional neural network to predict the position of 3D landmarks.
7. The method for estimating the three-dimensional posture of an animal according to any one of claims 1 to 6, characterized in that: The method of using the trained model to perform key point recognition and posture estimation on the experimental animal includes: The trained model is used to predict the key points of each frame of the experimental animal, and the relative position change relationship of the key points is used to estimate the animal's posture.
8. An animal three-dimensional posture estimation device, characterized in that: The animal three-dimensional posture estimation device comprises: A collection module is used to collect experimental animal video data based on the built multi-view shooting platform; A construction module, used to construct a training data set and a prediction data set based on the experimental animal video data; A screening module, used for performing frame screening processing on the training data set; The construction module is also used to construct an animal 3D posture estimation model based on key point errors, wherein the animal 3D posture estimation model includes a multi-view volume 3D posture estimation network and a framework designed for different key point errors; A training module, used for inputting a time block of a training data set and a prediction data set after frame screening into the animal 3D posture estimation model for training, and adjusting the weight of the loss during training by using a framework designed for different key point errors; The recognition module is used to use the trained model to perform key point recognition and posture estimation on experimental animals.
9. An animal three-dimensional posture estimation device, characterized in that: An animal three-dimensional posture estimation program is stored on the animal three-dimensional posture estimation device, and when the animal three-dimensional posture estimation program is executed by the processor, the steps of the animal three-dimensional posture estimation method according to any one of claims 1 to 7 are implemented.
10. A storage medium, characterized in that: The storage medium stores an animal three-dimensional posture estimation program, and when the animal three-dimensional posture estimation program is executed by the processor, the steps of the animal three-dimensional posture estimation method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Monocular camera object pose estimation method and system based on template matching
CN111768447A
Semi-supervised animal three-dimensional attitude estimation method and device and storage medium
CN118334755A
Lightweight human body posture estimation method based on key frame selection
CN118351565A
Intelligent image recognition method based on computer vision and machine learning
CN118537816A
Three-dimensional human body posture estimation method and system based on video sequence spatio-temporal context
CN118823833A
Cited By
Limb three-dimensional key point recognition model training method based on biological characteristics
CN121354170A