A method for animal surface motion capture based on differentiable rendering
By constructing high-precision grid templates through CT scanning, multi-view image acquisition and deep learning supervised information extraction, combined with the grid registration method of differentiable rendering, the problem of capturing the surface shape and color texture of animals in motion is solved, and accurate animal behavior analysis is achieved.
Patent Information
- Application Number
- CN202411605056.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-12
- Publication Date
- 2025-10-17
- Estimated Expiration
- 2044-11-12
AI Technical Summary
Existing technologies make it difficult to accurately capture the surface shape, color texture, and motion state of animals while they are in motion. Especially in the case of occlusion and rapid deformation, traditional key point estimation methods lack detailed information.
A mesh registration method based on high-precision animal mesh template construction based on CT scanning, multi-view synchronous image acquisition, deep learning supervised information extraction and differentiable rendering is used to reconstruct the three-dimensional mesh of the animal in motion state, and mesh registration is performed by combining data loss, regularization loss and timing loss.
It achieves accurate surface capture of animals in motion, reflecting motion status, surface shape, and color texture. It is robust and versatile, suitable for a variety of animal experiments, and provides detailed behavioral analysis support.
Smart Images

Figure CN119477987B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application belongs to the field of computational biology, and particularly relates to an animal surface motion capture method based on differentiable rendering. BACKGROUND
[0002] Accurate perception of animal behavior is of great significance in basic science and engineering applications. In basic science, accurate, intelligent, and efficient animal motion perception methods can help researchers easily obtain individual and social behaviors of animals, provide new research ideas for solving important scientific problems in disciplines such as neuroscience, animal behavior, and bionics, and help reveal the neural mechanisms of animal communication and mate selection behaviors. In engineering applications, in-depth research on animal motion behavior can help achieve complex animal motion behavior regulation, promote the construction of a new generation of animal robots with simple system structure, superior motion capability, optimized system energy consumption, and excellent concealment performance, and have broad application prospects in scientific research, dangerous environment search and rescue, weather observation, industrial production, and many other fields.
[0003] With the development of computer vision and machine learning technologies, animal behavior perception technology is showing a development trend from artificial to intelligent, from single modality to multi-modality, from planar to stereoscopic, from low precision to high precision, from sparse representation to dense representation, and from single individual to multi-individual interaction. The current mainstream method focuses on two-dimensional and three-dimensional key point estimation in animal motion state. DeepLabCut is a representative method for two-dimensional key point estimation. This method is based on deep neural network and realizes accurate prediction of the position of animal body key points in images. Anipose is a representative method for three-dimensional key point estimation. This method uses multi-camera triangulation and space-time filtering based on DeepLabCut to effectively obtain the position of animal body key points in three-dimensional space.
[0004] However, key points are a sparse representation of the animal body, and the method of simplifying the animal body into several key points loses important information such as the shape, color texture, and motion state of the animal body surface. Three-dimensional mesh as a dense representation of the animal body effectively makes up for the defects of key point representation and can reflect animal information more finely. Since three-dimensional mesh has a more complex data form than key points, and animals have rapid surface deformation in motion state, and there are occlusion, motion blur, and other complex situations, existing methods are difficult to realize surface capture in animal motion state. SUMMARY
[0005] To solve the above problems, the application aims to provide an animal surface motion capture method based on differentiable rendering, which comprises: (1) high-precision animal grid template construction based on CT scanning, (2) multi-view synchronous animal motion image acquisition, (3) subsequent supervised information extraction based on deep learning, and (4) grid registration based on differentiable rendering. The method can reconstruct the surface grid of the animal in the motion state, capture the surface shape, color texture and motion state of the animal, and more accurately and in detail describe the behavior information of the animal compared with the traditional animal behavior capture method, thereby providing strong support for animal behavior analysis.
[0006] The high-precision animal grid template construction method based on CT scanning of step (1) comprises the following steps: scanning an animal specimen by using a high-precision CT scanner, and constructing a three-dimensional grid template based on the scanning result, wherein the grid template finely depicts the shape of the animal body, thereby providing strong prior for the surface capture of the animal in the motion state and improving the authenticity of the surface capture effect. Figure 2 As shown in the grid template; the built-in skeleton in the grid template is consistent with the anatomical structure of the animal, a vertex-skeleton binding matrix is obtained, and three-dimensional key points are implanted in the template to establish a vertex-key point mapping matrix for driving the grid deformation.
[0007] The multi-view synchronous animal motion image acquisition method of step (2) is based on an animal image acquisition system composed of multiple high-speed industrial cameras, a synchronous signal generator and an animal experiment platform. In the experiment, the animal is placed on the experiment platform, and multiple high-speed cameras are placed around the experiment platform to capture images of the animal motion process from different angles and provide animal motion information from different angles. The synchronous signal generator controls the synchronous shooting of the multiple cameras to ensure the inter-frame synchronization of each view. The acquired multi-view images record the behavior information of the animal from multiple different angles, and the effective information complementation between the different views can provide a full and comprehensive description of the animal motion state. An example of the system is shown in Figure 3 .
[0008] The supervised information extraction method based on deep learning of step (3) comprises the following steps: two-dimensional key point extraction based on DeepLabCut, contour mask extraction based on SimpleClick and optical flow extraction based on Raft. The two-dimensional key point detection is performed on the acquired images by using the two-dimensional posture tool DeepLabCut based on deep learning to obtain the two-dimensional pixel coordinates of the key points in each view, thereby providing supervision for the positions of the key points of the animal body, Figure 4Examples of the key point supervision obtained by the method are given; the SimpleClick segmentation tool is used to segment the collected images, and the key points obtained by DeepLabCut are used as prompt input into the SAM model, and the SAM model realizes image segmentation according to the prompt to obtain the pixel area of the animal in each image, that is, the contour mask of the animal, Figure 5 Examples of the contour mask map obtained by the above method are given; the Raft optical flow estimation model is used to extract the optical flow map between adjacent two frames under each view, which indicates the dense pixel correspondence relationship of the two frames of images, and provides supervision for the direction of animal movement, Figure 6 Examples of the optical flow map obtained by the above method are given. The extracted key points, contour mask maps and optical flow will be used as supervision information for the mesh registration algorithm based on differentiable rendering.
[0009] The mesh registration algorithm based on differentiable rendering of the step (4) comprises the following steps: first, a series of learnable mesh deformation parameters are introduced to describe the difference between the real animal and the template animal; then a non-rigid mesh deformation method is used to transform the template mesh to obtain a predicted mesh through the action of the parameters; finally, a loss function combining data loss, regularization loss and time loss is used to calculate the difference between the predicted mesh and the real motion state, and the gradient descent method is used to update the parameters, so as to realize the surface capture of the animal in the motion state. The flow chart of the algorithm is shown in Figure 7 .
[0010] The learnable mesh deformation parameters include {β, s, γ, φ, f, θ}. Wherein β represents a local scaling coefficient, which allows each bone to control the surface mesh around it through a vertex-skeleton binding matrix to perform three degrees of freedom scaling, so as to describe the scale difference between different parts of the real animal and the template animal. s represents a global scaling coefficient, which is used to control the overall size scaling of the template. γ and φ represent displacement parameters, which are used to represent the translation and rotation of the model in the world coordinate system respectively, so as to capture the overall motion of the animal. f represents a color texture parameter, which is represented by the RGB color value corresponding to each mesh vertex, and is used to reflect the color texture of the animal's body surface. θ represents the pose parameter, wherein each element θ b represents the rotation angle of bone b relative to its father bone in the kinematic tree, and the rotation angle of the root bone represents the global rotation.
[0011] The non-rigid mesh deformation method introduces a local scaling factor β into the traditional linear blend skinning algorithm. This allows the built-in skeleton to not only control the mesh for rigid pose transformations but also for non-rigid local scaling, thereby capturing the shape differences between the real animal and the template mesh. This mesh deformation function is denoted as D. Similar to the linear blend skinning algorithm, the calculation of this function involves two transformation matrices: the world-local transformation matrix A corresponding to the initial state of each bone b, and b and the world-local transformation matrix T that controls the mesh deformation b To calculate A b , starting from bone b, traverse its ancestors in the kinematic tree from bottom to top until the root bone. During the traversal process, merge the transformation matrix of each bone relative to its father bone to obtain A b , the process is formulated as:
[0012]
[0013] in and Indicates the rotation angle and starting position of bone b in its father's coordinate system in the initial state, R(.) represents the function that converts a rotation angle into a rotation matrix, and Ψ(b) represents all ancestors of bone b. b In the same traversal process, we need to consider the effect of local scaling, a non-rigid transformation. When calculating, we first consider the translation effect caused by the local scaling of the father of bone b, ψ(b), and get T′ b Then consider the scaling effect of the local scaling of bone b itself (given by the scaling matrix (represented), we get T b :
[0014]
[0015] The coordinates of vertex i on the mesh in the world coordinate system in the initial state are V i * , first use (A b ) -1 Transform it to the local coordinate system of bone b, and then use T b Get its coordinates in the world coordinate system after being transformed by bone b. If the total number of bones is B, then the above operation can get B different coordinates for each vertex. Using the vertex-bone binding weight to perform weighted summation of these coordinates, the coordinate V' of the vertex in the world coordinate system can be obtained. i Then, the global scaling, translation, and rotation represented by s,γ,φ are applied to the network V' i , we get the final coordinate V of the vertex in the world coordinate system i, the mathematical form of the process is:
[0016]
[0017] V i =sR φ V' i +γ
[0018] The loss function contains three different losses, namely: data loss E data , regularization loss E reg and timing loss E temp These three different losses use actual observation data, prior regularization constraints, and temporal supervision as learning targets, respectively, to achieve alignment between the model and the animal's real behavior and complete surface capture of the animal's motion state. The data items use key points, contour masks, and original color images obtained by the deep learning-based supervised information extraction method as supervision, and their expression is (in the following formula, T represents the total number of recorded video frames, and t represents each frame in the video):
[0019] It contains three items: E kp , E sil , E rgb and their corresponding weights λ kp ,λ sil ,λ rgb . These three items represent the key point loss item, the contour mask loss item, and the image color loss item respectively. The key point loss item encourages the matching between the model key points and the actual key points, thereby promoting the model to capture the true movement state of the animal. The specific calculation method of the key point loss item is to map the model vertices after the mesh deformation function D into three-dimensional key point coordinates through the vertex-key point mapping matrix, and project the three-dimensional key points to each perspective through the camera projection function to obtain the two-dimensional key point coordinates generated by the model. And calculate its two-dimensional key point supervision y under each perspective (t) The contour mask loss term encourages the model contour to coincide with the actual observed animal contour, thereby promoting the model to capture the surface shape of the animal. The specific calculation method of the contour mask loss term is to use a differentiable renderer to render the model after the mesh deformation function D into a binary contour mask under each view angle. And calculate its correlation with the supervised contour mask S (t) The image color loss term encourages the model’s color to match the color in the actual observed image, thereby promoting the model to capture the surface color texture of the animal. The specific calculation method of the image color loss term is to use a differentiable renderer to render the model and its surface texture f after the mesh deformation function D into color images under different viewing angles. and calculate the pixel color error between them and the real color image I (t) .
[0020] The regularization loss introduces prior constraints to limit the pose and shape parameters, encouraging the model to learn more realistic and reasonable parameters. Its expression is:
[0021] E reg = λ β1 E β1 + λ β2 E β2 + λ θ E θ
[0022] where there are three terms, shape symmetry constraint E β1 , shape specification constraint E β2 and pose specification constraint E θ and their corresponding weights λ β1 , λ β2 , λ θ . The shape symmetry constraint encourages anatomically symmetric body parts to have the same local scaling factor. The specific calculation method is to penalize the difference of local scaling factors of symmetric parts. The shape specification constraint limits the shape parameters within the allowed range. The specific calculation method is to penalize the shape parameters that exceed the allowed range. Similar to the shape specification constraint, the pose specification constraint encourages the pose to change within the allowed range and penalizes the pose parameters that exceed the allowed range.
[0023] The optical flow, an important temporal information, is introduced in the temporal loss to enhance the robustness of the method in the case of key point supervision failure. In actual experiments, due to factors such as fast movement of animals and environmental occlusion, there are often a large number of errors or missing in key point supervision, which seriously affects the surface motion capture effect. Optical flow supervision can provide temporal information, indicating the dense pixel correspondence between two frames of images, effectively reflecting the motion direction of the animal between two frames, and making up for the influence of key point missing. The expression of the temporal loss is:
[0024] where E mor is the multi-view optical flow loss, and λ mor is its weight. The specific calculation method is as follows: first, calculate the displacement vector between the vertex projection of the t+1 frame and the vertex projection of the t frame under each view, to form the optical flow feature of each vertex, then input the mesh and optical flow feature into the differentiable renderer to render the optical flow map under each view, and calculate the pixel-wise loss with the supervised optical flow map O (t) , so as to encourage the model to conform the surface motion between two frames to the optical flow supervision information.
[0025] By utilizing the gradient generated in the loss function calculation process, the numerical value of the learnable parameters can be updated in a gradient descent manner, and iteration is continuously performed, so as to realize the alignment from the template mesh to the real animal, and thus realize the surface capture of the animal in the motion state.
[0026] The present application has the beneficial effects of:
[0027] (1) The present application can reconstruct a three-dimensional mesh of an animal in a motion state, reflect the motion state, surface shape and color texture of the animal, and accurately and in detail describe the behavior information of the animal, thereby providing strong support for fine animal behavior analysis.
[0028] (2) The present application has robustness and can effectively deal with environmental occlusion, key point detection errors and the like in animal experiments, thereby helping to realize animal behavior perception in complex experimental scenarios.
[0029] (3) The present application has universality and can be applied to various animal experiments including bumblebees, fruit fly larvae and mice, without the need to design specific algorithms for different animals, and can adapt to various experimental requirements in the field of biomedical research. BRIEF DESCRIPTION OF DRAWINGS
[0030] Figure 1 It is a schematic diagram of the implementation steps of the present method.
[0031] Figure 2 It is a high-precision animal mesh template constructed by CT scanning.
[0032] Figure 3 It is a multi-view synchronous animal image acquisition system for bumblebees.
[0033] Figure 4 It is the key point supervision obtained by the deep learning-based supervision information extraction method.
[0034] Figure 5 It is the contour mask supervision obtained by the deep learning-based supervision information extraction method.
[0035] Figure 6 It is the optical flow supervision obtained by the deep learning-based supervision information extraction method.
[0036] Figure 7 It is a flowchart of the mesh registration method based on differentiable rendering.
[0037] Figure 8 It is the bumblebee key point definition used in the implementation case.
[0038] Figure 9 It is the surface motion capture effect of bumblebees in the implementation case. DETAILED DESCRIPTION
[0039] The application will be further described in conjunction with the accompanying drawings and examples.
[0040] Example 1 uses the method of the application to capture the surface motion of bumblebees in motion.
[0041] As Figure 1 , step (1), first based on CT scan to build a high-precision bumblebee grid template, which finely depicts the three-dimensional grid model of the animal body shape, the template as Figure 2 shown, get the final grid template, according to the bumblebee body anatomy structure, in the grid template, the built-in bone is planted, according to the connection relationship between the bones, the bone kinematics tree is constructed, and the vertex-skeleton binding matrix is obtained, which is used to realize the grid deformation driven by the skeleton; In the template, three-dimensional key points are implanted, and according to the relative position relationship between each key point and the grid surface vertex, a vertex-key point mapping matrix is established, which is used to map the grid vertex to the three-dimensional coordinates of the key point.
[0042] Step (2), a multi-view bumblebee motion image acquisition system as shown in Figure 3 is constructed, which is equipped with 6 high-speed industrial cameras (Hikrobot MV-CS016-10UC area array camera, Hikrobot MVL-HF2524M-10MP lens). In the experiment, a synchronous signal generator is used to control all cameras to shoot synchronously in an external trigger mode, the shooting frame rate of this experiment is set to 200 frames, and the image resolution is 800*800. Before placing the experimental bumblebee, first calibrate the multiple cameras, and use the multi-view calibration method based on Anipose to get the parameters of each camera. After calibration, the head of the experimental bumblebee is fixed and placed on the air pump suspension device, and an LED display screen is placed in front of it, and through the pattern change on the LED screen, it stimulates it to make forward, backward, turning and other actions, while recording its motion process from 6 different angles simultaneously, and the recorded video is saved in the host computer.
[0043] Step (3), using a deep learning-based supervised information extraction method to extract information from multi-view images. 46 key points on the bumblebee body are defined in advance, and the key point definition is as Figure 8The 6 key points on each leg are shown: coxa-trochanter, trochanter-femur, femur-tibia, tibia-basitarsus, basitarsus-telotarsus, telotarsus-claw; the 3 key points on each antenna: head-antenna, antenna node, antenna tip; and the 4 key points on the trunk: mouth, head-thorax, thorax-abdomen, tail. A total of 18000 frames of experimental data were collected, from which a continuous 2000-frame segment was selected as the data set for mesh reconstruction. The supervision information extraction method based on deep learning described above was used to extract supervision information from the video. To obtain the 2D key point supervision corresponding to each image, the pre-trained DeepLabCut model was used to estimate the pose of the collected images, and a total of 46 key points were detected on each image, as shown in Figure 4 To obtain the contour mask data corresponding to each image, the obtained key points were input into SimpleClick, and the image was automatically segmented using the key points as a prompt to obtain the mask of the corresponding part of the bumblebee in the image, as shown in Figure 5 To obtain the optical flow supervision, the Raft algorithm was used to calculate the optical flow map between adjacent two frames, as shown in Figure 6
[0044] Step (4), surface motion capture of the bumblebee was performed using a mesh registration algorithm based on differentiable rendering, and the algorithm flow is as shown in Figure 7 The algorithm implementation is based on PyTorch and PyTorch3D framework. The objective function introduced above is adopted to learn the parameters, combined with data loss, regularization loss and temporal loss. The data loss uses the extracted key points, contour mask and color image in the bumblebee motion image to reconstruct the motion state, surface shape and color texture of the bumblebee. The regularization loss constrains the shape parameters and pose parameters of the bumblebee to improve the rationality of parameter learning. The temporal loss introduces multi-view optical flow supervision to enhance the stability of the algorithm when the key point detection is wrong. To balance the influence of each term in the objective function on the final learning result, the corresponding weight coefficients of each term need to be adjusted. In this experiment, the weight values of each term determined through actual debugging are shown in Table 1. When performing parameter optimization, the Adam optimizer in PyTorch is used to update the parameters in the gradient descent manner. The initial learning rate of the optimizer is set to 0.001, 200 iterations are performed for each frame, and the cosine annealing learning rate decay strategy is used to gradually reduce the learning rate during the iteration process. The final learning rate is 0.00001. After the parameter learning process is completed, the 3D key point Euclidean distance index and the overlap degree IOU index of the contour mask are used to evaluate the reconstruction effect. Among the total 2000 frames of data, 50 frames are manually labeled as the test set. The average value of the 3D key point Euclidean distance error obtained by the method on the test set is less than 5% of the length of the bee, and the contour overlap degree IOU is 0.801. This result proves that the surface motion capture result of the method has very high precision. The mesh reconstruction result on one frame is shown in Figure 9
[0045] Table 1
[0046] kp ]]> sil ]]> rgb ]]> β1 ]]> β2 ]]> θ ]]> mor ]]> 10.0 5.0 1.0 0.3 0.5 0.5 1.0
Claims
1. A method for capturing animal surface motion based on differentiable rendering, characterized in that: This is achieved through the following steps: (1) high-precision animal mesh template construction based on CT scans, (2) followed by multi-view synchronous animal motion image acquisition, (3) followed by supervised information extraction based on deep learning, and (4) finally mesh registration based on differentiable rendering; The step (1) is a method for constructing a high-precision animal mesh template based on CT scanning, which uses CT scanning of animal specimens to reconstruct a mesh template as a shape prior for capturing the surface of the animal in motion, and implants bones and key points into the template to obtain a vertex-bone binding matrix and a key point-vertex mapping matrix as a driver for template deformation; The step (3) is based on a deep learning-based supervised information extraction method, which uses a deep learning method to extract two-dimensional key points, contour masks, and optical flows from the acquired multi-view images as supervised information used in the subsequent mesh registration method based on differentiable rendering; The step (4) is based on a mesh registration method for differentiable rendering, which first introduces a series of learnable mesh deformation parameters to characterize the difference between the real animal and the template animal; Then, a non-rigid mesh deformation method is used to transform the template mesh through the action of parameters to obtain a predicted mesh; finally, a loss function that combines data loss, regularization loss and timing loss is used to calculate the difference between the predicted mesh and the actual animal behavior, and the parameters are updated using the gradient descent method to achieve surface motion capture.
2. The method for capturing animal surface motion based on differentiable rendering according to claim 1, characterized in that: The step (2) uses a multi-view synchronous image acquisition system to capture images of animals in motion from multiple camera positions, providing animal motion information from different perspectives.
3. The method for capturing animal surface motion based on differentiable rendering according to claim 1, wherein: The learnable grid deformation parameters include ,in Represents the local scaling factor, which is used to depict the proportional differences between different parts of the body of the real animal and the template; Represents the global scaling factor, which is used to control the overall size scaling of the template; and Represents the displacement parameter, which is used to represent the translation and rotation of the template in the world coordinate system; Represents the color texture parameter, which is used to reflect the color texture of the animal's body surface; Represents the posture parameters, which are used to control the posture of the model.
4. The method for capturing animal surface motion based on differentiable rendering according to claim 1, wherein: The non-rigid mesh deformation method introduces a local scaling factor in the linear blend skinning algorithm. , local scaling is performed by controlling the surrounding mesh vertices through bones, thereby capturing the shape differences between the real animal and the template mesh.
5. The method for capturing animal surface motion based on differentiable rendering according to claim 1, wherein: The loss function includes data loss, regularization loss and timing loss. The data loss uses key points, contour masks and color image data as supervision to achieve alignment between the model and actual observation data; the regularization loss introduces shape symmetry constraints, shape normalization constraints and posture normalization constraints to limit the parameters within a realistic and reasonable range; the timing loss introduces multi-view optical flow information to improve the robustness of the method in the case of key point errors or missing.
Citation Information
Patent Citations
Three-dimensional face and eyeball movement modeling and capturing method and system
CN110807364A
Multi-view mouse dynamic three-dimensional reconstruction method for mouse
CN111105486A