Livestock three-dimensional reconstruction system construction method based on monocular vision

By constructing a 3D reconstruction system for livestock based on monocular vision, the problem that general models cannot accurately capture the real anatomical structure of livestock is solved. This system enables high-fidelity 3D model generation and automated body size measurement, thereby improving the precision management capabilities of smart animal husbandry.

CN121767559APending Publication Date: 2026-03-31NORTHWEST A & F UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-24
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

In existing technologies, general quadrupedal parametric models cannot accurately capture the anatomical proportions and physiological structural details of real livestock, resulting in distortion of the three-dimensional reconstruction model, affecting the reliability of body size parameters, and limiting the precision management of smart animal husbandry.

Method used

A monocular vision-based approach is used to reconstruct a 3D mesh model of livestock by fusing multi-angle RGB video and depth video. A model with parameterized skeletal structure is constructed, and personalized livestock models are generated using orthogonal shape basis vectors and morphological mixing functions. Finally, a standardized pose is unified through inverse linear hybrid skinning technology, thus constructing a complete 3D livestock reconstruction system.

Benefits of technology

It achieves end-to-end automatic generation from monocular images to 3D models with skeletons, significantly improving the model's fidelity and the representation accuracy of key physiological structures, ensuring the visual consistency, geometric accuracy, and biomechanical rationality of the reconstructed model, avoiding abnormal phenomena such as limb penetration, and providing a reliable basis for automated body size measurement.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121767559A_ABST
    Figure CN121767559A_ABST
Patent Text Reader

Abstract

The invention discloses a livestock three-dimensional reconstruction system construction method based on monocular vision, and belongs to the technical field of computer three-dimensional vision and intelligent animal husbandry. Comprising the following steps: fusing and reconstructing a plurality of livestock three-dimensional grid models according to a plurality of livestock multi-angle RGB videos and depth videos, and constructing a parameterized skeleton model in combination with preset skeleton levels and skin weights; unifying the personalized model into a standard attitude through reverse linear hybrid skin, calculating a shape deviation matrix, extracting an orthogonal shape basis vector through principal component analysis, and establishing a parameterized morphological hybrid function; and finally, constructing a monocular vision reconstruction system, and generating a three-dimensional livestock model with a skeleton structure from a single picture by adjusting the weight coefficient of the shape base vector. According to the invention, automatic generation from a two-dimensional image to a livestock three-dimensional model with a complete skeleton structure can be realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of computer 3D vision and smart animal husbandry technology, and in particular to a method for constructing a 3D reconstruction system for livestock based on monocular vision. Background Technology

[0002] In the intersection of smart animal husbandry and computer vision, digital modeling and automated body size measurement of livestock based on 3D vision technology have become key means to achieve precision breeding management. However, the widespread application of this technology is still limited by a core bottleneck: the lack of high-fidelity parametric 3D models that can accurately represent the true anatomical structure and body shape variations of specific livestock breeds.

[0003] In existing technologies, researchers often use general-purpose parametric models for tetrapods, such as the SMAL (Skinned Multi-Animal Linear) model. The shape space of such models is constructed based on a limited number of animal toy scans, and their training sets severely lack data on real, live livestock, especially specific breeds with significant economic value, such as the Saanen dairy goat. Therefore, general-purpose models cannot accurately capture the anatomical proportions, trunk morphology, and details of key physiological structures crucial for assessing milk production performance in real livestock. When these inherently biased models are used for 3D reconstruction or body shape analysis from monocular images, the inaccuracies in their shape representation are propagated and amplified, leading to severe distortions in both overall proportions and local morphology in the reconstructed model. Consequently, key body size parameters such as body length, height, and chest circumference calculated based on the model become unreliable.

[0004] This fundamental inaccuracy at the model level directly limits the accuracy ceiling of all subsequent automated processing steps. Whether it's 3D shape reconstruction from monocular RGB images or non-contact body size measurement based on 3D models, the accuracy of the results highly depends on the underlying parametric model's ability to depict real biological morphology. Therefore, learning from multi-view scanning data of real livestock and constructing a breed-specific, high-fidelity parametric model, and then establishing a complete technical system from data acquisition and model building to automated measurement, has become an essential requirement for overcoming current technological bottlenecks and promoting the development of precision animal husbandry. Summary of the Invention

[0005] The purpose of this invention is to provide a method for constructing a three-dimensional reconstruction system for livestock based on monocular vision, which can realize the automatic generation of a three-dimensional livestock model with a complete skeletal structure from a two-dimensional image.

[0006] To address the aforementioned technical problems, embodiments of the present invention provide a method for constructing a three-dimensional reconstruction system for livestock based on monocular vision, comprising the following steps: Based on multiple RGB and depth videos of livestock from multiple angles, multiple 3D mesh models of livestock are fused and reconstructed; based on the preset bone joint hierarchy and the corresponding skin weight matrix, a livestock model with parameterized skeletal structure is constructed. Multiple personalized livestock models are generated based on multiple 3D mesh models of livestock and livestock models with parameterized skeletal structures. Multiple personalized livestock models are uniformly restored to a standardized livestock posture model through inverse linear hybrid skinning technology. The shape deviation vectors of each standardized livestock posture model and the livestock model with parameterized skeletal structure are calculated, and all shape deviation vectors are combined into a shape deviation matrix. By performing principal component analysis on the shape deviation matrix, orthogonal shape basis vectors that characterize the main body shape variation patterns of livestock are extracted, and a parameterized morphological mixing function is constructed based on the shape basis vectors. A monocular vision-based 3D livestock reconstruction system is constructed. The system generates a 3D livestock model containing skeletal structure by adjusting the weight coefficients of the shape basis vectors in the parameterized morphological mixing function for unidentified livestock images.

[0007] In some optional embodiments, fusing and reconstructing the RGB video and depth video into multiple livestock 3D mesh models includes the following steps: The initial camera pose is obtained through checkerboard calibration, and the initial camera pose is refined using a time-constrained ICP optimization framework. The foreground mask of the livestock is extracted from the RGB video using the Track-Anything tool. The refined initial camera pose, depth video, and foreground mask are then input into the DynamicFusion algorithm for fusion and reconstruction to generate a 3D mesh model of the livestock.

[0008] In some optional embodiments, the method further includes optimizing the livestock model with parameterized skeletal structure, including the following steps: Based on anatomical data, the livestock model with parameterized skeletal structure was subjected to skeletal optimization processing, including joint position adjustment and the addition of joints in key physiological regions; the optimized model was then subjected to cyclic subdivision processing to improve the mesh resolution.

[0009] In some optional embodiments, the parameterized morphological mixing function is formulated as follows: In the formula, For parameterized morphological blending functions, These are the weight coefficients of the orthogonal shape basis vectors. These are orthogonal shape basis vectors; K The number of orthogonal shape basis vectors.

[0010] In some optional embodiments, generating multiple personalized livestock models based on multiple three-dimensional livestock mesh models and livestock models with parameterized skeletal structures includes the following steps: Establish the correspondence between the multiple three-dimensional mesh models of livestock and the livestock model with parameterized skeletal structure, and generate multiple driveable digital models through optimal rigid registration; Based on the preset posture parameters of each drivable digital model, the surface deformation of each drivable digital model is driven by kinematic tree and linear hybrid skinning technology. The deformed drivable digital models are projected onto a two-dimensional image plane, and the difference loss is calculated with the contour, key points and depth information extracted from the corresponding RGB video and depth video. The posture parameters of each drivable digital model are iteratively adjusted according to the difference loss. When the difference loss of each drivable digital model is less than the preset threshold, multiple personalized livestock models with consistent shape and skeletal driving ability are output.

[0011] In some optional embodiments, a three-dimensional reconstruction system for livestock using monocular vision is also included to obtain a skeletal model of the livestock; by calculating the Euclidean distance between corresponding key points of the skeletal model of the livestock, the body length, body height, chest width, chest circumference, hip width, and hip height of the livestock are automatically obtained.

[0012] Embodiments of the present invention also provide a computer device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the above-described method for constructing a three-dimensional livestock reconstruction system based on monocular vision.

[0013] Embodiments of the present invention also provide a computer-readable storage medium storing a computer program, which, when run by a processor, is capable of executing the above-described method for constructing a three-dimensional livestock reconstruction system based on monocular vision.

[0014] The method for constructing a three-dimensional livestock reconstruction system based on monocular vision provided by this invention has at least the following beneficial effects: This invention achieves end-to-end automatic generation from monocular images to skeletal 3D models. By constructing a parametric morphological mixing function based on real scan data, the system input is simplified to ordinary monocular images, significantly lowering the technical application threshold. Unlike general models based on non-real data, the orthogonal shape basis vectors extracted in this invention can accurately characterize the body shape variation patterns of livestock, significantly improving model fidelity, especially the characterization accuracy of key physiological structures such as the mammary gland, providing a reliable foundation for automated body size measurement. Through optimal rigid registration and linear hybrid skinning technology, this invention achieves integrated modeling of geometric morphology and skeletal driving capabilities, generating a drivable digital twin that can be directly used for dynamic analysis. Furthermore, through multi-dimensional constraint joint optimization, the reconstructed model is ensured to maintain visual consistency, geometric accuracy, and biomechanical rationality, effectively avoiding abnormal phenomena such as limb penetration, ultimately generating a 3D model that is both realistic and conforms to physiological laws. Attached Figure Description

[0015] The accompanying drawings, which are included to provide a further understanding of the invention and form part of this invention, illustrate exemplary embodiments of the invention and are used to explain the invention, but do not constitute an undue limitation of the invention. In the drawings:

[0016] Figure 1 This is a flowchart of a method for constructing a three-dimensional reconstruction system for livestock based on monocular vision, according to an embodiment of the present invention; Figure 2 This is a schematic diagram of three-dimensional data of Saanen dairy goats of different ages provided according to an embodiment of the present invention; Figure 3 This is a spatial representation diagram of livestock based on orthogonal shape basis vectors according to an embodiment of the present invention; Figure 4 This is a schematic diagram of the reconstruction result of a three-dimensional livestock reconstruction system based on monocular vision according to an embodiment of the present invention; Figure 5 This is a comparison diagram of manual measurement and parametric model key point measurement of livestock body size based on monocular vision, provided according to an embodiment of the present invention. Figure 6 This is a schematic diagram of the construction process of a three-dimensional livestock reconstruction system based on monocular vision according to an embodiment of the present invention. Detailed Implementation

[0017] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below in conjunction with specific embodiments and corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this invention, and not all of them. Based on the embodiments of this invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this invention.

[0018] One embodiment of the present invention relates to a method for constructing a three-dimensional reconstruction system for livestock based on monocular vision. The implementation details of the method for constructing a three-dimensional reconstruction system for livestock based on monocular vision in this embodiment are described in detail below. The following content is only for the convenience of understanding and is not necessary for implementing this solution.

[0019] The specific process of the livestock 3D reconstruction system construction method based on monocular vision in this embodiment can be described as follows: Figure 1 As shown, it includes: Step 101: Based on multiple RGB and depth videos of livestock from multiple angles, fuse and reconstruct multiple three-dimensional mesh models of livestock; construct livestock models with parameterized skeletal structures based on preset bone joint levels and corresponding skinning weight matrices. A corridor-style acquisition setup was constructed, comprising an RGBD camera array with eight perspectives (one top-view and seven surround-view). Synchronous RGB and depth video of the target livestock was captured using eight synchronized Microsoft Azure Kinect DK RGBD cameras.

[0020] The initial camera pose was obtained through checkerboard calibration, and refined using a time-constrained ICP optimization framework. The foreground mask of the livestock was extracted from the RGB video using the Track-Anything tool. The refined initial camera pose, depth video, and foreground mask were then input into the DynamicFusion algorithm for fusion and reconstruction, generating a 3D livestock mesh model. Radius filtering was used to further suppress noise in the 3D livestock mesh model, resulting in a high-fidelity 3D scanning model of the livestock.

[0021] Three-dimensional data of Saanen dairy goats of different ages, such as Figure 2 As shown in the figure, a three-dimensional model of Saanen dairy goats at 6, 10, 14, and 18 months of age is displayed, showing the changes in their growth and development morphology as they age.

[0022] Construct a livestock model with a parameterized skeletal structure, wherein the livestock model with a parameterized skeletal structure includes a predefined skeletal joint hierarchy and a corresponding skinning weight matrix; Based on anatomical data, the livestock model with parameterized skeletal structure was subjected to skeletal optimization processing, including joint position adjustment and the addition of joints in key physiological regions; the optimized model was then subjected to cyclic subdivision processing to improve the mesh resolution.

[0023] Step 102: Based on multiple 3D mesh models of livestock and livestock models with parameterized skeletal structures, generate multiple personalized livestock models. Establish the correspondence between the multiple three-dimensional mesh models of livestock and the livestock model with parameterized skeletal structure, and generate multiple driveable digital models through optimal rigid registration; The optimized parametric template is registered with the 3D scanned model. First, 32 identical anatomical keypoints are defined on both the scanned model and the template. Optimal rigidity transformation parameters are calculated using singular value decomposition to achieve initial spatial alignment of the models. The purpose of this step is to establish a precise correspondence, generating a drivable digital model that retains the complete skeletal drive capability of the template while maintaining a high degree of geometric consistency with the actual scanned data through optimal rigidity registration. This model forms the basis for subsequent pose and shape optimization, as well as monocular reconstruction.

[0024] Based on the preset posture parameters of each drivable digital model, the surface deformation of each drivable digital model is driven by kinematic tree and linear hybrid skinning technology. The deformed drivable digital model is projected onto a two-dimensional image plane, and the difference loss is calculated with the contour, key points and depth information extracted from the corresponding RGB video and depth video. The posture parameters of each drivable digital model are iteratively adjusted according to the difference loss. When the difference loss of each drivable digital model is less than the preset threshold, multiple personalized livestock models with consistent shape and skeletal driving ability are output. Attitude fitting: Based on the initial attitude parameters, global joint transformations are recursively calculated through a kinematic tree, and linear hybrid skinning technology is used to drive the surface deformation of the drivable digital model.

[0025] Joint optimization: The deformed model is projected onto a 2D image plane and compared in multiple dimensions with contour, keypoint, and depth information extracted from the original RGB-D video. Pose and shape parameters are iteratively optimized by minimizing a joint loss function that includes contour alignment, keypoint matching, depth consistency, and biomechanical constraints to ensure the fitting results are visually, geometrically, and physiologically reasonable and reliable. Once the loss function converges, a personalized livestock model that matches the morphology of a real individual and maintains skeletal drive capabilities is output.

[0026] Step 103: Multiple personalized livestock models are uniformly restored to a standardized livestock posture model through inverse linear hybrid skinning technology. The shape deviation vectors of each standardized livestock posture model and the livestock model with parameterized skeletal structure are calculated, and all shape deviation vectors are combined into a shape deviation matrix. Principal component analysis is performed on the shape deviation matrix to extract orthogonal shape basis vectors that characterize the main body shape variation patterns of livestock. A parameterized morphological mixing function is constructed based on the shape basis vectors. Using the pose parameters estimated by the driveable digital model, all personalized livestock models are uniformly restored to a standard T-pose pose through inverse linear hybrid skinning technology.

[0027] The shape deviation vector between each normalized model and the standard template is calculated and combined into a shape deviation matrix. Principal component analysis is performed on this matrix to extract a set of orthogonal shape basis vectors representing the main body shape variation patterns of livestock. Based on these shape basis vectors, a parametric morphological mixing function is constructed.

[0028] The formula for the parameterized morphological mixing function is as follows: In the formula, For parameterized morphological blending functions, These are the weight coefficients of the orthogonal shape basis vectors. These are orthogonal shape basis vectors; K The number of orthogonal shape basis vectors.

[0029] The size of the generated model can be controlled by adjusting the weighting coefficients. This function is the core of the subsequent construction of the monocular vision reconstruction system.

[0030] A spatial representation of livestock based on orthogonal shape basis vectors is shown below. Figure 3 As shown in the figure, the diverse 3D forms of livestock generated by adjusting the weight coefficients of different orthogonal shape basis vectors demonstrate the variability and flexibility of the parametric model in shape space.

[0031] Step 104: Construct a three-dimensional livestock reconstruction system based on monocular vision. The system generates a three-dimensional livestock model containing skeletal structure by adjusting the weight coefficients of the shape basis vectors in the parameterized morphological mixing function for unidentified livestock images.

[0032] Based on the parametric model built in the preceding steps (combining shape space and skeletal driving capabilities), an extended optimization framework (such as SMALify) is used for fully automated monocular livestock mesh reconstruction. During optimization, shape and pose parameters are adjusted to align the model projection with the input monocular image in terms of contours and key points. Innovative loss functions, such as 3D geometric constraints (e.g., monocular depth) and biomechanical regularization, are introduced to resolve inherent ambiguities in monocular reconstruction and ensure the accuracy and rationality of the reconstructed model. This ultimately constructs a complete monocular vision-based 3D livestock reconstruction system.

[0033] The reconstruction results of the livestock 3D reconstruction system based on monocular vision are as follows: Figure 4 As shown in the figure, the process of generating corresponding 3D models from a multi-view image sequence of goats using algorithms, and then unifying different pose models into a standard T-Pose, is illustrated. This demonstrates the reconstruction and normalization process of livestock 3D models from multiple poses to a standardized pose.

[0034] A monocular vision-based 3D reconstruction system for livestock is used to obtain a skeletal model of livestock. By calculating the Euclidean distance between corresponding key points in the skeletal model of livestock, the body length, body height, chest width, chest circumference, rump width, and rump height of livestock are automatically obtained. Comparison of manual measurement and parametric model key point measurement of livestock body size is shown in the figure below. Figure 5 As shown in the figure, the left side of the image demonstrates the manual measurement of a goat's body length, height, hip height, chest width, and chest circumference, while the right side presents a schematic diagram of the corresponding body size key point measurement based on a parametric 3D model. This provides a direct comparison between the traditional manual measurement and digital model measurement methods for obtaining livestock body size.

[0035] A schematic diagram illustrating the construction process of a monocular vision-based 3D livestock reconstruction system is shown below. Figure 6 As shown in the figure, the complete construction process of the livestock 3D model is presented, which starts from octagonal RGBD data acquisition, generates point cloud model through multi-view Dynamic Fusion, and then obtains T-pose registration results through template initialization and template registration.

[0036] The steps of the various methods described above are only for clarity. In practice, they can be combined into one step or some steps can be split into multiple steps. As long as they include the same logical relationship, they are all within the protection scope of this invention. Adding insignificant modifications or introducing insignificant designs to the algorithm or process, without changing the core design of the algorithm and process, are also within the protection scope of this invention.

[0037] Another embodiment of the present invention relates to a computer-readable storage medium storing a computer program. When executed by a processor, the computer program implements the method embodiments described above.

[0038] That is, those skilled in the art will understand that all or part of the steps in the methods of the above embodiments can be implemented by a program instructing related hardware. This program is stored in a storage medium and includes several instructions to cause a device (which may be a microcontroller, chip, etc.) or processor to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0039] Those skilled in the art will understand that the above embodiments are specific embodiments for implementing the present invention, and in practical applications, various changes can be made to them in form and detail without departing from the spirit and scope of the present invention.

Claims

1. A monocular vision-based livestock three-dimensional reconstruction system construction method, characterized in that, The method comprises: According to the multi-angle RGB video and depth video of the livestock, a three-dimensional grid model of the livestock is reconstructed by fusion; and a livestock model with a parameterized skeletal structure is constructed according to a preset skeletal joint level and a corresponding skin weight matrix. Based on the three-dimensional grid model of the livestock and the livestock model with the parameterized skeletal structure, a plurality of individualized livestock models are generated. The plurality of individualized livestock models are restored to a livestock standardized posture model by reverse linear mixed skin technology, the shape deviation vectors of the livestock standardized posture model and the livestock model with the parameterized skeletal structure are calculated, all the shape deviation vectors are combined into a shape deviation matrix, the orthogonal shape basis vectors representing the main body shape variation mode of the livestock are extracted by performing principal component analysis on the shape deviation matrix, and a parameterized shape mixing function is constructed based on the shape basis vectors. A livestock three-dimensional reconstruction system based on monocular vision is constructed, the livestock three-dimensional model containing the skeletal structure is generated by adjusting the weight coefficients of the shape basis vectors in the parameterized shape mixing function about the unrecognized livestock picture.

2. The monocular vision-based livestock 3D reconstruction system construction method of claim 1, wherein, The RGB video and the depth video are fused to reconstruct the three-dimensional grid model of the livestock, which comprises the following steps: The initial pose of the camera is obtained through the chessboard calibration, the initial pose of the camera is refined using a time-constrained ICP optimization framework, the foreground mask of the livestock is extracted from the RGB video through a Track-Anything tool, and the refined initial pose of the camera, the depth video and the foreground mask are jointly input into a dynamic multi-view DynamicFusion algorithm for fusion reconstruction to generate the three-dimensional grid model of the livestock.

3. The monocular vision-based livestock 3D reconstruction system construction method of claim 1, wherein, The livestock model with the parameterized skeletal structure is also optimized, which comprises the following steps: The livestock model with the parameterized skeletal structure is subjected to a skeletal optimization process based on anatomical data, including joint position adjustment and joint addition in key physiological regions, and the model after the skeletal optimization is subjected to a loop subdivision process to improve the grid resolution.

4. The monocular vision-based livestock 3D reconstruction system construction method of claim 1, wherein, The three-dimensional grid model of the livestock and the livestock model with the parameterized skeletal structure are established, and the corresponding relationship between the three-dimensional grid model of the livestock and the livestock model with the parameterized skeletal structure is established, and a plurality of drivable digital models are generated through optimal rigid registration; Based on the preset posture parameters of each drivable digital model, the surface of each drivable digital model is deformed through a kinematic tree and a linear mixed skin technology, each deformed drivable digital model is projected onto a two-dimensional image plane, and a difference loss is calculated with the contour, key point and depth information extracted from the corresponding RGB video and depth video, the posture parameters of each drivable digital model are iteratively adjusted according to the difference loss, and when the difference loss of each drivable digital model is less than a preset threshold, a plurality of individualized livestock models with consistent shapes and maintaining skeletal driving ability are output. The parameterized shape mixing function formula is as follows:

5. The monocular vision-based livestock 3D reconstruction system construction method of claim 1, wherein, ​ wherein is a parametric morphing function, is a weight coefficient of the orthogonal shape basis vector, is an orthogonal shape basis vector; K is a number of orthogonal shape basis vectors.

6. The monocular vision-based livestock 3D reconstruction system construction method of claim 1, wherein, The method further comprises: using a livestock three-dimensional reconstruction system based on monocular vision to obtain a livestock skeleton model; and automatically obtaining a body length, a body height, a chest width, a chest girth, a hip width and a hip height of the livestock by calculating Euclidean distances between corresponding key points of the livestock skeleton model.

7. A computer system, characterized by The method comprises: at least one processor; and a memory connected to the at least one processor in communication; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the livestock three-dimensional reconstruction system construction method based on monocular vision as claimed in any one of claims 1 to 6.

8. A computer-readable storage medium, characterized in that, The computer readable storage medium stores a computer program, and the computer program, when executed by the processor, can perform the livestock three-dimensional reconstruction system construction method based on monocular vision as defined in any one of claims 1 to 6.