Pig three-dimensional reconstruction method based on sparse multiple views

By introducing adjacent Gaussian anti-pooling strategy and geometric depth constraints in the 3DGS model, the overfitting problem of three-dimensional reconstruction technology from sparse perspective is solved, and a more stable and accurate three-dimensional reconstruction effect of pigs is achieved.

CN120014170APending Publication Date: 2025-05-16SOUTH CHINA AGRICULTURAL UNIVERSITY
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510105711.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-23
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

Existing 3D Gaussian Splatting (3DGS)-based three-dimensional reconstruction technology is prone to overfitting under sparse viewing conditions, resulting in unstable and inaccurate reconstruction results in new perspectives.

Method used

The training and rendering process of 3DGS models is optimized by introducing adjacent Gaussian anti-pooling strategies and geometric depth constraints. Specific measures include increasing the Gaussian distribution density during the three-dimensional scene optimization stage, and improving the model's understanding of scene geometric structure through deep joint optimization of pseudo-view and training view.

Benefits of technology

The stability and accuracy of the three-dimensional reconstruction model at different perspectives is improved, the performance ability of scene geometric details is enhanced, the geometric blur problem under sparse perspective conditions is alleviated, and the reconstruction quality is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014170A_ABST
    Figure CN120014170A_ABST
Patent Text Reader

Abstract

The invention discloses a pig three-dimensional reconstruction method based on sparse multiple views, and the method carries out the optimization of the following two aspects on the basis of a 3DGS three-dimensional model: (1), introducing an adjacent Gaussian anti-pooling strategy, which increases the density of Gaussian distribution in a scene in a more reasonable manner, effectively improves the geometric detail expression capability of the scene, and improves the robustness of the scene; the defects of the 3DGS model in processing complex details are overcome; (2) geometric depth constraint is introduced, depth information of a real or pseudo view is utilized, consistency of rendering depth and estimated depth is optimized through a Pearson's correlation coefficient, the understanding ability of the model to a scene geometric structure is improved, the problem of geometric blur under a sparse view angle condition is effectively relieved, and the method is suitable for being used in a scene. And the reconstruction quality of the sparse multi-view 3DGS model is further improved. According to the method, efficient and accurate pig three-dimensional reconstruction can be realized under the condition of a sparse image visual angle, and the reconstructed pig three-dimensional model is better ensured to still have relatively high stability and precision under different visual angles.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to the technical field of three-dimensional reconstruction, and in particular to a three-dimensional reconstruction method for pigs based on sparse multi-views. Background Art

[0002] my country's pig farming industry has developed rapidly, with significant growth in pig inventory, pork production and consumption, making China the world's largest pork producer and consumer. With the improvement of living standards, consumers' requirements for pork quality and safety continue to increase. In the daily breeding and feeding management of pigs, in order to improve production performance and meat quality, scientific breeding models and the health of pigs are particularly critical. Large-scale and standardized breeding requires the selection of excellent breeds and health monitoring. The body size parameters of pigs are crucial for selecting pigs with excellent body shape and monitoring their health. The efficient and accurate acquisition of the three-dimensional model of pigs is the core technology for obtaining pig body size parameters and optimizing breeding management.

[0003] The sparse multi-view based 3D reconstruction technology can effectively restore the 3D shape of the target by inputting a small number of images from different perspectives while reducing the data acquisition cost. Compared with the traditional dense multi-view method, the sparse multi-view reconstruction technology has significant advantages such as convenient data shooting, low computational overhead and high data processing efficiency. It is especially suitable for application scenarios with limited resources or complex scenes. Figure 3 The 3D reconstruction technology has strong practicality and applicability.

[0004] Among them, 3D Gaussian Splatting (3DGS) technology provides a feasible basic solution. As an innovative technology, 3DGS has rapidly emerged in the fields of explicit radiation field and computer graphics. Figure 1 As shown in the figure, 3DGS generates 3D scenes from a set of 2D images, which is suitable for scene data captured from a set of photos or videos. Unlike implicit neural network models such as Neural Radiance Field (NeRF), 3DGS explicitly represents the scene using millions of 3D Gaussian ellipsoids and achieves efficient rendering through rasterization technology. This method accelerates the rendering process of new perspective synthesis and brings the possibility of real-time rendering.

[0005] However, when the 3DGS model only has sparse viewpoint data, the reduced constraints of the training data make the 3DGS model prone to overfitting of the training viewpoint, resulting in unsatisfactory performance in new viewpoint synthesis. In other words, the performance of the 3DGS model depends largely on the number and accuracy of the initial sparse points. Sparse input views may not only lead to insufficient initialization of Gaussian points, which in turn leads to reconstruction collapse, but also increase the risk of overfitting, resulting in oversmoothing results, that is, it is difficult for the 3DGS model to ensure that the reconstructed 3D model of the pig still has high stability and accuracy under different viewpoints. Therefore, it is challenging to obtain sparse viewpoint data in many complex scene environments in daily life and use sparse data to generate accurate 3D scenes. Summary of the invention

[0006] 1. Technical issues to be solved

[0007] In view of the shortcomings of the prior art, the present invention provides a sparse multi-view based three-dimensional reconstruction method for pigs, which can achieve efficient and accurate three-dimensional reconstruction of pigs under sparse image viewing conditions. The present invention reduces data dependence and optimizes model training strategies. While reducing data acquisition and computing costs, the present invention better ensures that the reconstructed three-dimensional pig model still has high stability and accuracy under different viewing angles.

[0008] (II) Technical solution

[0009] In order to solve the above technical problems, the present invention provides the following technical solutions: a 3D reconstruction method of a pig based on sparse multi-views, comprising the following steps:

[0010] S1, non-contactly photographing a sparse multi-view image of a pig scene with a camera, and obtaining camera parameters of the sparse multi-view image and a sparse point cloud of the three-dimensional scene;

[0011] S2, 3D scene initialization representation: initialize 3D Gaussian based on sparse point cloud;

[0012] S3, 3D scene optimization: Update 3D Gaussian parameters and introduce neighboring Gaussian anti-pooling strategy to increase Gaussian distribution density of the scene;

[0013] S4: Rasterization rendering generates new perspective images and introduces geometric depth constraints to optimize the geometric structure of the model.

[0014] Preferably, in step S3, the neighboring Gaussian de-pooling strategy includes: calculating the average distance of the K nearest neighbors of the existing Gaussian as a proximity score, specifically connecting each existing Gaussian with its nearest K neighbors, representing the Gaussian at the head of the line as the source Gaussian and the Gaussian at the tail as the target Gaussian; if the proximity score exceeds a set threshold, two new Gaussians are grown on each line connecting the source Gaussian and the target Gaussian; wherein the two new Gaussians are respectively at one-third and two-thirds of the line, and the opacity, scaling, and rotation attribute settings of the new Gaussian close to the source Gaussian are consistent with those of the source Gaussian, and the opacity, scaling, and rotation attribute settings of the new Gaussian close to the target Gaussian are consistent with those of the target Gaussian.

[0015] Preferably, in step S4, the geometric depth constraint comprises performing joint depth optimization on the pseudo view and the training view.

[0016] Preferably, the geometric depth constraint specifically includes the following steps:

[0017] The pseudo view and the training view are rendered by differentiable rasterization to generate a rendering depth map as the model prediction value;

[0018] The pseudo view and the training view are input into the pre-trained DPT model to generate the corresponding estimated depth map as the reference depth for optimization;

[0019] The relative loss is used to measure the difference in depth distribution between the rendered depth map and the estimated depth map.

[0020] Preferably, the relative loss is a Pearson correlation coefficient.

[0021] Preferably, in step S1, the camera parameters of each image in the sparse multi-view image and the sparse point cloud of the three-dimensional scene are obtained specifically by using the COLMAP tool.

[0022] Preferably, step S2 further includes inputting camera parameters of the sparse multi-view image and a sparse point cloud of the three-dimensional scene to further initialize the sparse point cloud into a 3D Gaussian.

[0023] Preferably, step S3 further comprises adaptive density control to optimize the Gaussian density of the scene.

[0024] (III) Beneficial effects

[0025] Compared with the prior art, the present invention provides a 3D reconstruction method for pigs based on sparse multi-views, which has the following beneficial effects: The present invention performs the following two optimizations on the basis of the 3DGS model: (1) a neighboring Gaussian anti-pooling strategy is introduced, which effectively improves the geometric detail expression capability of the scene by increasing the density of Gaussian distribution in the scene, thus making up for the deficiency of the 3DGS model in processing complex details; (2) a geometric depth constraint is introduced, which utilizes the depth information of the real or pseudo view and optimizes the depth distribution consistency between the rendered depth and the estimated depth through the Pearson correlation coefficient, thereby improving the model's ability to understand the geometric structure of the scene, effectively alleviating the geometric blur problem under sparse viewing conditions, and further improving the sparse multi-view. Figure 3 Reconstruction quality of DGS model. Through the above methods, the present invention can achieve efficient and accurate three-dimensional reconstruction of pigs under sparse image viewing conditions. By reducing data dependence and optimizing model training strategies, the present invention can reduce data acquisition and computing costs while better ensuring that the reconstructed three-dimensional pig model still has high stability and accuracy under different viewing angles. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] Figure 1 It is a flow chart of the steps of the prior art three-dimensional reconstruction technology method based on 3DGS;

[0027] Figure 2 A flowchart of one step of a method for three-dimensional reconstruction of pigs based on sparse multi-views of the present invention;

[0028] Figure 3 Another step flow chart of a method for three-dimensional reconstruction of pigs based on sparse multi-views of the present invention;

[0029] Figure 4 This is a three-dimensional reconstruction effect diagram of a pig based on sparse multi-views according to the present invention. DETAILED DESCRIPTION

[0030] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0031] The present invention provides a method for three-dimensional reconstruction of a pig based on sparse multi-views, comprising the following steps:

[0032] S1. Sparse multi-view images of a pig scene are captured by a camera in a non-contact manner, and camera parameters of the sparse multi-view images and a sparse point cloud of the three-dimensional scene are obtained.

[0033] S2. Initialization representation of three-dimensional scene: Initialize 3D Gaussian (i.e., 3D Gaussian distribution) based on sparse point cloud.

[0034] S3, 3D scene optimization: Update 3D Gaussian parameters and introduce a neighboring Gaussian anti-pooling strategy to increase the Gaussian distribution density of the scene.

[0035] S4: Rasterization rendering generates new perspective images and introduces geometric depth constraints to optimize the geometric structure of the model.

[0036] The present invention is optimized on the basis of the 3DGS of the prior art. First, a multi-perspective image of a scene is captured by a camera in a non-contact manner. Then, a three-dimensional reconstruction technology method based on 3DGS is used to perform high-quality three-dimensional modeling of the scene. The present invention is introduced and explained in conjunction with 3DGS below.

[0037] The flow chart of the steps of the 3DGS-based 3D reconstruction technology method is as follows: Figure 1 As shown in FIG, the whole method consists of four stages, namely, data preprocessing, 3D scene initialization, 3D scene optimization, and rasterization rendering to generate new perspective images.

[0038] (1) Data preprocessing is the process of preprocessing multi-view images into the data format required by the model, that is, Figure 1 Part (1) of the paper. In this process, the camera first takes multi-view images of the scene non-contactly, and then uses the COLMAP tool to perform three steps: feature extraction, feature matching, and sparse reconstruction to obtain camera parameters such as camera pose and sparse point cloud of the three-dimensional scene. That is, the motion-based structure recovery algorithm (SfM) in COLMAP is used to perform pose estimation and sparse three-dimensional point cloud acquisition. The SFM module in COLMAP adopts an incremental reconstruction method. First, unordered images are selected for feature matching, and matching optimization is performed based on geometric conditions. The sparse structure of the point cloud is restored through triangulation and relative pose estimation is performed. Then, the bundle adjustment method is used to optimize the data structure. Finally, the camera parameters of the multi-view images and the sparse point cloud of the three-dimensional scene are output.

[0039] It can be understood that the above-mentioned step S1 of the present invention corresponds to the data preprocessing stage of 3DGS, in which the sparse multi-view images of the pig scene are photographed non-contactly by the camera, and the camera parameters of each image in the sparse multi-view images and the sparse point cloud of the three-dimensional scene are obtained by the COLMAP tool.

[0040] (2) Three-dimensional scene initialization. The 3DGS model uses a set of 3D Gaussian distributions to represent the three-dimensional scene. These 3D Gaussian distributions are used as basic units. Each 3D Gaussian distribution (referred to as Gaussian distribution, Gaussian, Gaussian point or 3D Gaussian point) is defined by the following parameters: the mean μ is used to represent the center position of the Gaussian; the rotation quaternion q and the scaling vector s form the three-dimensional covariance matrix ∑, which controls the shape and direction of the Gaussian distribution; the opacity α describes the degree of blocking of light propagation by each 3D Gaussian point; and the spherical harmonic coefficients sh used to describe the Gaussian color; these parameters can not only flexibly describe the geometric structure and texture information of the three-dimensional scene, but also can be optimized through back-propagation, thereby achieving automatic learning during the optimization process. This design allows the scene to be parameterized as a set of optimizable Gaussian functions. This provides an efficient and accurate representation for three-dimensional reconstruction.

[0041] Each 3D Gaussian distribution is defined by the complete three-dimensional covariance matrix ∑ in the world coordinate system, with its center point (mean) as μ. The formula of the 3D Gaussian distribution is defined as follows:

[0042]

[0043] Among them, G(x) represents the probability density value of point x under Gaussian distribution G; x represents the coordinates of a point in three-dimensional space, and the dimension is μ represents the center point (mean) position of the 3D Gaussian in three-dimensional space, and its dimension is ∑ is the covariance matrix of 3D Gaussian points (∑ -1 is the inverse matrix of the covariance matrix), which defines the shape and direction of the Gaussian distribution, and its dimension is It is composed of a scaling factor and a rotation matrix: ∑ = RSS T R T ; R is the rotation matrix, which is an orthogonal matrix used to describe the rotation direction of the Gaussian point in three-dimensional space. The dimension is S is a scaling vector used to describe the scale of the Gaussian distribution in different directions, with a dimension of

[0044] It can be understood that the above-mentioned step S2 of the present invention corresponds to the three-dimensional scene initialization representation stage of 3DGS; step S2 also includes inputting camera parameters of the multi-view image and the sparse point cloud of the three-dimensional scene to further initialize the sparse point cloud into a 3D Gaussian distribution.

[0045] (3) 3D scene optimization: Starting from the initial sparse point set generated by the SfM above, an adaptive density control strategy is used to adaptively adjust the number of Gaussians and their density according to the unit volume, thereby optimizing the sparse Gaussian set to a denser Gaussian set that can more accurately represent the scene and has the correct 3D Gaussian parameters. The adaptive density control strategy specifically targets two key issues in 3D Gaussian scenes: under-reconstructed areas and over-reconstructed areas; these two types of areas usually show significant view space position gradients, indicating that these areas have not been well reconstructed. In under-reconstructed areas, new geometry is covered by cloning Gaussians and moving their positions along the gradient direction; in over-reconstructed areas with high variance, large Gaussians are split into smaller Gaussians and their scale is reduced by a factor of 1.6, thereby increasing geometric details while ensuring coverage.

[0046] like Figure 1 As shown in part (3) of , densification is performed every 100 iterations and any Gaussian distribution that is essentially transparent is removed, that is, any distribution whose opacity α is less than a threshold ∈ α In addition, to manage the number and size of Gaussians, the optimization process periodically resets the opacity value α of the Gaussians to near zero, allowing density adjustment while removing low-utilization Gaussians. To optimize the Gaussian parameters, a Sigmoid activation function is used for opacity to constrain its range to [0,1) to obtain smooth gradients, and an exponential activation function is used for covariance scaling.

[0047] It can be understood that the above-mentioned step S3 of the present invention corresponds to the three-dimensional scene optimization stage of 3DGS. The step S3 of the present invention also includes the Gaussian density (i.e., Gaussian distribution density) of the optimized scene by adaptive density control as mentioned in 3DGS, and deletes Gaussians whose opacity is less than a threshold.

[0048] In addition, in sparse view modeling, limited 3D scene coverage is a key issue affecting the reconstruction effect. In order to improve the quality of 3D scene reconstruction, step S3 of the present invention also introduces a proximity Gaussian unpooling strategy based on 3DGS, which improves the model performance by increasing the density of Gaussian distribution in the scene, and can effectively increase the accuracy and details of the 3D scene representation.

[0049] Specifically, in step S3, the neighboring Gaussian unpooling strategy includes: calculating the average distance of the K (K=3) nearest neighbors of the existing Gaussian (i.e., 3D Gaussian distribution) as the proximity score, specifically connecting each existing Gaussian with its K nearest neighbors, and representing the Gaussian at the head of the connection, i.e., the starting Gaussian, as the source Gaussian and the Gaussian at the tail as the target Gaussian. The target Gaussian is one of the K neighbors of the source Gaussian, and these target Gaussians are determined by formula (2):

[0050]

[0051] in, represents the i-th source Gaussian G i K (K=3) minimum distances of the Euclidean distances to nearby Gaussians; K-min() represents the smallest K values, and min() is the minimum function; where d ij =||μ i -μ j || represents Gaussian G i Center and Gaussian G j Euclidean distance between centers, μ i is the source Gaussian G i The center point position, μ j is the target Gaussian G j The center point position; i represents the source Gaussian G i Subscript index, j represents the target Gaussian G j Subscript index.

[0052] If the proximity score of the Gaussian exceeds the set threshold, two new Gaussians are grown on each line connecting the source Gaussian and the target Gaussian, that is, on each edge. i The proximity score P i The calculation formula is shown in formula (3):

[0053]

[0054] Among them, P i is the source Gaussian G i The proximity score is calculated as the average distance of the K (K=3) smallest distances to the nearby Gaussian distances; i represents the source Gaussian G i Subscript index, j represents the target Gaussian G j Subscript index.

[0055] Among them, the centers of the two new Gaussians are respectively at one-third and two-thirds of the connecting line, and the opacity, scaling, and rotation attribute settings of the new Gaussian close to the source Gaussian are consistent with the source Gaussian, and the opacity, scaling, and rotation attribute settings of the new Gaussian close to the target Gaussian are consistent with the target Gaussian; the spherical harmonic coefficients of the two new Gaussians are initialized to 0.

[0056] (4) Rasterization rendering generates a new perspective image. The 3D Gaussian distribution is projected onto the 2D image plane during the rendering process and accumulated through α blending. Given the Jacobian determinant of the affine approximation of the observation transformation W and the projection transformation J, the 2D covariance matrix in the camera coordinates is given by equation (4):

[0057] ∑′=(JWΣW T J T )1:2,1:2 (4)

[0058] The covariance matrix Σ of a 3D Gaussian distribution is similar to describing the configuration of an ellipsoid. Given a scaling matrix S and a rotation matrix R, the corresponding covariance matrix Σ can be found, and the formula is defined as follows:

[0059] ∑=RSS T R T (5)

[0060] Therefore, the pixel color corresponding to the rasterized rendered image is calculated by blending 3D Gaussians that overlap at a given pixel and sort them according to their depth, pixel color The formula definition is as follows:

[0061]

[0062] Among them, α′ i represents the opacity learned by weighting the probability density of the 2D Gaussian projected to the target pixel location. c i Represents the view-dependent color computed from the stored spherical harmonic coefficients f.

[0063] c i is the learned color, opacity α′ i is the learned 3D Gaussian opacity α i The product result with Gaussian distribution, where x′ represents the projection coordinates of the 3D point and μ′ i Represents the projection coordinates of the center point of the 3D Gaussian. Opacity α′ i The formula definition is as follows:

[0064]

[0065] In addition, if Figure 1 As shown in part (4) of , after rasterizing the rendered image, 3DGS further constructs a loss function for the rendered image to back-propagate the optimization model. The loss function combines the L1 loss and the D-SSIM term to improve the balance between the geometric structure and the image quality, where the balance weight λ = 0.2. The loss function formula is defined as formula (8):

[0066] L=(1-λ)L 1 +λL D-SSIM (8)

[0067] Step S4 of the present invention also uses the above 3DGS method to rasterize and render the image. In addition, image data with limited training perspectives can easily lead to overfitting of the model to the existing perspectives. Figure 3The reconstruction quality of the DGS model is improved. Therefore, the present invention introduces a geometric depth constraint mechanism, which includes deep joint optimization of pseudo views and training views. The pseudo views are introduced to enhance the model's unseen perspective images, so as to achieve more accurate synthesis of new perspectives. The pseudo views are sampled from the two closest training views in Euclidean space, the average camera direction is calculated and the virtual direction is inserted between them, random noise is applied to the 3-DOF camera position, and the calculation formula of the pseudo view camera position P′ is shown in formula (9):

[0068]

[0069] Where t∈P represents the camera position and q is the quaternion representing the average rotation of the two cameras; represents a random noise ε that follows a normal distribution with a mean of 0 and a standard deviation of δ.

[0070] The above geometric depth constraint specifically includes the following steps:

[0071] First, the pseudo view and the training view are rendered by differentiable rasterization to generate a rendered depth map as the model prediction value. The 3D Gaussian distribution is used to generate a depth map per pixel x by using a point-based rendering method. p The depth value D(x) of the corresponding pixel is calculated by mixing and superimposing the depth of N ordered Gaussians. p ), the rendering opacity of N ordered Gaussians 2D Gaussian with opacity α and its projection on the image plane Calculate and render opacity as shown in formula (10):

[0072]

[0073] For rendering depth map, the depth D(x p ) can be represented by the Gaussian center μ i It is represented by the α blending of the distance from the camera center o, and the formula is defined as follows:

[0074]

[0075] Furthermore, the pseudo view and the training view are input into the pre-trained DPT (Dense Prediction Transformer) model to generate the corresponding estimated depth map as the reference depth for optimization. The DPT model is a depth estimation method based on the Vision Transformer (ViT), which uses the global attention mechanism to capture long-range dependencies in the scene to achieve high-quality depth map prediction. The core idea of ​​the DPT model is to generate a dense depth map consistent with the resolution of the input image through feature extraction and context modeling.

[0076] Specifically, during the application process, the input image is first adjusted to the format required by the model through preprocessing, including size scaling, pixel normalization, and channel order adjustment. The preprocessed image is used as input and feature extraction is performed through the backbone network of the model. The backbone network of DPT is based on the Vision Transformer structure, which divides the image into patches of fixed size, converts the patches into vector sequences through embedding operations, and inputs them into multi-layer attention modules. These modules model the global dependencies within the image through the self-attention mechanism to extract depth information. After feature extraction, DPT gradually upsamples the features through a specific decoder module. The decoder combines convolution operations with attention mechanisms to capture local details while retaining global context information, and finally generates a high-resolution depth map. In the output stage, the model converts the feature map generated by the decoder into depth values, and the grayscale value of each pixel represents the depth information of the corresponding position. In this way, the depth distribution of the scene can be accurately estimated from a single image through the DPT model.

[0077] In the geometric depth constraint, in order to alleviate the ratio ambiguity between the real scene depth and the estimated depth, the present invention uses a relative loss to measure the depth distribution difference between the rendered depth map and the estimated depth map. The relative loss is preferably the Pearson correlation coefficient, that is, the present invention introduces a relaxed relative loss on the rendered depth map and the estimated depth map: Pearson correlation coefficient It is used to measure the distribution difference between 2D depth maps, and its formula is defined as follows:

[0078]

[0079] in, Represents a rendered depth map and estimated depth map The covariance of and Represents the rendered depth map and estimated depth map The variance of .

[0080] In addition, the total loss function of the present invention includes not only the loss function constructed by the rendered image as described above in 3DGS, but also the Pearson L Pearson Loss: Specifically, the rendered image of the model and construct L between the real image C color Loss, which corresponds to the L1 loss of the above loss function of 3DGS, as shown in formula (13):

[0081]

[0082] Rendered images of models and construct L between the real image C D-SSIM Loss, which corresponds to the D-SSIM term of the above loss function of 3DGS, as shown in formula (14):

[0083]

[0084] In formula (14), SSIM (Structural Similarity) is an indicator for measuring the similarity between two images; D-SSIM (Depth-aware Structural Similarity) is a loss function based on structural similarity (SSIM) and takes depth information into account.

[0085] Construct L for the rendered depth map and estimated depth map of the training view, as well as the rendered depth map and estimated depth map of the pseudo view Pearson The loss is shown in formula (15):

[0086]

[0087] Among them, formula (15) corresponds to formula (12).

[0088] In the whole optimization process, the loss function needs to be back-propagated to optimize the three-dimensional scene, that is, the total loss function formula of the whole model of the present invention is shown in formula (16):

[0089] L=λ 1 L color +λ 2 L D-SSIM +λ 3 L Pearson (16)

[0090] Among them, λ 1 , 2 , 3 is the weight coefficient of each loss term, λ 1 , 2 ,3 The values ​​can be 0.5, 0.2, and 0.05 respectively.

[0091] The present invention proposes a sparse multi-view based 3D reconstruction method for pigs, aiming to achieve high-quality 3D reconstruction of pigs through a small amount of viewing angle input. Compared with the current mainstream 3D reconstruction benchmarks and methods, the present invention evaluates its reconstruction effect under different viewing angles through rendering test comparison. In order to quantify the performance of the model, the commonly used evaluation indicators, including PSNR (peak signal-to-noise ratio), SSIM (structural similarity index) and LPIPS (perceptual loss) are used below to comprehensively evaluate the quality of the rendered image. These indicators can effectively measure the image's detail restoration ability, structural fidelity and perceptual quality, ensuring the visual effect and accuracy of the 3D reconstruction results.

[0092] In the verification process, the pig data used came from the actual 3D model of pigs. As can be seen from Table 1 below, the proposed method (Ours) has better reconstruction effect on the pig dataset than the existing method, and the rendered image of the reconstructed pig model leads in the main performance evaluation indicators PSNR, SSIM and LPIPS.

[0093] Table 1 Comparison of sparse reconstruction effects of the proposed method and existing methods on pig datasets

[0094]

[0095]

[0096] It can be understood that the present invention is an optimization based on the 3DGS model. First, the multi-view image data of the pigs is obtained by non-contact camera shooting, and the data is pre-processed by the COLMAP tool to extract the camera parameters of the multi-view image and the sparse point cloud of the three-dimensional scene. Then, an initial 3D Gaussian distribution is generated with each point as the center, and these Gaussian points are projected to the screen space according to the camera parameters, and rasterization rendering is used to generate a new perspective image. In the three-dimensional scene optimization process, the properties of the 3D Gaussian distribution are optimized by back propagation, and adaptive density control is performed alternately. The 3D Gaussian model generated after the optimization is completed can render high-quality images from any perspective with the help of rasterization technology. The technical problems to be solved by the present invention are the following two points.

[0097] (1) In the 3D reconstruction of pig scenes under sparse viewing angles, limited 3D scene coverage is a key factor affecting the reconstruction quality. In order to improve the reconstruction effect, the present invention introduces a neighboring Gaussian anti-pooling strategy in the 3D scene optimization process, densifies existing Gaussian points to generate new Gaussian points, increases the Gaussian distribution density, and thus improves the scene detail performance and reconstruction accuracy.

[0098] (2) In addition, image data with limited viewing angles can easily lead to overfitting of the model to the existing viewing angles. Therefore, in the process of rasterization rendering to generate images, the present invention introduces geometric depth constraints to optimize the geometric structure of the model. First, the pseudo view and the training view are used to generate a rendering depth map as a model prediction value through differentiable rasterization rendering. Subsequently, the pseudo view and the training view are input into the pre-trained DPT model to generate the corresponding estimated depth map as a reference depth for optimization. In order to alleviate the scale ambiguity problem between the rendered depth map and the estimated depth map, the present invention introduces a loose relative loss-Pearson correlation coefficient between the rendered depth map and the estimated depth map to measure the distribution difference between the 2D depth maps, thereby optimizing the geometric structure of the reconstructed scene model.

[0099] Finally, the present invention uses the current mainstream 3D reconstruction benchmarks and methods to train and test the model to verify the improvement of the present invention in 3D reconstruction quality.

[0100] It should be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the sentence "comprises a ..." does not exclude the existence of other identical elements in the process, method, article or device including the element.

[0101] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A 3D reconstruction method for pigs based on sparse multi-views, characterized in that: The following steps are involved: S1. non-contactly photographing a sparse multi-view image of a pig scene with a camera, and obtaining camera parameters of the sparse multi-view image and a sparse point cloud of the three-dimensional scene; S2, 3D scene initialization representation: initializing 3D Gaussian according to the sparse point cloud; S3, 3D scene optimization: Update 3D Gaussian parameters and introduce neighboring Gaussian anti-pooling strategy to increase Gaussian distribution density of the scene; S4: Rasterization rendering generates new perspective images and introduces geometric depth constraints to optimize the geometric structure of the model.

2. The method for pig 3D reconstruction based on sparse multi-views according to claim 1, characterized in that: In step S3, the neighboring Gaussian de-pooling strategy includes: calculating the average distance of the K nearest neighbors of the existing Gaussian as a proximity score, specifically connecting each of the existing Gaussian with its nearest K neighbors, representing the Gaussian at the head of the line as the source Gaussian and the Gaussian at the tail as the target Gaussian; if the proximity score exceeds a set threshold, growing two new Gaussians on each line connecting the source Gaussian and the target Gaussian; wherein the two new Gaussians are respectively located at one-third of the line and two-thirds of the line, and the opacity, scaling, and rotation attribute settings of the new Gaussian close to the source Gaussian are consistent with those of the source Gaussian, and the opacity, scaling, and rotation attribute settings of the new Gaussian close to the target Gaussian are consistent with those of the target Gaussian.

3. The pig 3D reconstruction method based on sparse multi-views according to claim 1, characterized in that: In the step S4, the geometric depth constraint includes performing joint depth optimization on the pseudo view and the training view.

4. The method for pig 3D reconstruction based on sparse multi-views according to claim 3, characterized in that: The geometric depth constraint specifically includes the following steps: The pseudo view and the training view are rendered by differentiable rasterization to generate a rendering depth map as a model prediction value; Inputting the pseudo view and the training view into a pre-trained DPT model to generate a corresponding estimated depth map as a reference depth for optimization; The depth distribution difference between the rendered depth map and the estimated depth map is measured using a relative loss.

5. The method for pig 3D reconstruction based on sparse multi-views according to claim 4, characterized in that: The relative loss is the Pearson correlation coefficient.

6. The method for pig 3D reconstruction based on sparse multi-views according to claim 1, characterized in that: In the step S1, the camera parameters of each image in the sparse multi-view image and the sparse point cloud of the three-dimensional scene are obtained specifically by using the COLMAP tool.

7. The method for pig 3D reconstruction based on sparse multi-views according to claim 1, characterized in that: The step S2 also includes inputting camera parameters of the sparse multi-view image and a sparse point cloud of the three-dimensional scene to further initialize the sparse point cloud into a 3D Gaussian.

8. The pig 3D reconstruction method based on sparse multi-views according to claim 1, characterized in that: The step S3 also includes adaptive density control to optimize the Gaussian density of the scene.

Citation Information

Cited By

  • Gradient control-based Gaussian rendering image degradation processing method

    CN120580335A

  • Cow body condition automatic scoring method and device

    CN121120922A

  • Urban building agent reconstruction method based on aerial images

    CN122156490A