A realistic 3D color texture reconstruction method
By precalibrating the system parameters of the three-dimensional sensor and color texture camera, and optimizing texture fusion using mapping relationships and composite weights, the color problem in three-dimensional geometric measurement is solved, real three-dimensional color texture reconstruction is achieved, and image clarity and efficiency are improved.
Patent Information
- Application Number
- CN201910687176.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2019-07-29
- Publication Date
- 2025-08-26
- Estimated Expiration
- 2039-07-29
AI Technical Summary
The existing three-dimensional geometric measurement technology cannot effectively solve the color problem, resulting in texture artifacts such as blur, ghosting and color discontinuity, and the existing methods require manual participation, reducing efficiency.
By precalibrating the system parameters of the three-dimensional sensor and color texture camera, multi-view three-dimensional images and two-dimensional color texture images are obtained, texture fusion is used to use mapping relationships, composite weights and bidirectional similarity functions are introduced to optimize texture reconstruction, and real three-dimensional color textures are generated.
Realistic three-dimensional color texture reconstruction is realized, eliminating texture discontinuity, improving image clarity and efficiency, and reducing manual participation.
Smart Images

Figure CN110599578B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of electronic technology, and more particularly, relates to a realistic three-dimensional color texture reconstruction method. Background Art
[0002] 3D measurement technology has been widely used in various industries and disciplines, including urban surveying, anthropometrics, and prototyping. Optical 3D measurement technology provides a flexible method for acquiring 3D images. With the development of high-performance optoelectronic devices such as charge-coupled devices (CCDs) and digital light processing (DLP) projectors, optical 3D measurement can obtain highly sensitive and high-speed data. Structured light 3D measurement technology, which uses dynamic, spatially varying patterned illumination, has been widely used in 3D measurement systems. The pattern can be periodic stripes, two-dimensional grids, or random spots. The object's geometry is encoded in a distorted structured light pattern in order to accurately decode it from the captured image.
[0003] However, 3D geometric measurements cannot solve the color problem. Typically, multi-view images are captured and used for mapping onto geometric surfaces to generate color information independent of the geometric reconstruction. Any errors in geometry or camera pose can make the mapping difficult to align, and inconsistent lighting between different views can lead to unrealistic colors. These problems will lead to texture artifacts such as blurring, ghosting, and color discontinuities.
[0004] To address these issues, various texture reconstruction methods have been proposed. Image fusion in the spatial or frequency domain can improve color consistency, but these methods do so at the expense of image clarity. Image stitching using Markov random field optimization methods is another approach to avoid image degradation, but visible seams cannot be completely eliminated. Post-processing methods, such as Poisson blending, heat diffusion, and other color adjustment methods, are often used to adjust the color of texture patches at the seams. Optimizing camera poses can be used to correct for misalignment, such as manual camera calibration, mutual information-based methods, and methods that maximize color consistency. Some methods address misalignment by correcting the input images using non-rigid calibration techniques. For example, optical flow methods for image warping have been introduced to address image misalignment. Furthermore, super-resolution methods have been proposed to overcome blurring. A recent method proposes a patch-based optimization method for texture mapping across multiple images. While various approaches have been employed to eliminate various artifacts, they all require manual intervention, reducing efficiency.
[0005] To solve at least one of the above problems, this paper proposes a realistic 3D color texture reconstruction method. Summary of the Invention
[0006] To solve the above problems, the present invention provides a realistic three-dimensional color texture reconstruction method, which is characterized by including: pre-calibrating system parameters of a three-dimensional sensor and a color texture camera; obtaining multi-perspective three-dimensional images and two-dimensional color texture images; generating a three-dimensional mesh model using the multi-perspective three-dimensional images; establishing a mapping relationship between the two-dimensional color texture image and the three-dimensional mesh model at each perspective based on the system parameters; performing texture fusion based on the mapping relationship to obtain a fused image to achieve color texture reconstruction of the entire three-dimensional model; and generating a corresponding texture map according to the mapping relationship.
[0007] In some embodiments, the texture fusion is to evaluate the confidence of each texture pixel by defining a composite weight based on depth data, and to calculate the fusion result by weighted average based on the confidence under each viewing angle. The composite weight is calculated by the following formula:
[0008] f(x k )=f norm (x k )·f depth (x k )·f edge (x k )
[0009] Among them, the normal weight Depth Weight Edge weight Each weight is normalized to the range [0, 1]. The coefficients for the normal weight are: a = 0.1, b = 50°, the coefficients for the depth weight are: a = 0.4, b = 50 mm, d0 = 55 mm, and the coefficients for the edge weight are: a = -0.08, b = 50 mm.
[0010] In some embodiments, a target image is introduced between the original two-dimensional color texture image and the fused image, and a bidirectional similarity function E is used according to the displacement of the fused image at each viewing angle. BDS (S, T) reconstructs the original two-dimensional color texture image and calculates an energy function, and minimizes the energy function to cause the global target image to shift and reduce the image blur of the overall model texture fusion, ultimately obtaining a new high-resolution target image.
[0011] In some embodiments, the energy function further includes an optical measure consistency function E C :
[0012]
[0013] Among them, M i represents the fused image under the i-th perspective, xk represents the pixel position of the image, P(·) represents the projection function, and N represents the number of viewing angles. j Represents the composite weight of the image at the j-th perspective.
[0014] In some embodiments, the energy function E is ultimately constructed as:
[0015]
[0016]
[0017] E=E1+λE2
[0018] where λ is the scaling factor between the two energy functions E1 and E2.
[0019] In some embodiments, the objective function is solved by a two-step alternating optimization strategy, that is, the following two steps are iteratively solved: S1: fix the fused image M i , optimize the target image T i ; S2: fix the target image T i , optimize the fused image M i .
[0020] In step S1, the target image Ti is expressed as follows:
[0021]
[0022] In step S2, the objective function Mi is expressed as follows:
[0023]
[0024] In some embodiments, the iterative calculation process utilizes a multi-scale optimization method. Specifically, in a low-scale phase, all images are downsampled to a low resolution and the iterative calculation is performed. After the energy function E converges, the target image and the fused image are upsampled to a larger scale, while the original two-dimensional color texture image is still downsampled. This is to inject high-frequency information from the original two-dimensional color texture image into the target image and the fused image.
[0025] The present invention proposes a method for reconstructing realistic 3D color textures. This method uses a composite weight parameter to evaluate the confidence level of texture color. By weighted averaging the projected texture image, texture discontinuities can be eliminated. For inaccurate geometric shapes, a bidirectional similarity (BDS) function, representing the structural similarity between two images, is introduced to correct for inconsistencies and generate realistic textures. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] Figure 1 2 is a schematic diagram of a method for reconstructing realistic three-dimensional color texture according to an embodiment of the present invention.
[0027] Figure 2 Schematic diagram of various weight curves according to an embodiment of the present invention.
[0028] Figure 3 4 is a flow chart of a color texture fusion algorithm based on BSF according to an embodiment of the present invention. DETAILED DESCRIPTION
[0029] The present invention will be further described in detail below in conjunction with specific embodiments and with reference to the accompanying drawings. It should be emphasized that the following description is merely illustrative and is not intended to limit the scope of the present invention and its application.
[0030] Figure 1 FIG2 is a schematic diagram of a method for reconstructing realistic three-dimensional color texture according to an embodiment of the present invention.
[0031] In step 101, the system parameters of the 3D sensor and color texture camera are pre-calibrated. The 3D sensor is used to acquire depth images, i.e., 3D images. The 3D sensor can be a binocular vision 3D sensor based on structured light technology, such as a binocular vision 3D sensor composed of a digital fringe projector and dual color cameras. Of course, any 3D sensor capable of acquiring 3D images can be used, such as a monocular structured light 3D sensor or a time-of-flight 3D sensor. The color texture camera is used to acquire high-resolution object texture information, and can be, for example, a high-resolution SLR camera. Calibration involves calibrating the internal parameters of the 3D sensor and color texture camera, as well as their relative external parameters.
[0032] In one embodiment, a 3D sensor contains three cameras: two black-and-white industrial cameras for generating depth data, and the third is a color camera for acquiring texture images. Calibration of the 3D sensor is a prerequisite for subsequent depth data generation and texture fusion. Starting from the mathematical model of a single camera, the following describes the calibration principles of a binocular sensor composed of black-and-white cameras, as well as the calibration principles of a color texture camera.
[0033] Camera model.
[0034] If the diffraction effect of the imaging system is ignored and the camera lens is assumed to strictly meet the paraxial condition, the camera imaging can be equivalent to pinhole imaging, and the imaging process satisfies the perspective projection transformation. The object point is marked as X in the world coordinate system. w =(X w ,Y w ,Z w ) T , the ideal image point in the image coordinate system is mc =(u,v) T , then the imaging process can be expressed as
[0035]
[0036] Among them, the superscript represents homogeneous coordinates; X c Represents the coordinates of the object point in the camera coordinate system; It's X c Projection on the camera image plane; R c , t c The rotation matrix and translation vector from the world coordinate system to the camera coordinate system are called the external parameters of the camera; K c is the camera internal parameter matrix, including the equivalent focal length along the image coordinate axis (f u ,f v ) T , the projection of the optical center on the image plane, that is, the principal point (u0,v0) of the image plane T , and the image tilt factor γ; M c =K c [R c |t c ] is called the projection matrix, which also includes the internal and external parameters of the camera; λ = Z c is the scale factor.
[0037] In real optical imaging systems, due to factors such as the processing technology and structural assembly of the imaging lens, there will inevitably be deviations between the actual imaging surface and the ideal imaging surface mentioned above, which is called camera lens distortion. The classic Brown-Conrady model is currently the most widely used lens distortion model, and its distortion can be expressed as:
[0038] x′ c =x c +Δ(x c ),
[0039]
[0040] Where x′ c =(x′ c ,y′ c ) T Indicates a distorted image point; Δ(x c ) represents the distortion term, including radial distortion and decentering distortion; is the distance from the undistorted image point to the principal point; (k1, k2, k3, ...) and (p1, p2, p3, ...) are the radial distortion and centrifugal distortion parameters, respectively. Usually, three radial distortions and two centrifugal distortions are sufficient to meet the accuracy requirements. Let k = (k1, k2, k3, p1, p2) T Represents the distortion parameter vector, then the camera model including lens distortion is expressed as
[0041]
[0042] In the nonlinear camera model expressed in Equation (3), k and K c Represents the intrinsic parameters of the camera, R c With t c Represents the external parameters of the camera. The camera calibration process is usually to minimize the reprojection error from the target reference point to the actual image point, that is, X c →m c , so by analyzing the expression x′ c =x c +Δ(x c k) The ideal image point can be distorted, that is, the undistorted image point x c Calculate the distorted image point x′ c The 3D reconstruction process is to construct accurate object points through actual image points, namely X c →m c The 3D reconstruction process is to reconstruct the spatial coordinates of the object point through the actual image points, that is, m c →X c , this process needs to remove the distortion, that is, from the actual image point x′ with distortion c Calculate the undistorted image point x c Since Equation (2) is a complex nonlinear function, its inverse function cannot be obtained by analytical expression. Considering that the distortion term is relatively small, the numerical solution of the undistorted image point can be obtained by recursive approximation:
[0043]
[0044] Distorted image points are used to approximate initial undistorted image points, and more accurate undistorted image points can be obtained by controlling the number of iterations.
[0045] Binocular sensor calibration.
[0046] The left and right cameras can form a three-dimensional sensor based on binocular stereo vision. The world coordinate system is usually set up on the camera. Taking the left camera as an example, the mathematical model of the binocular sensor is expressed as
[0047] X l =R l X w +t l(5)
[0048]
[0049] Where I is the identity matrix, R s and t s is the rotation matrix and translation vector from the left camera coordinate system to the right camera coordinate system, [R s |t s ] represents the structural parameters of the sensor, satisfying
[0050]
[0051] The goal of dual-target calibration is to determine the internal parameters of the two cameras and the structural parameters between the two cameras.
[0052] The camera calibration process usually uses the reprojection error between the projected image points of the target reference points and the actual measured image points as the optimization objective function to obtain the optimal estimate of the internal and external parameters from the target image data. The objective function of the binocular sensor calibration is expressed as
[0053]
[0054] Among them, m′ l and m′ r are the real image coordinates, and The reprojected image coordinates of the reference points calculated based on the model can be optimized using the Gauss-Newton or Levenberg-Marquardt algorithms.
[207] The above formula is optimized and solved to finally obtain the system parameters.
[0055] Color texture camera calibration.
[0056] The function of a color texture camera is to capture color 2D images of objects. By establishing a mapping relationship between a 3D geometric model and a 2D color image, the color information of the 3D mesh is obtained, ultimately achieving color 3D imaging. To achieve this mapping, the internal and structural parameters of the color texture camera must be determined in advance, i.e., the color texture camera must be calibrated. Typically, there are two ways to acquire color images of objects using a color 3D sensor: 1. The two cameras in a binocular sensor are color cameras, capable of simultaneously capturing both depth and color texture information. In this case, the camera internal and external parameters are already determined through binocular sensor calibration, eliminating the need for additional calibration. 2. The binocular camera is separate from the color camera, requiring additional calibration of the color camera. In practical applications, both approaches have advantages and disadvantages. The first approach offers a simple structure and low cost, but due to the Bayer filter imaging principle of color cameras, the grayscale image converted from the color image has lower grayscale accuracy than that obtained directly from a black and white camera, thus affecting the accuracy of the depth data. Furthermore, given the speed of 3D measurement, the resolution of the industrial camera used in the binocular sensor cannot be too high, thus limiting the resolution of the color texture image. The second method has a relatively complex structure and a higher cost, but since it is not restricted by depth data acquisition, a dedicated color camera, such as a professional SLR camera, can be selected according to needs to achieve very high resolution and color reproduction.
[0057] Assuming the left camera coordinate system is the three-dimensional sensor coordinate system, the structural parameters of the three cameras are:
[0058]
[0059] Among them, R t and t t are the rotation matrix and translation vector from the world coordinate system to the color camera coordinate system, R p and t p is the rotation matrix and translation vector between the left camera and the color texture camera. In order to obtain more accurate structural parameters, we add the transformation matrix to the nonlinear objective function of the three cameras and minimize the objective function by the Gauss-Newton or Levenberg-Marquardt method to achieve camera parameter estimation:
[0060]
[0061] where τ={K l ,K r ,K c ,k l ,k r ,k c ,R s ,t s ,Rp ,t p},K c 、k c are the internal parameters of the color camera respectively.
[0062] General process of 3D sensor calibration
[0063] Taking the plane target with circular reference point as an example, the specific calibration process is as follows:
[0064] (1) Target image reference point extraction: Capture multiple sets of target images, extract the coordinates of the circle center on the target image and match them with the known three-dimensional coordinates of the reference point. Use the three-dimensional coordinates of the reference point and the image coordinates as input parameters to solve and optimize the sensor system parameters.
[0065] (2) Obtaining the initial values of camera parameters: Ignoring lens distortion, a linear camera model is used to estimate the internal and external parameters of the camera. To prevent overfitting of the objective function in the next step and to accelerate the convergence of the objective function, the least squares algorithm is used to further optimize the estimated parameters. The obtained results are used as the initial values of the camera parameters and sensor structure parameters.
[0066] (3) Nonlinear optimization of sensor parameters: Lens distortion is added to the camera model, and the nonlinear camera model and sensor structure parameters are used to construct the optimization objective function. The optimal estimation of the sensor parameters is achieved by minimizing the objective function.
[0067] Back to Figure 1 In step 102, the object is captured from multiple perspectives using a 3D sensor and a color texture camera to obtain a multi-perspective 3D image and a 2D color texture image. In one embodiment, the object can be placed on a rotating stage. As the stage rotates, the 3D sensor and the color texture camera capture the object, thereby capturing 3D images and 2D color texture images from multiple perspectives containing 360-degree information about the object.
[0068] Step 103 generates a three-dimensional mesh model using the multi-view three-dimensional image.
[0069] Step 104 establishes a mapping relationship between the two-dimensional color texture image and the three-dimensional mesh model at each viewing angle based on the system parameters to achieve mesh parameterization.
[0070] Step 105 performs texture fusion based on the mapping relationship to obtain a fused image to achieve color texture reconstruction of the entire 3D model. In one embodiment, the multi-view 2D color texture images can be projected onto the 3D network model based on the mapping relationship. Once the 2D color texture images from all viewpoints are projected, a color 3D texture model can be obtained, thus achieving color texture reconstruction of the entire 3D model.
[0071] Step 106 generates a corresponding texture map based on the mapping relationship. For easy storage, multiple texture maps are generated based on the mapping relationship, and the three-dimensional mesh model and texture maps are saved in formats such as obj, ply, and wrl. It is understood that in actual applications, it is necessary to frequently perform new mesh parameterization on the model and generate new texture maps.
[0072] Performing a mean operation on all texture images that establish a mapping relationship is a direct and simple global texture fusion method. However, in reality, due to factors such as uneven lighting and changes in the object's morphology, the brightness of the surface image captured by the color camera is also inconsistent, resulting in a relatively obvious color jump in the fusion result. In order to achieve realistic three-dimensional color texture reconstruction, this patent proposes that the use of texture fusion based on composite weights is an effective method to solve texture color jumps and achieve a natural transition of texture boundaries from different perspectives. That is, the confidence of each texture pixel is evaluated by defining a composite weight through depth data, and the fusion result is calculated by weighted average based on the confidence at each perspective.
[0073] Here we introduce the Sigmoid kernel function (also called logistic function) to distribute the composite weights. The Sigmoid kernel function is defined as follows:
[0074]
[0075] Where f(·)∈(0,1), coefficients a and b are real numbers that control the distribution of the Sigmoid curve. The Sigmoid kernel function is introduced into the composite weights to flexibly control the weight curve by adjusting the coefficients to meet practical needs. The composite weights include normal weights, depth weights, and edge weights.
[0076] The normal weight is assigned based on the angle between the normal of the object surface and the camera's line of sight. According to the classic bidirectional reflectance (BRDF) model, the brightness of the object surface captured by the camera is related to parameters such as the incident angle of the light source, the normal of the object surface, the camera's line of sight, and the surface reflectivity. However, in practical applications, it is difficult to obtain the exact values of these parameters, such as the specific spatial position of the light source, the actual reflectivity of the object surface, and so on. Therefore, we approximate the normal weight as a Sigmoid function. Let the angle between the normal of the object surface and the camera's line of sight in a certain valid area of the image be Δθ k , then the normal weight satisfies:
[0077]
[0078] where x k is a pixel in the image. The larger the angle, the lower the weight. The normal weight curve is as follows: Figure 2As shown in (a), the coefficients in the curve are a=0.1 and b=50°.
[0079] Depth weight is assigned based on the distance from the object surface to the camera imaging plane. Since the camera imaging model is limited by the depth of field (DOF), when certain areas of the object surface exceed the depth of field range of the lens, they will become blurred due to defocus, which will have a negative impact on the quality of texture fusion. The depth weight is based on the optimal imaging distance. The greater the deviation, the smaller the weight. The weight decays slowly within the depth of field range, decays quickly near the depth of field boundary, and is quickly cut off beyond the depth of field range. The depth weight is defined as follows:
[0080]
[0081] Where D(·) represents the shortest distance from the object surface point d to the optimal imaging reference plane d0. The depth weight curve is as follows Figure 2 As shown in (b), the coefficients in the curve are a=0.4, b=50mm, and d0=55mm.
[0082] Edge weights are assigned based on the shortest Euclidean distance between the target pixel in the image and the edge contour of the valid area. Since texture fusion from different perspectives can easily cause brightness jumps at the edge, the closer the target pixel is to the edge contour of the valid area, the lower the weight. The edge weights in this paper are defined as follows:
[0083]
[0084] Where D(·) represents the point x k The shortest distance to the edge contour line. The edge weight can significantly reduce the color jump at the texture boundary and is a very important weight function in texture fusion. The edge weight curve is as follows Figure 2 As shown in (c), coefficient a = -0.08, b = 50 mm.
[0085] The composite weight is constructed by multiplying the above three weights:
[0086] f(x k )=f norm (x k )·f depth (x k )·f edge (x k ) (15)
[0087] Each weight is normalized and its value range is [0,1].
[0088] If the texture mapping relationship under each perspective is accurate enough, and the reconstructed geometric model is fine enough, texture fusion based on composite weights can achieve good texture reconstruction effects. However, due to the combined influence of various links in three-dimensional imaging such as system calibration, single-view depth data reconstruction, and ICP matching, it is actually difficult to meet these assumptions, resulting in texture dislocation and blurring, which reduces the quality of texture reconstruction. To this end, this patent also proposes a texture fusion algorithm that combines composite weights with a bidirectional similarity (BDS) function. The main idea of the algorithm is to introduce a target image between the original image and the fused image, and use a bidirectional similarity function to reconstruct the original image according to the displacement of the fused image at each perspective to calculate and generate an energy function. By minimizing the energy function, the global target image is displaced and the image blur of the overall model texture fusion is reduced to obtain a new target image. The introduction of the bidirectional similarity function is to deform the target image according to the fused image during reconstruction, while including the original image information as much as possible.
[0089] In 2008, Simakov et al. defined the bidirectional similarity function as:
[0090]
[0091] Where S represents the original image, T represents the target image, s and t represent the blocks of the original and target images, respectively. D(s, t) represents the sum of the squared differences between blocks s and t in RGB color space. α is the scaling parameter between the two terms. L represents the number of pixels in each block. For example, for a 7×7 block, L = 49. The first term on the right side of Equation (16) is the completeness term, which indicates the completeness of the target image's information contained in the original image. Lower values indicate greater completeness. The second term is the coherence term, which indicates the presence of new visual structure in the target image relative to the original image (e.g., caused by artifacts). Lower values indicate less new visual structure. By minimizing this function, the target image is guaranteed to contain the maximum amount of information from the original image, subject to visual coherence constraints.
[0092] However, relying solely on the bidirectional similarity function cannot effectively improve the quality of texture fusion. It is also necessary to make the target image and the fused image under multiple views photometrically consistent. For this purpose, another energy function is introduced:
[0093]
[0094] Among them, M i represents the fused image under the i-th perspective, x krepresents the pixel position of the image, and P(·) represents the projection function, for example, P i (T j ) represents the projection of the target image of the jth perspective onto the target image of the ith perspective, and N represents the number of perspectives. j Represents the composite weight of the image at the jth viewpoint, and the weight value is generated according to formula (5). Expand formulas (6) and (7) from a single viewpoint to a global viewpoint, and construct the final energy function:
[0095]
[0096] Where λ is the scaling factor between the two energy functions E1 and E2. By minimizing the energy function E, the target image T under each viewing angle can be generated. i , so that the target image satisfies two constraints: Similarity constraint, that is, it contains the information of the original image as much as possible (corresponding to energy function E1); Consistency constraint, that is, it maintains consistency with the fused image (corresponding to energy function E2).
[0097] Texture alignment and blending
[0098] From the energy function (8), we can see that the objective function T1,...,T N and fused images M1,...,M N are all variables. In order to obtain the optimal solution of the energy function (8), this paper uses a two-step alternating optimization strategy. The basic idea is that when optimizing the target image, all fused images remain unchanged, and when generating the fused image, all target images remain unchanged. First, initialize the original image as the target image and the initialization image of the fused image, that is, T i =S i ,M i =S i , the two-step alternating optimization method is as follows:
[0099] Step S1: Fix the fused image M i , optimize the target image T i At this stage, the fused images M1,...,M N According to formula (18), the target image is associated with both energy functions E1 and E2, so the target image T is solved separately. i For Equation (16), according to Simakov's method, block search is performed by minimizing D(s, t) to determine the correspondence between all blocks in the target image and the blocks in the original image with the smallest error. In order to describe the solution process more clearly, Equation (16) is rewritten as:
[0100]
[0101] Where E1(i,x k ) is the target image pixel x under the i-th viewing angle k The energy function at s u and s v Cover the target image pixel x in the complete term and correlation term of the BSF function respectively k The block of the original image corresponding to the block is determined by block search, y u and y v They are block s u and s v , and corresponds to the pixel in the target image block x k The pixel position of , U and V correspond to the number of blocks in the complete term and the related term respectively. For example, if the block size is 7×7, then U and V are less than or equal to 49. From formula (19), we can see that the energy function E1(i,x k ) is about T i (x k ), we can differentiate (19) and set the derivative to 0, and get
[0102]
[0103] Thus we get the expression of the target image:
[0104]
[0105] From Equation (21), we can see that the target image in the first term is reconstructed based on the information of the original image. Note that 1 / L is retained in the formula in order to merge it with Equation (14).
[0106] The method for minimizing the energy function E2 is similar. Considering P j (P i (T j ))=T j , rewrite the expression of energy function E2:
[0107]
[0108] Taking the derivative of equation (22) and setting the derivative to 0, we get
[0109]
[0110] To keep consistent with the expression of E1, swap the symbols i and j to obtain T i The expression:
[0111]
[0112] Similarly, there is no wi (x k ) is eliminated in order to merge with Equation (21). From Equation (24), we can see that the target image is solved by weighted averaging the texture images under all current relevant perspectives. This constraint reflects that the target image will be aligned according to the result of the fused image.
[0113] Finally, the energy function E is derived and the derivative is set to 0, that is, Combining equations (21) and (24), we can get the expression of the target image:
[0114]
[0115] Step S2: Fix the texture image T i , optimize the fused image M i At this stage, the fused images M1,...,M N is the optimization parameter. According to formula (18), the fused image is only related to the energy function E2, so a similar method can be used to obtain the generation formula of the texture image:
[0116]
[0117] As can be seen from Equation (26), the fused image is obtained by weighted averaging the target images at each relevant perspective. At the beginning of the iterative operation, if the target images at each perspective are misaligned, the fused image will produce ghosting and blurring. During the iterative operation, the target images at each perspective will be aligned according to the fused image, thereby continuously reducing the ghosting and blurring of the image during the fused image reconstruction process. Until the energy function E is less than the set critical value c, it is judged to be converged.
[0118] During the iterative operation, in order to avoid falling into the local optimum and accelerate the convergence speed, we adopt a multi-scale optimization method. In the low-scale stage, all images are downsampled to a low resolution and the above-mentioned iterative operation is performed. After the energy function E converges, the target image and the fused image are upsampled to a larger scale, while the original image is still downsampled, in order to inject the high-frequency information of the original image into the target image and the fused image. In some embodiments, the fused image is blurred in the initial stage and there are obvious ripples in the grayscale curve. However, with the iterative operation of different scales, the image becomes clearer and the contrast of the grayscale curve is enhanced. In the iterative operation of the highest scale, the resolution of all images is adjusted to the initial resolution of the original image, and the target images T1,...,T under all viewing angles are obtained at the same time. N and fused images M1,...,M N In one embodiment, a 10-level scale is used for multi-scale optimization, where the image size of any one dimension of the i-th level image is li The calculation formula is:
[0119] l i =(l0 / 8)·8 (i-1) / 9 (27)
[0120] l0 is the initial image size of the original image. For example, if the initial resolution of the original image is 5520 pixels × 3680 pixels, the image resolution at the first scale is 690 pixels × 460 pixels.
[0121] After the iterative operation of the above two-step method, the new highest resolution target image under all viewing angles is finally obtained, and the target image has been aligned and optimized. At this time, the composite weighted fusion algorithm is used to fuse all the target images to obtain the final texture map of the 3D model. The overall process of the color texture fusion algorithm based on BSF is as follows: Figure 3 shown.
[0122] The above is a further detailed description of the present invention in conjunction with specific / preferred embodiments, and the specific implementation of the present invention should not be considered to be limited to these descriptions. Those skilled in the art to which the present invention relates may make various substitutions or modifications to the described embodiments without departing from the scope of the present invention, and such substitutions or modifications should be considered to fall within the scope of protection of the present invention.
Claims
1. A realistic three-dimensional color texture reconstruction method, characterized in that: include: Pre-calibrate the system parameters of the 3D sensor and color texture camera; Acquire multi-view 3D images and 2D color texture images; generating a three-dimensional mesh model using the multi-view three-dimensional images; Establishing a mapping relationship between the two-dimensional color texture image and the three-dimensional mesh model at each viewing angle based on the system parameters; Performing texture fusion based on the mapping relationship to obtain a fused image to achieve color texture reconstruction of the entire three-dimensional model; Generate a corresponding texture map according to the mapping relationship; The texture fusion is to evaluate the confidence of each texture pixel by defining a composite weight based on depth data, and to calculate the fusion result by weighted average according to the confidence under each viewing angle; The target image is introduced between the original two-dimensional color texture image and the fused image, and a bidirectional similarity function E is used according to the displacement of the fused image at each viewing angle. BDS (S, T) reconstructs the original two-dimensional color texture image and calculates an energy function, and minimizes the energy function to cause the global target image to shift, thereby reducing the image blur of the overall model texture fusion, and finally obtaining a new high-resolution target image; the target image and the fused image under multiple perspectives maintain optical measurement consistency.
2. The method for reconstructing realistic three-dimensional color texture according to claim 1, characterized in that: The composite weight is calculated by the following formula: f(x k )=f norm (x k )·f depth (x k )·f edge (x k ) Among them, the normal weight Depth Weight Edge weight And each weight is normalized and its value range is [0,1].
3. The method for reconstructing realistic three-dimensional color texture according to claim 2, characterized in that: The coefficient values in the normal weight are: a=0.1, b=50°, the coefficient values in the depth weight are: a=0.4, b=50mm, d0=55mm, and the coefficient values in the edge weight are: a=-0.08, b=50mm.
4. The method for reconstructing realistic three-dimensional color texture according to claim 1, characterized in that: The energy function also includes an optical measure consistency function E C : Among them, M i represents the fused image under the i-th perspective, x k represents the pixel position of the image, P(·) represents the projection function, N represents the number of viewing angles, and w j Represents the composite weight of the image at the j-th perspective.
5. The method for reconstructing realistic three-dimensional color texture according to claim 4, characterized in that: The energy function E is finally constructed as: E=E1+λE2 where λ is the scaling factor between the two energy functions E1 and E2.
6. The method for reconstructing realistic three-dimensional color texture according to claim 5, characterized in that: The objective function is solved by a two-step alternating optimization strategy, which iteratively solves the following two steps: S1: Fix the fused image M i , optimize the target image T i ; S2: Fix the target image T i , optimize the fused image M i .
7. The method for reconstructing realistic three-dimensional color texture according to claim 6, characterized in that: In step S1, the target image Ti is expressed as follows: In step S2, the fused image Mi is expressed as follows:
8. The method for reconstructing realistic three-dimensional color texture according to claim 7, characterized in that: A multi-scale optimization method is used in the iterative operation process, that is, in the low-scale stage, all images are downsampled to a low resolution and the above-mentioned iterative operation is performed; when the energy function E converges, the target image and the fused image are upsampled to a larger scale, while the original two-dimensional color texture image is still downsampled, in order to inject the high-frequency information of the original two-dimensional color texture image into the target image and the fused image.
Citation Information
Patent Citations
Depth and color imaging integrated handheld three-dimensional modeling device
CN106530395A