Three-dimensional model reconstruction method and system for weak texture scene, intelligent terminal and storage medium
By identifying weakly textured regions and utilizing neural radiation field processing and optimization strategies, the problem of 3D reconstruction of weakly textured regions was solved, achieving higher accuracy and consistency in 3D model reconstruction.
Patent Information
- Application Number
- CN202511029004.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-25
- Publication Date
- 2025-11-04
- Estimated Expiration
- 2045-07-25
AI Technical Summary
Existing technologies have limitations in the 3D reconstruction of weakly textured regions, making it difficult to obtain complete 3D attribute data, resulting in poor reconstruction results.
By acquiring multi-view images, identifying weak texture regions, obtaining three-dimensional attribute data through neural radiation field processing, and employing a preset optimization strategy to regularize the volume density, including center focusing constraint, non-surface sparsity constraint, and depth diffusion control constraint, the volume density distribution is optimized.
It improves the integrity and effect of 3D reconstruction in areas with weak textures, and enhances the accuracy and consistency of geometric reconstruction.
Smart Images

Figure CN120526064B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of three-dimensional modeling processing, in particular to a three-dimensional model reconstruction method and system for weak texture scenes, an intelligent terminal and a storage medium. BACKGROUND
[0002] The reconstruction from two-dimensional to three-dimensional scenes is one of the key core tasks in the field of computer vision. Compared with the reconstruction based on single-view images, multi-view images can capture scene information at different angles and angles. The three-dimensional reconstruction based on multi-view images has a more realistic and delicate reconstruction effect.
[0003] At present, the three-dimensional scene reconstruction is usually based on the multi-view stereo matching (MVS, Multi-View Stereo) method. The MVS method can obtain high-precision depth estimation, but has limitations in weak texture areas. Specifically, the MVS method has fewer feature points in the weak texture area, resulting in the inability to obtain complete three-dimensional scene reconstruction. Therefore, the method of the prior art is not conducive to better obtaining three-dimensional attribute data in the weak texture area, and thus is not conducive to improving the completeness and reconstruction effect of three-dimensional reconstruction in the scene existing in the weak texture area.
[0004] Therefore, the related technology needs to be improved and developed. SUMMARY
[0005] The main purpose of the present application is to provide a three-dimensional model reconstruction method and system for weak texture scenes, an intelligent terminal and a storage medium, which aims to solve the technical problem that the MVS method is not conducive to better obtaining three-dimensional attribute data in the weak texture area when used for three-dimensional scene reconstruction in the related technology, and thus is not conducive to improving the completeness and reconstruction effect of three-dimensional reconstruction in the scene existing in the weak texture area.
[0006] In order to achieve the above purpose, the first aspect of the present application provides a three-dimensional model reconstruction method for weak texture scenes, wherein the three-dimensional model reconstruction method for weak texture scenes comprises:
[0007] Obtaining multi-view images corresponding to a three-dimensional scene;
[0008] According to the above multi-view images, determining the corresponding weak texture area in the three-dimensional scene;
[0009] According to the above multi-view images, determining the radiation field input data corresponding to the three-dimensional scene, and according to the radiation field input data, obtaining the three-dimensional attribute data corresponding to the three-dimensional scene through neural radiation field processing, wherein the radiation field input data comprises the spatial coordinates and the viewing angle direction of each sampling point, and the three-dimensional attribute data comprises the color value and the body density of each sampling point in the viewing angle direction;
[0010] optimizing the volume density corresponding to the weak texture region according to a preset optimization strategy, to obtain target optimized volume density corresponding to the weak texture region, wherein the preset optimization strategy comprises a regularization constraint for the volume density;
[0011] generating a three-dimensional reconstruction model corresponding to the three-dimensional scene according to the target optimized attribute data corresponding to the weak texture region and three-dimensional attribute data corresponding to other regions in the three-dimensional scene, wherein the target optimized attribute data corresponding to the weak texture region comprises color values and target optimized volume density corresponding to the weak texture region, and the other regions comprise regions in the three-dimensional scene other than the weak texture region.
[0012] Optionally, the determining of the weak texture region in the three-dimensional scene according to the multi-view image comprises:
[0013] performing a preprocessing operation on the multi-view image to obtain a processed image, wherein the preprocessing operation comprises grayscale processing and smoothing processing;
[0014] calculating gradient values corresponding to each pixel point in the processed image, and determining a first region in the processed image according to the gradient values and a preset gradient threshold;
[0015] calculating a local variance corresponding to the processed image based on a sliding window, and determining a second region in the processed image according to the local variance and a preset variance threshold;
[0016] determining the weak texture region in the three-dimensional scene according to the first region and the second region.
[0017] Optionally, the determining of the weak texture region in the three-dimensional scene according to the first region and the second region comprises:
[0018] obtaining a candidate region according to a union set of the first region and the second region;
[0019] determining a weak texture region in the three-dimensional scene corresponding to a position of the candidate region according to a position correspondence relationship between the multi-view image and the three-dimensional scene.
[0020] Optionally, the determining of the weak texture region in the three-dimensional scene corresponding to the position of the candidate region according to the position correspondence relationship between the multi-view image and the three-dimensional scene comprises:
[0021] generating an initial weak texture mask region corresponding to the candidate region;
[0022] performing a median filtering process on the initial weak texture mask region to obtain a processed initial weak texture mask region;
[0023] obtain a foreground mask region, and obtain a target weak texture mask region according to the foreground mask region and the processed initial weak texture mask region;
[0024] determine a weak texture region in the three-dimensional scene corresponding to the target weak texture mask region according to the position correspondence between the multi-view image and the three-dimensional scene.
[0025] Optionally, the regularization constraint on the volume density includes a center focus constraint, a non-surface sparsity constraint, and a depth diffusion control constraint.
[0026] The center focus constraint is used to control the depth of the center of gravity of the weight distribution of the volume density to the neural radiance field.
[0027] The non-surface sparsity constraint is used to suppress the volume density in a low weight region.
[0028] The depth diffusion control constraint is used to limit the diffusion range of the volume density along the ray direction of the neural radiance field.
[0029] Optionally, the preset optimization strategy further includes a neighborhood volume density smoothing constraint, and the neighborhood volume density smoothing constraint is used to suppress the difference in volume density between adjacent sampling points.
[0030] The second aspect of the present application provides a three-dimensional model reconstruction system for a weak texture scene, wherein the three-dimensional model reconstruction system for the weak texture scene includes:
[0031] a data acquisition module configured to acquire a multi-view image corresponding to a three-dimensional scene;
[0032] a weak texture region determination module configured to determine a corresponding weak texture region in the three-dimensional scene according to the multi-view image;
[0033] a processing module configured to determine radiance field input data corresponding to the three-dimensional scene according to the multi-view image, and obtain three-dimensional attribute data corresponding to the three-dimensional scene by processing the radiance field input data through a neural radiance field, wherein the radiance field input data includes spatial coordinates and a viewing direction corresponding to each sampling point, and the three-dimensional attribute data includes a color value and a volume density corresponding to each sampling point in the viewing direction;
[0034] an optimization module configured to optimize a volume density corresponding to the weak texture region according to a preset optimization strategy to obtain a target optimized volume density corresponding to the weak texture region, wherein the preset optimization strategy includes a regularization constraint on the volume density.
[0035] a three-dimensional reconstruction module configured to generate a three-dimensional reconstruction model corresponding to the three-dimensional scene according to target optimized attribute data corresponding to the weak-textured region and three-dimensional attribute data corresponding to other regions in the three-dimensional scene, wherein the target optimized attribute data corresponding to the weak-textured region comprises color values and target optimized volume densities corresponding to the weak-textured region, and the other regions comprise regions in the three-dimensional scene other than the weak-textured region.
[0036] The third aspect of the present application provides an intelligent terminal, which comprises a memory, a processor, and a computer program stored in the memory and executable on the processor, and the computer program, when executed by the processor, implements the steps of any one of the three-dimensional model reconstruction methods for weak-textured scenes.
[0037] The fourth aspect of the present application provides a computer-readable storage medium, which stores a computer program, and the computer program, when executed by a processor, implements the steps of any one of the three-dimensional model reconstruction methods for weak-textured scenes.
[0038] As can be seen, in the scheme of the present application, multi-view images corresponding to a three-dimensional scene are acquired; a weak-textured region in the three-dimensional scene is determined according to the multi-view images; three-dimensional attribute data corresponding to the three-dimensional scene is obtained through neural radiance field processing according to radiance field input data corresponding to the three-dimensional scene determined according to the multi-view images, wherein the radiance field input data comprises spatial coordinates and viewing direction of each sampling point, and the three-dimensional attribute data comprises color values and volume densities corresponding to each sampling point in the viewing direction; a target optimized volume density corresponding to the weak-textured region is obtained by optimizing the volume density corresponding to the weak-textured region according to a preset optimization strategy, wherein the preset optimization strategy comprises a regularization constraint for the volume density; a three-dimensional reconstruction model corresponding to the three-dimensional scene is generated according to target optimized attribute data corresponding to the weak-textured region and three-dimensional attribute data corresponding to other regions in the three-dimensional scene, wherein the target optimized attribute data corresponding to the weak-textured region comprises color values and target optimized volume densities corresponding to the weak-textured region, and the other regions comprise regions in the three-dimensional scene other than the weak-textured region.
[0039] Compared with the prior art, in the three-dimensional model reconstruction method for a weak texture scene provided in the application, three-dimensional attribute data is obtained based on neural radiance field processing, so that the subsequent three-dimensional reconstruction process is realized. Moreover, after obtaining the three-dimensional attribute data based on neural radiance field processing, the volume density in the three-dimensional attribute data corresponding to the weak texture region is further optimized according to the weak texture region determined based on multi-view images, so that the optimization of the volume density can be realized. In this way, it is beneficial to better obtain the three-dimensional attribute data of the weak texture region, and thus it is beneficial to improve the integrity and reconstruction effect of the three-dimensional reconstruction in the scene existing in the weak texture region. BRIEF DESCRIPTION OF DRAWINGS
[0040] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed to be used in the embodiments or prior art description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.
[0041] Figure 1 is a flowchart of a three-dimensional model reconstruction method for a weak texture scene provided by an embodiment of the present application;
[0042] Figure 2 is a body density regularization diagram based on weight guidance provided by an embodiment of the present application;
[0043] Figure 3 is a general framework logic diagram of a three-dimensional model reconstruction method for a weak texture scene provided by an embodiment of the present application;
[0044] Figure 4 is a component module diagram of a three-dimensional model reconstruction system for a weak texture scene provided by an embodiment of the present application;
[0045] Figure 5 is an internal structure principle diagram of an intelligent terminal provided by an embodiment of the present application. DETAILED DESCRIPTION
[0046] In the following description, specific details such as specific system structures, techniques, etc. are presented in order to thoroughly understand the embodiments of the present application. However, it should be clear to those skilled in the art that the present application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits and methods are omitted to avoid unnecessary details that hinder the description of the present application.
[0047] It should be understood that the word "comprising" when used in the specification and claims herein, specifies the presence of stated features, integers, steps, operations, elements, and / or components, but does not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0048] It should also be understood that the terminology used in the description herein is for the purpose of describing particular embodiments only and is not intended to be limiting. As used in this specification and the appended claims, the singular forms "a," "an," and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise.
[0049] It should further be understood that the term "and / or" as used in the specification and in the claims, if and when used, means any one of the items, or combinations of items, listed is possible and all possible combinations and permutations of these items are encompassed as well.
[0050] As used in this specification and claims, the terms "if' and "when" can each be interpreted to mean "upon determination" or "in response to a determination" or "in response to a classification into" depending on the context. Similarly, the phrase "if determined" or "if classified into [described condition or event]" can be interpreted to mean "upon determination" or "in response to a determination" or "upon classification into [described condition or event]" or "in response to a classification into [described condition or event]" depending on the context.
[0051] The technical solutions in the embodiments of the present application are clearly and completely described below with reference to the drawings of the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work fall within the scope of protection of the present application.
[0052] In the following description, a large number of specific details are set forth in order to provide a thorough understanding of the present application, but the present application can also be implemented in other ways different from those described herein, and those skilled in the art can make similar generalizations without departing from the connotation of the present application, therefore the present application is not limited by the specific embodiments disclosed below.
[0053] In oblique photogrammetry, the MVS method can obtain a higher precision depth estimation, but there is a limitation in weak texture areas. Specifically, the MVS method has fewer feature points in the weak texture area, which leads to the inability to obtain complete three-dimensional scene reconstruction. In some application scenarios, the MVS variant method can be used, which has achieved certain results in weak texture areas, but the completeness and geometric details of three-dimensional reconstruction in weak texture areas need to be improved.
[0054] In some application scenarios, a neural radiance field (NeRF) can be used for processing. Compared with traditional tilt photography methods, NeRF has superior performance in three-dimensional reconstruction of weak texture areas, semi-transparent objects, reflective and refractive surfaces, and the like. Although NeRF-based methods have made significant progress in three-dimensional reconstruction tasks, they still have problems such as discontinuous geometric structure, surface cracks, and local missing in weak texture areas, which affect the structural consistency and integrity of the reconstructed model.
[0055] In some specific application scenarios, a geometric and texture decoupled modeling framework and increasing the sampling density of weak texture areas can be used to improve the reconstruction accuracy of weak texture areas. However, the accurate restoration of geometric details and the maintenance of surface continuity still face significant difficulties, especially in weak texture areas. Therefore, improving the geometric reconstruction quality of weak texture areas is still one of the key problems that need to be broken through in the process of high-precision explicit representation of NeRF.
[0056] To solve at least one of the above technical problems, the present application proposes a three-dimensional model reconstruction method for weak texture scenes. The method effectively enhances the expression ability of complex geometric structures by optimizing the spatial distribution of volume density in weak texture areas, and significantly improves the accuracy and consistency of geometric reconstruction.
[0057] Specifically, in the present application, a plurality of view images corresponding to a three-dimensional scene are obtained. According to the plurality of view images, a weak texture area in the three-dimensional scene is determined. According to the plurality of view images, radiation field input data corresponding to the three-dimensional scene is determined. According to the radiation field input data, three-dimensional attribute data corresponding to the three-dimensional scene is obtained by processing a neural radiance field. The radiation field input data includes spatial coordinates and viewing direction of each sampling point. The three-dimensional attribute data includes color values and volume density of each sampling point in the viewing direction. According to a predetermined optimization strategy, the volume density corresponding to the weak texture area is optimized to obtain target optimized volume density corresponding to the weak texture area. The predetermined optimization strategy includes a regularization constraint for volume density. According to the target optimized attribute data corresponding to the weak texture area and the three-dimensional attribute data corresponding to other areas in the three-dimensional scene, a three-dimensional reconstruction model corresponding to the three-dimensional scene is generated. The target optimized attribute data corresponding to the weak texture area includes color values and target optimized volume density corresponding to the weak texture area. The other areas include areas in the three-dimensional scene except the weak texture area.
[0058] Compared with the prior art, the three-dimensional model reconstruction method for a weak texture scene provided in the application obtains three-dimensional attribute data based on neural radiation field processing, thereby realizing the subsequent three-dimensional reconstruction process. Moreover, after obtaining the three-dimensional attribute data based on neural radiation field processing, the volume density in the three-dimensional attribute data corresponding to the weak texture region is further optimized according to the weak texture region determined based on the multi-view image, and the optimization of the volume density can be realized. In this way, it is beneficial to better obtain the three-dimensional attribute data of the weak texture region, and thus it is beneficial to improve the integrity and reconstruction effect of the three-dimensional reconstruction in the scene existing in the weak texture region.
[0059] As shown in Figure 1 The embodiment of the application provides a three-dimensional model reconstruction method for a weak texture scene, and specifically, the above method comprises the following steps:
[0060] Step S100, acquiring multi-view images corresponding to a three-dimensional scene.
[0061] The three-dimensional scene is a scene that needs to be three-dimensionally reconstructed, and the multi-view images include two-dimensional images acquired based on a plurality of different angles.
[0062] Step S200, determining a weak texture region in the three-dimensional scene according to the multi-view images.
[0063] Specifically, the determination of the weak texture region in the three-dimensional scene according to the multi-view images comprises:
[0064] performing a preprocessing operation on the multi-view images to obtain a processed image, wherein the preprocessing operation comprises grayscale processing and smoothing processing;
[0065] calculating gradient values corresponding to each pixel point in the processed image, and determining a first region in the processed image according to the gradient values and a preset gradient threshold;
[0066] calculating a local variance corresponding to the processed image based on a sliding window, and determining a second region in the processed image according to the local variance and a preset variance threshold;
[0067] determining the weak texture region in the three-dimensional scene according to the first region and the second region.
[0068] Further, the determination of the weak texture region in the three-dimensional scene according to the first region and the second region comprises:
[0069] obtaining a candidate region according to the union of the first region and the second region;
[0070] According to the position correspondence relationship between the multi-view image and the three-dimensional scene, a weak texture region in the three-dimensional scene corresponding to the position of the candidate region is determined.
[0071] Specifically, the weak texture region in the three-dimensional scene corresponding to the position of the candidate region is determined according to the position correspondence relationship between the multi-view image and the three-dimensional scene, including:
[0072] An initial weak texture mask region corresponding to the candidate region is generated.
[0073] The initial weak texture mask region is subjected to median filtering processing to obtain a processed initial weak texture mask region.
[0074] A foreground mask region is obtained, and a target weak texture mask region is obtained according to the foreground mask region and the processed initial weak texture mask region.
[0075] According to the position correspondence relationship between the multi-view image and the three-dimensional scene, a weak texture region in the three-dimensional scene corresponding to the position of the target weak texture mask region is determined.
[0076] In a three-dimensional reconstruction task based on volume rendering, a weak texture region is difficult to provide an effective supervision signal for volume density learning due to the lack of significant image gradient and texture features, resulting in difficulty in accurately focusing the volume density distribution on the real surface position, thereby causing the blur, fracture or even loss of local geometric structure in the reconstruction result. To improve the geometric expression ability of the model in such a region, a weak texture region identification strategy that fuses image edges and local texture features is proposed in the present application, which can be used to pixel-level discriminate the potential weak texture region during the training process and provide structure perception support for the spatial regularization of volume density.
[0077] First, the multi-view image is taken as an input image, the input image is subjected to grayscale processing, converted into a gray value, and smoothed by using a Gaussian filter to suppress noise images. Further, the Sobel operator is used to calculate the Sobel gradient of the image , and the regions with weak edge features are screened as the first region based on a preset gradient threshold . At the same time, the local variance of the image is estimated by using a sliding window , and the texture flat region with a local variance lower than a preset variance threshold is taken as the second region.
[0078] In the present application, the union of the first region and the second region is taken as the candidate region, and an initial weak texture mask region corresponding to the candidate region is generated To further enhance the spatial consistency, a median filter is applied to smooth the generated mask. In addition, to ensure that the extracted region is located on the foreground object, a foreground mask region generated based on the transparency channel is obtained , and an edge intensity gradient map obtained by combining the normalized Sobel operator , the degree of participation of the mask at the edge is controlled to avoid misjudging the real boundary as a weak texture region. The calculation formula of the final target weak texture mask region is as follows:
[0079] (1);
[0080] wherein, represents the initial weak texture mask region, which is 1 only when the pixel position meets the condition, otherwise it is 0. and respectively represent the Sobel gradient of the image and the local variance of the image estimated by the sliding window, and are the corresponding threshold condition parameters. represents the final target weak texture mask region. is an edge intensity gradient map extracted based on the Sobel operator, which is used to control the participation degree of the high gradient structure boundary in the mask construction, so as to avoid misjudging the real edge region as a weak texture region, wherein, and are the length and width of the image respectively. Further, after obtaining the target weak texture mask region, the weak texture region in the three-dimensional scene corresponding to the position of the target weak texture mask region and all the sampling points included in the weak texture region can be determined according to the position correspondence between the multi-view image and the three-dimensional scene. Thus, subsequent optimization processing can be performed on the sampling points in the weak texture region. Step S300, according to the above multi-view image, determine the corresponding radiation field input data of the three-dimensional scene, according to the above radiation field input data, through neural radiation field processing, obtain the corresponding three-dimensional attribute data of the three-dimensional scene, wherein, the radiation field input data includes the spatial coordinates and the viewing direction corresponding to each sampling point, the three-dimensional attribute data includes the color value and the body density corresponding to each sampling point in the viewing direction.
[0081]
[0082] Step S300, according to the above multi-view image, determine the corresponding radiation field input data of the three-dimensional scene, according to the above radiation field input data, through neural radiation field processing, obtain the corresponding three-dimensional attribute data of the three-dimensional scene, wherein, the radiation field input data includes the spatial coordinates and the viewing direction corresponding to each sampling point, the three-dimensional attribute data includes the color value and the body density corresponding to each sampling point in the viewing direction.
[0083] It should be noted that the data set corresponding to the multi-view image includes the internal and external parameters and coordinates corresponding to each image, and the radiation field input data is determined based on the internal and external parameters and coordinates. Alternatively, based on a preset image information processing tool, the multi-view image is obtained. For example, through the open source computer vision library (COLMAP, Common Library for Multi-view Applications with Photogrammetry), the internal and external parameters, coordinates and other information corresponding to the image are obtained, and then the radiation field input data is obtained.
[0084] Specifically, the color value and the body density of each sampling point in the corresponding view direction are obtained by tensor decomposition in the neural radiation field.
[0085] In the embodiment of the application, the vector-matrix tensor decomposition mode (Vector-Matrix Tensor Decomposition Mode) in the tensor-based neural radiation field (TensoRF) is used to perform tensor decomposition on the three-dimensional position and direction of the input multi-view image to obtain the corresponding body density and color value, and the three-dimensional attribute data of the corresponding sampling point.
[0086] Step S400, according to the preset optimization strategy, the body density corresponding to the weak texture region is optimized to obtain the target optimized body density corresponding to the weak texture region, wherein the preset optimization strategy includes a regularization constraint for the body density.
[0087] In the embodiment of the application, the target optimized body density corresponding to the weak texture region close to the real surface is obtained based on optimization.
[0088] Specifically, the regularization constraint for the body density includes a center focusing constraint, a non-surface sparse constraint and a depth diffusion control constraint.
[0089] The center focusing constraint is used to control the depth aggregation of the weight distribution of the body density to the center of gravity of the neural radiation field.
[0090] The non-surface sparse constraint is used to suppress the body density of the low weight region.
[0091] The depth diffusion control constraint is used to limit the diffusion range of the body density along the ray direction of the neural radiation field.
[0092] In this embodiment, the above-mentioned regularization constraint on volume density is based on a weight-guided volume density regularization method. By utilizing the weight distribution generated during the rendering process, the focus, sparsity, and thickness of volume density in space are jointly constrained, guiding the volume density distribution to focus more closely on the geometric surface area in space, thereby improving the geometric fidelity and structural restoration capability of weak texture areas.
[0093] Figure 2 This is a schematic diagram of a weight-guided volume density regularization provided in an embodiment of this application, such as... Figure 2 As shown, the above weight-guided volume density regularization includes three parts: center focusing constraint, non-surface sparsity constraint, and depth diffusion control constraint.
[0094] The aforementioned center-focusing constraint is used to control the centroid depth clustering of the weight distribution of the volume density into the neural radiation field. First, the weights on the rays are calculated. Centroid depth of distribution Based on this, a control function that decays with depth and distance is constructed. This function takes a value close to 1 near the depth center and decays rapidly with increasing distance, effectively distinguishing between weight clustering regions and edge regions. The center focusing constraint term is defined as follows: :
[0095] (2);
[0096] In the above formula (2), and Representing the first The depth and weight values of each sampling point. The hyperparameters used to control the degree of decay can be set and adjusted according to actual needs; the smaller the value, the faster the decay. Indicates the presence of rays Expected value Representing the Volume density at each sampling point.
[0097] The aforementioned non-surface sparsity constraints are used to suppress the volume density of low-weight regions. In this embodiment, an adaptive gating mechanism based on weight distribution is provided to construct a volume density suppression strategy for low-weight regions. First, a dynamic threshold based on the maximum weight of the current ray is set. Then, construct a mask function that decays exponentially with the weights. This mask takes a value close to 1 in high-weight regions and rapidly decays to 0 in low-weight regions, thus suppressing density in unstructured regions. Finally, the non-surface sparsity regularization term... Defined as:
[0098] (3);
[0099] In the above formula (3), Representing the The volume rendering weight corresponding to each sampling point on a ray. The coefficient for adjusting the gate threshold can be set and adjusted according to actual needs. It represents the maximum weight among all sampling points on the current ray. To prevent division by zero errors, a small constant can be set and adjusted according to actual needs.
[0100] The aforementioned depth diffusion control constraint is used to limit the diffusion range of the aforementioned volume density along the ray direction of the neural radiation field. In this embodiment, a volume density-based constraint is introduced. The depth diffusion constraint term for the distribution variance. It should be noted that... Represents volume density, Represents the specific The volume density of a sampling point, and the two have the same meaning in a general sense, which will not be repeated below. First, define The depth-weighted expectation is Then, its variance in the depth direction is calculated to measure... The degree of depth diffusion constitutes the regularization term. This item as The penalty index for thickness distribution encourages a clear and compact response of volume density around the geometric surface region, thereby significantly improving the clarity of structural boundaries and suppressing fuzzy redundancy in the depth direction. Specifically, it is shown in formula (4) below:
[0101] (4);
[0102] In the above formula (4), This is a regularization term for the depth diffusion control constraint.
[0103] Furthermore, to synergistically improve the geometric alignment, regional sparsity, and depth focus of volume density, embodiments of this application combine the above three regularization terms into a unified density regularization objective:
[0104] (5);
[0105] in, , and These represent the relative contribution weight coefficients for controlling the three sub-objectives, which can be set and adjusted according to actual needs. Furthermore, to achieve targeted optimization of weakly textured regions, this paper introduces region masks. The mask is the weak texture mask identified in formula (1). The mapping result in the ray dimension. The regularization term is applied to weakly textured regions under the masking effect to avoid over-constraining textured regions and to maintain the integrity and flexibility of its expression. , and The relative weights of the three regularized sub-objectives are controlled separately. This represents the overall density regularization loss. For ray-level guidance based on weakly textured region masks, This represents the final region-guided density regularization term.
[0106] Furthermore, the aforementioned preset optimization strategy also includes a neighborhood volume density smoothing constraint, which is used to suppress volume density differences between adjacent sampling points.
[0107] In the 3D modeling process, to alleviate the volume density in weakly textured areas To address the discontinuities and high-frequency oscillations in the distribution, this application introduces a neighborhood volume density smoothing constraint term to enhance the local continuity of volume density along the depth sampling dimension, thereby improving the consistency of the geometric structure and the stability of surface reconstruction. This regularization term guides the smoothing of volume density by suppressing drastic fluctuations between adjacent depth sampling points. This results in a smoother spatial distribution, especially in regions with sparse texture information, effectively reducing the generation of pseudo-structures and density instability, thereby improving the spatial consistency and geometric fidelity of the reconstruction results. Specifically, the smoothing term corresponding to the aforementioned neighborhood volume density smoothing constraint... As shown in the following formula (6):
[0108] (6);
[0109] in, This indicates the index of the sampling point along each ray direction. This represents the number of sampling points on a ray. Representing the The volume density of each sampling point. It should be noted that this constraint is only applied to rays in weakly textured regions to ensure the structural coherence of the volume density in the absence of texture guidance, while avoiding excessive smoothing in other regions.
[0110] The above volume density neighborhood volume density smoothing regularization term It only affects the density gradient along the line of sight and does not directly interfere with the color reconstruction loss. The optimization path is optimized. Therefore, while maintaining the stability of the main task optimization, this mechanism can assist the model in generating a more continuous and geometrically consistent volume representation. To further improve the geometric representation accuracy and structural stability of weakly textured regions, this application constructs a joint optimization framework oriented towards volume density, comprehensively considering image reconstruction loss, volume density regularization loss, and local smoothing constraints to jointly optimize the spatial structure of volume density. Its total loss function is... As shown in the following formula (7):
[0111] (7);
[0112] in, This represents the original image supervision loss within the neural radiation field. and The weights of the two regularization sub-terms control the relative contributions of different constraints during training; their values range from [value range missing]. The settings can be adjusted according to actual needs. This joint loss design effectively enhances the model's ability to model the structure and its robustness to geometric representation of weakly textured regions without significantly interfering with the color regression task.
[0113] Step S500: Based on the target optimization attribute data corresponding to the weak texture region and the three-dimensional attribute data corresponding to other regions in the three-dimensional scene, generate a three-dimensional reconstruction model corresponding to the three-dimensional scene. The target optimization attribute data corresponding to the weak texture region includes the color value and target optimization volume density corresponding to the weak texture region. The other regions include regions in the three-dimensional scene other than the weak texture region.
[0114] Specifically, based on the aforementioned color and optimized volume density, a NeRF-based volume rendering method is used for rendering. The obtained rendered image and the real image are used to construct an image supervision loss. Simultaneously, a total loss function as shown in formula (7) is constructed by combining the volume density regularization loss for iterative optimization to improve the 3D reconstruction effect. Specifically, the optimized volume density is densely sampled using a regular grid, discretized into voxel volumes, and then the corresponding 3D grid is obtained based on the moving cube algorithm. Methods such as the moving cube algorithm, Poisson reconstruction, and voxel grid extraction can be used, without specific limitations here.
[0115] In this embodiment of the application, the three-dimensional attribute data used in the three-dimensional reconstruction is optimized three-dimensional attribute data for weak texture areas, so as to better realize the reconstruction of weak texture areas.
[0116] It can be seen that in the three-dimensional model reconstruction method for a weak texture scene provided in the application, the three-dimensional attribute data is obtained based on the neural radiation field processing, so as to realize the subsequent three-dimensional reconstruction process. Moreover, after the three-dimensional attribute data is obtained based on the neural radiation field processing, the volume density in the three-dimensional attribute data corresponding to the weak texture region is further optimized according to the weak texture region determined based on the multi-view image, so that the volume density can be optimized. In this way, it is beneficial to better obtain the three-dimensional attribute data of the weak texture region, and thus it is beneficial to improve the integrity and reconstruction effect of the three-dimensional reconstruction in the scene existing in the weak texture region.
[0117] Figure 3 is a general framework logic schematic diagram of a three-dimensional model reconstruction method for a weak texture scene provided by an embodiment of the application. As shown in Figure 3As shown, in view of the problems of geometric structure blurring and surface discontinuity in the reconstruction of weak texture areas for the NeRF volume rendering method, in the embodiment of the present application, a high-precision three-dimensional reconstruction method for weak texture areas of a three-dimensional scene is proposed based on the framework of TensoRF. It mainly includes three parts: (1) three-dimensional scene weak texture mask identification: based on the input multi-view image, combined with Sobel operator, local variance and filtering algorithm, the edge features and texture distribution of the image are comprehensively analyzed, the area with weak texture change but with stable geometric structure in the multi-view image is extracted, that is, the weak texture mask area of the image, which provides structural guidance for subsequent volume density distribution modeling. (2) Volume density regularization of weak texture area: in view of the problem that due to insufficient illumination or texture features in the weak texture area of the three-dimensional scene (such as flat surface, low gradient area and repeated texture area), the rendering weight distribution is not concentrated, the volume density is difficult to focus on the real surface, and the three-dimensional modeling quality is affected, the embodiment of the present application proposes a volume density regularization strategy based on rendering weight distribution. First, according to the matrix-vector tensor decomposition (VM tensor decomposition, Vector-Matrix decomposition) mode of the TensoRF method, the three-dimensional position and direction of the input multi-view image are tensor decomposed to obtain the corresponding volume density and color value. On this basis, the strategy speculates the concentrated position of the rendering weight based on the volume density and color calculation, and constructs a local density adjustment mechanism based on the rendering weight, so that the volume density is closer to the possible surface position in space, thereby improving the geometric accuracy and surface consistency in the weak texture area. In addition, by designing the density sparsity constraint and the distribution control mechanism of the depth direction, the expansion range of the volume density in the non-surface area is effectively limited, and the structure clarity and spatial continuity of the reconstructed surface are enhanced. (3) Joint loss constraint: to solve the problem of uneven distribution or local incoherence of volume density along the ray direction in the weak texture area, the embodiment of the present application introduces a neighborhood density smoothness regularization term based on image reconstruction supervision and volume density regularization, and constructs a joint loss optimization function. Through the continuous iteration optimization of the joint loss function, a high-quality image is finally generated, and a three-dimensional grid can also be extracted based on the volume density.
[0118] In summary, in the present application, a three-dimensional modeling method for weak texture areas is proposed, which combines a region perception mechanism based on texture distribution and a volume density geometric alignment strategy. This method constructs a spatial perception discrimination mechanism in the weak texture area, guides the volume density distribution to focus on the real geometric surface, and thus realizes high-fidelity modeling and reconstruction of complex structures in low-texture areas.
[0119] Meanwhile, a volume density regularization and joint loss optimization strategy for weak texture regions is designed, which comprehensively considers the focusing property of volume density at the depth center, the sparsity constraint of unstructured regions, and the thickness control and local continuity constraint of density distribution. The strategy jointly regulates the density distribution by introducing a spatial guidance mechanism based on rendering weights and a smoothing regularization term of volume density sequence, so that the density distribution is more accurately fitted to the real geometric surface, thereby effectively improving the geometric positioning accuracy and surface structure consistency of weak texture regions.
[0120] As shown in Figure 4 corresponding to the above-mentioned three-dimensional model reconstruction method for weak texture scenes, the embodiment of the application also provides a three-dimensional model reconstruction system for weak texture scenes, the three-dimensional model reconstruction system for weak texture scenes comprises:
[0121] The data acquisition module 410 is configured to acquire multi-view images corresponding to a three-dimensional scene.
[0122] The weak texture region determination module 420 is configured to determine a weak texture region in the three-dimensional scene according to the multi-view images.
[0123] The processing module 430 is configured to determine radiance field input data corresponding to the three-dimensional scene according to the multi-view images, and obtain three-dimensional attribute data corresponding to the three-dimensional scene by neural radiance field processing according to the radiance field input data, wherein the radiance field input data comprises spatial coordinates and a viewing direction corresponding to each sampling point, and the three-dimensional attribute data comprises a color value and a volume density corresponding to each sampling point in the viewing direction.
[0124] The optimization module 440 is configured to optimize the volume density corresponding to the weak texture region according to a preset optimization strategy to obtain target optimized volume density corresponding to the weak texture region, wherein the preset optimization strategy comprises a regularization constraint for the volume density.
[0125] The three-dimensional reconstruction module 450 is configured to generate a three-dimensional reconstruction model corresponding to the three-dimensional scene according to target optimized attribute data corresponding to the weak texture region and three-dimensional attribute data corresponding to other regions in the three-dimensional scene, wherein the target optimized attribute data corresponding to the weak texture region comprises a color value and target optimized volume density corresponding to the weak texture region, and the other regions comprise regions in the three-dimensional scene except the weak texture region.
[0126] Further, the weak texture region determination module 420 specifically comprises:
[0127] The preprocessing unit is configured to perform a preprocessing operation on the multi-view images to obtain a processed image, wherein the preprocessing operation comprises grayscale processing and smoothing processing.
[0128] a first region determining unit, configured to calculate gradient values corresponding to each pixel point in the processed image, and determine a first region in the processed image according to the gradient values and a preset gradient threshold;
[0129] a second region determining unit, configured to calculate local variances corresponding to the processed image based on a sliding window, and determine a second region in the processed image according to the local variances and a preset variance threshold;
[0130] a weak texture region determining unit, configured to determine a corresponding weak texture region in the three-dimensional scene according to the first region and the second region.
[0131] In this way, the three-dimensional attribute data is obtained based on the neural radiance field processing, so that the subsequent three-dimensional reconstruction process is realized. Moreover, after the three-dimensional attribute data is obtained based on the neural radiance field processing, the volume density in the three-dimensional attribute data corresponding to the weak texture region is further optimized according to the weak texture region determined based on the multi-view image, so that the optimization of the volume density can be realized. In this way, it is beneficial to better obtain the three-dimensional attribute data of the weak texture region, and further beneficial to improve the integrity and reconstruction effect of the three-dimensional reconstruction in the scene existing in the weak texture region.
[0132] It should be noted that the specific structure and implementation manner of the three-dimensional model reconstruction system for a weak texture scene and each module or unit thereof can refer to the corresponding description in the above method embodiments, which will not be described here.
[0133] It should be noted that the division manner of each module of the three-dimensional model reconstruction system for a weak texture scene is not unique, and is not specifically limited here.
[0134] Based on the above embodiments, the present application further provides an intelligent terminal, and a principle block diagram thereof can be as shown in Figure 5 The intelligent terminal includes a processor, a memory, a network interface and a display screen connected through a system bus. The processor of the intelligent terminal is configured to provide computing and control capabilities. The memory of the intelligent terminal includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operating system and the computer program in the non-volatile storage medium to run. The network interface of the intelligent terminal is configured to communicate with external terminals through network connection. The computer program is executed by the processor to realize the steps of any one of the three-dimensional model reconstruction methods for a weak texture scene. The display screen of the intelligent terminal can be a liquid crystal display screen or an electronic ink display screen.
[0135] Those skilled in the art can understand that Figure 5The principle block diagram shown in the figure is only a block diagram of part of the structure related to the scheme of the present application, and does not constitute a limitation on the intelligent terminal to which the scheme of the present application is applied. The specific intelligent terminal can include more or fewer components than those shown in the figure, or combine certain components, or have a different component arrangement.
[0136] In an embodiment, an intelligent terminal is provided, which includes a memory, a processor, and a computer program stored in the memory and executable on the processor, and when the computer program is executed by the processor, the steps of any of the three-dimensional model reconstruction methods for a weak-texture scene provided in the embodiments of the present application are implemented.
[0137] The embodiments of the present application also provide a computer readable storage medium, which stores a computer program, and when the computer program is executed by a processor, the steps of any of the three-dimensional model reconstruction methods for a weak-texture scene provided in the embodiments of the present application are implemented.
[0138] It should be understood that the sequence numbers of the steps in the above embodiments do not mean the order of execution, and the execution order of the processes should be determined according to their functions and inherent logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.
[0139] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the division of the above functional units and modules is taken as an example for illustration, and in actual application, the above functions can be completed by different functional units and modules according to needs, that is, the internal structure of the above device is divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated in one processing unit, or each unit can exist physically, or two or more units can be integrated in one unit, and the integrated unit can be realized in the form of hardware or in the form of software functional unit. In addition, the specific names of the functional units and modules are only for easy distinction, and do not limit the protection scope of the present application. The specific working process of the units and modules in the above device can refer to the corresponding process in the foregoing method embodiments, which will not be repeated here.
[0140] In the above embodiments, the description of each embodiment has its own emphasis, and the parts not described or recorded in detail in a certain embodiment can refer to the relevant description of other embodiments.
[0141] Those skilled in the art can understand that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be realized in electronic hardware or a combination of computer software and electronic hardware. Whether the functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different ways to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.
[0142] In the embodiments provided in the present application, it should be understood that the disclosed system / terminal device and method can be implemented in other ways. For example, the above-described system / terminal device embodiments are only schematic, for example, the division of the above modules or units is only a logical function division, and an actual implementation can be different, for example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed.
[0143] The integrated modules / units described above, if realized in the form of software function units and sold or used as independent products, can be stored in a computer readable storage medium. Based on such understanding, all or part of the processes in the above-described embodiment methods can also be completed by computer programs instructing related hardware, and the above computer programs can be stored in a computer readable storage medium. The computer programs are executed by the processor, and the steps of the above various method embodiments can be implemented. The computer programs include computer program codes, which can be in the form of source code, object code, executable files or some intermediate forms. The computer readable medium can include any entity or device capable of carrying the above computer program codes, recording medium, U disk, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal and software distribution medium, etc. It should be noted that the computer readable storage medium contains contents which can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction.
[0144] The above-described embodiments are only used to illustrate the technical solutions of the present application, rather than limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand; it can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement for part of the technical features; and these modifications or replacements, which do not deviate from the spirit and scope of the technical solutions of the embodiments of the present application, should be included in the protection scope of the present application.
Claims
1. A method for reconstructing 3D models for weakly textured scenes, characterized in that, The method includes: Acquire multi-view images corresponding to a 3D scene; Based on the multi-view images, determine the corresponding weak texture regions in the three-dimensional scene; The radiation field input data corresponding to the three-dimensional scene is determined based on the multi-view image. Based on the radiation field input data, the three-dimensional attribute data corresponding to the three-dimensional scene is obtained through neural radiation field processing. The radiation field input data includes the spatial coordinates and viewing direction corresponding to each sampling point, and the three-dimensional attribute data includes the color value and volume density corresponding to each sampling point under the viewing direction. The volume density corresponding to the weak texture region is optimized according to a preset optimization strategy to obtain the target optimized volume density corresponding to the weak texture region. The preset optimization strategy includes regularization constraints on the volume density. Based on the target optimization attribute data corresponding to the weak texture region and the three-dimensional attribute data corresponding to other regions in the three-dimensional scene, a three-dimensional reconstruction model corresponding to the three-dimensional scene is generated. The target optimization attribute data corresponding to the weak texture region includes the color value and target optimization volume density corresponding to the weak texture region. The other regions include regions in the three-dimensional scene other than the weak texture region. The step of determining the corresponding weak texture region in the 3D scene based on the multi-view image includes: The multi-view image is preprocessed to obtain a processed image, wherein the preprocessing operation includes grayscale processing and smoothing processing; The gradient value corresponding to each pixel in the processed image is calculated, and the first region in the processed image is determined based on the gradient value and a preset gradient threshold. The local variance of the processed image is obtained by calculating based on a sliding window, and a second region in the processed image is determined based on the local variance and a preset variance threshold. Candidate regions are obtained based on the union of the first region and the second region; Generate the initial weak texture mask region corresponding to the candidate region; The initial weak texture mask region is subjected to median filtering to obtain the processed initial weak texture mask region. Obtain the foreground mask region, and based on the foreground mask region and the processed initial weak texture mask region, obtain the target weak texture mask region; Based on the positional correspondence between the multi-view image and the three-dimensional scene, the weak texture region in the three-dimensional scene corresponding to the position of the target weak texture mask region is determined.
2. The 3D model reconstruction method for weakly textured scenes according to claim 1, characterized in that, The regularization constraints on volume density include center-focusing constraints, non-surface sparsity constraints, and depth diffusion control constraints. The central focusing constraint is used to control the centroid depth clustering of the weighted distribution of the volume density toward the neural radiation field. The non-surface sparsity constraint is used to suppress the volume density of low-weight regions; The depth diffusion control constraint is used to limit the diffusion range of the volume density along the ray direction of the neural radiation field.
3. The method for reconstructing 3D models for weakly textured scenes according to claim 1 or 2, characterized in that, The preset optimization strategy also includes a neighborhood volume density smoothing constraint, which is used to suppress the volume density difference between adjacent sampling points.
4. A 3D model reconstruction system for weakly textured scenes, characterized in that, The system includes: The data acquisition module is used to acquire multi-view images corresponding to the 3D scene; The weak texture region determination module is used to determine the corresponding weak texture region in the three-dimensional scene based on the multi-view image. The processing module is used to determine the radiation field input data corresponding to the three-dimensional scene based on the multi-view image, and to obtain the three-dimensional attribute data corresponding to the three-dimensional scene through neural radiation field processing based on the radiation field input data. The radiation field input data includes the spatial coordinates and viewing direction corresponding to each sampling point, and the three-dimensional attribute data includes the color value and volume density corresponding to each sampling point under the viewing direction. An optimization module is used to optimize the volume density corresponding to the weak texture region according to a preset optimization strategy to obtain the target optimized volume density corresponding to the weak texture region, wherein the preset optimization strategy includes regularization constraints on volume density. The 3D reconstruction module is used to generate a 3D reconstruction model corresponding to the 3D scene based on the target optimization attribute data corresponding to the weak texture region and the 3D attribute data corresponding to other regions in the 3D scene. The target optimization attribute data corresponding to the weak texture region includes the color value and target optimization volume density corresponding to the weak texture region, and the other regions include regions in the 3D scene other than the weak texture region. The weak texture region determination module includes: A preprocessing unit is used to perform preprocessing operations on the multi-view image to obtain a processed image, wherein the preprocessing operations include grayscale processing and smoothing processing; The first region determination unit is used to calculate the gradient value corresponding to each pixel in the processed image, and determine the first region in the processed image based on the gradient value and a preset gradient threshold. The second region determination unit is used to calculate the local variance corresponding to the processed image based on a sliding window, and determine the second region in the processed image based on the local variance and a preset variance threshold. A weak texture region determination unit is configured to: obtain candidate regions based on the union of the first region and the second region; generate an initial weak texture mask region corresponding to the candidate regions; perform median filtering on the initial weak texture mask region to obtain a processed initial weak texture mask region; obtain a foreground mask region; obtain a target weak texture mask region based on the foreground mask region and the processed initial weak texture mask region; and determine the weak texture region in the three-dimensional scene corresponding to the position of the target weak texture mask region based on the positional correspondence between the multi-view image and the three-dimensional scene.
5. A smart terminal, characterized in that, The smart terminal includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the computer program is executed by the processor, it implements the steps of the three-dimensional model reconstruction method for weakly textured scenes as described in any one of claims 1 to 3.
6. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the steps of the three-dimensional model reconstruction method for weakly textured scenes as described in any one of claims 1 to 3.
Citation Information
Patent Citations
Three-dimensional model reconstruction method and system for local dark texture scene
CN119048710A
3D Gaussian weak texture compensation and density control reconstruction method
CN120070752A