Semantic and posture coupled image reconstruction method, device, equipment and storage medium
By combining depth information and semantic information, using a two-dimensional Gaussian function to adjust the camera pose and perform global calibration, the difficulty of obtaining accurate camera poses in the prior art is solved, and efficient and accurate three-dimensional image reconstruction is achieved.
Patent Information
- Application Number
- CN202510139411.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-08
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-02-08
AI Technical Summary
The prior art requires expensive high-precision sensors or complex preprocessing steps when obtaining accurate camera positioning, which makes it time-consuming and susceptible to noise, limiting the application of neural rendering technology in different scenarios.
By obtaining the depth information and semantic information of each frame image, an initial local point cloud is generated, and it is converted into an initial local Gaussian point cloud based on a two-dimensional Gaussian function, the camera pose is adjusted, and the calibrated camera pose is then obtained through global calibration.
It realizes the acquisition of accurate camera poses without relying on expensive sensors or complex preprocessing steps, improving the speed and quality of 3D image reconstruction.
Smart Images

Figure CN119579807B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image processing technology, and in particular to a semantic and posture coupled image reconstruction method, apparatus, computer equipment, computer-readable storage medium and computer program product. Background Art
[0002] In recent years, image reconstruction technology based on real-life images for 3D reconstruction has gradually developed, and the demand and application of this technology in the fields of cultural relics protection, virtual reality, augmented reality, and game modeling are becoming more and more extensive. Currently, neural rendering technology has made significant progress in scene reconstruction and new perspective synthesis. These technologies usually rely on accurate pre-calculated camera poses as a priori conditions.
[0003] However, obtaining accurate camera poses often requires expensive high-precision sensors to obtain the camera poses, or requires the use of complex preprocessing steps such as motion recovery structure (SFM), which is time-consuming and susceptible to noise. In practical applications, especially in dynamic scenes or large-scale data scenarios, this dependence limits the applicability of methods such as neural radiance field rendering (NeRF) and 3D Gaussian Splatting (3DGS). This limits the application of the new technology of neural rendering in different scenarios, making it unable to meet the current requirements for speed and quality of 3D image reconstruction of real scenes. Summary of the invention
[0004] Based on this, it is necessary to provide a semantic and posture coupled image reconstruction method, device, computer equipment, computer-readable storage medium and computer program product that can balance speed and accuracy in order to address the above technical problems.
[0005] In a first aspect, the present application provides a semantic and pose coupled image reconstruction method, comprising:
[0006] Obtain the depth information and semantic information of each frame image;
[0007] Generate an initial local point cloud corresponding to each frame of the image according to the depth information;
[0008] Based on a two-dimensional Gaussian function, the initial local point cloud is converted into an initial local Gaussian point cloud, and according to the initial local Gaussian point cloud, the camera pose corresponding to each frame of the image is adjusted to obtain an initial camera pose;
[0009] Performing global calibration on the initial camera pose according to the semantic information to obtain a calibrated camera pose;
[0010] A three-dimensional image is reconstructed according to the calibrated camera posture.
[0011] In one embodiment, adjusting the camera pose corresponding to each frame of the image according to the initial local Gaussian point cloud to obtain the initial camera pose includes:
[0012] Adjusting the initial local Gaussian point cloud according to the local attributes corresponding to the initial local point cloud to obtain an adjusted local Gaussian point cloud;
[0013] Determining adjusted local semantic information corresponding to the adjusted local Gaussian point cloud through semantic attributes and local semantic information corresponding to the initial local point cloud;
[0014] According to the adjusted local Gaussian point cloud and the adjusted local semantic information, the camera pose corresponding to each frame of the image is adjusted to obtain an initial camera pose.
[0015] In one embodiment, the local attribute includes color value and transparency; the semantic attribute includes transparency; the local semantic information includes semantic features; the initial local point cloud includes the i-th point; the initial local Gaussian point cloud includes a two-dimensional Gaussian calculation result of the i-th point;
[0016] The step of adjusting the initial local Gaussian point cloud by using the local attributes corresponding to the initial local point cloud to obtain an adjusted local Gaussian point cloud includes:
[0017] Combine the color value and transparency corresponding to the i-th point and the two-dimensional Gaussian calculation result of the i-th point to obtain the Gaussian combination result of the i-th point; determine the i-th Gaussian weight that is negatively correlated with the depth value order of the i-th point in the initial local point cloud; adjust the Gaussian combination result of the i-th point according to the i-th Gaussian weight to obtain the Gaussian value to be accumulated at the i-th point; accumulate the Gaussian value to be accumulated at the i-th point and the Gaussian cumulative value of the previous (i-1) points to obtain the Gaussian cumulative value of the previous i points; the Gaussian cumulative value of the previous i points is the value of the adjusted local Gaussian point cloud at the i-th point; wherein i is a positive integer greater than 2;
[0018] The method of determining the adjusted local semantic information corresponding to the adjusted local Gaussian point cloud through the semantic attributes and local semantic information corresponding to the initial local point cloud includes: combining the semantic feature, transparency and semantic feature corresponding to the i-th point to obtain the semantic combination result of the i-th point; adjusting the semantic combination result of the i-th point according to the i-th Gaussian weight to obtain the semantic value to be accumulated of the i-th point; accumulating the semantic value to be accumulated of the i-th point with the semantic accumulation values of the first (i-1) points to obtain the semantic accumulation values of the first i points; the semantic accumulation values of the first i points are the adjusted local semantic information corresponding to the i-th point; wherein i is a positive integer greater than 2.
[0019] In one embodiment, the image of each frame includes images of a current frame and a next frame, and the value of the adjusted local Gaussian point cloud of the current frame and the adjusted local semantic information are parameters of a single frame loss;
[0020] The step of adjusting the camera pose corresponding to each frame of the image according to the adjusted local Gaussian point cloud and the adjusted local semantic information to obtain an initial camera pose includes:
[0021] Under the condition that the initial camera pose of the current frame remains unchanged, the local attribute is adjusted until the single frame loss of the current frame satisfies a convergence condition, thereby obtaining the adjusted local attribute of the current frame;
[0022] When the adjusted local properties remain unchanged, the camera pose difference between the current frame and the next frame is optimized until the single frame loss of the next frame meets the convergence condition, and the initial camera pose of the next frame is obtained according to the optimized camera pose difference.
[0023] In one of the embodiments, the single-frame loss includes a local Gaussian loss with the value of the adjusted local Gaussian point cloud as an independent variable, an L1 regularized loss with the adjusted local semantic information corresponding to the adjusted local Gaussian point cloud as an independent variable, and an L1 regularized loss with the depth information corresponding to the adjusted local Gaussian point cloud as an independent variable; the local Gaussian loss includes an L1 regularized loss with the value of each block of the adjusted local Gaussian point cloud as an independent variable, and a pixel similarity loss of the value of each block of the adjusted local Gaussian point cloud;
[0024] The adjusting the local attribute includes:
[0025] The local attribute, the value of the adjusted local Gaussian point cloud, the adjusted local semantic information corresponding to the adjusted local Gaussian point cloud, and the depth information corresponding to the adjusted local Gaussian point cloud are adjusted.
[0026] In one embodiment, the globally calibrating the initial camera pose according to the semantic information to obtain a calibrated camera pose includes:
[0027] Generating an initial global point cloud of each frame of the image according to the depth information;
[0028] Based on the initial camera pose and a global Gaussian model corresponding to the initial global point cloud, adjusting the initial global point cloud to obtain an adjusted global point cloud;
[0029] Determining adjusted global semantic information through semantic attributes and semantic information corresponding to the initial global point cloud;
[0030] The initial camera pose is optimized until the global loss corresponding to the adjusted global semantic information meets a convergence condition, thereby obtaining a calibrated camera pose.
[0031] In a second aspect, the present application also provides a semantic and posture coupled image reconstruction device, comprising:
[0032] An information acquisition module is used to acquire the depth information and semantic information of each frame image;
[0033] A point cloud conversion module, used to generate an initial local point cloud corresponding to each frame of the image according to the depth information;
[0034] A Gaussian processing module, used for converting the initial local point cloud into an initial local Gaussian point cloud based on a two-dimensional Gaussian function, and adjusting the camera pose corresponding to each frame of the image according to the initial local Gaussian point cloud to obtain an initial camera pose;
[0035] A global calibration module, used for globally calibrating the initial camera pose according to the semantic information to obtain a calibrated camera pose;
[0036] An image reconstruction module is used to reconstruct a three-dimensional image according to the calibrated camera posture.
[0037] In a third aspect, the present application further provides a computer device, which includes a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the steps of semantic and posture coupled image reconstruction in any of the above embodiments are implemented.
[0038] In a fourth aspect, the present application further provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of semantic and posture coupled image reconstruction in any of the above embodiments.
[0039] In a fifth aspect, the present application further provides a computer program product, which includes a computer program, and when the computer program is executed by a processor, the steps of semantic and posture coupled image reconstruction in any of the above embodiments are implemented.
[0040] The semantic and posture coupled image reconstruction method, device, computer equipment, computer readable storage medium and computer program product obtain the depth information and semantic information of each frame of the image, and obtain two types of information that have not yet been coupled. The initial local point cloud corresponding to each frame of the image is generated according to the depth information, which can omit the image preprocessing step, avoid the problem of long time consumption of the preprocessing step, and avoid the problem of noise introduced by the preprocessing step; then based on the two-dimensional Gaussian function, the initial local point cloud is converted into an initial local Gaussian point cloud, and the camera posture corresponding to each frame of the image is adjusted according to the initial local Gaussian point cloud to obtain the initial camera posture; thanks to the consistency of the two-dimensional Gaussian under multiple perspectives, when the initial Gaussian point cloud uses the two-dimensional Gaussian function as the basic expression method, the camera posture corresponding to each frame of the image can be made more accurate. Finally, according to the semantic information, the initial camera pose is globally calibrated to obtain the calibrated camera pose; thereby eliminating the accumulated error of the initial camera pose in the process of gradual adjustment of each frame image, realizing the coupling of depth information and semantic information, and making the calibrated camera pose more accurate; finally, image reconstruction is performed according to the calibrated camera pose. In this case, the reconstructed three-dimensional image surface quality is better and the three-dimensional reconstructed surface mesh quality is higher. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the drawings required for use in the embodiments of the present application or related technical descriptions will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without paying creative work.
[0042] Figure 1 A diagram showing an application environment of a semantics and posture coupling image reconstruction method in one embodiment;
[0043] Figure 2 Schematic diagram of a flow chart of a semantic and pose coupled image reconstruction method in one embodiment;
[0044] Figure 3 A schematic diagram of a process for obtaining an initial camera pose in one embodiment;
[0045] Figure 4 Schematic diagram of the stages of obtaining two camera postures in one embodiment;
[0046] Figure 5 is a schematic diagram of an image in one embodiment;
[0047] Figure 6 is a schematic diagram of depth information in an embodiment;
[0048] Figure 7 is a structural block diagram of a semantics and posture coupled image reconstruction device in one embodiment;
[0049] Figure 8 FIG. 4 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION
[0050] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0051] The semantic and posture coupled image reconstruction method provided in the embodiment of the present application can be applied to Figure 1 In the application environment shown, the terminal 102 communicates with the server 104 via a network. The data storage system can store data that the server 104 needs to process. The data storage system can be integrated on the server 104, or placed on the cloud or other network servers.
[0052] The terminal 102 may be, but is not limited to, various personal computers, laptops, smart phones, tablet computers, IoT devices, and portable wearable devices. The IoT devices may be smart speakers, smart TVs, smart air conditioners, smart car devices, projection devices, etc. Portable wearable devices may be smart watches, smart bracelets, head-mounted devices, etc. Head-mounted devices may be virtual reality (VR) devices, augmented reality (AR) devices, smart glasses, etc. The server 104 may be an independent physical server, a server cluster or distributed system consisting of multiple physical servers, or a cloud server that provides cloud computing services.
[0053] In an exemplary embodiment, Figure 2 As shown in FIG, a semantic and pose coupled image reconstruction method is provided, and the method is applied to Figure 1 The server 104 in the example is used as an example to illustrate, including the following steps 202 to 210. Among them:
[0054] Step 202: Obtain depth information and semantic information of each frame image.
[0055] Depth information represents the distance information between the object in the image and the camera. Depth information can be a depth feature predicted based on a depth detection model, and the depth feature can be in the form of a depth map or a vector. The process of acquiring depth information omits complex preprocessing steps such as motion recovery structure (SFM), and does not have the sparse point cloud provided by algorithms such as motion recovery structure. Instead, the sparse point cloud is replaced by depth information. Exemplarily, the depth information can be a monocular depth map, and the monocular depth prediction model used to obtain the monocular depth map can use DPT or DepthAnything.
[0056] Semantic information is the meaning of an object in an image. Semantic information indicates what the object in the image is, or what a certain part of the object in the image is. For example, an image represents object A through pixel values, and the name or identifier of object A is the semantic information of object A; part A1 of object A is also in the image, and the name or identifier of part A1 is the semantic information of part A1. Semantic information can be determined by labels, or it can be recognized based on models such as recurrent neural networks, attention mechanisms, and GPT.
[0057] Step 204: Generate an initial local point cloud corresponding to each frame image according to the depth information.
[0058] The initial local point cloud includes the point cloud of each local area of the image. The multiple local areas divided into each frame of the image can be determined according to the size or number of the local areas, and the depth information of each local area is used to obtain the initial local point cloud of each local area. For example, the depth information of the i-th local area can be used as the initial local point cloud of the i-th local area, or information fusion can be performed based on the depth information of the i-th local area and the depth information of the adjacent local areas of the i-th local area, and the corresponding fusion result can be used as the initial local point cloud of the i-th local area.
[0059] Step 206 , based on a two-dimensional Gaussian function, the initial local point cloud is converted into an initial local Gaussian point cloud, and the camera pose corresponding to each frame image is adjusted according to the initial local Gaussian point cloud to obtain an initial camera pose.
[0060] The two-dimensional Gaussian function is a two-dimensional expression of the Gaussian distribution. Unlike the three-dimensional Gaussian, where the geometric meaning of each Gaussian is equivalent to an ellipsoid, the geometric shape of the two-dimensional Gaussian is approximately a disk. Unlike the three-dimensional Gaussian function, the shape of the two-dimensional Gaussian function remains consistent under perspective transformations such as rotation or projection, but the direction or angle changes, thus being able to present multi-perspective consistency. Each two-dimensional Gaussian is defined by a center point and two principal tangent vectors, and the rotation information of the two-dimensional Gaussian function in three-dimensional space can be defined by a rotation matrix.
[0061] The initial local Gaussian point cloud is a two-dimensional Gaussian point cloud initialized in each local space for each frame. The initial local Gaussian point cloud includes the calculation result of substituting the initial local point cloud into the two-dimensional Gaussian function, and the attributes carried by each point in the initial local point cloud. The attributes carried by each point can be used to obtain parameters such as weights to adjust the camera pose corresponding to each frame image, thereby obtaining the initial camera pose. For example, the initial local Gaussian point cloud includes multiple Gaussian points, each Gaussian point includes the calculation result of substituting the pixel value of the initial local point cloud into the two-dimensional Gaussian function, and each Gaussian point also contains additional attributes, including the opacity of the Gaussian point, the color of the Gaussian point, and the semantic information of the Gaussian point.
[0062] Camera pose refers to the position and posture of the camera in space, which forms a multi-view image. Camera pose can be represented by the rotation matrix and translation matrix of the camera pose. The camera pose corresponding to each frame of the image is the camera pose estimated based on the initial local Gaussian point cloud of each frame of the image. When using each frame of the image, the camera pose can be calculated and adjusted based on each frame of the image, so as to obtain the initial camera pose more accurately.
[0063] The initial camera pose is the result of the gradual adjustment of the camera pose of each frame. The initial camera pose is calculated based on the initial local Gaussian point cloud of each frame. Due to the overlapping characteristics between frames, the process of determining the initial camera pose is the process of pose recovery frame by frame. Optionally, in the process of pose recovery frame by frame, after optimizing the initial camera pose of two frames, the local initialization two-dimensional Gaussian will not be retained, that is, it will not occupy video memory.
[0064] Step 208: globally calibrate the initial camera pose according to the semantic information to obtain a calibrated camera pose.
[0065] The calibrated camera pose is obtained by coupling semantic information with the initial camera pose. Compared with the initial camera pose, semantic information can provide higher-dimensional geometric feature information and eliminate the gradually accumulated errors in each frame image, thereby obtaining the calibrated camera pose.
[0066] The position deviation of the initial camera pose can be determined according to the semantic information, so as to perform global calibration according to the position deviation to obtain the calibrated camera pose. The position deviation includes point cloud deviation or image deviation.
[0067] The change of camera pose will lead to the change of the content of the captured image. Therefore, when performing global Gaussian training, the camera pose will affect the generation of the image and the distribution of features in the image. Therefore, the semantic information can be matched with the initial global point cloud to substitute into the corresponding restoration model for initial pose restoration; wherein, the initial global point cloud includes the point cloud of the image in all areas. For example, the depth information of the image in all areas can be used as the initial global point cloud, or the depth information of the image in all areas can be filtered, deformed, etc. to obtain the initial global point cloud.
[0068] Step 210: Reconstruct a three-dimensional image according to the calibrated camera posture.
[0069] Through the calibrated camera pose, the pixel value or depth information of each frame image can be converted into a point in three-dimensional space, and then the three-dimensional scene can be reconstructed by fusing multiple depth maps. Since the calibrated camera pose is a pose coupled with semantics and pose, the reconstructed three-dimensional image can be more accurate.
[0070] Exemplarily, a Gaussian function at a corresponding viewing angle is determined for the calibrated camera pose, and the pixel values or depth values of each frame image are mapped according to the Gaussian function to obtain a three-dimensional mesh of the object surface. Optionally, when the pixel values or depth values of each frame image have corresponding attributes, the three-dimensional mesh of the object surface can also be adjusted according to attributes such as color and transparency to obtain a three-dimensional image.
[0071] In the above semantic and posture coupled image reconstruction method, the depth information and semantic information of each frame image are obtained to obtain two uncoupled information. The initial local point cloud corresponding to each frame image is generated according to the depth information, which can omit the image preprocessing step, avoid the problem of long time-consuming preprocessing steps, and avoid the problem of noise introduced in the preprocessing step; then based on the two-dimensional Gaussian function, the initial local point cloud is converted into an initial local Gaussian point cloud, and the camera posture corresponding to each frame image is adjusted according to the initial local Gaussian point cloud to obtain the initial camera posture; thanks to the consistency of the two-dimensional Gaussian under multiple perspectives, when the initial Gaussian point cloud uses the two-dimensional Gaussian function as the basic expression method, the camera posture corresponding to each frame image can be made more accurate. Finally, according to the semantic information, the initial camera posture is globally calibrated to obtain the calibrated camera posture; in this way, the accumulated error of the initial camera posture in the gradual adjustment process of each frame image is eliminated, the coupling of depth information and semantic information is realized, and the calibrated camera posture is more accurate; finally, the image is reconstructed according to the calibrated camera posture. In this case, the surface quality of the reconstructed three-dimensional image is better and the quality of the three-dimensional reconstructed surface mesh is higher.
[0072] In an exemplary embodiment, Figure 3As shown, according to the initial local Gaussian point cloud, the camera pose corresponding to each frame image is adjusted to obtain the initial camera pose, including steps 302 to 306, wherein:
[0073] Step 302 , adjusting the initial local Gaussian point cloud according to the local attributes corresponding to the initial local point cloud to obtain an adjusted local Gaussian point cloud.
[0074] The local attributes are the attributes corresponding to each point in the initial local point cloud. The local attributes include at least the color value and transparency of each point, so as to adjust the corresponding value of the initial local Gaussian point cloud through the attributes directly related to the three-dimensional image display; the local attributes may also include the rotation matrix of the two-dimensional Gaussian function, and the semantic information of the two-dimensional Gaussian function. The rotation matrix of the two-dimensional Gaussian function is used to represent the orientation of the two-dimensional Gaussian in space, so as to determine the initial local Gaussian point cloud through the attributes related to the camera posture.
[0075] The adjusted local Gaussian point cloud is obtained by adjusting the value of the initial local Gaussian point cloud based on the local attributes. Exemplarily, the depth map of the image is used as the value of each point in the initial local point cloud, and the value of each point in the initial local point cloud is substituted into the two-dimensional Gaussian function for calculation to obtain the result of the two-dimensional Gaussian function; at the same time, the weight of the two-dimensional Gaussian function result is determined according to the local attributes, and data fusion is performed based on the two-dimensional Gaussian function result and the weight of the two-dimensional Gaussian function result to obtain the adjusted local Gaussian point cloud.
[0076] Step 304 : determining adjusted local semantic information corresponding to the adjusted local Gaussian point cloud through semantic attributes and local semantic information corresponding to the initial local point cloud.
[0077] The semantic attribute corresponding to the initial local point cloud is an attribute related to the local semantic information among the local attributes. The semantic attribute at least includes the transparency of each point, so as to adjust the local semantic information corresponding to the initial local Gaussian point cloud through the attribute related to the semantic information.
[0078] Local semantic information is the semantic information corresponding to each point in the initial local point cloud. Multiple points in a region can correspond to the same local semantic information, but each point has its own local attributes. Therefore, by adjusting the semantic attributes and local semantic information, each point in the adjusted local Gaussian point cloud can have relatively accurate adjusted local semantic information. The adjusted local semantic information is obtained by adjusting the local semantic information based on the semantic attributes corresponding to the initial local point cloud.
[0079] Step 306 , adjusting the camera pose corresponding to each frame image according to the adjusted local Gaussian point cloud and the adjusted local semantic information to obtain an initial camera pose.
[0080] For example, the camera pose of each frame image may be adjusted until the adjusted local Gaussian point cloud and the adjusted local semantic information meet their respective convergence conditions to obtain the initial camera pose. The camera pose corresponding to each frame image may also be adjusted according to parameters related to the adjusted local Gaussian point cloud and the adjusted local semantic information to obtain the initial camera pose.
[0081] In this embodiment, the initial local Gaussian point cloud is adjusted by the local attributes corresponding to the initial local point cloud to obtain the adjusted local Gaussian point cloud; thus, the adjusted local Gaussian point cloud is adapted to the local attributes of the initial local point cloud to form the information of the local space in the posture dimension; at the same time, the adjusted local semantic information corresponding to the adjusted local Gaussian point cloud is determined by the semantic attributes and semantic information corresponding to the initial local point cloud, so as to form the information of the local space in the semantic dimension; further, according to the coupled adjusted local Gaussian point cloud and the adjusted local semantic information, the camera posture of each frame image is adjusted according to the posture and semantic information of the local level, so as to form the relative posture recovery between frames and obtain the initial camera posture more accurately.
[0082] In an exemplary embodiment, the local attributes include color values and transparency; the semantic attributes include transparency, and the local semantic information includes semantic features; the initial local point cloud includes the i-th point; and the initial local Gaussian point cloud includes the two-dimensional Gaussian calculation result of the i-th point.
[0083] The color value includes the color channel value of the pixel corresponding to the i-th point in the image, and the color channel value can be an RGB color channel value. Transparency is the transparency value of the pixel corresponding to the i-th point, and the transparency value can be an alpha value; transparency is both a local attribute and a semantic attribute. The semantic feature can be a semantic label in the image area, and the i-th point has a corresponding label category. For example, the image includes pixel A, and the depth information of pixel A is converted into pixel A in the initial local point cloud; the value of pixel A is substituted into the two-dimensional Gaussian function to obtain the two-dimensional Gaussian calculation result of the i-th point; at the same time, the color value and transparency of the i-th pixel are used as the local attributes of the i-th point, and the transparency of the i-th pixel is also used as the semantic attribute of the i-th point; the semantic feature of the i-th pixel is the local semantic information of the i-th point.
[0084] The initial local Gaussian point cloud is adjusted by the local attributes corresponding to the initial local point cloud to obtain the adjusted local Gaussian point cloud, including: combining the color value and transparency corresponding to the i-th point and the two-dimensional Gaussian calculation result of the i-th point to obtain the Gaussian combination result of the i-th point; determining the i-th Gaussian weight negatively correlated with the depth value order according to the depth value order of the i-th point in the initial local point cloud; adjusting the Gaussian combination result of the i-th point according to the i-th Gaussian weight to obtain the Gaussian value to be accumulated at the i-th point; accumulating the Gaussian value to be accumulated at the i-th point and the Gaussian cumulative value of the previous (i-1) points to obtain the Gaussian cumulative value of the previous i points; the Gaussian cumulative value of the previous i points is the value of the adjusted local Gaussian point cloud at the i-th point; wherein i is a positive integer greater than 2.
[0085] Correspondingly, the adjusted local semantic information corresponding to the adjusted local Gaussian point cloud is determined through the semantic attributes and local semantic information corresponding to the initial local point cloud, including: combining the semantic features, transparency and semantic features corresponding to the i-th point to obtain the semantic combination result of the i-th point; adjusting the semantic combination result of the i-th point according to the i-th Gaussian weight to obtain the semantic value to be accumulated of the i-th point; accumulating the semantic value to be accumulated of the i-th point with the semantic accumulation values of the previous (i-1) points to obtain the semantic accumulation values of the previous i points; the semantic accumulation values of the previous i points are the adjusted local semantic information corresponding to the i-th point; wherein i is a positive integer greater than 2.
[0086] The Gaussian combination result of the i-th point is the adjusted information of the two-dimensional Gaussian calculation result of the i-th point. The color value and transparency corresponding to the i-th point are used to adjust the two-dimensional Gaussian calculation result of the i-th point, in a manner similar to weighting, so that the adjusted local Gaussian point cloud can be closer to the corresponding information of the actual scene. Optionally, the color value corresponding to the i-th point, the transparency corresponding to the i-th point, and the two-dimensional Gaussian calculation result of the i-th point can be multiplied to obtain the Gaussian combination result of the i-th point. The color value corresponding to the i-th point and the transparency corresponding to the i-th point can also be substituted into the corresponding function for calculation to obtain the color and transparency mapping result of the i-th point; the color and transparency mapping result of the i-th point is multiplied by the two-dimensional Gaussian calculation result of the i-th point to obtain the Gaussian combination result of the i-th point.
[0087] Since the depth information is related to the color value and transparency, and the adjusted local Gaussian point cloud is obtained by performing multiple processing conversions based on the depth information, the accuracy can be increased by adjusting the color value and transparency.
[0088] The depth value order is the order from large to small. Since the larger the depth value of each point is, the farther it is from the camera, the more information is lost; the smaller the depth value of each point is, the closer it is to the camera, the less information is lost, so the i-th Gaussian weight negatively correlated with the depth value order is determined to ensure the accuracy of the adjustment process.
[0089] The i-th Gaussian weight is the Gaussian weight of the i-th point; the i-th Gaussian weight is used to weight the Gaussian combination result of the i-th point and the semantic value to be combined of the i-th point, respectively, to obtain the Gaussian value to be accumulated at the i-th point.
[0090] Optionally, the product of the transparency corresponding to the j-th point and the two-dimensional Gaussian calculation result of the j-th point can be calculated first to obtain the j-th product; the difference between the preset value and the j-th product can be determined to obtain the j-th difference; the j-th difference can be multiplied continuously until j is equal to (i-1), to obtain the i-th Gaussian weight; wherein j is a positive integer less than i, and the maximum value of j is (i-1).
[0091] The Gaussian value to be accumulated at the i-th point is a value determined based on the data of the i-th point itself. Since each frame of image is only data from a single perspective, the data determined based on each frame of image will have deviations. Therefore, the Gaussian value to be accumulated at the i-th point is accumulated with the Gaussian cumulative value of the previous (i-1) point to improve accuracy. The previous (i-1) point is every point before the i-th point.
[0092] The Gaussian cumulative value of the first i points is the result of sequentially accumulating the Gaussian values to be accumulated of each point in the initial local point cloud. The Gaussian cumulative value of the first i points can be obtained by accumulation or by adjusting the result of normalizing the semantic values to be accumulated of the first i points.
[0093] The semantic combination result of the i-th point is the adjusted information of the semantic feature of the i-th point. The transparency corresponding to the i-th point is used to adjust the semantic feature of the i-th point, in a manner similar to weighting, so that the adjusted local semantic information is closer to the real information of the actual scene. Optionally, the transparency corresponding to the i-th point and the semantic feature of the i-th point can be multiplied to obtain the semantic combination result of the i-th point. The transparency corresponding to the i-th point can also be substituted into the corresponding function for calculation to obtain the transparency mapping result of the i-th point; the transparency mapping result of the i-th point is multiplied by the semantic feature of the i-th point to obtain the semantic combination result of the i-th point.
[0094] The semantic value to be accumulated at the i-th point is the semantic value determined based on the data of the i-th point itself. Since each frame of image is only data from a single perspective, the semantic value determined based on each frame of image has deviations. Therefore, the semantic value to be accumulated at the i-th point is accumulated with the semantic accumulated value of the previous (i-1) point to improve accuracy. The previous (i-1) point is every point before the i-th point.
[0095] The semantic accumulation value of the first i points is the result of sequentially accumulating the semantic values to be accumulated of each point in the initial local point cloud. The semantic accumulation value of the first i points can be obtained by accumulation or by adjusting the result of normalizing the semantic values to be accumulated of the first i points.
[0096] In this embodiment, the adjusted local Gaussian point cloud and the adjusted local semantic information are both obtained by adjusting with the same weight and different attributes; since the depth information is related to the color value and transparency, and the adjusted local Gaussian point cloud is obtained by performing multiple processing and conversion based on the depth information, the adjusted local Gaussian point cloud can be made closer to the actual scene by adjusting the color value and transparency; and since the degree to which transparency affects the local semantic information is relatively large, transparency is used to adjust the local semantic information so that the adjusted local semantic information is closer to the actual scene. Determine the i-th Gaussian weight that is negatively correlated with the depth value sequence to ensure the accuracy of the adjustment process; finally, accumulate the two accumulated values of the i-th point and the (i-1) points before the i-th point respectively, so that the accuracy of relative posture recovery between frames is improved.
[0097] In an exemplary embodiment, each frame image includes the current frame and the next frame image, and the value of the adjusted local Gaussian point cloud and the adjusted local semantic information of the current frame belong to the parameters of the single frame loss. The current frame and the next frame are two consecutive frames of images. When the current frame is the kth frame, the next frame of the kth frame is the (k+1)th frame, and so on, and the result of optimizing each frame image in turn can be obtained; wherein k is a positive integer. The initial camera pose of the first frame is a differentiable Lie group SE-3 affine transformation matrix, thereby forming the initial initial camera pose.
[0098] Single-frame loss is the loss in the optimization process for each frame. Single-frame loss can be a loss function or a loss value calculated by a loss function. Single-frame loss is used to determine whether the two stages of processing the current frame and the next frame are completed; the two stages refer to the stage of adjusting the local attributes of the current frame and the stage of adjusting the camera pose difference between the current frame and the next frame. These two stages are performed frame by frame to gradually obtain the initial camera pose of the last frame; the initial camera pose of the last frame is the camera pose used for global calibration.
[0099] According to the adjusted local Gaussian point cloud and the adjusted local semantic information, the camera pose corresponding to each frame image is adjusted to obtain the initial camera pose, including: when the initial camera pose of the current frame remains unchanged, the local attributes are adjusted until the single frame loss of the current frame meets the convergence condition, and the adjusted local attributes of the current frame are obtained; when the adjusted local attributes remain unchanged, the camera pose difference between the current frame and the next frame is optimized until the single frame loss of the next frame meets the convergence condition, and the initial camera pose of the next frame is obtained according to the camera pose difference obtained by optimization.
[0100] The camera pose difference may be a camera pose residual between the current frame and the next frame, which is used to indicate the difference degree of the initial camera pose between the current frame and the next frame. The camera pose residual may include a translation residual, a rotation residual, and a homogeneous coordinate matrix residual composed of the translation residual and the rotation residual.
[0101] In the stage of adjusting the camera pose difference, the local attributes include the color value and transparency of each point, the rotation matrix of the two-dimensional Gaussian function, the rotation matrix of the two-dimensional Gaussian function, and the scaling factor of the two-dimensional Gaussian function.
[0102] The adjusted local attributes are the results of adjusting the local attributes. Each attribute in the local attributes can be adjusted as an independent variable in turn, or multiple attributes in the local attributes can be adjusted as independent variables according to certain corresponding relationships. Since the local attributes of each frame are used to adjust the initial camera pose of the next frame, the initial camera pose of the next frame is more and more accurate, and at the same time, the local attributes of the next frame are also more and more accurate, until they gradually get closer to the real scene.
[0103] The convergence condition of the single frame loss is a condition that the loss value of the single frame loss reaches a certain interval. The convergence condition can be a minimization condition, that is, when the loss value of the single frame loss is less than a certain threshold.
[0104] The difference in camera pose between the current frame and the next frame can be determined based on the adjusted local attributes, and the initial camera pose of the next frame can also be determined by matching feature points. Exemplarily, the initial camera pose and depth information of the current frame are used to transform the feature points in the current frame image from the camera coordinate system of the current frame to the world coordinate system. At the same time, the corresponding feature points in the next frame image are determined based on the feature points in the current frame image, so that the feature points in the current frame image and the corresponding feature points in the next frame image form a 3D-2D point pair; the 3D-2D point pair matching process is implemented based on the Gauss-Newton method or the g2o library.
[0105] The process of optimizing the difference between the current frame and the next frame's camera pose is the process of gradually adjusting the initial camera pose. Since the local attributes and the initial local Gaussian point cloud are used to obtain the adjusted local Gaussian point cloud, the adjusted local Gaussian point cloud can be made closer and closer to reality through the adjustment process of the local attributes; correspondingly, when the camera pose changes, the position of the camera changes, so the initial camera pose can affect the position of the pixel points of the image, thereby changing the point corresponding to the pixel point in the initial local point cloud, thereby affecting the initial local Gaussian point cloud; therefore, optimizing the initial camera pose can also make the adjusted local Gaussian point cloud closer and closer to reality.
[0106] After obtaining the camera pose difference and the initial camera pose of the current frame, the initial camera pose of the current frame may be adjusted based on the camera pose difference to obtain an adjustment result; the adjustment result is the initial camera pose of the next frame.
[0107] In this embodiment, during the training process using a single frame loss, the first stage involves the adjustment process of local attributes, and the second stage involves the adjustment process of camera pose differences. Through the adjustment process of these two stages, the adjusted local Gaussian point cloud can be made closer and closer to reality, thereby making the initial pose of the next frame more accurate than the initial pose of the current frame.
[0108] In an exemplary embodiment, the single-frame loss includes a local Gaussian loss with the value of the adjusted local Gaussian point cloud as an independent variable, an L1 regularized loss with the adjusted local semantic information corresponding to the adjusted local Gaussian point cloud as an independent variable, and an L1 regularized loss with the depth information corresponding to the adjusted local Gaussian point cloud as an independent variable; the local Gaussian loss includes an L1 regularized loss with the value of each block of the adjusted local Gaussian point cloud as an independent variable, and a pixel similarity loss of the value of each block of the adjusted local Gaussian point cloud. Adjusting the local attributes includes: adjusting the local attributes, the value of the adjusted local Gaussian point cloud, the adjusted local semantic information corresponding to the adjusted local Gaussian point cloud, and the depth information corresponding to the adjusted local Gaussian point cloud.
[0109] The L1 regularized loss is the sum of the absolute values of the parameter vector; when the adjusted local semantic information is used as the independent variable, the L1 regularized loss includes the sum of the absolute values of the adjusted local semantic information of each block; when the depth information is used as the independent variable, the L1 regularized loss includes the L1 regularized loss of the depth information of each block; when the value of the adjusted local Gaussian point cloud of each block is used as the independent variable, the L1 regularized loss includes the sum of the absolute values of the values of the adjusted local Gaussian point cloud. Among them, the L1 regularized loss with different independent variables has different hyperparameters.
[0110] The local Gaussian loss includes two types of losses, namely, the L1 regularization loss with the value of the adjusted local Gaussian point cloud of each block as the independent variable, and the pixel similarity loss; these two losses have their own hyperparameters, and the sum of the hyperparameters of these two losses is 1. The pixel similarity loss can be structural similarity (SSIM) or other similarities. Exemplarily, the value of the adjusted local Gaussian point cloud of each block and the reference value of the adjusted local Gaussian point cloud of each block are substituted into the pixel similarity loss function to form the pixel similarity loss.
[0111] In this example, during the process of adjusting local attributes, the numerical value of the adjusted local Gaussian point cloud, the adjusted local semantic information and the depth information are adjusted, so that the adjustment process of local attributes involves the optimization process of more parameters. Since the value of the adjusted local Gaussian point cloud, the adjusted local semantic information corresponding to the adjusted local Gaussian point cloud, and the depth information corresponding to the adjusted local Gaussian point cloud can directly or indirectly affect the adjusted local Gaussian point cloud, the optimized local attributes of the current frame can be closer to the real scene.
[0112] In an exemplary embodiment, the initial camera pose is globally calibrated according to the semantic information to obtain the calibrated camera pose, including: generating an initial global point cloud for each frame image according to the depth information; adjusting the initial global point cloud based on the global Gaussian model corresponding to the initial camera pose and the initial global point cloud to obtain the adjusted global point cloud; determining the adjusted global semantic information through the semantic attributes and semantic information corresponding to the initial global point cloud; optimizing the initial camera pose until the global loss corresponding to the adjusted global semantic information meets the convergence condition, thereby obtaining the calibrated camera pose.
[0113] The initial global point cloud includes the point cloud of the image in all areas. For example, the initial global point cloud can be based on the depth information of the image in all areas, or the depth information of the image in all areas can be filtered, deformed, etc. to obtain the initial global point cloud.
[0114] The global Gaussian model is a Gaussian function used to perform Gaussian processing on the initial global point cloud. When the initial camera pose remains unchanged, the initial global Gaussian model can be more accurately determined by the mean, covariance, and weight of the initial global point cloud. The initial global Gaussian model can be adjusted until the optimal global Gaussian model is obtained. When Gaussian sputtering is performed based on the optimal global Gaussian model, the three-dimensional image created by Gaussian sputtering can be made more consistent with the real scene. Exemplarily, the difference between the initial global point cloud and the reference global point cloud can be calculated, and the gradient calculation of each parameter of the global Gaussian model can be performed based on the difference, and the parameters of the global Gaussian model can be updated according to the gradient until the difference is less than the preset value, or the number of updates reaches the preset number, and the global Gaussian model is obtained.
[0115] The adjusted global point cloud is the point cloud output by the global Gaussian model after the value of the initial global point cloud is substituted into the global Gaussian model. The adjusted global point cloud can be considered as the global Gaussian point cloud.
[0116] The semantic attribute corresponding to the initial global point cloud is an attribute related to the global semantic information among the global attributes. The semantic attribute at least includes the transparency of each point, so as to adjust the global semantic information corresponding to the initial global Gaussian point cloud through the attribute related to the semantic information.
[0117] The global semantic information is the semantic information corresponding to each point in the initial global point cloud. Multiple points in a region can correspond to the same global semantic information, and each point has its own global attributes. Therefore, according to the semantic attributes and the global semantic information, adjustments can be made so that each point in the adjusted global Gaussian point cloud has relatively accurate adjusted global semantic information. The adjusted global semantic information is obtained by adjusting the global semantic information based on the semantic attributes corresponding to the initial global point cloud.
[0118] The process of optimizing based on the differences in camera poses between different frames can adjust the initial camera pose. The initial camera pose can affect the position of the pixels in the image, thereby changing the points corresponding to the pixels in the initial global point cloud, and thus affecting the initial global Gaussian point cloud; therefore, optimizing the initial camera pose can also make the adjusted global Gaussian point cloud closer to reality.
[0119] The convergence condition of the global loss is the condition that the loss value of the global loss reaches a certain interval. The convergence condition can be a minimization condition, that is, when the loss value of the global loss is less than a certain threshold. The global loss takes the adjusted global semantic information as the independent variable.
[0120] Adjusting the loss with the adjusted global semantic information as the independent variable can provide higher-dimensional geometric feature information and eliminate the error accumulated frame by frame to ensure the accuracy of the global Gaussian model when used for Gaussian sputtering processing. The accuracy means that the reconstructed three-dimensional image is closer to the real scene.
[0121] In this embodiment, the pixel points obtained by adjusting the depth information under the initial camera pose are used to determine the initial global point cloud, and some preprocessing steps are omitted, so that the image preprocessing step can be omitted, the problem of long time consumption in the preprocessing step can be avoided, and the problem of noise introduced in the preprocessing step can be avoided; at the same time, based on the global Gaussian model corresponding to the initial camera pose and the initial global point cloud, the initial global point cloud is adjusted to obtain an adjusted global point cloud, so that the global Gaussian model injected with the error of the initial camera pose adjusts the initial global point cloud, so that the adjusted global point cloud can reflect the error of the initial camera pose; further, through the semantic attributes and semantic information corresponding to the initial global point cloud, the adjusted global semantic information is more accurately determined, and the initial camera pose is optimized until the global loss corresponding to the adjusted global semantic information meets the convergence condition, and the calibrated camera pose is obtained to ensure the accuracy of the camera pose through the global loss from the perspective of semantic information.
[0122] In one embodiment, an example is given from the perspective of a specific application field. Taking images of real scenes and reconstructing them in three dimensions has a wide range of needs and applications in the fields of cultural relics protection, virtual reality, augmented reality, and game modeling. At present, neural rendering technology has made significant progress in scene reconstruction and new perspective synthesis. These technologies usually rely on accurate pre-calculated camera poses as a priori conditions. However, obtaining accurate camera poses often requires expensive high-precision sensors to obtain camera poses, or requires the use of complex pre-processing steps such as motion recovery structure (SFM), which is time-consuming and susceptible to noise. In practical applications, especially in dynamic scenes or large-scale data scenarios, this dependence limits the applicability of methods such as neural radiance field rendering (NeRF) and 3D Gaussian Splatting (3DGS). This limits the application of the new technology of neural rendering in different scenarios, making it unable to meet the current requirements for speed and quality of three-dimensional reconstruction of real scenes.
[0123] The main purpose of this embodiment is to solve the problem of achieving high-quality new perspective synthesis and accurate 3D geometric reconstruction without preprocessing the camera pose. By combining the geometric representation of explicit 2D Gaussians, this study avoids the complexity of traditional SFM algorithm preprocessing and further expands the application of 3D reconstruction from Gaussian point clouds to generating structured 3D meshes. By coupling 2D semantics with pose optimization, this embodiment aims to provide a novel and efficient 3D reconstruction paradigm that meets the current requirements for speed and accuracy of 3D reconstruction of real scenes.
[0124] 3D Gaussian-based rendering technology is an emerging explicit geometric representation method that directly describes scene geometry in the form of point clouds. It has a natural advantage of explicit representation, making it more robust under geometric structure optimization and multi-view consistency constraints. 3D Gaussian splashing technology greatly improves the quality of 3D reconstruction and new view synthesis. However, existing 3D Gaussian distribution methods still rely on the preprocessing of camera poses, so it is difficult to completely get rid of the reliance on traditional algorithms such as SFM to restore camera poses or use sensors to obtain camera poses in practical applications. At the same time, because the 3D Gaussian expression used by 3DGS is not consistent under multiple views, 3DGS cannot accurately express the surface of objects, resulting in poor quality of the 3D mesh of the object surface extracted by the 3DGS-based method. These two shortcomings limit the further application of 3DGS in the field of 3D reconstruction.
[0125] In the field of new perspective synthesis, a variety of 3D scene representation methods have been proposed in recent years, including planes (such as multi-perspective stereo reconstruction), meshes, point clouds, and methods that use neural networks. Among them, neural radiation fields and 3D Gaussian splashing technology based on point cloud representation have become mainstream technologies in this field due to their excellent lighting realistic rendering capabilities. However, most existing methods still rely on pre-calculated camera pose parameters obtained through the SFM algorithm, and this limitation affects their application.
[0126] In order to eliminate the dependence of NeRF and 3DGS on pre-processing camera poses and obtaining camera poses from high-precision sensors, some research works in recent years are gradually trying NeRF methods without SFM pre-processing. However, most methods still assume that the camera motion is small, the range of pose changes is limited, or rely on prior knowledge. These methods often perform poorly when faced with complex camera trajectories, especially large camera motions (such as 360° scenes or large data sets). Therefore, some works that do not require camera pose initialization have been born. For example, Nope-NeRF combines the dedistorted depth prior from monocular depth estimation during training, improves the performance of the network, and no longer relies on the initialization of camera poses, but its training time is very long, requiring more than ten hours. Colmap-Free 3DGS (CF-3DGS) proposes a method for jointly optimizing camera parameters and three-dimensional Gaussian (3DGS), combining local and global 3DGS strategies, aiming to overcome the limitations of existing technologies and improve the efficiency and accuracy of new perspective synthesis and three-dimensional reconstruction. CF-3DGS successfully utilizes the explicit expression of three-dimensional Gaussian technology to render high-resolution scenes without camera pose input, but this method cannot reconstruct a smooth three-dimensional surface mesh.
[0127] This embodiment proposes a new method for 3D reconstruction of surface meshes to solve the problem of high camera pose requirements of current neural rendering technology. First, this embodiment uses two-dimensional Gaussian as the basic expression method, and uses both frame-by-frame pixel information and global semantic information in the process of restoring poses. After pose restoration is completed, a high-precision 3D surface mesh can be reconstructed to achieve 3D image reconstruction.
[0128] In a specific embodiment, the existing 3D reconstruction technology based on neural rendering has the following defects: First, in the existing solutions, before using the neural radiation field NeRF or the 3D Gaussian splattering technology, the traditional SFM\MVS algorithm needs to be used to preprocess the camera pose. The preprocessing step is time-consuming and susceptible to noise, especially in scenes with weak textures, the algorithm can easily fail. In order to solve this problem, this embodiment couples pose optimization and neural rendering algorithms, eliminating the current 3D reconstruction technology's dependence on pose preprocessing.
[0129] In a specific embodiment, compared with the techniques of coupling pose optimization and neural rendering, these techniques are only for rendering new perspectives of the scene, and there is little work focused on improving the quality of the reconstructed scene surface mesh, which is also a major problem in the application of 3DGS technology. In order to solve this problem, this embodiment replaces the existing three-dimensional Gaussian-based expression method and uses a two-dimensional Gaussian with more consistency in multiple perspectives to improve the quality of the reconstruction of the three-dimensional reconstructed surface mesh.
[0130] In a specific embodiment, in the existing solutions, such as Colmap-Free 3DGS, only the local Gaussian connection between frames is used to restore the posture during the posture restoration process, which is a solution for gradually restoring the posture. However, this solution is prone to accumulating frame-by-frame errors. In order to solve this problem, this embodiment uses two-dimensional semantics to provide global semantic information and perform global correction, and also solves the problem of high video memory consumption of the current Colmap-Free 3DGS.
[0131] In an exemplary embodiment, Figure 4 As shown, this embodiment is divided into two stages:
[0132] The first stage is the step of recovering the relative pose between frames, which corresponds to the above steps 202-206; it specifically includes: first, in the initial stage of training, a pre-trained depth prediction model is used to obtain a predicted depth map, and the predicted depth is used to initialize each frame of Gaussian point cloud. Gaussian point cloud uses two-dimensional Gaussian as the basic expression method. Thanks to the consistency of two-dimensional Gaussian under multiple perspectives, the camera pose recovered between frame images is more accurate. In the process of recovering the relative pose between frames, the pixel loss of the two-dimensional image, the depth map rendering loss and the two-dimensional semantic loss are used as supervision. The pixel loss of the two-dimensional image is the above-mentioned local Gaussian loss; the depth map rendering loss is the L1 regularized loss with the depth information corresponding to the adjusted local Gaussian point cloud as the independent variable; the two-dimensional semantic loss is the L1 regularized loss with the adjusted local semantic information corresponding to the adjusted local Gaussian point cloud as the independent variable. Among them, the monocular depth map is as follows: Figure 5 As shown, the depth map corresponding to the monocular depth map is as follows Figure 6 shown.
[0133] The first stage is the step of global camera correction using semantic information, which corresponds to the above steps 208-210; it specifically includes: after the camera pose is restored in the first stage, there is an accumulated error in restoring the camera pose frame by frame. Then, semantic information is used as global camera correction to make the pose learning more accurate and improve the surface reconstruction quality.
[0134] In one embodiment, step 202, generating an initial local point cloud corresponding to each frame image according to the depth information; and step 204, converting the initial local point cloud into an initial local Gaussian point cloud based on a two-dimensional Gaussian function, is a two-dimensional Gaussian initialization process.
[0135] The process of two-dimensional Gaussian initialization includes: using a monocular depth map as a local initialization point cloud of a two-dimensional Gaussian. The input is a two-dimensional image. The current Gaussian splashing technology, whether it is a two-dimensional Gaussian or a three-dimensional Gaussian, requires an explicit point cloud as initialization before training. However, this embodiment abandons the process of restoring the camera pose using the traditional SFM algorithm, so there is no sparse point cloud provided by the SFM algorithm. Therefore, a pre-trained depth model is used to provide a monocular depth map as the local Gaussian point cloud initialization of a single frame. The pre-trained monocular depth prediction model can use DPT or DepthAnything.
[0136] Different from the three-dimensional Gaussian, each Gaussian has the same geometric meaning as an ellipsoid, and the two-dimensional Gaussian is similar to a disk. , two main tangent vectors and the corresponding scaling factor The rotation information of a two-dimensional Gaussian in three-dimensional space can be obtained by cross-producting two vectors to obtain the orthogonal tangent vector So the rotation matrix can be defined as Each Gaussian point also contains additional attributes, including opacity ,color And the 64-dimensional semantic features that will be used next To perform the 2D Gaussian splatting algorithm, any point on the 2D Gaussian disk can be defined by the following equation:
[0137]
[0138] in, is the homogeneous transformation matrix of the two-dimensional Gaussian.
[0139] In addition, for any point on the initial local point cloud , whose value can be obtained using the standard two-dimensional Gaussian function:
[0140]
[0141] With the basic expression of two-dimensional Gaussian, we can use the Gaussian splash algorithm to render the final image.
[0142] The rendering formula of the two-dimensional Gaussian function is as follows:
[0143]
[0144] in, is the value of the adjusted local Gaussian point cloud; i is the order of the point cloud, is the color value of the i-th point; is the transparency of the i-th point; is the two-dimensional Gaussian function mentioned above; is the i-th Gaussian weight, which is negatively correlated with the depth value order.
[0145] Similarly, semantic features can be rendered:
[0146]
[0147] in, is the adjusted local semantic information; i is the order of the point cloud, is the semantic feature of the i-th point; is the transparency of the i-th point; is the two-dimensional Gaussian function mentioned above; is the i-th Gaussian weight, which is negatively correlated with the depth value order.
[0148] The above step 206 also includes a step for obtaining the initial camera pose, which can be called the stage of relative pose recovery between frames. The specific goal of this stage is to preliminarily restore the camera pose as a basis for global pose recovery. The specific method of this stage is: After having the above-mentioned two-dimensional Gaussian point cloud initialized in the local space of a single frame, the overlapping characteristics between frames can be used to perform pose recovery between frames. Take two frames as an example: First, optimize the local two-dimensional Gaussian point cloud of one of the frames according to the above-mentioned rendering equation:
[0149]
[0150] in, It’s the color; is the rotation matrix of the two-dimensional Gaussian, which is used to represent the orientation of the two-dimensional Gaussian in space; is the i-th semantic feature; It is transparency; is the image feature; is to minimize the single frame loss; the single frame loss includes , and ; is the local Gaussian loss; It is the L1 regularized loss with the adjusted local semantic information corresponding to the adjusted local Gaussian point cloud as the independent variable; It is the L1 regularized loss with the depth information corresponding to the adjusted local Gaussian point cloud as the independent variable.
[0151] While optimizing pixel information, semantic information and depth information will also be optimized. , as well as The calculation formula is as follows:
[0152]
[0153]
[0154]
[0155] Among them, L1 is the L1 regularization loss, is the pixel similarity loss; yes Hyperparameters of yes Hyperparameters of yes Hyperparameters of .
[0156] After optimizing the local 2D Gaussian of a single frame, we can use this local Gaussian information to optimize the relative camera pose between this frame and the next frame. The camera pose is constructed as a differentiable Lie group SE-3 affine transformation matrix. We restore the camera pose by minimizing the image, depth, and semantic information of the next frame. Relative pose This is achieved by optimizing the following single frame losses:
[0157]
[0158] The independent variable is The matrix of the initial camera pose is formed. is the translation matrix, is the rotation matrix.
[0159] It should be noted that in this process, after optimizing the pose between two frames, the local initialized two-dimensional Gaussian will not be retained, that is, it will not occupy video memory.
[0160] Finally, step 208 is a step of using semantic information to correct the camera pose from a global perspective. The goal is to use semantic information to eliminate the accumulated inter-frame errors that are initially restored in the above steps. The idea of the method is to first use the monocular depth map and the corresponding steps to obtain a global Gaussian point cloud, that is, the adjusted global point cloud mentioned above; then, use the above restored initial camera pose to train the global Gaussian model of all images. After training 7000 iterations, the initial camera pose is restored using semantic information, semantic attributes and the corresponding global loss:
[0161]
[0162] The independent variable is The matrix of the initial camera pose is formed. is the translation matrix, is the rotation matrix; is the global loss, .
[0163] Semantic information can improve higher-dimensional geometric feature information and eliminate the accumulated inter-frame errors. After the training is completed through the above steps, a high-precision surface mesh can be extracted, thereby obtaining step 210.
[0164] Based on this, firstly, the method of this embodiment uses two-dimensional Gaussian as the basic expression method, which is more consistent with multiple perspectives, which is conducive to the optimization of camera pose and the extraction of high-precision three-dimensional surface meshes. Secondly, the method of this embodiment provides global high-dimensional semantics in the entire optimization process by using the assistance of two-dimensional semantic information. The cumulative error of restoring pose using inter-frame pixel information can be eliminated, making pose recovery more accurate and the quality of reconstruction results higher. The above two advantages enable the neural rendering pose coupling three-dimensional reconstruction technology of this embodiment to directly restore the camera pose in the three-dimensional reconstruction process without the need for complex preprocessing steps of the camera pose, and reconstruct high-precision object surface meshes and high-quality rendered images.
[0165] This embodiment has proved through a large number of experiments that the semantic and posture coupling 3D reconstruction technology based on neural rendering of this embodiment is applicable to various data types. This embodiment uses Hausdorff distance as a geometric error metric, peak signal-to-noise ratio PSNR and structural similarity SSIM as image quality evaluation indicators, which shows that the reconstruction technology of this embodiment is of better quality than the existing posture coupling optimization technology.
[0166] In addition, the semantic information used in this embodiment to provide a basis for global posture correction can be replaced by different current semantic prediction models to obtain more accurate prediction results and improve the correction quality.
[0167] It should be understood that, although the various steps in the flowcharts involved in the above-mentioned embodiments are displayed in sequence according to the indication of the arrows, these steps are not necessarily executed in sequence according to the order indicated by the arrows. Unless there is a clear explanation in this article, the execution of these steps does not have a strict order restriction, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-mentioned embodiments can include multiple steps or multiple stages, and these steps or stages are not necessarily executed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a part of the steps or stages in other steps.
[0168] Based on the same inventive concept, the embodiment of the present application also provides a semantic and posture coupled image reconstruction device for implementing the semantic and posture coupled image reconstruction method involved above. The implementation solution provided by the device to solve the problem is similar to the implementation solution recorded in the above method, so the specific limitations in one or more embodiments of the semantic and posture coupled image reconstruction device provided below can refer to the limitations of the semantic and posture coupled image reconstruction method above, and will not be repeated here.
[0169] In an exemplary embodiment, Figure 7 As shown, a semantic and posture coupled image reconstruction device is provided, comprising:
[0170] The information acquisition module 702 is used to acquire the depth information and semantic information of each frame image;
[0171] A point cloud conversion module 704 is used to generate an initial local point cloud corresponding to each frame of the image according to the depth information;
[0172] A Gaussian processing module 706 is used to convert the initial local point cloud into an initial local Gaussian point cloud based on a two-dimensional Gaussian function, and adjust the camera pose corresponding to each frame of the image according to the initial local Gaussian point cloud to obtain an initial camera pose;
[0173] A global calibration module 708, configured to perform global calibration on the initial camera pose according to the semantic information to obtain a calibrated camera pose;
[0174] The image reconstruction module 710 is used to reconstruct a three-dimensional image according to the calibrated camera posture.
[0175] In one embodiment, the Gaussian processing module 706 is used to:
[0176] Adjusting the initial local Gaussian point cloud according to the local attributes corresponding to the initial local point cloud to obtain an adjusted local Gaussian point cloud;
[0177] Determining adjusted local semantic information corresponding to the adjusted local Gaussian point cloud through semantic attributes and local semantic information corresponding to the initial local point cloud;
[0178] According to the adjusted local Gaussian point cloud and the adjusted local semantic information, the camera pose corresponding to each frame of the image is adjusted to obtain an initial camera pose.
[0179] In one embodiment, the local attribute includes color value and transparency; the semantic attribute includes transparency; the local semantic information includes semantic features; the initial local point cloud includes the i-th point; the initial local Gaussian point cloud includes a two-dimensional Gaussian calculation result of the i-th point;
[0180] The Gaussian processing module 706 is used to:
[0181] Combine the color value and transparency corresponding to the i-th point and the two-dimensional Gaussian calculation result of the i-th point to obtain the Gaussian combination result of the i-th point; determine the i-th Gaussian weight that is negatively correlated with the depth value order of the i-th point in the initial local point cloud; adjust the Gaussian combination result of the i-th point according to the i-th Gaussian weight to obtain the Gaussian value to be accumulated at the i-th point; accumulate the Gaussian value to be accumulated at the i-th point and the Gaussian cumulative value of the previous (i-1) points to obtain the Gaussian cumulative value of the previous i points; the Gaussian cumulative value of the previous i points is the value of the adjusted local Gaussian point cloud at the i-th point; wherein i is a positive integer greater than 2;
[0182] Combine the semantic feature corresponding to the i-th point, the transparency and the semantic feature of the i-th point to obtain the semantic combination result of the i-th point; adjust the semantic combination result of the i-th point according to the i-th Gaussian weight to obtain the semantic value to be accumulated of the i-th point; accumulate the semantic value to be accumulated of the i-th point with the semantic accumulation value of the previous (i-1) points to obtain the semantic accumulation value of the previous i points; the semantic accumulation value of the previous i points is the adjusted local semantic information corresponding to the i-th point; wherein i is a positive integer greater than 2.
[0183] In one embodiment, each frame of the image includes an image of a current frame and a next frame, and the value of the adjusted local Gaussian point cloud of the current frame and the adjusted local semantic information belong to parameters of single frame loss; the Gaussian processing module 706 is used to:
[0184] Under the condition that the initial camera pose of the current frame remains unchanged, the local attribute is adjusted until the single frame loss of the current frame satisfies a convergence condition, thereby obtaining the adjusted local attribute of the current frame;
[0185] When the adjusted local properties remain unchanged, the camera pose difference between the current frame and the next frame is optimized until the single frame loss of the next frame meets the convergence condition, and the initial camera pose of the next frame is obtained according to the optimized camera pose difference.
[0186] In one of the embodiments, the single-frame loss includes a local Gaussian loss with the value of the adjusted local Gaussian point cloud as an independent variable, an L1 regularized loss with the adjusted local semantic information corresponding to the adjusted local Gaussian point cloud as an independent variable, and an L1 regularized loss with the depth information corresponding to the adjusted local Gaussian point cloud as an independent variable; the local Gaussian loss includes an L1 regularized loss with the value of each block of the adjusted local Gaussian point cloud as an independent variable, and a pixel similarity loss of the value of each block of the adjusted local Gaussian point cloud;
[0187] In terms of adjusting the local attributes, the Gaussian processing module 706 is used to:
[0188] The local attribute, the value of the adjusted local Gaussian point cloud, the adjusted local semantic information corresponding to the adjusted local Gaussian point cloud, and the depth information corresponding to the adjusted local Gaussian point cloud are adjusted.
[0189] In one embodiment, the global calibration module 708 is used to:
[0190] Generating an initial global point cloud of each frame of the image according to the depth information;
[0191] Based on the initial camera pose and a global Gaussian model corresponding to the initial global point cloud, adjusting the initial global point cloud to obtain an adjusted global point cloud;
[0192] Determining adjusted global semantic information through semantic attributes and semantic information corresponding to the initial global point cloud;
[0193] The initial camera pose is optimized until the global loss corresponding to the adjusted global semantic information meets a convergence condition, thereby obtaining a calibrated camera pose.
[0194] Each module in the semantic and posture coupled image reconstruction device can be implemented in whole or in part by software, hardware, or a combination thereof. Each module can be embedded in or independent of a processor in a computer device in the form of hardware, or can be stored in a memory in a computer device in the form of software, so that the processor can call and execute operations corresponding to each module.
[0195] In an exemplary embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as shown in FIG. Figure 8As shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, referred to as I / O) and a communication interface. The processor, the memory and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a semantic and posture coupled image reconstruction method is implemented.
[0196] Those skilled in the art will understand that Figure 8 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.
[0197] In one embodiment, a computer device is further provided, including a memory and a processor, wherein a computer program is stored in the memory, and the processor implements the steps in the above method embodiments when executing the computer program.
[0198] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.
[0199] In one embodiment, a computer program product is provided, including a computer program, which implements the steps in the above method embodiments when executed by a processor.
[0200] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant regulations.
[0201] A person of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiment method can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to the memory, database or other medium used in the embodiments provided in the present application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. As an illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The database involved in each embodiment provided in this application may include at least one of a relational database and a non-relational database. Non-relational databases may include distributed databases based on blockchains, etc., but are not limited to this. The processor involved in each embodiment provided in this application may be a general-purpose processor, a central processing unit, a graphics processor, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, an artificial intelligence (AI) processor, etc., but are not limited to this.
[0202] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.
[0203] The above-described embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the present application. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the attached claims.
Claims
1. A semantic and pose coupled image reconstruction method, characterized in that: The method comprises: Obtain the depth information and semantic information of each frame image; Generate an initial local point cloud corresponding to each frame of the image according to the depth information; Based on a two-dimensional Gaussian function, the initial local point cloud is converted into an initial local Gaussian point cloud; the initial local point cloud includes the i-th point; the initial local Gaussian point cloud includes a two-dimensional Gaussian calculation result of the i-th point; The color value and transparency corresponding to the i-th point and the two-dimensional Gaussian calculation result of the i-th point are combined to obtain the Gaussian combination result of the i-th point; according to the depth value sequence of the i-th point in the initial local point cloud, the i-th Gaussian weight negatively correlated with the depth value sequence is determined; according to the i-th Gaussian weight, the Gaussian combination result of the i-th point is adjusted to obtain the Gaussian value to be accumulated at the i-th point; the Gaussian value to be accumulated at the i-th point is accumulated with the Gaussian accumulation value of the previous (i-1) points to obtain the Gaussian accumulation value of the previous i points; the Gaussian accumulation value of the previous i points is the value of the adjusted local Gaussian point cloud at the i-th point; The semantic feature corresponding to the i-th point, the transparency and the semantic feature of the i-th point are combined to obtain the semantic combination result of the i-th point; the semantic combination result of the i-th point is adjusted according to the i-th Gaussian weight to obtain the semantic value to be accumulated of the i-th point; the semantic value to be accumulated of the i-th point is accumulated with the semantic accumulation value of the previous (i-1) points to obtain the semantic accumulation value of the previous i points; the semantic accumulation value of the previous i points is the adjusted local semantic information corresponding to the i-th point; wherein i is a positive integer greater than 2; According to the adjusted local Gaussian point cloud and the adjusted local semantic information, adjusting the camera pose corresponding to each frame of the image to obtain an initial camera pose; Performing global calibration on the initial camera pose according to the semantic information to obtain a calibrated camera pose; A three-dimensional image is reconstructed according to the calibrated camera posture.
2. The method according to claim 1, characterized in that: The images of each frame include images of the current frame and the next frame, and the value of the adjusted local Gaussian point cloud of the current frame and the adjusted local semantic information belong to parameters of single frame loss; The step of adjusting the camera pose corresponding to each frame of the image according to the adjusted local Gaussian point cloud and the adjusted local semantic information to obtain an initial camera pose includes: Under the condition that the initial camera pose of the current frame remains unchanged, the local attributes are adjusted until the single frame loss of the current frame meets the convergence condition, thereby obtaining the adjusted local attributes of the current frame; When the adjusted local attributes remain unchanged, the camera pose difference between the current frame and the next frame is optimized until the single frame loss of the next frame meets the convergence condition, and the initial camera pose of the next frame is obtained according to the optimized camera pose difference.
3. The method according to claim 2, characterized in that The single-frame loss includes a local Gaussian loss with the value of the adjusted local Gaussian point cloud as an independent variable, an L1 regularized loss with the adjusted local semantic information corresponding to the adjusted local Gaussian point cloud as an independent variable, and an L1 regularized loss with the depth information corresponding to the adjusted local Gaussian point cloud as an independent variable; the local Gaussian loss includes an L1 regularized loss with the value of each block of the adjusted local Gaussian point cloud as an independent variable, and a pixel similarity loss of the value of each block of the adjusted local Gaussian point cloud; The adjusting the local attribute includes: The local attribute, the value of the adjusted local Gaussian point cloud, the adjusted local semantic information corresponding to the adjusted local Gaussian point cloud, and the depth information corresponding to the adjusted local Gaussian point cloud are adjusted.
4. The method according to claim 1, characterized in that: The globally calibrating the initial camera pose according to the semantic information to obtain a calibrated camera pose includes: Generating an initial global point cloud of each frame of the image according to the depth information; Based on the initial camera pose and a global Gaussian model corresponding to the initial global point cloud, adjusting the initial global point cloud to obtain an adjusted global point cloud; Determining adjusted global semantic information through semantic attributes and semantic information corresponding to the initial global point cloud; The initial camera pose is optimized until the global loss corresponding to the adjusted global semantic information meets a convergence condition, thereby obtaining a calibrated camera pose.
5. A semantic and pose coupled image reconstruction device, characterized in that: The device comprises: An information acquisition module is used to acquire the depth information and semantic information of each frame image; A point cloud conversion module, used to generate an initial local point cloud corresponding to each frame of the image according to the depth information; A Gaussian processing module is used to convert the initial local point cloud into an initial local Gaussian point cloud based on a two-dimensional Gaussian function; the initial local point cloud includes the i-th point; the initial local Gaussian point cloud includes the two-dimensional Gaussian calculation result of the i-th point; the color value and transparency corresponding to the i-th point and the two-dimensional Gaussian calculation result of the i-th point are combined to obtain the Gaussian combination result of the i-th point; according to the depth value sequence of the i-th point in the initial local point cloud, the i-th Gaussian weight negatively correlated with the depth value sequence is determined; according to the i-th Gaussian weight, the Gaussian combination result of the i-th point is adjusted to obtain the Gaussian value to be accumulated at the i-th point; the Gaussian value to be accumulated at the i-th point is accumulated with the Gaussian cumulative value of the previous (i-1) points to obtain the Gaussian cumulative value of the previous i points; the Gaussian cumulative value of the previous i points is the value of the adjusted local Gaussian point cloud at the i-th point; The Gaussian processing module is used to combine the semantic features and transparency corresponding to the i-th point with the semantic features of the i-th point to obtain the semantic combination result of the i-th point; adjust the semantic combination result of the i-th point according to the i-th Gaussian weight to obtain the semantic value to be accumulated of the i-th point; accumulate the semantic value to be accumulated of the i-th point with the semantic accumulation value of the first (i-1) points to obtain the semantic accumulation value of the first i points; the semantic accumulation value of the first i points is the adjusted local semantic information corresponding to the i-th point; wherein i is a positive integer greater than 2; adjust the camera pose corresponding to the image of each frame according to the adjusted local Gaussian point cloud and the adjusted local semantic information to obtain an initial camera pose; A global calibration module, used for globally calibrating the initial camera pose according to the semantic information to obtain a calibrated camera pose; An image reconstruction module is used to reconstruct a three-dimensional image according to the calibrated camera posture.
6. The device according to claim 5, characterized in that The images of each frame include images of the current frame and the next frame, and the value of the adjusted local Gaussian point cloud of the current frame and the adjusted local semantic information belong to the parameters of the single frame loss; the Gaussian processing module is used to: Under the condition that the initial camera pose of the current frame remains unchanged, the local attributes are adjusted until the single frame loss of the current frame meets the convergence condition, thereby obtaining the adjusted local attributes of the current frame; When the adjusted local attributes remain unchanged, the camera pose difference between the current frame and the next frame is optimized until the single frame loss of the next frame meets the convergence condition, and the initial camera pose of the next frame is obtained according to the optimized camera pose difference.
7. The device according to claim 6, characterized in that The single-frame loss includes a local Gaussian loss with the value of the adjusted local Gaussian point cloud as an independent variable, an L1 regularized loss with the adjusted local semantic information corresponding to the adjusted local Gaussian point cloud as an independent variable, and an L1 regularized loss with the depth information corresponding to the adjusted local Gaussian point cloud as an independent variable; the local Gaussian loss includes an L1 regularized loss with the value of each block of adjusted local Gaussian point cloud as an independent variable, and a pixel similarity loss of the value of each block of adjusted local Gaussian point cloud; the Gaussian processing module is used to: The local attribute, the value of the adjusted local Gaussian point cloud, the adjusted local semantic information corresponding to the adjusted local Gaussian point cloud, and the depth information corresponding to the adjusted local Gaussian point cloud are adjusted.
8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 4 are implemented.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 4 are implemented.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 4 are implemented.
Citation Information
Patent Citations
Coupled indoor three-dimensional semantic mapping and modeling method
CN112347550A
Indoor mobile robot three-dimensional map construction and maintenance method
CN115993121A