Three-dimensional reconstruction model training method and device, equipment and medium
By training a 3D reconstruction model that includes a vignette reconstruction network layer and a 3D reconstruction network layer, the problem of vignette modeling with fisheye cameras is solved, achieving high-fidelity 3D scene reconstruction, reducing floaters, and improving rendering effects.
Patent Information
- Application Number
- CN202511474366.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-15
- Publication Date
- 2026-02-24
AI Technical Summary
Existing technologies cannot effectively model fisheye camera vignetting, resulting in a large number of floaters in 3D scene reconstruction, which affects the rendering effect.
By training a 3D reconstruction model that includes a vignetting reconstruction network layer and a 3D reconstruction network layer, the vignetting and non-vignetting regions of fisheye camera images are processed separately. The vignetting reconstruction network layer is used to reconstruct the vignetting region and generate a vignetting rendering map and opacity, while the 3D reconstruction network layer performs 3D reconstruction on the non-vignetting region. Finally, the two layers are mixed to generate the target rendering map.
It effectively decouples camera vignetting from 3D scene reconstruction, reduces the occurrence of floaters, and improves the fidelity of the reconstructed scene.
Smart Images

Figure CN121564459A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of three-dimensional reconstruction technology, and in particular to a three-dimensional reconstruction model training method, apparatus, equipment and medium. Background Technology
[0002] Fisheye-GS is a lightweight, scalable Gaussian sputtering module for fisheye cameras. It renders 3D Gaussian images directly onto a unit sphere using fisheye projection (such as isometric projection) to alleviate the distortion and efficiency problems of traditional 3DGS under wide field of view (FoV) conditions. This method achieves 3D GS scene reconstruction of fisheye images. Based on the isometric projection model, it projects the mean (the spatial center (centroid / value) of the 3D Gaussian, used to locate the Gaussian's position in the scene) and rot (the rotation parameter of the 3D Gaussian, which determines its orientation in space; common representations include rotation matrix, quaternion, or Euler angles) of the 3D Gaussian onto the fisheye unit sphere, thus achieving the purpose of rendering fisheye images. This can solve the projection distortion problem of traditional 3DGS under extreme curvature and wide field of view (FoV) conditions of fisheye lenses.
[0003] The above-mentioned solution can only handle standard fisheye images. However, when the fisheye image is more complex, such as when the edges of the fisheye lens often contain blurred camera lens vignetting, if this part is not decoupled and reconstructed, it will result in a large number of black floaters in the 3D GS scene reconstruction, severely degrading the reconstruction quality of the fisheye image. This method only solves the projection of 3D GS primitives onto the fisheye model, but it cannot model the fisheye camera vignetting, resulting in a large number of floaters and affecting the final rendering effect. Summary of the Invention
[0004] This application provides a method, apparatus, device, and medium for training a 3D reconstruction model to solve the problem that existing methods cannot model fisheye camera vignetting, resulting in a large number of floaters and affecting the final rendering effect.
[0005] To solve the above-mentioned technical problems, the embodiments of this application are implemented as follows: In a first aspect, embodiments of this application provide a method for training a three-dimensional reconstruction model, the method comprising: Acquire fisheye camera images, camera parameters, 3D point cloud information, and vehicle pose information. The fisheye camera images are images of the same target scene captured by a fisheye camera on the vehicle from different perspectives. The camera parameters are the parameters of the fisheye camera. The camera parameters, the 3D point cloud information, and the vehicle pose information are input into the 3D reconstruction model to be trained. The 3D reconstruction model to be trained includes: a vignette reconstruction network layer and a 3D reconstruction network layer. The vignette reconstruction network layer is used to reconstruct the vignette region of the fisheye camera image based on the camera parameters, thereby obtaining a vignette rendering image and opacity. The 3D reconstruction network layer is used to perform 3D reconstruction of the non-vague region of the fisheye camera image based on the 3D point cloud information, the camera parameters and the vehicle pose information to obtain a background rendering image. The background rendering image, the vignette rendering image, and the opacity are blended to obtain the target rendering image; The loss value of the 3D reconstruction model to be trained is calculated based on the fisheye camera image and the target rendering image. The model parameters of the 3D reconstruction model to be trained are adjusted based on the loss value until the termination training condition of the 3D reconstruction model to be trained is reached, so as to obtain the final 3D reconstruction model.
[0006] Secondly, embodiments of this application provide a three-dimensional reconstruction model training device, the device comprising: The information acquisition module is used to acquire fisheye camera images, camera parameters, 3D point cloud information and vehicle pose information. The fisheye camera images are images of the same target scene captured by the fisheye camera on the vehicle from different perspectives, and the camera parameters are the parameters of the fisheye camera. The information input module is used to input the camera parameters, the 3D point cloud information and the vehicle pose information into the 3D reconstruction model to be trained. The 3D reconstruction model to be trained includes: a vignette reconstruction network layer and a 3D reconstruction network layer. The vignette reconstruction module is used to reconstruct the vignette region of the fisheye camera image based on the camera parameters using the vignette reconstruction network layer, so as to obtain a vignette rendering image and opacity. The non-vague reconstruction module is used to use the 3D reconstruction network layer to perform 3D reconstruction of the non-vague region of the fisheye camera image based on the 3D point cloud information, the camera parameters and the vehicle pose information, to obtain a background rendering image. The target rendering image acquisition module is used to blend the background rendering image, the vignette rendering image, and the opacity to obtain the target rendering image; The loss value calculation module is used to calculate the loss value of the three-dimensional reconstruction model to be trained based on the fisheye camera image and the target rendering image. The 3D reconstruction model acquisition module is used to adjust the model parameters of the 3D reconstruction model to be trained based on the loss value until the termination training condition of the 3D reconstruction model to be trained is reached, so as to obtain the final 3D reconstruction model.
[0007] Thirdly, embodiments of this application provide an electronic device, including: The system includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program, when executed by the processor, implements the aforementioned three-dimensional reconstruction model training method.
[0008] Fourthly, embodiments of this application provide a readable storage medium that, when the instructions in the storage medium are executed by the processor of an electronic device, enables the electronic device to perform the aforementioned three-dimensional reconstruction model training method.
[0009] In this embodiment of the application, by training a 3D reconstruction model that includes a vignette reconstruction network layer and a 3D reconstruction network layer, the vignettes around the fisheye camera can be decoupled from the 3D scene reconstruction. The camera vignettes can be modeled separately into a 2D space, so that the reconstruction of the 3D GS field and sky texture map is not affected by the camera vignettes, greatly reducing the occurrence of floaters and improving the fidelity of the reconstructed scene.
[0010] The above description is only an overview of the technical solution of this application. In order to better understand the technical means of this application and to implement it in accordance with the contents of the specification, and to make the above and other objects, features and advantages of this application more obvious and understandable, the following are specific embodiments of this application. Attached Figure Description
[0011] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments of this application will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0012] Figure 1 A flowchart illustrating the steps of a three-dimensional reconstruction model training method provided in this application embodiment; Figure 2 A flowchart illustrating the steps of a multi-source data preprocessing method provided in this application embodiment; Figure 3 A flowchart illustrating the steps of a method for obtaining a vignette rendering image provided in this application embodiment; Figure 4 A flowchart illustrating the steps of a method for obtaining a background rendering image provided in this application embodiment; Figure 5 A flowchart illustrating the steps of a background image acquisition method provided in this application embodiment; Figure 6 A flowchart illustrating the steps of a loss value calculation method provided in this application embodiment; Figure 7 A schematic diagram of an adaptive decoupling reconstruction process based on a complex fisheye camera provided for an embodiment of this application; Figure 8 This is a schematic diagram of the structure of a three-dimensional reconstruction model provided in an embodiment of this application; Figure 9 This is a schematic diagram of the structure of a three-dimensional reconstruction model training device provided in an embodiment of this application; Figure 10 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0013] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0014] Reference Figure 1 The diagram illustrates a flowchart of a three-dimensional reconstruction model training method provided in an embodiment of this application. Figure 1 As shown, the training method for the three-dimensional reconstruction model may include steps 101 to 107.
[0015] Step 101: Acquire fisheye camera images, camera parameters, 3D point cloud information, and vehicle pose information. The fisheye camera images are images of the same target scene captured by the fisheye camera on the vehicle from different perspectives. The camera parameters are the parameters of the fisheye camera.
[0016] In this embodiment, a fisheye camera refers to a special camera with an ultra-large field of view (usually exceeding 180°). When imaging, it is prone to producing halos with brightness attenuation and color distortion at the edges of the image. It is often placed around the vehicle to achieve panoramic visual perception.
[0017] 3D point cloud information refers to a dataset containing the 3D coordinates (x, y, z) of a large number of spatial points in a target scene, collected by point cloud acquisition devices (such as LiDAR), which can be used to characterize the 3D structure of the scene.
[0018] Vehicle pose information refers to information describing the vehicle's position (such as x, y, z coordinates) and attitude (such as roll angle, pitch angle, yaw angle) in three-dimensional space, which can provide a reference for the spatial alignment of multi-view fisheye camera images.
[0019] Fisheye camera images are images of the same target scene captured by a fisheye camera from different perspectives. Fisheye cameras can be mounted on vehicles. 3D point cloud information is point cloud information collected by point cloud acquisition devices on the vehicle.
[0020] Among them, multi-view fisheye camera images: images captured from different angles by multiple fisheye cameras, used to capture different perspectives of a scene. These images provide a wider field of view and greater coverage. Fisheye camera images often have significant distortion, thus requiring subsequent image processing.
[0021] LiDAR point cloud: Point cloud data from LiDAR (Light Detection and Ranging) provides precise depth information about a scene. Point clouds obtained through LiDAR scanning can help accurately reconstruct the position and structure of various objects in space.
[0022] Camera parameters include intrinsic and extrinsic parameters. The intrinsic parameters (such as focal length, principal point, etc.) and extrinsic parameters (the position and orientation of the camera in the vehicle coordinate system) are crucial parameters in the reconstruction process.
[0023] Vehicle pose information refers to the position and orientation of onboard equipment in space. It provides spatial location information of the scene, helping to transform image data and LiDAR point cloud data into a unified world coordinate system.
[0024] In practical implementation, multi-source data can be acquired. The acquisition process can be as follows: First, determine the data acquisition scenario (such as urban roads, parks, etc.). Install at least two fisheye cameras on the vehicle (e.g., four cameras in the front, rear, left, and right to achieve 360° coverage), ensuring that the viewpoints of each camera can cover the same target scene and that there is a certain degree of overlap. Simultaneously, deploy point cloud acquisition devices (such as LiDAR) and pose acquisition devices (such as a GPS (Global Positioning System) + IMU (Inertial Measurement Unit) integrated navigation system) on the vehicle. When the vehicle moves in the target scene, the fisheye cameras are simultaneously triggered to capture images, the point cloud acquisition devices acquire 3D point clouds, and the pose acquisition devices record the vehicle's real-time pose. Finally, the data is aggregated to obtain multi-source data including multi-view fisheye camera images, camera intrinsic parameters (such as focal length and distortion coefficients) and extrinsic parameters (such as the camera's installation position and angle on the vehicle), 3D point cloud information, and vehicle pose information.
[0025] After obtaining multi-source data (i.e., fisheye camera images, camera parameters, 3D point cloud information, and vehicle pose information), data preprocessing can be performed to obtain preprocessed multi-source data. Data preprocessing methods include distortion correction, semantic segmentation, and dynamic object detection. The data preprocessing process can be combined with... Figure 2 The following is a detailed description.
[0026] Reference Figure 2The diagram illustrates a flowchart of a multi-source data preprocessing method provided in an embodiment of this application. Figure 2 As shown, the multi-source data preprocessing method may include steps 201, 202, 203 and 204.
[0027] Step 201: Acquire multi-source data, which includes: initial fisheye camera image, camera parameters, initial 3D point cloud information, and vehicle pose information.
[0028] In this embodiment, when training the 3D reconstruction model, multi-source information can be acquired first, including: initial fisheye camera images, camera parameters, initial 3D point cloud information, and vehicle pose information. The specific acquisition process can be referred to the description in step 101 above, and will not be repeated here.
[0029] Step 202: Based on the distortion coefficients of the camera intrinsic parameters in the camera parameters, perform radial distortion correction on the initial fisheye camera image to obtain the corrected fisheye camera image.
[0030] Camera intrinsic parameters are parameters that describe the optical characteristics of a camera itself, including focal length, principal point coordinates, and distortion coefficients, and are used to establish the mapping relationship between three-dimensional spatial points and two-dimensional image pixels.
[0031] Radial distortion refers to image distortion caused by the optical lens characteristics of fisheye cameras. It manifests as pixels at the edge of the image shifting towards the center or edge, causing straight lines to appear as curves.
[0032] After obtaining multi-source data, the initial fisheye camera image can be radially distorted by using the distortion coefficients of the camera intrinsic parameters in the camera parameters to obtain a corrected fisheye camera image.
[0033] In practical applications, fisheye cameras often exhibit slight radial distortion, causing image edge stretching. This distortion needs to be corrected using distortion correction algorithms. The specific correction process can be as follows: Based on the distortion coefficients in the camera's intrinsic parameters (such as radial distortion coefficients k1, k2, and k3), a fisheye camera distortion correction model (such as a fisheye distortion model extended from a pinhole camera model) is used for correction. First, the distortion coordinates of each pixel in the original image need to be determined. Then, the ideal coordinates of that pixel in the distortion-free state are calculated based on the distortion coefficients. An interpolation algorithm (such as bilinear interpolation) is then used to map the original pixel values to the ideal coordinates, generating a corrected fisheye camera image that eliminates image distortion caused by radial distortion.
[0034] Step 203: Perform pixel-level classification on the corrected fisheye camera image to obtain the semantic label image corresponding to the corrected fisheye camera image.
[0035] Pixel-level classification refers to the process of labeling each pixel in an image with a category (such as sky, ground, vehicles, pedestrians, etc.) to generate a semantically labeled image with the same size as the original image.
[0036] A semantically labeled image is an image that contains annotations of the category to which each pixel in the image belongs. Each pixel value corresponds to a specific semantic category (such as using 0 to represent the sky and 1 to represent the ground), which is used to distinguish different scene elements in the image.
[0037] After obtaining the calibrated fisheye camera image, pixel-level classification can be performed to obtain a semantic label image corresponding to the calibrated fisheye camera image. In this example, semantic segmentation of the fisheye image aims to label each pixel in the image as a certain category (such as road, building, pedestrian, vehicle, etc.). This segmentation helps in scene understanding and provides semantic information for subsequent modeling.
[0038] Specifically, semantic segmentation can be achieved by using a semantic segmentation model (such as U-Net, DeepLab series, etc.) to perform pixel-level classification on the corrected fisheye camera image: 1. Input the corrected fisheye image, and extract multi-scale features (such as edges, textures, and contextual information) through the model's encoder. 2. The decoder maps the features back to the original image size and outputs the probability of each pixel belonging to different categories through the softmax function. 3. Take the category with the highest probability as the semantic label for that pixel, and generate a semantic label image of the same size as the corrected image, which includes category labels for scene elements such as sky, ground, and moving objects.
[0039] Of course, in practical applications, other semantic segmentation methods can also be used, such as threshold segmentation and cluster segmentation, and this embodiment does not limit them.
[0040] Step 204: Based on the initial 3D point cloud information, determine the dynamic target detection result, the aggregated dynamic point cloud data, and the static point cloud data, and use the dynamic target detection result, the dynamic point cloud data, and the static point cloud data as the 3D point cloud information.
[0041] Dynamic target detection results refer to the detection of the 3D bounding box of dynamic targets from single-frame point cloud data using point cloud target detection algorithms.
[0042] Aggregated dynamic and static point cloud data: Using the dynamic target detection results described above, the point cloud data within the dynamic target bounding boxes of each frame's original point cloud are extracted and aggregated into the same coordinate system to obtain the aggregated dynamic target point cloud. The static point cloud data after extracting the dynamic target point cloud from each frame's original point cloud is then aggregated to obtain the aggregated static target point cloud. The above aggregation steps only require the vehicle's pose information (specifically, the lidar pose corresponding to each frame's point cloud) and do not require any camera parameters.
[0043] This application employs a targeted preprocessing procedure, first using camera intrinsic parameters to correct radial distortion in fisheye images, and then obtaining semantic style data and dynamic target pose data through pixel-level classification and dynamic point cloud detection. This provides high-quality input for independent modeling of dual network layers in subsequent models, reduces floaters, and enhances scene fidelity.
[0044] Step 102: Input the camera parameters, the 3D point cloud information and the vehicle pose information into the 3D reconstruction model to be trained. The 3D reconstruction model to be trained includes: a vignette reconstruction network layer and a 3D reconstruction network layer.
[0045] The camera vignetting reconstruction network layer is a network module in the 3D reconstruction model to be trained that is specifically used to process camera vignetting data. Its core function is to model and reconstruct the vignetting generated when the fisheye camera is imaging, and to decouple the fisheye camera vignetting region.
[0046] The 3D reconstruction network layer is a network module in the model responsible for processing scene data other than camera vignetting (such as scene object structure, sky texture, etc.), and is used to reconstruct background information that is not affected by vignetting and output rendered images.
[0047] After obtaining the preprocessed multi-source data (i.e., fisheye camera images, camera parameters, 3D point cloud information, and vehicle pose information), the camera parameters, 3D point cloud information, and vehicle pose information can be input into the 3D reconstruction model to be trained. The 3D reconstruction model to be trained includes a vignette reconstruction network layer and a 3D reconstruction network layer.
[0048] Step 103: Using the vignette reconstruction network layer, the vignette region of the fisheye camera image is reconstructed according to the camera parameters to obtain a vignette rendering image and opacity.
[0049] After inputting the camera parameters of a fisheye camera into the 3D reconstruction model to be trained, a vignette reconstruction network layer can be used to reconstruct the vignette region of the fisheye camera image based on the camera parameters, obtaining a vignette rendering image and opacity. In this example, the core purpose of the vignette reconstruction network layer is to decouple and reconstruct the camera vignette region into a vignette texture image based on a decoupling reconstruction loss, then render the camera vignette image (describing the RGB values of the vignette region) through a texture rendering process, and predict the opacity of the vignette region using a camera vignette opacity network. This implementation process can be combined with... Figure 3 The following is a detailed description.
[0050] Reference Figure 3 The diagram illustrates a flowchart of a method for obtaining a vignette rendering image according to an embodiment of this application. Figure 3As shown, the method for obtaining the vignette rendering image may include steps 301 and 302.
[0051] Step 301: Using the 2D vignette texture prediction module, based on the fisheye camera texture coordinates in the camera parameters, perform 2D texture sampling on the learned fisheye camera vignette texture image to obtain the vignette rendering image.
[0052] In this embodiment, the vignette reconstruction network layer may include: a two-dimensional vignette texture prediction module and a vignette opacity prediction module.
[0053] Among them, the two-dimensional vignette texture prediction module is the core module in the vignette reconstruction network layer responsible for generating vignette color information. It samples vignette data based on fisheye camera texture coordinates and optimizes the sampling features through unit spherical projection, and finally outputs an image containing the vignette RGB color distribution.
[0054] The vignette opacity prediction module is a dedicated module in the network layer for calculating the transparency of vignettes. It combines the positional encoding information of the fisheye camera texture coordinates with the camera embedding vector to output an opacity map that represents the transparency / occlusion of the vignette region, which is used to optimize the superposition effect of vignettes and background.
[0055] Fisheye camera texture coordinates can be used to describe the normalized pixel coordinates of the position of fisheye camera image pixels on a two-dimensional texture plane (usually with the upper left corner of the image as the origin, the horizontal axis as the u axis and the vertical axis as the v axis). They can accurately correspond to the spatial position of each pixel in the fisheye image and provide a positioning reference for vignette data sampling.
[0056] The vignette rendering is the output of the 2D vignette texture prediction module. It represents the color distribution of the fisheye camera vignette in RGB three-channel form (such as red attenuation in the edge area and blue shift in the corner). The pixel value corresponds to the color intensity of the vignette at that location.
[0057] After inputting the camera parameters into the 3D reconstruction model to be trained, the 2D vignette texture prediction module can be invoked to perform 2D texture sampling on the learned fisheye camera vignette texture image based on the fisheye camera texture coordinates in the fisheye camera parameters, thereby obtaining a vignette rendering image. The core objective of this step is to generate a vignette brightness distribution (vignette rendering image) adapted to the fisheye viewpoint by sampling from the learned vignette texture image using the fisheye camera texture coordinates.
[0058] Step 302: Use the vignette opacity prediction module to process the position encoding information of the texture coordinates of the fisheye camera and the camera embedding vector of the fisheye camera to obtain the opacity.
[0059] Simultaneously, the vignette opacity prediction module can be used to process the positional encoding information of the fisheye camera texture coordinates and the camera embedding vector of the fisheye camera to obtain the opacity. The core objective of this step is to predict the probability (opacity) of each pixel belonging to the vignette region by fusing the spatial location features of the fisheye texture coordinates with camera-specific features. The specific implementation process may include: 1. Location encoding of texture coordinates in fisheye cameras. This involves converting discrete 2D texture coordinates into continuous feature vectors, enhancing the model's ability to perceive spatial location (especially edge regions).
[0060] High-dimensional feature vectors can be generated from standardized coordinates using a preset encoding method (such as sine or cosine encoding). A positional encoding vector is then output, preserving the spatial distance and orientation information of the coordinates.
[0061] 2. Construction of fisheye camera embedding vectors. This involves encoding camera parameters (intrinsic parameters, optical characteristics, etc.) into fixed-dimensional vectors to enable the model to adapt to the differences in vignetting characteristics of different fisheye cameras (such as the vignetting range and attenuation rate of different lenses).
[0062] The camera parameters are mapped to low-dimensional vectors through a fully connected network, and the resulting vectors are obtained by an activation function.
[0063] 3. Feature fusion and opacity prediction. This involves concatenating the location encoding vector (d-dimensional) with the camera embedding vector (32-dimensional) to form a fused feature vector of d+32 dimensions.
[0064] The fused features can be processed using a lightweight MLP (Multilayer Perceptron) and then processed by a sigmoid activation function to map the output value to the range of [0, 1]. The closer the value is to 1, the higher the probability that the pixel belongs to the vignetting region.
[0065] 4. Obtain the opacity.
[0066] The opacity values of each pixel are arranged according to the resolution of the original fisheye image to form an opacity of the same size as the vignette rendering image, which is used for subsequent blending with the background rendering image.
[0067] Step 104: Using the 3D reconstruction network layer, based on the 3D point cloud information, the camera parameters, and the vehicle pose information, perform 3D reconstruction on the non-vague region of the fisheye camera image to obtain a background rendering image.
[0068] After inputting 3D point cloud information, camera parameters, and vehicle pose information into the 3D reconstruction model to be reconstructed, a 3D reconstruction network layer can be used to perform 3D reconstruction of the non-hallucinated regions of the fisheye camera image based on the 3D point cloud information, camera parameters, and vehicle pose information, obtaining a background rendering image. That is, based on the 3D point cloud, camera parameters, and vehicle pose, the scene structure of the non-hallucinated region is reconstructed, and a background rendering image (including visual features of the sky, static background, and dynamic objects) is output. This implementation process can be combined with... Figure 4 The following is a detailed description.
[0069] Reference Figure 4 The diagram illustrates a flowchart of the steps involved in obtaining a background rendering image according to an embodiment of this application. Figure 4 As shown, the method for obtaining the background rendering image may include steps 401, 402, and 403.
[0070] Step 401: Using the three-dimensional reconstruction network layer, a three-dimensional sky cube map is constructed based on the six-face texture map. The isometric projection model is determined using the camera parameters of the fisheye camera. The input rays of the sky cube map are projected onto a unit sphere through the model. The projection data on the unit sphere is sampled in two dimensions to obtain a sky texture map adapted to the fisheye view.
[0071] In this embodiment, the main function of the equidistant projection module is to map the sky node data in three-dimensional space onto a unit sphere through the equidistant projection algorithm, and then convert it into a two-dimensional feature map through two-dimensional image sampling, so as to achieve accurate conversion of the sky region from three-dimensional to two-dimensional.
[0072] Sky node data is used to reconstruct the sky region in an image, and can be modeled using texture map sampling.
[0073] The unit sphere is only used to achieve the fisheye image effect. It maps planar coordinates to spherical coordinates before sampling to realize the fisheye effect of the image.
[0074] A 3D reconstruction network layer can be used to construct a 3D sky cube map based on a six-faceted texture map. An isometric projection model is determined using the camera parameters of a fisheye camera. The input rays of the sky cube map are projected onto a unit sphere through this model, and the projection data on the unit sphere is sampled in 2D to obtain a sky texture map adapted to the fisheye view. In other words, the 3D sky cube map is mapped onto a unit sphere using an isometric projection model from a fisheye camera, and a sky texture adapted to the fisheye distortion viewpoint is generated through sampling. The specific implementation process can be as follows: 1. Construction of a 3D sky cube map.
[0075] Input texture maps of 6 faces (front, back, left, right, top, and bottom, covering a 360° panoramic sky scene) and stitch them together to form a 3D sky cube (a cube centered on the camera, with each face corresponding to a sky texture in one direction).
[0076] Cube parameter standardization: the side length is set to unit length, and the center coincides with the optical center of the virtual camera to ensure that the texture coordinates correspond one-to-one with the spatial direction.
[0077] 2. Determination of the equidistant projection model.
[0078] Analyze the parameters of the fisheye camera (intrinsic parameters: focal length f, principal point (c_x, c_y), distortion coefficients k1, k2; imaging model: equidistant projection θ=r, where θ is the field of view angle and r is the normalized radial distance).
[0079] The projection function is constructed based on parameters: the input is the direction of the three-dimensional ray, and the output is the pixel coordinates of the corresponding fisheye image, thus clarifying the mapping relationship between "spatial direction → distorted image position".
[0080] 3. Input ray generation and unit spherical projection.
[0081] Generate an input ray covering all pixels of the fisheye image: starting from the camera optical center, each pixel (u,v) corresponds to a ray, and the direction vector is inferred from the pixel coordinates through the camera intrinsic parameters (the ideal direction before distortion elimination).
[0082] Ray projection onto the unit sphere: Normalize the direction vector of each ray to obtain the point (x,y,z) on the unit sphere (spherical coordinates (θ,φ), where θ is the angle between the ray and the optical axis, and φ is the azimuth angle), thus realizing the conversion from "fisheye pixel to spherical coordinates".
[0083] 4. Unit spherical sampling and sky texture generation.
[0084] Spherical coordinates mapped to sky cube: convert the (x,y,z) coordinates of a unit sphere into the sampled coordinates of a sky cube (determining which face of the cube is hit and the specific texture location).
[0085] Two-dimensional image sampling: Perform bilinear interpolation sampling on the texture map of the corresponding face of the cube to obtain the sky color (RGB value) in that direction.
[0086] Output: Arrange the sampling results of all pixels according to the fisheye image resolution to obtain a sky texture map adapted to the fisheye distortion viewpoint.
[0087] Step 402: Using the 3D reconstruction network layer, the static background node data and dynamic target node data are processed according to the 3D point cloud information and the vehicle pose information to obtain a background image.
[0088] Static background node data is a subset of 3D point clouds representing fixed background elements in a scene, such as point cloud data corresponding to long-term static objects like roads, buildings, and trees. It has the characteristic of stable spatial position.
[0089] Dynamic target node data refers to a subset of three-dimensional point clouds in the target detection point cloud information that represents the elements of the core dynamic targets in the scene, such as moving cars and pedestrians.
[0090] In practical applications, static background nodes represent static objects in a scene, such as buildings, roads, and trees. These objects are represented using 3D GS (Gaussian point cloud) models during training, and their three-dimensional shapes and structures are accurately constructed by fusing them with LiDAR point cloud data.
[0091] Dynamic target nodes include moving objects, such as moving cars and pedestrians. Unlike static backgrounds, modeling dynamic targets requires greater real-time performance and flexibility. These targets are identified and labeled using dynamic target detection methods, and then appropriate models are used for 3D reconstruction.
[0092] The background image refers to the image output by the 3D Gaussian sputtering module after processing static and dynamic target node data.
[0093] Simultaneously, a 3D reconstruction network layer can be used to process static background node data and dynamic target node data based on 3D point cloud information and vehicle pose information to obtain a background image. That is, the geometric features of the static and dynamic scenes are initialized using 3D point clouds, and combined with a unified coordinate system of vehicle pose to render and generate a background image containing both static background and dynamic targets. The specific implementation process may include: 1. Input data preprocessing.
[0094] 3D point cloud classification: The original point cloud is divided into static point clouds (roads, buildings, etc.) and dynamic target point clouds (pedestrians, vehicles, etc.), and the category and initial pose of the dynamic targets are labeled.
[0095] Vehicle pose analysis: Extract the vehicle's world coordinate system pose (rotation matrix R, translation vector T) for transforming dynamic target coordinates.
[0096] 2. Static background node processing.
[0097] Gaussian element initialization: Initialize the 3D Gaussian sphere (GS) field with a static point cloud in the world coordinate system. Each point cloud point corresponds to a Gaussian element (containing position (x,y,z), covariance matrix σ, and color (r,g,b)).
[0098] Geometric optimization: By learning through the network, the position, covariance (direction and magnitude of Gaussians), and color (fitting scene texture) of primitives are adjusted to enhance the continuous representation of static scenes.
[0099] 3. Dynamic target node processing.
[0100] Target center point cloud initialization: For each dynamic target, initialize the Gaussian elements with the point cloud in its local coordinate system (with the target center as the origin).
[0101] Coordinate system transformation: Combining the vehicle pose and the target's own pose (position (dx,dy,dz) and attitude (rx,ry,rz) relative to the vehicle), the Gaussian elements of the dynamic target are transformed from the local coordinate system to the world coordinate system.
[0102] 4. Dynamic and static rendering integration.
[0103] Unified projection: Based on the isometric projection model of a fisheye camera, static and dynamic Gaussian primitives in the world coordinate system are projected onto the fisheye image plane (calculating the color contribution of each primitive to the pixel).
[0104] Occlusion handling: Primitives are sorted by their depth value (distance from the camera), and nearby primitives occlude distant primitives to ensure spatial consistency in rendering.
[0105] Output: Generates a background image (scene texture of non-sky areas) containing static background and dynamic objects.
[0106] Next, combined Figure 5 The detailed processing procedures for dynamic target nodes and static background nodes are described in detail.
[0107] Reference Figure 5 The diagram illustrates a flowchart of a background image acquisition method provided in an embodiment of this application. Figure 5 As shown, the method for obtaining the background image may include steps 501 and 502.
[0108] Step 501: Using the 3D reconstruction network layer, the static background node data is initialized in the initial world coordinate system based on the 3D point cloud information to obtain static background node parameters; and the dynamic target node data is initialized with the dynamic target as the center to obtain dynamic target node parameters.
[0109] In this embodiment, a 3D reconstruction network layer can be used to initialize the static background node data in the initial world coordinate system based on the 3D point cloud information, obtaining the static background node parameters; and the dynamic target node data can be initialized with the dynamic target as the center, obtaining the dynamic target node parameters. Specifically, static point clouds (such as non-moving target point clouds like roads, buildings, and guardrails) can be extracted from the 3D point cloud, noise points can be filtered (isolated points can be removed through statistical filtering or radius filtering), and effective geometric information can be retained. A 3D Gaussian sphere (GS) field is used to model the static background nodes, assigning a Gaussian sphere element to each static point cloud point. The initial parameters include: position: directly reusing the world coordinate system coordinates (x_w, y_w, z_w) of the static point cloud; covariance (composed of a rotation matrix and a scaling matrix, describing the shape of the Gaussian ellipse, i.e., ...). ), and spherical harmonic coefficients. Then, the static Gaussian nodes are initialized: containing the parameter set (position, covariance, color) of all static Gaussian primitives, and all primitives are in the world coordinate system.
[0110] For dynamic target node data: individual dynamic target point clouds (such as pedestrians and vehicles) can be segmented from 3D point clouds. The bounding box and center coordinates of each dynamic target are determined using target detection algorithms (such as PointPillars and YOLOv8-3D). A local coordinate system is established with the center coordinates of the dynamic target as the origin (x-axis along the target's direction of travel, y-axis horizontal and perpendicular to the direction of travel, z-axis vertically upward). Gaussian primitives are assigned to the point cloud of each dynamic target, with initial parameters set based on the local coordinate system: position, covariance, color, etc. The Gaussian primitives for each dynamic target node are initialized: a set of Gaussian primitive parameters (position, covariance, color) for each dynamic target, with all primitives located in the target's local coordinate system.
[0111] Step 502: During frame-by-frame rendering, the Gaussian parameters of the dynamic target node and the Gaussian parameters of the static background node are transformed to the world coordinate system through the dynamic target pose in that frame. Then, the dynamic target Gaussian points and the static background Gaussian points in the world coordinate system are projected and rendered using the camera extrinsic parameters and vehicle pose information to obtain the rendered background image.
[0112] After obtaining the dynamic and static Gaussian parameters through initialization in step 501, frame-by-frame rendering can be performed. During rendering, the Gaussian parameters of the dynamic target nodes and the Gaussian parameters of the static background nodes can be transformed to the world coordinate system using the dynamic target pose in that frame. Then, the dynamic and static Gaussian points in the world coordinate system are projected and rendered using camera extrinsic parameters and vehicle pose information to obtain the rendered background image.
[0113] The embodiments of this application realize accurate geometric modeling and viewpoint adaptation of static background and dynamic target, and finally generate a background image with complete structure and fit to fisheye view, which effectively ensures the scene realism and coordinate consistency of the subsequent background rendering image.
[0114] Step 403: Combine the background image and the sky texture image to obtain the background rendering image.
[0115] After obtaining the background image and sky texture image, they can be combined to obtain the background rendering image. Specifically, the sky and dynamic / static backgrounds can be merged based on the background opacity information to generate a complete scene rendering image without vignetting.
[0116] This application's embodiments address the stretching distortion problem in traditional sky rendering at wide-angle views by adapting to fisheye distortion using cube-spherical projection, thus improving the accuracy of sky region reconstruction. Based on point cloud initialization and coordinate system one, accurate geometric modeling of static scenes and dynamic targets is achieved, ensuring the consistency of scene structure under different viewpoints. Background rendering image combination: Through depth-aware fusion, a non-halo region rendering result containing complete scene elements (sky + dynamic and static background) is generated, providing a high-quality foundation for subsequent blending with the halo region.
[0117] Step 105: Blend the background rendering image, the vignette rendering image, and the opacity to obtain the target rendering image.
[0118] After obtaining the background rendering image, vignette rendering image, and opacity, these elements can be blended to obtain the target rendering image. Specifically, the effect of overlapping vignette and non-vignette areas in a realistic fisheye image can be simulated to generate the final target rendering image.
[0119] Step 106: Calculate the loss value of the 3D reconstruction model to be trained based on the fisheye camera image and the target rendering image.
[0120] After obtaining the fisheye camera image and the target rendering image, the loss value of the 3D reconstruction model to be trained can be calculated based on these images. The calculation process for the loss value can be combined with... Figure 6 The following is a detailed description.
[0121] Reference Figure 6 The diagram illustrates a flowchart of a loss value calculation method provided in an embodiment of this application. Figure 6 As shown, the loss value calculation method may include steps 601, 602 and 603.
[0122] Step 601: Calculate the mean absolute error loss and structural similarity loss based on the normalized RGB parameters of the target rendering image and the normalized RGB parameters of the fisheye camera image, and then perform a weighted summation of the mean absolute error loss and structural similarity loss to obtain the reprojection error loss.
[0123] In this embodiment, the normalized RGB parameter of the target rendering image refers to the normalized value of the RGB value of each pixel in the target rendering image output by the model (range [0,1], obtained by dividing the original RGB value (0-255) by 255).
[0124] The normalized RGB parameter of a fisheye camera image refers to the normalized RGB value of each pixel in a real fisheye camera image (range [0,1], used as a supervision benchmark).
[0125] Mean absolute error loss (L1) refers to the loss that measures the pixel-level brightness difference between the target rendered image and the real fisheye image. It is calculated as the average of the absolute difference of the normalized RGB values for each pixel.
[0126] Structural similarity loss (SSIM loss) is a loss that measures the overall similarity of structure, brightness, and contrast between two images. It is obtained by subtracting the structural similarity index (SSIM, range [0,1]) from 1. The smaller the value, the more similar the structures of the two images are.
[0127] After obtaining the target rendered image, the mean absolute error loss and structural similarity loss can be calculated based on the normalized RGB parameters of the target rendered image and the normalized RGB parameters of the fisheye camera image. The mean absolute error loss and structural similarity loss are then weighted and summed to obtain the reprojection error loss. In other words, L1 loss captures pixel brightness differences, SSIM loss captures image structural similarity, and the weighted fusion comprehensively evaluates the overall consistency between the rendered image and the real image.
[0128] Step 602: Calculate the decoupling loss of the non-vague scene area based on the mask parameters and vignette opacity of the non-vague scene area in the target rendering image.
[0129] Non-vagina scene region mask parameters: binary mask (value is 0 or 1), where 1 represents pixels in the real fisheye image that belong to the non-vagina region, and 0 represents pixels in the vagina region (used to mark non-vagina regions that need to be optimized).
[0130] Simultaneously, the decoupling loss for non-vague areas can be calculated based on the mask parameters of the non-vague scene areas and the vignette opacity in the target rendered image. That is, the vignette opacity is monitored through the non-vague mask to ensure that the opacity of the non-vague areas is as close to 0 as possible (to avoid interference from the vignette on the non-vague areas). In this example, to better decouple the fisheye vignette and reduce the floaters of the non-vague areas, a non-vague area decoupling loss L_rgb_pure based on vignette opacity is designed, with the following form: L_rgb_pure=mean(abs(r_s*(1-vignette_opacity)-rgb_gt* (1-vignette_opacity))).
[0131] Here, vignette_opacity is the predicted vignette opacity, and (1-vignette_opacity) can be understood as the pure scene area that does not contain vignettes.
[0132] Step 603: Calculate the loss value of the 3D reconstruction model to be trained based on the reprojection error loss and the non-halo region decoupling loss.
[0133] After obtaining the reprojection error loss and the decoupling loss of the non-halo region, the loss value of the 3D reconstruction model to be trained can be calculated based on the reprojection error loss and the decoupling loss of the non-halo region. Specifically, the reprojection error loss and the decoupling loss of the non-halo region can be directly added together to obtain the final loss value. Alternatively, a weighted sum of the reprojection error loss and the decoupling loss of the non-halo region can be performed to obtain the final loss value.
[0134] This application embodiment ensures that the target rendered image and the real fisheye image are highly consistent in pixel brightness and structure through reprojection error loss, and forces the vignetting effect of non-vignetting regions to be minimized through non-vignetting region decoupling loss. The synergistic optimization of the two enables the model to accurately restore the overall appearance of the scene and clearly separate the vignetting and non-vignetting regions, ultimately improving the visual realism of 3D reconstruction and the accuracy of region modeling.
[0135] Step 107: Adjust the model parameters of the 3D reconstruction model to be trained based on the loss value until the termination training condition of the 3D reconstruction model to be trained is reached, and the final 3D reconstruction model is obtained.
[0136] After calculating the loss value, the model parameters of the 3D reconstruction model to be trained can be adjusted based on the loss value until the termination condition of the 3D reconstruction model is reached, resulting in the final 3D reconstruction model. Specifically, based on the calculated total loss value, the weights, biases, and other parameters of the vignette reconstruction network layer and the 3D reconstruction network layer in the 3D reconstruction model to be trained can be updated using a backpropagation algorithm (such as the Adam optimizer). The process of data input, dual-network layer reconstruction, loss calculation, and parameter adjustment is repeated. After each round of training, the reconstruction accuracy of the model on the validation set (such as scene structure similarity and vignette restoration error) is calculated. When the validation set accuracy does not improve for N consecutive rounds (such as N=10), or the number of training rounds reaches a preset maximum value, or the loss value is lower than a preset threshold, the termination condition is determined, and the final 3D reconstruction model is output.
[0137] This application embodiment trains a 3D reconstruction model that includes a vignette reconstruction network layer and a 3D reconstruction network layer. This decouples the vignettes around the fisheye camera from the 3D scene reconstruction. The camera vignettes can be modeled separately in a 2D space, so that the reconstruction of the 3D GS field and sky texture map is not affected by the camera vignettes. This greatly reduces the occurrence of floaters and improves the fidelity of the reconstructed scene.
[0138] Next, combined Figure 7 and Figure 8 The adaptive decoupling and reconstruction process is described in detail. For example... Figure 7 As shown, the process may include: 1. Input: Multi-view fisheye camera images; LiDAR point clouds; camera intrinsic and extrinsic parameters; vehicle pose, etc.
[0139] 2. Data preprocessing: Radial distortion removal using fisheye camera; 2D semantic segmentation based on image prediction; aggregation of lidar point clouds and dynamic object detection.
[0140] 3. Creating a Reconstruction Model: The scene is divided into sky nodes, camera vignette nodes, static background nodes, and dynamic target nodes. This means the entire scene is divided into different types of nodes, each representing a part of the scene. Sky Nodes: The sky is usually unaffected by external objects and can be modeled using a sky texture map model. This model can capture the overall effect of the sky area well, making the rendered sky more natural and realistic. Camera Vignette Nodes: Fisheye lenses often produce a vignette effect in images, where the center of the image is brighter than the edges. For this part, a vignette blending model can be used to handle the lighting changes at the image edges by modeling camera distortion. Static Background Nodes: Static background nodes represent static objects in the scene, such as buildings, roads, and trees. These objects are represented using a 3D GS (Gaussian point cloud) model during training, and their 3D shape and structure are accurately constructed by fusing with LiDAR point cloud data. Dynamic Target Nodes: Dynamic targets include moving objects, such as moving cars and pedestrians. Unlike static backgrounds, modeling dynamic targets requires higher real-time performance and flexibility. These targets are identified and calibrated using dynamic target detection methods, and then 3D reconstruction is performed using appropriate models.
[0141] 4. Scene Rendering: Sky nodes use a sky texture map model, camera vignette nodes use a vignette blending model, and static background nodes and dynamic target nodes are rendered using 3D GS models. Sky Node Rendering: Sky nodes are rendered using a sky texture map. Textures are applied to a unit spherical model using mapping technology to generate a realistic fisheye model sky effect. Fisheye Camera Vignette Rendering: A vignette prediction network is designed for fisheye camera vignette rendering. It outputs the RGB values and opacity of the camera vignette and then blends them with the background rendering image. Static Background and Dynamic Target Rendering: Static background and dynamic target nodes are reconstructed using 3D GS models. The static background uses LiDAR point clouds to build a 3D model in the world coordinate system, and is rendered based on the vehicle's position in the world coordinate system. The 3D model of the dynamic target is built at the target's center position, and combined with the dynamic target detection information to transform it into the world coordinate system for combined rendering.
[0142] 5. Scene Optimization: Based on semantic segmentation, different loss functions are applied to child nodes and the global scene to optimize the entire scene. Scene optimization aims to improve reconstruction accuracy and overall scene consistency. Child node optimization: Different types of nodes in the scene (such as sky, static background, dynamic objects, etc.) have different processing requirements. Different loss functions can be applied to these child nodes to ensure that nodes do not confuse with each other. Global scene optimization: The entire scene is optimized using a global loss function to ensure smooth transitions between different nodes and reduce unnecessary errors. The goal of global optimization is to achieve high realism in both visual and spatial aspects of the reconstructed scene.
[0143] The structure of the 3D reconstruction model consists of two parts: the first part is the reconstruction of the non-camera vignetting area, and the second part is the reconstruction of the camera vignetting.
[0144] 1. Non-camera vignette reconstruction: This part reconstructs the sky, background, and dynamic objects in 3D space. The final scene RGB (r_s) is obtained through combined rendering. Each part is described below: Sky Node: A 3D sky cube map composed of 6 texture maps. In order to be reconstructed under a fisheye camera, its input rays are projected onto a unit sphere using an equidistant projection model before 2D image sampling. This ensures that the sky cube map can be learned to a standard space (non-fisheye distortion). During sampling, it is then projected onto the distorted fisheye model, which greatly improves the accuracy of sky reconstruction.
[0145] Background and dynamic target nodes: Both of these are modeled using a 3D GS field. The Gaussian distribution of the background nodes is initialized using a lidar point cloud in the initial world coordinate system, while the dynamic target nodes are initialized using a point cloud centered on the target, and then transformed to the world coordinate system using the target's pose. After being unified to the same world coordinate system, the positions and covariances of the 3D GS primitives are projected onto a fisheye model for rendering according to an isometric projection model, ultimately resulting in the rendered image of the fisheye model. The output of this part is combined with the output of the sky nodes to obtain the final scene RGB (r_s).
[0146] 2. Camera Vignette Reconstruction: This part mainly aims to reconstruct the camera vignette with varying opacity. The network primarily consists of two components: 2D vignette texture prediction and vignette opacity prediction. 2D Vignette Texture Map: Camera vignetting is a shadow at the edge of a camera image with gradient opacity. It is an image with high-frequency fixed elements in 2D space. Therefore, modeling camera vignetting in 2D space using 2D texture maps is beneficial for fast and accurate vignetting reconstruction. This part consists of 2D texture maps for n cameras to better adapt to different camera models. The model samples the UV coordinates of the fisheye camera onto a texture map at a fixed resolution to obtain the final RGB image (r_v) of the camera vignetting under the fisheye model.
[0147] Vignette Opacity Network: Since camera vignetting is semi-transparent, a vignette opacity prediction is also needed to blend it with the scene's RGB image. The prediction of vignette opacity is implemented using an MLP network. The model's input consists of the positional encoding of the uv coordinates (uv_emb) and the embedding vectors from different cameras (c_emb). The output is an opacity map within the range [0,1], which is then used for feature map blending.
[0148] This application's embodiments adaptively decouple camera vignetting from scene reconstruction in complex fisheye cameras. This ensures the reconstruction of complex fisheye camera vignetting while reducing floaters in the scene. Simultaneously, an adaptive decoupling reconstruction algorithm is developed, utilizing 2D texture maps and opacity maps for both vignetting prediction and decoupling, and employing a decoupling loss designed to further decouple vignetting.
[0149] Reference Figure 9 The diagram shows a structural schematic of a three-dimensional reconstruction model training device provided in an embodiment of this application. Figure 9 As shown, the 3D reconstruction model training device 900 may include the following modules: The information acquisition module 910 is used to acquire fisheye camera images, camera parameters, 3D point cloud information and vehicle pose information. The fisheye camera images are images of the same target scene captured by the fisheye camera on the vehicle from different perspectives. The camera parameters are the parameters of the fisheye camera. The information input module 920 is used to input the camera parameters, the three-dimensional point cloud information and the vehicle pose information into the three-dimensional reconstruction model to be trained. The three-dimensional reconstruction model to be trained includes: a vignette reconstruction network layer and a three-dimensional reconstruction network layer. The vignette reconstruction module 930 is used to reconstruct the vignette region of the fisheye camera image based on the camera parameters using the vignette reconstruction network layer, so as to obtain a vignette rendering image and opacity. The non-vague reconstruction module 940 is used to use the three-dimensional reconstruction network layer to perform three-dimensional reconstruction of the non-vague region of the fisheye camera image based on the three-dimensional point cloud information, the camera parameters and the vehicle pose information, to obtain a background rendering image. The target rendering image acquisition module 950 is used to blend the background rendering image, the vignette rendering image, and the opacity to obtain the target rendering image; The loss value calculation module 960 is used to calculate the loss value of the three-dimensional reconstruction model to be trained based on the fisheye camera image and the target rendering image. The 3D reconstruction model acquisition module 970 is used to adjust the model parameters of the 3D reconstruction model to be trained based on the loss value until the termination training condition of the 3D reconstruction model to be trained is reached, so as to obtain the final 3D reconstruction model.
[0150] Optionally, the information acquisition module includes: The data acquisition unit is used to acquire multi-source data, which includes: initial fisheye camera image, camera parameters, initial 3D point cloud information and vehicle pose information; The image acquisition unit is configured to perform radial distortion correction on the initial fisheye camera image based on the distortion coefficients of the camera intrinsic parameters in the camera parameters, so as to obtain the corrected fisheye camera image. A fisheye image acquisition unit is used to perform pixel-level classification on the corrected fisheye camera image to obtain the fisheye camera image containing semantic labels; The three-dimensional point cloud acquisition unit is used to determine the dynamic target detection result, the aggregated dynamic point cloud data and static point cloud data based on the initial three-dimensional point cloud information, and to use the dynamic target detection result, the dynamic point cloud data and the static point cloud data as the three-dimensional point cloud information.
[0151] Optionally, the non-hallucination reconstruction module includes: The texture map acquisition unit is used to construct a three-dimensional sky cube map based on the six-face texture map using the three-dimensional reconstruction network layer, determine the isometric projection model using the camera parameters of the fisheye camera, project the input ray of the sky cube map onto a unit sphere through the model, and perform two-dimensional image sampling on the projection data on the unit sphere to obtain a sky texture map adapted to the fisheye view. The background image acquisition unit is used to process static background node data and dynamic target node data based on the three-dimensional point cloud information and the vehicle pose information using the three-dimensional reconstruction network layer to obtain a background image. The rendering image acquisition unit is used to combine the background image and the sky texture image to obtain the background rendering image.
[0152] Optionally, the background image acquisition unit includes: The parameter acquisition subunit is used to initialize the static background node data in the initial world coordinate system based on the three-dimensional point cloud information using the three-dimensional reconstruction network layer to obtain static background node parameters; and to initialize the dynamic target node data with the dynamic target as the center to obtain dynamic target node parameters. The background image acquisition sub-unit is used to transform the Gaussian parameters of the dynamic target node and the Gaussian parameters of the static background node into the world coordinate system through the dynamic target pose in the current frame during frame-by-frame rendering. Then, the dynamic target Gaussian points and static background Gaussian points in the world coordinate system are projected and rendered using camera extrinsic parameters and vehicle pose information to obtain the rendered background image.
[0153] Optionally, the vignette reconstruction network layer includes: a two-dimensional vignette texture prediction module and a vignette opacity prediction module. The vignette reconstruction module includes: The vignette image acquisition unit is used to use the two-dimensional vignette texture prediction module to perform 2D texture sampling on the learned fisheye camera vignette texture image based on the fisheye camera texture coordinates in the camera parameters, and during the sampling process, project the sampled features onto the unit sphere to obtain the vignette rendering image. The opacity acquisition unit is used to process the position encoding information of the texture coordinates of the fisheye camera and the camera embedding vector of the fisheye camera using the vignette opacity prediction module to obtain the opacity.
[0154] Optionally, the loss value calculation module includes: The first loss calculation unit is used to calculate the mean absolute error loss and structural similarity loss based on the normalized RGB parameters of the target rendering image and the normalized RGB parameters of the fisheye camera image, and to perform a weighted summation of the mean absolute error loss and structural similarity loss to obtain the reprojection error loss. The second loss calculation unit is used to calculate the decoupling loss of the non-vague scene area based on the mask parameters and vignette opacity of the non-vague scene area in the target rendering image. The loss value calculation unit is used to calculate the loss value of the three-dimensional reconstruction model to be trained based on the reprojection error loss and the non-halo region decoupling loss.
[0155] This application embodiment trains a 3D reconstruction model that includes a vignette reconstruction network layer and a 3D reconstruction network layer. This decouples the vignettes around the fisheye camera from the 3D scene reconstruction. The camera vignettes can be modeled separately in a 2D space, so that the reconstruction of the 3D GS field and sky texture map is not affected by the camera vignettes. This greatly reduces the occurrence of floaters and improves the fidelity of the reconstructed scene.
[0156] This application provides an electronic device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor. When the computer program is executed by the processor, it implements the above-described three-dimensional reconstruction model training method.
[0157] Figure 10 A schematic diagram of the structure of an electronic device 1000 according to an embodiment of the present invention is shown. Figure 10As shown, the electronic device 1000 includes a central processing unit (CPU) 1001, which can perform various appropriate actions and processes according to computer program instructions stored in read-only memory (ROM) 1002 or loaded from storage unit 1008 into random access memory (RAM) 1003. The RAM 1003 may also store various programs and data required for the operation of the electronic device 1000. The CPU 1001, ROM 1002, and RAM 1003 are interconnected via a bus 1004. An input / output (I / O) interface 1005 is also connected to the bus 1004.
[0158] Multiple components in electronic device 1000 are connected to I / O interface 1005, including: input unit 1006, such as keyboard, mouse, microphone, etc.; output unit 1007, such as various types of monitors, speakers, etc.; storage unit 1008, such as disk, optical disk, etc.; and communication unit 1009, such as network card, modem, wireless transceiver, etc. Communication unit 1009 allows electronic device 1000 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0159] The various processes and handling described above can be executed by processing unit 1001. For example, the methods of any of the above embodiments can be implemented as computer software programs tangibly contained in a computer-readable medium, such as storage unit 1008. In some embodiments, part or all of the computer program can be loaded and / or installed on electronic device 1000 via ROM 1002 and / or communication unit 1009. When the computer program is loaded into RAM 1003 and executed by CPU 1001, one or more actions of the methods described above can be performed.
[0160] This application provides a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, it implements the various processes of the above-described three-dimensional reconstruction model training method embodiments and achieves the same technical effect. To avoid repetition, it will not be described again here. The computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, etc.
[0161] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.
[0162] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of this application.
[0163] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.
[0164] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed in this application can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0165] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0166] In the embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative. For instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0167] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0168] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0169] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, ROM, RAM, magnetic disks, or optical disks.
[0170] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A method for training a three-dimensional reconstruction model, characterized in that, The method includes: Acquire fisheye camera images, camera parameters, 3D point cloud information, and vehicle pose information. The fisheye camera images are images of the same target scene captured by a fisheye camera on the vehicle from different perspectives. The camera parameters are the parameters of the fisheye camera. The camera parameters, the 3D point cloud information, and the vehicle pose information are input into the 3D reconstruction model to be trained. The 3D reconstruction model to be trained includes: a vignette reconstruction network layer and a 3D reconstruction network layer. The vignette reconstruction network layer is used to reconstruct the vignette region of the fisheye camera image based on the camera parameters, thereby obtaining a vignette rendering image and opacity. The 3D reconstruction network layer is used to perform 3D reconstruction of the non-vague region of the fisheye camera image based on the 3D point cloud information, the camera parameters and the vehicle pose information to obtain a background rendering image. The background rendering image, the vignette rendering image, and the opacity are blended to obtain the target rendering image; The loss value of the 3D reconstruction model to be trained is calculated based on the fisheye camera image and the target rendering image. The model parameters of the 3D reconstruction model to be trained are adjusted based on the loss value until the termination training condition of the 3D reconstruction model to be trained is reached, so as to obtain the final 3D reconstruction model.
2. The method according to claim 1, characterized in that, The acquisition of fisheye camera images, camera parameters, 3D point cloud information, and vehicle pose information includes: Acquire multi-source data, which includes: initial fisheye camera image, camera parameters, initial 3D point cloud information, and vehicle pose information; Based on the distortion coefficients of the camera intrinsic parameters in the camera parameters, radial distortion correction is performed on the initial fisheye camera image to obtain the corrected fisheye camera image; The corrected fisheye camera image is classified at the pixel level to obtain the fisheye camera image containing semantic labels; Based on the initial 3D point cloud information, the dynamic target detection result, the aggregated dynamic point cloud data, and the static point cloud data are determined, and the dynamic target detection result, the dynamic point cloud data, and the static point cloud data are used as the 3D point cloud information.
3. The method according to claim 2, characterized in that, The step of using the 3D reconstruction network layer to perform 3D reconstruction of the non-vague region of the fisheye camera image based on the 3D point cloud information, the camera parameters, and the vehicle pose information to obtain a background rendering image includes: Using the aforementioned 3D reconstruction network layer, a 3D sky cube map is constructed based on a six-face texture map. An equidistant projection model is determined using the camera parameters of the fisheye camera. The input rays of the sky cube map are projected onto a unit sphere through this model, and the projection data on the unit sphere is sampled in two dimensions to obtain a sky texture map adapted to the fisheye view. Using the aforementioned 3D reconstruction network layer, static background node data and dynamic target node data are processed based on the 3D point cloud information and the vehicle pose information to obtain a background image; The background image and the sky texture image are combined to obtain the background rendering image.
4. The method according to claim 3, characterized in that, The process of using the 3D reconstruction network layer to process static background node data and dynamic target node data based on the 3D point cloud information and the vehicle pose information to obtain a background image includes: Using the aforementioned 3D reconstruction network layer, the static background node data is initialized in the initial world coordinate system based on the 3D point cloud information to obtain static background node parameters; and the dynamic target node data is initialized with the dynamic target as the center to obtain dynamic target node parameters. During frame-by-frame rendering, the Gaussian parameters of the dynamic target node and the Gaussian parameters of the static background node are transformed into the world coordinate system using the dynamic target pose in that frame. Then, the dynamic target Gaussian points and the static background Gaussian points in the world coordinate system are projected and rendered using camera extrinsic parameters and vehicle pose information to obtain the rendered background image.
5. The method according to claim 2, characterized in that, The vignette reconstruction network layer includes: a two-dimensional vignette texture prediction module and a vignette opacity prediction module. The process of using the vignette reconstruction network layer to reconstruct the vignette region of the fisheye camera image based on the camera parameters, to obtain a vignette rendering image and opacity, includes: The two-dimensional vignette texture prediction module uses the fisheye camera texture coordinates in the camera parameters to perform 2D texture sampling on the learned fisheye camera vignette texture image, and during the sampling process, the sampled features are projected onto a unit sphere to obtain the vignette rendering image. The vignette opacity prediction module is used to process the positional encoding information of the fisheye camera texture coordinates and the camera embedding vector of the fisheye camera to obtain the opacity.
6. The method according to claim 1, characterized in that, The step of calculating the loss value of the 3D reconstruction model to be trained based on the fisheye camera image and the target rendering image includes: Based on the normalized RGB parameters of the target rendering image and the normalized RGB parameters of the fisheye camera image, the mean absolute error loss and structural similarity loss are calculated, and the mean absolute error loss and structural similarity loss are weighted and summed to obtain the reprojection error loss. The decoupling loss of the non-vague scene region is calculated based on the mask parameters and vignette opacity of the non-vague scene region in the target rendering image. The loss value of the 3D reconstruction model to be trained is calculated based on the reprojection error loss and the non-halo region decoupling loss.
7. A three-dimensional reconstruction model training device, characterized in that, The device includes: The information acquisition module is used to acquire fisheye camera images, camera parameters, 3D point cloud information and vehicle pose information. The fisheye camera images are images of the same target scene captured by the fisheye camera on the vehicle from different perspectives, and the camera parameters are the parameters of the fisheye camera. The information input module is used to input the camera parameters, the 3D point cloud information and the vehicle pose information into the 3D reconstruction model to be trained. The 3D reconstruction model to be trained includes: a vignette reconstruction network layer and a 3D reconstruction network layer. The vignette reconstruction module is used to reconstruct the vignette region of the fisheye camera image based on the camera parameters using the vignette reconstruction network layer, so as to obtain a vignette rendering image and opacity. The non-vague reconstruction module is used to use the 3D reconstruction network layer to perform 3D reconstruction of the non-vague region of the fisheye camera image based on the 3D point cloud information, the camera parameters and the vehicle pose information, to obtain a background rendering image. The target rendering image acquisition module is used to blend the background rendering image, the vignette rendering image, and the opacity to obtain the target rendering image; The loss value calculation module is used to calculate the loss value of the three-dimensional reconstruction model to be trained based on the fisheye camera image and the target rendering image. The 3D reconstruction model acquisition module is used to adjust the model parameters of the 3D reconstruction model to be trained based on the loss value until the termination training condition of the 3D reconstruction model to be trained is reached, so as to obtain the final 3D reconstruction model.
8. The apparatus according to claim 7, characterized in that, The information acquisition module includes: The data acquisition unit is used to acquire multi-source data, which includes: initial fisheye camera image, camera parameters, initial 3D point cloud information and vehicle pose information; The image acquisition unit is configured to perform radial distortion correction on the initial fisheye camera image based on the distortion coefficients of the camera intrinsic parameters in the camera parameters, so as to obtain the corrected fisheye camera image. A fisheye image acquisition unit is used to perform pixel-level classification on the corrected fisheye camera image to obtain the fisheye camera image containing semantic labels; The three-dimensional point cloud acquisition unit is used to perform dynamic target detection on the initial three-dimensional point cloud information based on the pose information and the camera extrinsic parameters in the camera parameters, so as to obtain the three-dimensional point cloud information.
9. An electronic device, characterized in that, include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program, when executed by the processor, implements the three-dimensional reconstruction model training method according to any one of claims 1 to 6.
10. A readable storage medium, characterized in that, When the instructions in the storage medium are executed by the processor of the electronic device, the electronic device is able to perform the three-dimensional reconstruction model training method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Vehicle-mounted all-around panoramic image generation method and system
CN118587309A
Method and system for enhancing visual positioning based on cross-domain three-dimensional Gaussian sputtering
CN120388074A
Self-supervised single-view 3D reconstruction via semantic consistency
US20210287430A1