Scene modeling method and device, electronic equipment and vehicle
Patent Information
- Application Number
- CN202610968448.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-30
- Publication Date
- 2026-09-25
AI Technical Summary
[0004]本申请实施例提供一种场景建模方法、装置、电子设备及车辆,旨在改善相关技术中因先验信息利用方式粗糙、几何网络初始化策略针对性不强以及监督信号本身可靠性存疑等问题导致复杂场景下的几何重建效果差的问题
[0005]本申请的第一个目的在于提出一种场景建模方法,包括:获取目标场景的多视角图像序列和多视角中每个视角的相机位姿,多视角与多视角图像序列对应;基于所述多视角图像序列和所述每个视角的相机位姿,生成所述目标场景的初始三维模型;基于所述初始三维模型对初始有符号距离函数网络进行训练,以得到目标有符号距离函数网络,其中,所述目标有符号距离函数网络用于计算所述初始三维模型所处空间中的任一空间点到所述目标场景表面的有符号距离值;基于所述目标有符号距离函数网络提供的有符号距离值,对初始场景建模模型进行训练,以得到目标场景建模模型;通过所述目标场景建模模型对所述目标场景进行建模,以得到所述目标场景的场景模型获取目标场景的多视角图像序列和对应的相机位姿;基于多视角图像序列和对应的相机位姿,生成目标场景的初始建模结果;基于初始建模结果对初始有符号距离函数网络进行训练,以得到目标有符号距离函数网络;基于目标有符号距离函数网络对第一场景建模模型进行训练,以得到目标场景建模模型;通过目标场景建模模型对目标场景进行建模,以得到目标场景的场景模型。
Smart Images

Figure CN122821040A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of scene modeling technology, and in particular to a scene modeling method, apparatus, electronic device, and vehicle. Background Technology
[0002] In the field of 3D reconstruction technology for autonomous driving scenarios, implicit scene representation methods based on neural radiation fields still have shortcomings in terms of geometric surface reconstruction accuracy. To improve this issue, related technical solutions attempt to introduce geometric prior information provided by a pre-trained monocular depth estimation model as an additional supervision signal during the training process.
[0003] However, its geometric supervision mechanism directly relies on the pseudo-true value of monocular depth estimation as the basis for constraint. Since monocular depth estimation itself has inherent properties of scale ambiguity and inconsistency in local region accuracy, using this pseudo-true value as the direct supervision signal in the training process will lead to the confidence distribution at different training stages being difficult to accurately match with the actual convergence requirements of the network. When neural radiation field reconstruction methods introduce geometric priors to improve reconstruction results, they generally suffer from problems such as crude utilization of prior information, weak targeting of geometric network initialization strategies, and questionable reliability of the supervision signal itself. As a result, the geometric reconstruction effect in complex scenes still needs to be improved. Summary of the Invention
[0004] This application provides a scene modeling method, apparatus, electronic device, and vehicle, aiming to improve the poor geometric reconstruction effect in complex scenes caused by problems such as the crude use of prior information, the weak targeting of geometric network initialization strategies, and the questionable reliability of the supervision signals themselves.
[0005] The first objective of this application is to propose a scene modeling method, comprising: acquiring a multi-view image sequence of a target scene and the camera pose of each viewpoint in the multi-view sequence, wherein the multi-view sequence corresponds to the multi-view image sequence; generating an initial 3D model of the target scene based on the multi-view image sequence and the camera pose of each viewpoint; training an initial signed distance function network based on the initial 3D model to obtain a target signed distance function network, wherein the target signed distance function network is used to calculate the signed distance value from any spatial point in the space where the initial 3D model is located to the surface of the target scene; and based on the signed distance value provided by the target signed distance function network... The initial scene modeling model is trained to obtain the target scene modeling model; the target scene is modeled using the target scene modeling model to obtain the scene model of the target scene; multi-view image sequences and corresponding camera poses of the target scene are obtained; based on the multi-view image sequences and corresponding camera poses, an initial modeling result of the target scene is generated; an initial signed distance function network is trained based on the initial modeling result to obtain a target signed distance function network; the first scene modeling model is trained based on the target signed distance function network to obtain the target scene modeling model; the target scene is modeled using the target scene modeling model to obtain the scene model of the target scene.
[0006] This application solidifies the geometric priors of the initial 3D model into a signed distance function network during the pre-training stage, and then connects this network to the scene modeling model for joint training. This allows the model to have a preliminary understanding of the scene's geometric structure from the beginning of training, effectively reducing the optimization difficulty of the scene modeling model. The final extracted scene model retains high-quality appearance rendering capabilities while significantly improving the accuracy and surface integrity of 3D geometric reconstruction.
[0007] According to one embodiment of this application, an initial scene modeling model is trained based on the signed distance value provided by the target signed distance function network to obtain a target scene modeling model. This includes: constructing a second training set based on the pixels of each frame in the multi-view image sequence, the corresponding real color value of the pixel, and the monocular depth estimation result; for any pixel in the second training set, determining multiple second spatial points in the three-dimensional space corresponding to any pixel based on the corresponding camera pose; inputting the coordinates of the second spatial points into the target signed distance function network to obtain the second predicted signed distance value corresponding to the second spatial point; and training the initial scene modeling model based on the second training set and the second predicted signed distance value to obtain the target scene modeling model.
[0008] According to one embodiment of this application, training an initial scene modeling model based on a second training set and a second predicted signed distance value to obtain a target scene modeling model includes: inputting the coordinates of a second spatial point into the initial scene modeling model to obtain a predicted color value corresponding to the second spatial point; generating a rendering depth value corresponding to any pixel based on the second predicted signed distance value, and generating a rendering color value corresponding to any pixel based on the second predicted signed distance value and the predicted color value; calculating a discrimination loss based on the rendering color value, the rendering depth value, and the second predicted signed distance value to obtain a second calculation result; and updating the model parameters of the initial scene modeling model using the second calculation result until the updated initial scene modeling model satisfies a second preset convergence condition to obtain the target scene modeling model.
[0009] According to one embodiment of this application, the calculation of a discrimination loss based on a rendered color value, a rendered depth value, and a second predicted signed distance value to obtain a second calculation result includes: calculating a discrimination loss based on the rendered color value and the real color value of the corresponding pixel in the second training set to obtain a color reconstruction loss; calculating a discrimination loss based on the rendered depth value and the corresponding monocular depth estimation result in the second training set to obtain a monocular depth loss; calculating a ray sparsity constraint loss based on the second predicted signed distance value; and weighted summing the color reconstruction loss, monocular depth loss, and ray sparsity constraint loss to obtain the second calculation result.
[0010] According to one embodiment of this application, training an initial signed distance function network based on an initial 3D model to obtain a target signed distance function network includes: converting the initial 3D model into a meshed surface model and aligning the meshed surface model to the target coordinate system where the initial scene modeling model is located; sampling multiple first spatial points in the target coordinate system and calculating the actual signed distance value between each first spatial point and the meshed surface model; and training the initial signed distance function network based on the actual signed distance value to obtain the target signed distance function network.
[0011] According to one embodiment of this application, training an initial signed distance function network based on actual signed distance values to obtain a target signed distance function network includes: constructing a first training set based on the coordinates of a first spatial point and the corresponding actual signed distance value; inputting the coordinates of any first spatial point in the first training set into the initial signed distance function network to obtain a first predicted signed distance value; calculating a discriminative loss based on the first predicted signed distance value and the corresponding actual signed distance value in the first training set to obtain a first calculation result; and updating the model parameters of the initial signed distance function network using the first calculation result until the updated initial signed distance function network satisfies a first preset convergence condition to obtain the target signed distance function network.
[0012] According to one embodiment of this application, an initial three-dimensional model of a target scene is generated based on a multi-view image sequence and the corresponding camera pose, including: acquiring the depth map of each frame in the multi-view image sequence; back-projecting the depth map of each frame to a three-dimensional space based on the camera pose to obtain the point cloud of each frame; and stitching together the point clouds of each frame to obtain the initial three-dimensional model.
[0013] The second objective of this application is to propose a scene modeling device, comprising: an acquisition module for acquiring a multi-view image sequence of a target scene and the camera pose of each view in the multi-view sequence, wherein the multi-view sequence corresponds to the multi-view image sequence; a first modeling module for generating an initial 3D model of the target scene based on the multi-view image sequence and the corresponding camera pose; a first training module for training an initial signed distance function network based on the initial 3D model to obtain a target signed distance function network; a second training module for training an initial scene modeling model based on the target signed distance function network to obtain a target scene modeling model; and a second modeling module for modeling the target scene using the target scene modeling model to obtain a scene model of the target scene.
[0014] The third objective of this application is to provide an electronic device, including a memory, a processor, and a scene modeling program stored in the memory and capable of running on the processor, wherein when the processor executes the scene modeling program, it implements the aforementioned scene modeling method.
[0015] The fourth objective of this application is to provide a vehicle that includes the aforementioned electronic equipment. Attached Figure Description
[0016] Figure 1 This is a schematic diagram of the network structure of StreetSurf according to some embodiments of this application; Figure 2 This is a flowchart of a scene modeling method according to some embodiments of this application; Figure 3 This is a block diagram of a scene modeling apparatus according to some embodiments of this application; Figure 4 This is a block diagram of an electronic device according to some embodiments of this application; Figure 5 This is a block diagram of a vehicle according to some embodiments of this application. Detailed Implementation
[0017] To make the technical problems, technical solutions, and beneficial effects solved by this application clearer, the following detailed description is provided in conjunction with embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0018] Based on the background, in the field of 3D reconstruction technology for autonomous driving, the mainstream technical routes can be divided into two main categories: active reconstruction based on LiDAR and passive reconstruction based on vision sensors. While LiDAR-based solutions can directly acquire high-precision scene structure information, their high hardware cost limits their large-scale application in mass-produced vehicles. Therefore, lower-cost pure vision-based 3D reconstruction solutions have become a current research hotspot, mainly including methods such as stitching reconstruction based on monocular depth estimation, explicit reconstruction based on 3D Gaussian splashing, and implicit reconstruction based on neural radiation fields. Among these, neural radiation fields, as an implicit scene representation technique, achieves new perspective synthesis and scene geometry restoration by learning continuous radiation field functions; however, in practical applications, it still has significant shortcomings in terms of geometric surface reconstruction accuracy.
[0019] To improve the geometric reconstruction of neural radiation fields in complex scenes, some related technical solutions attempt to introduce additional geometric prior information to enhance constraints. However, a common drawback of these solutions is that their geometric supervision mechanism often directly relies on the pseudo-true values output by the pre-trained monocular depth estimation model. Since monocular depth estimation inherently possesses scale ambiguity and inconsistencies in local region accuracy, using these pseudo-true values as direct supervision signals during training leads to a mismatch between the confidence distribution at different training stages and the actual convergence requirements of the network, thus limiting the geometric accuracy and surface detail integrity of the final reconstruction result.
[0020] In the reconstruction of unbounded street scenes, StreetSurf proposes a multi-view implicit surface reconstruction method to address the specific challenges posed by street scene images captured by long-distance, narrow-view camera trajectories. This method divides the unbounded scene into three regions: foreground, background, and sky, and represents each region separately. The foreground region is geometrically reconstructed using an implicit surface model incorporating a signed distance function, while the background and sky regions are supplemented using a radiation field model. Furthermore, during training, this method also incorporates geometric prior information estimated by a general monocular model as a supervision signal and attempts to accelerate the convergence of the foreground network by pre-training on the 3D positions of the road surface areas.
[0021] Specifically, the StreetSurf algorithm is a novel multi-view implicit surface reconstruction technique that addresses the unique challenges of unbounded street scene images captured by long, narrow camera trajectories that are not object-centric. This is achieved by extending previous object-centric neural surface reconstruction techniques. StreetSurf divides the unbounded space into three parts: close-range, distant-view, and sky, and uses a cube / hypercube hash grid aligned with the cube boundaries, along with a road surface initialization scheme, to achieve a more refined and decoupled representation. To further address geometric errors caused by insufficient texture areas and limited viewpoints, StreetSurf employs geometric priors estimated using a general monocular model. Combined with an efficient multi-stage ray-stepping strategy, StreetSurf achieves high-quality geometric and appearance reconstructions in just one to two hours of training per street scene sequence.
[0022] 3D reconstruction using the Nerf series is very difficult. The fundamental reason is that the Nerf series is, in principle, an underfitting system lacking geometric constraints. Although Neus improved upon this deficiency by introducing the SDF function, the learning of the SDF function is indirectly achieved through gradient iteration updates based on the color difference between predicted and actual pixels. This implicit learning is less effective than direct learning, which is much more difficult. In StreetSurf, the foreground is represented by Neus, and the background by Nerf. The network structure diagram of StreetSurf can be found here. Figure 1 Although StreetSurf pre-trains a near-field Neus based on the 3D location of a given road surface, this pre-training is inherently inaccurate because the road surface cannot directly represent the real possible 3D space. The real 3D space includes not only the road surface but also buildings, trees, and so on. Furthermore, StreetSurf uses a geometric prior estimated using a general monocular model, but monocular depth estimation itself cannot be accurate, so its loss function expression is problematic.
[0023] The StreetSurf solution still has the following technical shortcomings in practical applications: First, the learning mechanism of its near-field geometric network is essentially still an indirect optimization process. Although a signed distance function is introduced to strengthen geometric constraints, the parameter update of this function still mainly relies on gradient backpropagation based on the difference between the rendered pixel color and the real pixel color. This implicit learning path does not provide direct feedback on geometric information, making it difficult for the signed distance function network to fit complex surface morphologies and resulting in insufficient stability during training.
[0024] Second, its road surface pre-training strategy for the near-field geometric network has limitations. This scheme uses the 3D location information of the road surface to initialize and guide the near-field signed distance function network. However, the road surface region can only represent a very limited set of planar structures in the scene, failing to encompass the non-planar geometries such as building facades, tree outlines, and traffic facilities that are widely present in real-world driving environments. Therefore, road surface-based pre-training cannot provide the network with accurate prior knowledge of the complete 3D spatial geometric distribution, offering limited help in reducing the overall training difficulty.
[0025] Third, the reliability of its geometric prior supervision signal is insufficient. Similar to the other schemes mentioned above, the depth and normal information estimated by the general monocular model on which StreetSurf relies cannot avoid the scale blur and accuracy fluctuation problems inherent in monocular vision. Directly using such pseudo-true values with inherent errors as the basis for the loss function will introduce bias during training, interfere with the sign distance function network's learning of accurate geometric surfaces, and thus affect the geometric quality of the final reconstruction result.
[0026] In summary, neural radiation field reconstruction methods, when introducing geometric priors to improve reconstruction results, generally suffer from problems such as coarse utilization of prior information, weak targeting of geometric network initialization strategies, and questionable reliability of the supervision signals themselves. This results in the geometric reconstruction performance in complex scenes still needing improvement. Therefore, a more effective scene modeling method is urgently needed to enhance the geometric accuracy and stability of 3D reconstruction.
[0027] Based on this, before training the neural radiation field network, this application first uses a pre-reconstruction method based on monocular depth estimation to obtain the initial dense geometric structure of the target scene, and then uses this initial geometric structure to pre-train the symbolic distance function branch in the neural radiation field network. In this way, the symbolic distance function network obtains a preliminary understanding of the macroscopic geometric contours and approximate surface morphology of the scene before formally participating in joint rendering optimization, essentially solidifying the explicitly reconstructed geometric priors directly into the initial parameter state of the symbolic distance function network. On this basis, joint optimization training is then performed by combining the volume rendering equation of the neural radiation field with the introduced geometric prior supervision signal. This pre-training step effectively reduces the initial exploration difficulty of the symbolic distance function network in complex high-dimensional geometric spaces, significantly narrowing the parameter search range in subsequent training and resulting in a smoother and more stable convergence path. Simultaneously, the improved geometric prior supervision strategy further corrects local geometric deviations left over from the pre-training stage. Ultimately, this application effectively improves the accuracy and detail reproduction of 3D scene geometric structure reconstruction while retaining the high-quality appearance rendering capabilities of the neural radiation field.
[0028] The embodiments of this application are described in detail below. Examples of these embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain this application, and should not be construed as limiting this application.
[0029] The scene modeling method, apparatus, electronic device, and vehicle of this application embodiment will now be described in detail with reference to the accompanying drawings.
[0030] Figure 2 This is a flowchart of a scene modeling method according to some embodiments of this application. (Refer to...) Figure 2 The scene modeling method in this application embodiment may include the following steps: S110: Obtain the multi-view image sequence of the target scene and the camera pose of each view in the multi-view, with the multi-view corresponding to the multi-view image sequence.
[0031] Specifically, the multi-view image sequence is acquired by the same camera under different camera poses. This sequence refers to a set of ordered images acquired from different spatial positions and orientations around the target scene. Each frame of the image corresponds to a unique observation direction and viewpoint position. The viewpoints are covered in multiple directions by the overlapping area of the field of view. The camera pose refers to the position and orientation of the camera in space when each frame of the image is captured. A one-to-one mapping relationship is established between the two, that is, the i-th frame of the image is strictly bound to the i-th camera pose, so that each two-dimensional image is positioned as an observation sample of the scene from a specific viewpoint.
[0032] The specific acquisition process includes two stages: image acquisition and pose calculation. In the image acquisition stage, the target scene is continuously or fixedly photographed by rotating the vehicle-mounted camera to obtain an RGB image sequence with sufficient field-of-view overlap. In the pose calculation stage, the acquired image sequence is feature extracted and matched using motion recovery structures or simultaneous localization and mapping algorithms. The camera parameters and the positions of three-dimensional feature points are jointly optimized by bundle adjustment to calculate the camera intrinsic parameter matrix and extrinsic parameter pose corresponding to each frame of the image, and stored as a pose transformation file in JSON or text format adapted for neural radiation field training.
[0033] S120 generates an initial 3D model of the target scene based on a multi-view image sequence and the corresponding camera pose.
[0034] Specifically, during the training of the neural radiation field, if the signed distance function network starts learning from a randomly initialized state, it needs to explore a complex, high-dimensional geometric space without any spatial reference. This not only leads to slow training convergence but also makes it prone to fragmented or floating geometric artifacts due to getting trapped in local optima. Therefore, this application pre-establishes an initial 3D model of the target scene, providing a geometric prior reference close to the real surface morphology for the subsequent SDF network (signed distance function network), thereby narrowing the parameter search range of the network and reducing the initial exploration difficulty of geometric learning. Specifically, an initial 3D model of the target scene can be output by inputting multi-view image sequences and corresponding camera poses into the pre-set initial modeling model.
[0035] S130, the initial signed distance function network is trained based on the initial 3D model to obtain the target signed distance function network, wherein the target signed distance function network is used to calculate the signed distance value from any spatial point in the space where the 3D model is located to the surface of the target scene.
[0036] Specifically, in related neural radiation field methods, signed distance function networks are typically trained from a randomly initialized state. This means that the network needs to explore complex high-dimensional geometric spaces solely based on indirect feedback from color rendering errors, without any spatial reference information. This blind fitting approach not only leads to slow training convergence but also easily causes the network to get stuck in local optima, resulting in broken, hollow, or floating geometric artifacts in weakly textured regions or at occlusion boundaries, severely affecting the geometric integrity and surface quality of the final reconstruction result.
[0037] To address the aforementioned issues, this application pre-trains the SDF network using the initial 3D model obtained in the preceding steps before integrating it into the neural radiation field for joint optimization. This pre-trains the network with the macroscopic geometric contours and approximate surface morphology information of the target scene contained in the initial 3D model, pre-setting it as the initial parameter state of the SDF network. This allows the network to possess a preliminary understanding of the scene's geometric structure before participating in complex volumetric rendering optimization. Specifically, for any spatial point, the output signed distance value roughly reflects its position relative to the target scene surface. For example, a positive signed distance value indicates the point is located outside the target scene surface (i.e., an open area outside the object); a negative signed distance value indicates the point is located inside the target scene surface (i.e., an internal region of the object); and a zero signed distance value indicates the point is exactly on the target scene surface. This significantly narrows the search range of network parameters in the subsequent joint training phase, reduces the initial exploration difficulty of geometric learning, makes the convergence path smoother and more stable, and provides an optimized starting point for subsequent restoration of refined surface details.
[0038] The specific training process is as follows: First, the initial 3D model is converted into a surface mesh model and aligned to the coordinate system of the initial signed distance function network. Then, sampling is performed within the spatial range of the surface mesh model, and the true value of the signed distance function of each sampling point relative to the surface mesh model is calculated, and a training set is constructed accordingly. Finally, the initial signed distance function network is trained under supervision using the training set until the preset convergence condition is met, and the target signed distance function network is obtained.
[0039] S140, based on the signed distance value provided by the target signed distance function network, train the initial scene modeling model to obtain the target scene modeling model.
[0040] Specifically, in related technologies, the geometric surface and appearance color of a scene are usually jointly learned by the same model from a randomly initialized state. Since the geometric branch lacks any prior knowledge of the spatial structure of the scene, it can only rely on indirect signals of color rendering errors to blindly explore in the early stages of training. This results in a slow convergence process, geometric surfaces are prone to fragmentation or floating artifacts, and the geometric reconstruction quality is particularly poor for weakly textured areas.
[0041] To address the aforementioned issues, this application, after pre-training the signed distance function network, uses the target signed distance function network, which already possesses macroscopic geometric cognition of the scene, as a geometric prior and integrates it into the initial scene modeling model for joint optimization training. The purpose of this is to deeply integrate the geometric prior information (signed distance values) obtained from explicit reconstruction with the volume rendering optimization framework of the neural radiation field. This allows the model to accurately determine the approximate location of the scene surface from the very beginning of training, thereby significantly reducing the parameter search space of the geometric branches, lowering the optimization difficulty, and providing a stable and reliable geometric foundation for subsequent learning of appearance details.
[0042] Specifically, a training set is constructed based on multi-view image sequences. Next, a target signed distance function network is integrated into the initial scene modeling model to obtain the initial scene modeling model. Subsequently, the initial scene modeling model is iteratively trained using the training set, updating the model parameters by minimizing the loss function until the corresponding preset convergence condition is met, thus obtaining the target scene modeling model.
[0043] S150, the target scene is modeled using the target scene modeling model to obtain the scene model of the target scene.
[0044] Specifically, after completing joint training and obtaining the target scene modeling model, this model can be used to perform a complete 3D modeling of the target scene to extract the final scene model. Specifically, in the modeling phase, a series of spatial point coordinates densely sampled within the 3D space of the target scene are sequentially input into the trained target scene modeling model. The signed distance function network branch in the model outputs the predicted signed distance value corresponding to each spatial point. Subsequently, the moving cube algorithm is used to extract continuous 3D triangular mesh surfaces at all isosurfaces where the predicted signed distance value is zero. This mesh surface is the final scene model of the target scene.
[0045] This application pre-trains the geometric priors of the initial 3D model into a signed distance function network, and then integrates this network into the scene modeling model for joint training. This allows the model to possess a preliminary understanding of the scene's geometric structure from the beginning of training, effectively reducing the optimization difficulty of the scene modeling model. The final extracted scene model retains high-quality appearance rendering capabilities while significantly improving the accuracy and surface integrity of 3D geometric reconstruction.
[0046] In some embodiments, an initial scene modeling model is trained based on the signed distance values provided by the target signed distance function network to obtain a target scene modeling model. This includes: constructing a second training set based on the pixels of each frame in the multi-view image sequence, the corresponding true color values of the pixels, and the monocular depth estimation results; for any pixel in the second training set, determining multiple second spatial points in the three-dimensional space corresponding to that pixel based on the corresponding camera pose; inputting the coordinates of the second spatial points into the target signed distance function network to obtain the second predicted signed distance value corresponding to the second spatial point; and training the initial scene modeling model based on the second training set and the second predicted signed distance value to obtain the target scene modeling model. The scene modeling model can be StreetSurf, and no specific restrictions are imposed here.
[0047] In some embodiments, training an initial scene modeling model based on a second training set and a second predicted signed distance value to obtain a target scene modeling model includes: inputting the coordinates of a second spatial point into the initial scene modeling model to obtain a predicted color value corresponding to the second spatial point; generating a rendering depth value corresponding to any pixel based on the second predicted signed distance value, and generating a rendering color value corresponding to any pixel based on the second predicted signed distance value and the predicted color value; calculating a discriminative loss based on the rendering color value, the rendering depth value, and the second predicted signed distance value to obtain a second calculation result; and updating the model parameters of the initial scene modeling model using the second calculation result until the updated initial scene modeling model satisfies a second preset convergence condition to obtain the target scene modeling model.
[0048] Specifically, firstly, a second training set is constructed based on the pixel labels of each frame in the multi-view image sequence, the corresponding true color values of each pixel, and the reference depth values pre-inferred by the monocular depth estimation model. Then, the target signed distance function network, pre-trained in the previous steps, is injected into the corresponding network branch responsible for geometric representation in the initial scene modeling model, replacing the original random initialization state. This results in the injected scene modeling model, at which point the geometric branch of the model has a preliminary understanding of the macroscopic contours of the target scene. Finally, the injected scene modeling model is iteratively trained using the second training set. By minimizing a preset loss function, the network parameters of the geometric and color branches in the model are updated synchronously until a preset convergence condition is met. The trained model is then the target scene modeling model.
[0049] During the joint training phase, for any pixel sample in the second training set, a ray path traversing three-dimensional space is first determined from the camera optical center along the observation direction of the pixel based on the camera pose parameters corresponding to the pixel. Multiple discrete second spatial points are then sampled along the ray path in order from near to far, thereby transforming the two-dimensional pixel observation into a series of sampling positions in three-dimensional space.
[0050] The StreetSurf model employs a multi-branch design, consisting of three parallel sub-networks: a near-field geometry branch, a far-field radiation branch, and a sky branch. The near-field geometry branch is responsible for reconstructing the main structure of the scene. Its core is a signed distance function network, which uses a fully connected multilayer perceptron structure with a position-encoded input layer. This structure comprises eight hidden layers, each containing 256 neurons, with fully connected layers. The hidden layer activation function is the Modified Linear Unit (ReLU) function, and the output layer directly outputs a scalar distance value without an activation function. In this embodiment, the entire target signed distance function network is injected into the corresponding network branch to replace the original random initialization state. The far-field radiation branch and the sky branch use a radiation field network structure to supplement the modeling of the appearance information of distant background and sky regions. The outputs of the three branches are fused using a spatial partitioning strategy to ultimately generate a complete scene representation.
[0051] Regarding the preprocessing of training data, the construction of the second training set includes: normalizing the pixel values of each frame of RGB image in the multi-view image sequence to the range of zero to one, converting the camera pose into a spatial transformation matrix adapted to the StreetSurf coordinate system, aligning the monocular depth estimation results with the image pixels pixel by pixel, and randomly shuffling and batching all training samples.
[0052] During training, the 3D coordinates of each second spatial point are upsized through position encoding and then input into the injected scene modeling model. For near spatial points, the signed distance function network branch outputs the second predicted signed distance value, and the color network branch outputs the predicted color value; for spatial points in the distant and sky regions, the radiation field branch directly outputs the volume density and color value.
[0053] Based on this, the second predicted signed distance value of each second spatial point is converted into the spatial occupancy value corresponding to each second spatial point through a preset mapping function. This spatial occupancy value represents the probability that a ray encounters the object surface at that location. Subsequently, the spatial occupancy value of each second spatial point is normalized along the ray path to obtain the weight distribution of each sampling point. Finally, the actual distance values of each second spatial point from the camera's optical center are weighted and summed according to the above weight distribution. The resulting expected distance is the rendering depth value corresponding to that pixel. This rendering depth value reflects the model's best estimate of the distance between the object surface and the camera.
[0054] Furthermore, based on the spatial occupancy of each second spatial point, the contribution weight of each second spatial point to the final pixel color is calculated sequentially along the light path from near to far. This contribution weight is jointly determined by the spatial occupancy and the cumulative transmittance of the light before reaching that point. The predicted color values of each second spatial point are accumulated and integrated according to their contribution weights, and the final color value obtained is the rendered color value corresponding to that pixel. This rendered color value simulates the physical process of light propagating in the scene and interacting with the object surface before reaching the photosensitive element during real camera imaging.
[0055] After obtaining the rendered color and depth values of the current pixel sample, the discriminative loss is calculated. First, the rendered color value is compared with the corresponding real color value of the pixel in the second training set to calculate the color reconstruction loss. This loss measures the appearance difference between the model's rendered result and the real image. Second, the rendered depth value is compared with the corresponding monocular reference depth value (monocular depth estimation result) to calculate the monocular depth loss. This loss is calculated by aligning the two values using a learnable scaling factor and translation, and then obtaining the mean square error, providing a depth-scale supervision signal for the geometric network. Third, based on the spatial occupancy converted from the second predicted signed distance value of each second spatial point, the ray sparsity constraint loss is calculated. This loss constrains the overall occupancy along the entire ray path and the occupancy of the region after the monocular reference depth to approach zero, thereby suppressing fog noise and background floating objects. The weighted sum of the above color reconstruction loss, monocular depth loss, and ray sparsity constraint loss yields the second calculation result for the current batch of pixel samples in this iteration.
[0056] Subsequently, backpropagation is performed based on the second calculation result, calculating the gradient of the loss relative to each network parameter in the injected scene model layer by layer. Using the Adam optimizer, each parameter is updated along the gradient descent direction according to the current learning rate, completing one iteration. Next, the next batch of pixel samples is extracted from the second training set, and the complete process of spatial point sampling, forward propagation to obtain predicted values, rendering depth and color generation, weighted summation of the three losses, backpropagation gradient calculation, and parameter update is repeated. As the number of iterations increases, the model parameters are gradually adjusted towards minimizing the loss, and the accuracy of color rendering and depth estimation is continuously improved.
[0057] After each iteration, the average discrimination loss of all batches in that round is recorded and compared with the average loss of several previous rounds. When the average loss of multiple consecutive training rounds decreases by less than a preset threshold, it indicates that the model is close to convergence and the effect of parameter updates on improving the loss has become saturated. At this point, or when the number of iterations reaches the preset maximum number of iterations, it is considered to have met the second preset convergence condition, and training is stopped. The scene modeling model obtained at this point is the target scene modeling model. Its internal geometric branches have refined the macroscopic prior information injected during the pre-training stage, and the color branch has also learned to generate appearance rendering results that are highly consistent with the real scene.
[0058] In some embodiments, the calculation of a discriminative loss based on the rendered color value, the rendered depth value, and the second predicted signed distance value to obtain a second calculation result includes: calculating a discriminative loss based on the rendered color value and the real color value of the corresponding pixel in the second training set to obtain a color reconstruction loss; calculating a discriminative loss based on the rendered depth value and the corresponding monocular depth estimation result in the second training set to obtain a monocular depth loss; calculating a ray sparsity constraint loss based on the second predicted signed distance value; and weighted summing the color reconstruction loss, the monocular depth loss, and the ray sparsity constraint loss to obtain the second calculation result.
[0059] For example, the formula for calculating color reconstruction loss is as follows: L_color=Σ||C_pred-C_true||²; Where L_color represents the color reconstruction loss; C_pred represents the rendered color value; C_true represents the true color value; and Σ represents the summation of all rays in the training batch.
[0060] The formula for calculating monocular depth loss is as follows: L_depth=Σ||(w×D_pred+q)-D_mono||²; Where L_depth represents the monocular depth loss; D_pred represents the desired rendering depth; D_mono represents the monocular depth estimation result; w represents the scale factor; q represents the translation amount; and Σ represents the summation over all rays in the training batch. w and q can be learned and updated during training.
[0061] The formula for calculating the sparse constraint loss is as follows: ; Where L_sparse represents the sparse constraint loss; t_start represents the spatial occupancy obtained by converting the second predicted signed distance value; t_start represents the ray start point, t_end represents the ray end point, ε is a preset threshold constant, representing the tolerance range near the monocular depth estimate; D(r) represents the monocular depth estimate result.
[0062] It should be noted that the monocular depth estimation result indicates the approximate spatial location of the object's surface at its distance from the camera. However, due to the inherent properties of scale ambiguity and inconsistencies in local accuracy in monocular depth estimation, the predicted value is not absolutely accurate, and the actual surface of the object may lie within a certain interval before or after the predicted value. Therefore, this application sets a threshold ε to form a narrow band interval before and after the predicted depth value, namely the range from the predicted depth value minus ε to the predicted depth value plus ε. Within this narrow band interval, the model is allowed to retain a certain amount of space occupancy to fit the actual object surface position. Outside this narrow band interval, namely the region from the light source to the predicted depth value minus ε, and the region from the predicted depth value plus ε to the light source end point, the space occupancy is forced to approach zero. The physical meaning of this constraint is that the light should traverse free space without encountering any obstructions before reaching the object surface, and the light should not produce any extra transmission or floating artifacts behind the object after passing through the object surface. Therefore, the introduction of monocular predicted depth and its threshold offset in the integral formula in the form of upper and lower limits is essentially to use the rough positional information provided by monocular depth estimation to define the possible spatial range of the object surface and impose strict spatial constraints on the area outside the range, thereby effectively eliminating the fog noise and background floating object problems commonly found in neural radiation field reconstruction.
[0063] The formula for calculating the second result is as follows: L_total=λ1×L_color+λ2×L_depth+λ3×L_sparse; Where L_total represents the second calculation result; L_color represents the color reconstruction loss; λ1 represents the weight coefficient corresponding to L_color; L_depth represents the monocular depth loss; λ2 represents the weight coefficient corresponding to L_depth; L_sparse represents the ray sparsity constraint loss; and λ3 represents the weight coefficient corresponding to L_sparse. λ1, λ2, and λ3 can be calibrated according to actual conditions, and no specific restrictions are imposed here.
[0064] This application weights and sums the color reconstruction loss, monocular depth loss, and ray sparsity constraint loss according to preset weight coefficients to form a joint optimization total loss. This weighted design allows the three types of losses to work synergistically during training. Among them, the color reconstruction loss serves as the dominant supervision signal, driving the model to learn an appearance rendering capability that is highly consistent with the input image, ensuring the visual realism and fidelity of the reconstruction results; the monocular depth loss, in the form of auxiliary supervision, injects the coarse depth information estimated by the monocular model into the geometric branch, providing scale alignment and approximate surface position guidance for the signed distance function network, accelerating geometric convergence and preventing the network from getting stuck in ambiguous solutions of depth scale; the ray sparsity constraint loss, as a regularization term, forces the occupancy of non-surface regions on the ray path to approach zero, effectively suppressing fog noise, background floating objects, and transmission artifacts commonly found in neural radiation field reconstruction. By adjusting the weight coefficients corresponding to various losses, the emphasis between appearance reconstruction accuracy and geometric reconstruction quality can be flexibly balanced at different stages of training. This allows the model to retain high-quality appearance rendering capabilities while significantly improving the integrity and clarity of the geometric surface of the 3D scene, ultimately achieving a high-quality 3D reconstruction effect that balances appearance realism and geometric accuracy.
[0065] In some embodiments, training an initial signed distance function network based on an initial 3D model to obtain a target signed distance function network includes: converting the initial 3D model into a meshed surface model and aligning the meshed surface model to the target coordinate system of the initial scene modeling model; sampling multiple first spatial points in the target coordinate system and calculating the actual signed distance value between each first spatial point and the meshed surface model; and training the initial signed distance function network based on the actual signed distance value to obtain the target signed distance function network.
[0066] Specifically, the initial 3D model is a discrete set of 3D point clouds, which is difficult to directly use as the continuous supervision signal required for network training. Therefore, it is necessary to first use algorithms such as Poisson surface reconstruction to transform this discrete point cloud into a triangular mesh model with continuous surfaces to achieve a complete description of the scene's surface morphology. Subsequently, since this mesh model is built at a real physical scale or in an arbitrary world coordinate system, while the signed distance function network to be trained usually operates in a normalized unit space, it is necessary to transform the spatial scale and position of the mesh model to align it to the target coordinate system of the initial scene modeling model, so as to ensure the spatial consistency between the coordinates of subsequent sampling points and the network input domain.
[0067] Within the aligned target coordinate system, a strategy combining uniform sampling and near-surface dense sampling is employed to select a large number of first spatial points in 3D space. For each sampled first spatial point, its nearest distance to the meshed surface model is calculated, and a corresponding positive or negative sign is assigned based on whether the point is inside or outside the mesh model, thus obtaining the actual signed distance value for that point. For example, for each sampled point, the shortest Euclidean distance from that point to the nearest facet of the meshed surface model is first calculated. Subsequently, the point's inside / outside assignment relative to the mesh model is determined using ray casting: if the point is inside the mesh model, the distance value is assigned a positive sign, resulting in a positive signed distance value; if the point is outside the mesh model, the distance value is assigned a negative sign, resulting in a negative signed distance value; if the point is exactly on the surface of the mesh model, its signed distance value is zero. Thus, the actual signed distance value of each sampled point simultaneously encodes both the geometric distance information to the surface and the spatial inside / outside assignment information.
[0068] Finally, a model training set is constructed based on the first spatial point and the corresponding actual signed distance value. The initial signed distance function network is trained using the model training set to obtain the scene model of the target scene.
[0069] In some embodiments, training an initial signed distance function network based on actual signed distance values to obtain a target signed distance function network includes: constructing a first training set based on the coordinates of a first spatial point and the corresponding actual signed distance value; inputting the coordinates of any first spatial point in the first training set into the initial signed distance function network to obtain a first predicted signed distance value; calculating a discriminative loss based on the first predicted signed distance value and the corresponding actual signed distance value in the first training set to obtain a first calculation result; and updating the model parameters of the initial signed distance function network using the first calculation result until the updated initial signed distance function network satisfies a first preset convergence condition to obtain the target signed distance function network.
[0070] For example, the three-dimensional coordinates of the first spatial point obtained by sampling in the aforementioned steps are used as input features, and the actual signed distance values calculated by each sampling point relative to the meshed surface model are used as supervision labels to jointly construct the first training set. The training set is then preprocessed by random shuffling and batch partitioning.
[0071] The initial signed distance function network adopts a fully connected multilayer perceptron structure with a position-encoded input layer, consisting of eight hidden layers, each containing 256 neurons. The layers are fully connected. The activation function of the hidden layers is the modified linear unit, i.e., the ReLU function, to introduce non-linear fitting capability. The output layer has no activation function to directly output the scalar distance value.
[0072] During training, in each iteration, a batch of 1024 training samples is randomly selected from the first training set. The coordinates of each first spatial point in this batch are encoden and then input into the initial signed distance function network. The first predicted signed distance value corresponding to each first spatial point is obtained through layer-by-layer forward propagation. The L1 discriminant loss between each first predicted value and the corresponding actual value in this batch is calculated, which is the average absolute error between the predicted and actual values. Based on this discriminant loss, the gradient of the weights of each layer of the network is calculated using the Adam optimizer with a current learning rate of 0.0001, and the weight parameters of each layer are updated through backpropagation. Every 5000 iterations, the learning rate is adjusted according to the exponential decay strategy multiplied by the decay rate of 0.96. The above iterative process of batch extraction, forward propagation, loss calculation, backpropagation, and parameter update is repeated, traversing the first training set round by round, until the average decrease of the discriminant loss in five consecutive training rounds is less than a preset value, or the number of iteration rounds reaches the maximum of 100 iteration rounds, and the first preset convergence condition is met, at which point training stops. The network obtained at this point is the target signed distance function network, which has solidified the macroscopic geometric prior information of the scene represented by the initial 3D model in the parameter space.
[0073] In some embodiments, the calculation of a discriminative loss based on the rendered color value, the rendered depth value, and the second predicted signed distance value of the second spatial point to obtain a second calculation result includes: calculating a discriminative loss based on the rendered color value and the real color value of the corresponding pixel in the second training set to obtain a color reconstruction loss; calculating a discriminative loss based on the rendered depth value and the corresponding monocular depth estimation result in the second training set to obtain a monocular depth loss; calculating a ray sparsity constraint loss based on the second predicted signed distance value of the second spatial point; and weighted summing the color reconstruction loss, the monocular depth loss, and the ray sparsity constraint loss to obtain the second calculation result.
[0074] In some embodiments, an initial 3D model of a target scene is generated based on a multi-view image sequence and the corresponding camera pose, including: acquiring the depth map of each frame in the multi-view image sequence; back-projecting the depth map of each frame to a 3D space based on the camera pose to obtain the point cloud of each frame; and stitching together the point clouds of each frame to obtain the initial 3D model.
[0075] For example, firstly, the RGB images of each frame in the multi-view image sequence are input into a pre-trained monocular depth estimation model to infer the dense depth map corresponding to each frame image; then, based on the camera intrinsic parameter matrix and extrinsic parameter pose corresponding to each frame image, the effective pixels in the depth map are back-projected point by point into the three-dimensional spatial coordinate system to generate single-frame point clouds under each viewpoint; finally, the single-frame point clouds under the above different viewpoints are accumulated, stitched and fused in a unified world coordinate system to remove redundant points and smooth noise, thereby constructing a dense initial three-dimensional point cloud that can reflect the macroscopic geometric contour and surface morphology of the target scene. This dense initial three-dimensional point cloud is the initial three-dimensional model of the target scene.
[0076] In some embodiments, after obtaining the depth map of each frame in the multi-view image sequence, the method further includes: obtaining the depth value of the pixels in the depth map; and removing pixels with depth values greater than or equal to a preset depth threshold. The preset depth threshold can be determined according to actual conditions and is not specifically limited here.
[0077] Specifically, after obtaining the depth map of each frame in the multi-view image sequence, the depth map needs to be filtered. Monocular depth estimation models exhibit significant uncertainty in predicting depth values for distant regions. Including such unreliable pixels in subsequent point cloud stitching introduces noise and affects the initial modeling accuracy. Therefore, a preset depth threshold is set, and pixels with depth values greater than or equal to this threshold are removed, retaining only valid pixels with depth values less than the threshold for subsequent processing.
[0078] In this way, the interference of unreliable depth values on subsequent point cloud stitching and initial modeling is avoided to a certain extent, the geometric purity of the initial 3D point cloud is improved, and a more reliable geometric prior foundation is provided for the pre-training of the signed distance function network.
[0079] Corresponding to the above embodiments, this application also proposes a scene modeling device.
[0080] Reference Figure 3 The scene modeling device 200 includes: an acquisition module 210, a first modeling module 220, a first training module 230, a second training module 240, and a second modeling module 250.
[0081] The acquisition module 210 is used to acquire a multi-view image sequence of the target scene and the camera pose of each view in the multi-view sequence, with each view corresponding to a different image sequence. The first modeling module 220 is used to generate an initial 3D model of the target scene based on the multi-view image sequence and the corresponding camera pose. The first training module 230 is used to train an initial signed distance function network based on the initial 3D model to obtain a target signed distance function network, wherein the target signed distance function network is used to calculate the signed distance value from any spatial point in the space where the initial 3D model is located to the surface of the target scene. The second training module 240 is used to train the initial scene modeling model based on the signed distance values provided by the target signed distance function network to obtain a target scene modeling model. The second modeling module 250 is used to model the target scene using the target scene modeling model to obtain a scene model of the target scene.
[0082] According to one embodiment of this application, the second training module 240 is specifically used to: construct a second training set based on the pixels of each frame in the multi-view image sequence, the corresponding real color values of the pixels, and the monocular depth estimation results; for any pixel in the second training set, determine multiple second spatial points in the three-dimensional space corresponding to any pixel according to the corresponding camera pose; input the coordinates of the second spatial points into the target signed distance function network to obtain the second predicted signed distance value corresponding to the second spatial point; and train the initial scene modeling model based on the second training set and the second predicted signed distance value to obtain the target scene modeling model.
[0083] According to one embodiment of this application, the second training module 240 is specifically used to: input the coordinates of the second spatial point into the initial scene modeling model to obtain the predicted color value corresponding to the second spatial point; generate the rendering depth value corresponding to any pixel based on the second predicted signed distance value, and generate the rendering color value corresponding to any pixel based on the second predicted signed distance value and the predicted color value; calculate the discrimination loss based on the rendering color value, the rendering depth value, and the second predicted signed distance value to obtain the second calculation result; and update the model parameters of the initial scene modeling model using the second calculation result until the updated initial scene modeling model satisfies the second preset convergence condition to obtain the target scene modeling model.
[0084] According to one embodiment of this application, the second training module 240 is further configured to: calculate a discrimination loss based on the rendered color value and the real color value of the corresponding pixel in the second training set to obtain a color reconstruction loss; calculate a discrimination loss based on the rendered depth value and the corresponding monocular depth estimation result in the second training set to obtain a monocular depth loss; calculate a ray sparsity constraint loss based on the second predicted signed distance value; and perform a weighted summation of the color reconstruction loss, monocular depth loss, and ray sparsity constraint loss to obtain a second calculation result.
[0085] According to one embodiment of this application, the first training module 230 is specifically used to: convert the initial 3D model into a meshed surface model and align the meshed surface model to the target coordinate system where the initial scene modeling model is located; sample multiple first spatial points in the target coordinate system and calculate the actual signed distance value between each first spatial point and the meshed surface model; and train the initial signed distance function network based on the actual signed distance value to obtain the target signed distance function network.
[0086] According to one embodiment of this application, the first training module 230 is specifically configured to: construct a first training set based on the coordinates of a first spatial point and the corresponding actual signed distance value; input the coordinates of any first spatial point in the first training set into an initial signed distance function network to obtain a first predicted signed distance value; calculate the discrimination loss based on the first predicted signed distance value and the corresponding actual signed distance value in the first training set to obtain a first calculation result; update the model parameters of the initial signed distance function network using the first calculation result until the updated initial signed distance function network satisfies a first preset convergence condition to obtain a target signed distance function network.
[0087] According to one embodiment of this application, the first modeling module 220 is specifically used to: acquire the depth map of each frame of the multi-view image sequence; backproject the depth map of each frame of the image to a three-dimensional space based on the camera pose to obtain the point cloud of each frame of the image; and stitch together the point clouds of each frame of the image to obtain an initial three-dimensional model.
[0088] It should be noted that the above explanation of the embodiments and beneficial effects of the scene modeling method also applies to the scene modeling apparatus of the embodiments of this application. To avoid redundancy, it will not be elaborated in detail here.
[0089] Corresponding to the above embodiments, this application also proposes an electronic device.
[0090] See Figure 4As shown, the electronic device 300 of this application includes a memory 310, a processor 320, and a scene modeling program stored in the memory 310 and capable of running on the processor 320. When the processor executes the scene modeling program, it implements the aforementioned scene modeling method.
[0091] It should be noted that the above explanation of the embodiments and beneficial effects of the electronic device also applies to the scenario modeling method of the embodiments of this application. To avoid redundancy, it will not be elaborated in detail here.
[0092] Corresponding to the above embodiments, this application also proposes a vehicle.
[0093] See Figure 5 As shown, the vehicle 30 of this application includes the aforementioned electronic device 300.
[0094] Corresponding to the above embodiments, this application also proposes a computer-readable storage medium.
[0095] The computer-readable storage medium of this application stores a scene modeling program thereon, which, when executed by a processor, implements the aforementioned scene modeling method.
[0096] It should be noted that the above explanation of the embodiments and beneficial effects of the scene modeling method also applies to the computer-readable storage medium of the embodiments of this application. To avoid redundancy, it will not be elaborated in detail here.
[0097] In this application, "multiple" refers to two or more.
[0098] In this application, unless otherwise expressly defined, the terms "installation," "connection," and "linking" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal connection between two components. Those skilled in the art can understand the specific meaning of the above terms in this application based on the specific circumstances.
[0099] The terms “first,” “second,” “third,” “fourth,” etc., in this application (if present) are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence.
[0100] In this application, the term "and / or" is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. Additionally, in this application, the character " / " generally indicates that the preceding and following related objects have an "or" relationship.
[0101] Unless otherwise specified, all steps in this application may be performed sequentially or randomly. For example, if the method includes steps A and B, it means that the method may include steps A and B performed sequentially, or it may include steps B and A performed sequentially. For example, if the method may also include step C, it means that step C may be added to the method in any order. For example, the method may include steps A, B, and C, or it may include steps A, C, and B, or it may include steps C, A, and B, etc.
[0102] The above description is merely a preferred embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this application should be included within the protection scope of this application.
Claims
1. A scene modeling method, characterized in that, include: Acquire a multi-view image sequence of the target scene and the camera pose of each view in the multi-view, wherein the multi-view corresponds to the multi-view image sequence; Based on the multi-view image sequence and the camera pose of each view, an initial 3D model of the target scene is generated; The initial signed distance function network is trained based on the initial 3D model to obtain the target signed distance function network, wherein the target signed distance function network is used to calculate the signed distance value from any spatial point in the space where the initial 3D model is located to the surface of the target scene. Based on the signed distance values provided by the target signed distance function network, the initial scene modeling model is trained to obtain the target scene modeling model; The target scene is modeled using the target scene modeling model to obtain the scene model of the target scene.
2. The scene modeling method according to claim 1, characterized in that, The step of training the initial scene modeling model based on the signed distance values provided by the target signed distance function network to obtain the target scene modeling model includes: A second training set is constructed based on the pixels of each frame in the multi-view image sequence, the corresponding true color values of the pixels, and the monocular depth estimation results; For any pixel in the second training set, multiple second spatial points in the three-dimensional space corresponding to the pixel are determined according to the corresponding camera pose. The coordinates of the second spatial point are input into the target signed distance function network to obtain the second predicted signed distance value corresponding to the second spatial point; The initial scene modeling model is trained based on the second training set and the second predicted signed distance value to obtain the target scene modeling model.
3. The scene modeling method according to claim 2, characterized in that, The step of training the initial scene modeling model based on the second training set and the second predicted signed distance value to obtain the target scene modeling model includes: The coordinates of the second spatial point are input into the initial scene modeling model to obtain the predicted color value corresponding to the second spatial point; The rendering depth value corresponding to any pixel is generated based on the second predicted signed distance value, and the rendering color value corresponding to any pixel is generated based on the second predicted signed distance value and the predicted color value. The discrimination loss is calculated based on the rendered color value, the rendered depth value, and the second predicted signed distance value to obtain the second calculation result; The model parameters of the initial scene modeling model are updated using the second calculation result until the updated initial scene modeling model satisfies the second preset convergence condition, thereby obtaining the target scene modeling model.
4. The scene modeling method according to claim 3, characterized in that, The calculation of the discrimination loss based on the rendered color value, the rendered depth value, and the second predicted signed distance value to obtain the second calculation result includes: The color reconstruction loss is obtained by calculating the discrimination loss based on the rendered color value and the real color value of the corresponding pixel in the second training set. The discrimination loss is calculated based on the rendered depth value and the corresponding monocular depth estimation results in the second training set to obtain the monocular depth loss; Calculate the ray sparsity constraint loss based on the second predicted signed distance value; The color reconstruction loss, the monocular depth loss, and the ray sparse constraint loss are weighted and summed to obtain the second calculation result.
5. The scene modeling method according to claim 1, characterized in that, The step of training the initial signed distance function network based on the initial 3D model to obtain the target signed distance function network includes: The initial 3D model is converted into a meshed surface model, and the meshed surface model is aligned to the target coordinate system of the initial scene modeling model. Multiple first spatial points are sampled within the target coordinate system, and the actual signed distance value between each first spatial point and the meshed surface model is calculated respectively. The initial signed distance function network is trained based on the actual signed distance values to obtain the target signed distance function network.
6. The scene modeling method according to claim 5, characterized in that, The step of training the initial signed distance function network based on the actual signed distance values to obtain the target signed distance function network includes: The first training set is constructed based on the coordinates of the first spatial point and the corresponding actual signed distance value. The coordinates of any first spatial point in the first training set are input into the initial signed distance function network to obtain the first predicted signed distance value. The discrimination loss is calculated based on the first predicted signed distance value and the corresponding actual signed distance value in the first training set to obtain the first calculation result; The model parameters of the initial signed distance function network are updated using the first calculation result until the updated initial signed distance function network satisfies the first preset convergence condition, thereby obtaining the target signed distance function network.
7. The scene modeling method according to claim 1, characterized in that, The step of generating an initial 3D model of the target scene based on the multi-view image sequence and the corresponding camera pose includes: Obtain the depth map of each frame in the multi-view image sequence; Based on the camera pose, the depth map of each frame is back-projected into three-dimensional space to obtain the point cloud of each frame. The point clouds of each frame of the image are stitched together to obtain the initial three-dimensional model.
8. A scene modeling device, characterized in that, include: The acquisition module is used to acquire a multi-view image sequence of the target scene and the camera pose of each view in the multi-view, wherein the multi-view corresponds to the multi-view image sequence; The first modeling module is used to generate an initial 3D model of the target scene based on the multi-view image sequence and the camera pose of each view. The first training module is used to train the initial signed distance function network based on the initial 3D model to obtain the target signed distance function network, wherein the target signed distance function network is used to calculate the signed distance value from any spatial point in the space where the initial 3D model is located to the surface of the target scene; The second training module is used to train the initial scene modeling model based on the signed distance value provided by the target signed distance function network to obtain the target scene modeling model. The second modeling module is used to model the target scene using the target scene modeling model to obtain a scene model of the target scene.
9. An electronic device, characterized in that, The method includes a memory, a processor, and a scene modeling program stored in the memory and capable of running on the processor. When the processor executes the scene modeling program, it implements the scene modeling method according to any one of claims 1-7.
10. A vehicle, characterized in that, Including the electronic device as described in claim 9.