A road surface new perspective reconstruction method, device, equipment and storage medium

By combining multi-view image and three-dimensional point cloud, using MLP network and Gaussian spherical optimization technology, the problem of low reconstruction quality caused by perspective transformation in three-dimensional scene reconstruction is solved, and a new perspective reconstruction of pavement with higher accuracy and efficiency is achieved.

CN120182514BActive Publication Date: 2025-08-08CHONGQING CHANGAN AUTOMOBILE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510663780.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-22
Publication Date
2025-08-08
Estimated Expiration
2045-05-22

AI Technical Summary

Technical Problem

The prior art has the problem of low reconstruction quality when reconstructing new views of three-dimensional scenes, especially when dealing with viewing angle transformation, the texture is blurred and the details are difficult to maintain, and the inaccurate depth information leads to object geometric distortion, affecting the realism of rendering.

Method used

By obtaining multi-view pavement image sequences and three-dimensional point clouds, the point cloud position features are encoded using a multi-layer perceptron (MLP) network model, a Gaussian sphere is generated and its attribute information is iteratively optimized, combining pavement image sequences and symbol distance functions until the optimization conditions are met, and a high-quality rendered image sequence is generated.

Benefits of technology

It improves the accuracy and efficiency of new perspective reconstruction of pavement, reduces the differences between reconstruction results and actual scenes, generates more realistic rendered images, and meets the needs of high-quality scene reconstruction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120182514B_ABST
    Figure CN120182514B_ABST
Patent Text Reader

Abstract

The present application discloses a method, apparatus, device and storage medium for reconstructing a road surface from a new perspective. The method comprises: obtaining a road surface image sequence from multiple perspectives, and a three-dimensional point cloud corresponding to the road surface image sequence; encoding the position of each feature point in the three-dimensional point cloud to obtain a point cloud position feature of the three-dimensional point cloud; inputting the point cloud position feature into a trained multi-layer perceptron (MLP) network model for feature processing to obtain a road surface signed distance function for each feature point in the three-dimensional point cloud; generating a Gaussian sphere at the position of each feature point in the three-dimensional point cloud and initializing the attribute information of the Gaussian sphere; iteratively optimizing the attribute information of the Gaussian sphere based on the road surface image sequence and the road surface signed distance function of each feature point in the three-dimensional point cloud until a stopping condition for the Gaussian sphere optimization is reached; and using the rendered image sequence corresponding to the attribute information of the Gaussian sphere when the iterative optimization stops as the image sequence for reconstructing the road surface from a new perspective.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer vision technology, and in particular to a method, device, equipment, and storage medium for reconstructing a road surface from a new perspective. Background Art

[0002] 3D scene reconstruction and new perspective synthesis combine images from different perspectives to reconstruct a 3D scene, achieving realistic image synthesis of the same object from different viewpoints. This technology plays a vital role in a variety of application scenarios, including virtual reality (VR), autonomous driving, intelligent transportation, urban planning, and gaming.

[0003] In related technologies, when reconstructing a new view of a three-dimensional scene through methods such as Neural Radiance Fields (NeRF) and 3D Gaussian Splatting (3DGS), there are problems such as low reconstruction quality. Summary of the Invention

[0004] One of the purposes of this application is to provide a method, device, equipment and storage medium for reconstructing a new perspective of a road surface to solve the problem of low reconstruction quality when reconstructing a new view of a three-dimensional scene in related technologies.

[0005] In order to achieve the above-mentioned purpose, the technical solution of the embodiment of the present application is implemented as follows:

[0006] The technical solution of the embodiment of the present application is implemented as follows:

[0007] In a first aspect, the present application provides a method for reconstructing a road surface from a new perspective, the method comprising:

[0008] Obtain multi-view road surface image sequences and the three-dimensional point clouds corresponding to the road surface image sequences;

[0009] Encode the position of each feature point in the three-dimensional point cloud to obtain the point cloud position feature of the three-dimensional point cloud;

[0010] The point cloud position features are input into the trained Multilayer Perceptron (MLP) network model for feature processing to obtain the road surface signed distance function for each feature point in the 3D point cloud.

[0011] Generate a Gaussian sphere at the position of each feature point in the 3D point cloud and initialize the attribute information of the Gaussian sphere;

[0012] Based on the road surface signed distance function of each feature point in the road surface image sequence and the 3D point cloud, the attribute information of the Gaussian sphere is iteratively optimized until the Gaussian sphere optimization stop condition is reached;

[0013] The rendered image sequence generated corresponding to the attribute information of the Gaussian sphere when the iterative optimization stops is used as the image sequence for reconstructing the road surface from a new perspective.

[0014] The above technical approach first acquires a multi-view road image sequence and its corresponding 3D point cloud. The image information is then combined with the point cloud data, enabling subsequent processing to comprehensively describe the road scene, improving the accuracy and realism of reconstruction and providing a more reliable scene foundation for applications such as autonomous driving and virtual reality. Secondly, the positions of feature points in the 3D point cloud are encoded and input into a trained MLP network model. The MLP network automatically learns and extracts the complex relationship between point cloud position features and the road surface signed distance function, transforming raw point cloud data into key features. This not only reduces data redundancy but also accurately calculates the road surface signed distance function for each feature point, providing a crucial basis for subsequent scene modeling and rendering. Next, a Gaussian sphere is generated at each feature point in the 3D point cloud and its attributes are initialized. The Gaussian sphere can effectively approximate local regions in space. By appropriately setting its attributes, it can effectively model road scenes of varying complexity. This representation facilitates subsequent optimization and rendering operations, helping to improve reconstruction accuracy and efficiency. Finally, the attribute information of the Gaussian sphere is iteratively optimized based on the road surface image sequence and the road surface signed distance function until the stopping condition is met. Through multiple iterations, the attributes of the Gaussian sphere are continuously adjusted to make it more consistent with the real road scene. In this way, the difference between the reconstruction result and the actual scene can be gradually reduced, the accuracy of the new perspective reconstruction of the road surface can be improved, and a more realistic rendering image sequence can be generated to meet the needs of practical applications for high-quality scene reconstruction.

[0015] In a second aspect, the present application provides a device for reconstructing a road surface from a new perspective, the device comprising:

[0016] An acquisition module is used to acquire a multi-view road surface image sequence and a three-dimensional point cloud corresponding to the road surface image sequence;

[0017] An encoding module is used to encode the position of each feature point in the three-dimensional point cloud to obtain the point cloud position feature of the three-dimensional point cloud;

[0018] The first processing module is used to input the point cloud position features into the trained MLP network model for feature processing to obtain the road surface signed distance function of each feature point in the three-dimensional point cloud;

[0019] A generation module is used to generate a Gaussian sphere at the position of each feature point of the three-dimensional point cloud and initialize the attribute information of the Gaussian sphere;

[0020] The second processing module is used to iteratively optimize the attribute information of the Gaussian sphere based on the road surface image sequence and the road surface signed distance function of each feature point in the three-dimensional point cloud until the Gaussian sphere optimization stopping condition is reached; and the rendered image sequence corresponding to the attribute information of the Gaussian sphere when the iterative optimization stops is used as the image sequence reconstructed from the new perspective of the road surface.

[0021] In a third aspect, the present application provides an electronic device, the electronic device comprising:

[0022] at least one processor; and,

[0023] A memory communicatively connected to at least one processor; wherein the memory stores a computer program executable by at least one processor, and the computer program is executed by at least one processor to implement some or all of the steps in the method for reconstructing a road surface from a new perspective as described in any one of the first aspects.

[0024] In a fourth aspect, the present application provides a computer-readable storage medium, which stores one or more computer programs, and the one or more computer programs can be executed by one or more processors to implement some or all of the steps in the method for reconstructing a road surface from a new perspective as described in any one of the first aspects.

[0025] In a fifth aspect, the present application provides a computer program product, comprising a computer program or instructions, which, when executed by a processor, implements some or all of the steps in the method for reconstructing a road surface from a new perspective as described in any one of the first aspects.

[0026] Beneficial technical effects of the embodiments of the present application:

[0027] The MLP network learns the normal vector and depth information of the point cloud, enhancing the learning ability of geometric information. After the iterative training of the MLP network is completed, the MLP network outputs the road surface signed distance function of each point to accurately describe the distance between the point cloud and the road surface, and initializes the point cloud into a series of Gaussian spheres. During the Gaussian sphere attribute optimization process, the rendering loss function is calculated based on the road surface image sequence, the rendered image sequence, and the road surface signed distance function of each point in the point cloud. The Gaussian sphere attributes are optimized based on the rendering loss function. In this way, normal constraints (gradient of the road surface signed distance function) and depth constraints (road surface signed distance function) are added to the Gaussian sphere, so that the rendered image retains more details and texture. At the same time, the depth constraint is also beneficial to improve the depth distortion and perspective distortion problems of the rendered image. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] The drawings herein are incorporated into and constitute a part of the specification. These drawings illustrate embodiments consistent with the present application and, together with the specification, are used to illustrate the technical solutions of the present application.

[0029] Figure 1 A schematic diagram of a process for reconstructing a road surface from a new perspective provided in an embodiment of the present application Figure 1 ;

[0030] Figure 2 A schematic diagram of the processing process of an optional method for reconstructing a road surface from a new perspective provided in an embodiment of the present application;

[0031] Figure 3 A schematic diagram of the structure of an optional MLP network model provided in an embodiment of the present application;

[0032] Figure 4 A schematic diagram comparing the reconstruction effect of the new road perspective reconstruction method provided by the embodiment of the present application with the reconstruction effect of the existing 3DGS;

[0033] Figure 5 A schematic diagram of a process for reconstructing a road surface from a new perspective provided in an embodiment of the present application Figure 2 ;

[0034] Figure 6 A schematic diagram of a process for reconstructing a road surface from a new perspective provided in an embodiment of the present application Figure 3 ;

[0035] Figure 7 A schematic diagram of a process for reconstructing a road surface from a new perspective provided in an embodiment of the present application Figure 4 ;

[0036] Figure 8 A schematic structural diagram of an optional device for reconstructing a road surface from a new perspective provided in an embodiment of the present application;

[0037] Figure 9 A schematic structural diagram of an optional electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0038] In order to make the purpose, technical solutions and advantages of this application clearer, the technical solutions of this application are further elaborated in detail below with reference to the accompanying drawings and embodiments. The described embodiments should not be regarded as limiting this application. All other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.

[0039] In the following description, references to "some embodiments" describe a subset of all possible embodiments. However, it is understood that "some embodiments" may be the same subset or different subsets of all possible embodiments, and may be combined with each other without conflict. The terms "first / second / third" are merely used to distinguish similar objects and do not represent a specific ordering of the objects. It is understood that the specific order or sequence of "first / second / third" may be interchanged where permitted, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.

[0040] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this application pertains. The terms used herein are for the purpose of describing this application only and are not intended to limit this application.

[0041] Three-dimensional scene reconstruction and novel view synthesis are long-standing and complex tasks in computer vision, aiming to recover the three-dimensional structure of a scene from a set of two-dimensional images. This technology plays a vital role in a variety of applications, including VR and autonomous driving. In recent years, implicit representation methods, especially neural radiance fields and their derivative models, have made significant progress in generating realistic images from arbitrary viewpoints. Although NeRF excels in creating realistic images, there is a growing demand for faster and more efficient rendering methods, especially for applications requiring real-time performance. 3DGS technology is a new three-dimensional scene representation and view synthesis method proposed in recent years. 3DGS explicitly represents the scene using millions of three-dimensional Gaussian functions, which are projected onto the image plane during rendering. Compared with implicit representation methods, 3DGS takes advantage of the differentiable rendering pipeline and point-based rendering techniques to provide an efficient computation and rendering process, avoiding the high computational cost of traditional NeRF methods.

[0042] While 3DGS has made significant progress in rendering efficiency and scene dynamics, most roads are curved, and the perspective changes (i.e., new perspective synthesis) involved when capturing images. Existing 3DGS technology, when handling perspective changes, especially those at large angles, can easily blur road surface textures and make it difficult to maintain clear details. This is especially true in distant areas or where point cloud data is sparse, where original details are easily lost, affecting the realism of the rendering. Furthermore, during the multi-perspective synthesis process, the lack of or inaccurate depth information causes the geometry of objects such as road markings to become distorted or lose depth in the new perspective, affecting the perspective effect of the scene.

[0043] To address one or more of the issues mentioned in the related art, embodiments of the present application provide a new perspective reconstruction method for roads. First, a multi-perspective road image sequence and its corresponding 3D point cloud are acquired. The image information is combined with the point cloud data, enabling subsequent processing to comprehensively describe the road scene, improving the accuracy and realism of the reconstruction and providing a more reliable scene foundation for applications such as autonomous driving and virtual reality. Second, the positions of feature points in the 3D point cloud are encoded and input into a trained MLP network model. The MLP network automatically learns and extracts the complex relationship between point cloud position features and the road surface signed distance function, transforming raw point cloud data into key features. This not only reduces data redundancy but also accurately calculates the road surface signed distance function for each feature point, providing an important basis for subsequent scene modeling and rendering. Next, a Gaussian sphere is generated at each feature point in the 3D point cloud and its attribute information is initialized. The Gaussian sphere can effectively approximate local regions in space. By properly setting its attributes, it can effectively model road scenes of varying complexity. This representation facilitates subsequent optimization and rendering operations, helping to improve the accuracy and efficiency of reconstruction. Finally, the attribute information of the Gaussian sphere is iteratively optimized based on the road surface image sequence and the road surface signed distance function until the stopping condition is met. Through multiple iterations, the attributes of the Gaussian sphere are continuously adjusted to make it more consistent with the real road scene. In this way, the difference between the reconstruction result and the actual scene can be gradually reduced, the accuracy of the new perspective reconstruction of the road surface can be improved, and a more realistic rendering image sequence can be generated to meet the needs of practical applications for high-quality scene reconstruction.

[0044] The method for reconstructing a road surface from a new perspective provided in the embodiments of the present application can be executed by an electronic device, which can be a laptop, tablet computer, desktop computer, in-vehicle device, set-top box, mobile device (e.g., mobile phone, portable music player, personal digital assistant, dedicated messaging device, portable gaming device), or other terminal. The electronic device can also be implemented as a server. The server can be a standalone physical server, a server cluster or distributed system consisting of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms.

[0045] Below, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the drawings in the embodiments of the present application.

[0046] Figure 1 A schematic diagram of an implementation flow of an optional method for reconstructing a road surface from a new perspective provided in an embodiment of the present application is shown as follows: Figure 1As shown, the method may include the following steps:

[0047] Step 101: Acquire a multi-view road surface image sequence and a three-dimensional point cloud corresponding to the road surface image sequence.

[0048] In embodiments of the present application, a road surface image sequence may include multiple road surface images. The multiple road surface images may be one or more images captured by different cameras at different camera angles. For example, the multiple road surface images may be captured by cameras at different locations on a vehicle, such as the front, rear, left, and right sides of the vehicle, respectively. There may be overlapping areas between the multiple road surface images.

[0049] In embodiments of the present application, the 3D point cloud corresponding to the road surface image sequence may include at least one feature point. The 3D point cloud corresponding to the road surface image sequence may be acquired using any suitable method. For example, a 3D point cloud corresponding to the area in the road surface image sequence may be acquired using a LiDAR or a Red, Green, Blue-Depth (RGB-D) camera. Another example is the estimation of the corresponding 3D point cloud based on the road surface image sequence using Structure from Motion (SFM) technology. Another example is the generation of the 3D point cloud using a generative model. The generative model may be any suitable model capable of generating point clouds, such as a neural network model. In some embodiments, the 3D cloud may include road surface point clouds and non-road surface point clouds. Non-road surface may include, but is not limited to, sky and trees. Non-road surface point clouds may include, but are not limited to, sky point clouds and tree point clouds. Road surface may include roads, lane markings, road signs, etc., and the road surface point cloud may be further divided into road point clouds, lane marking point clouds, and road sign point clouds. In some embodiments, the 3D point cloud may be divided into multiple pre-defined categories to obtain point clouds of each category. During implementation, the definition, division and number of categories can be set independently according to the actual scenario, and are not limited in the embodiments of the present application.

[0050] In some embodiments, the electronic device can also obtain the corresponding posture data when each camera shoots a sequence of road images. The posture data includes the external parameters and internal parameters of the camera. The external parameters and internal parameters of the camera are used to represent the viewing angle information of the image taken by the camera. The same viewing angle may include more than one road image.

[0051] In the embodiments of this application, the electronic device acquires a multi-view road image sequence and the corresponding 3D point cloud, fusing two different types of data: image information and point cloud information. The multi-view image sequence contains rich visual information such as texture and color, while the 3D point cloud provides precise spatial position and geometric structure information. This combination enables subsequent processing to more comprehensively describe the road scene, laying the foundation for accurate reconstruction.

[0052] Step 102: Encode the position of each feature point in the three-dimensional point cloud to obtain the point cloud position feature of the three-dimensional point cloud.

[0053] In an embodiment of the present application, since it is necessary to subsequently generate a Gaussian sphere for the position of each feature point in the three-dimensional point cloud, encoding the position of each feature point in the three-dimensional point cloud is equivalent to encoding the center coordinates of the Gaussian sphere, thereby obtaining a high-dimensional point cloud position feature after the three-dimensional point cloud position is encoded.

[0054] In some embodiments, step 102 encodes the position of each feature point in the three-dimensional point cloud, and obtaining the point cloud position features of the three-dimensional point cloud can be achieved in the following manner: through a position encoding function, the position of each feature point in the three-dimensional point cloud is converted into a high-dimensional vector to obtain multiple high-dimensional vectors; and the multiple high-dimensional vectors are fused to obtain point cloud position features with high dimensions.

[0055] In the embodiment of the present application, since it is necessary to generate a Gaussian sphere for the position of each feature point in the three-dimensional point cloud, encoding the position of each feature point in the three-dimensional point cloud is equivalent to encoding the center coordinates of the Gaussian sphere. The electronic device can convert the position of each feature point in the three-dimensional point cloud into a high-dimensional vector through the position encoding function. Figure 2 As shown, the position encoding function is used to encode the position of each feature point in the three-dimensional point cloud 201, which is equivalent to encoding the center coordinates of all Gaussian spheres, such as Figure 2 The dotted box 202 in the figure is converted into a high-dimensional vector, thereby obtaining multiple high-dimensional vectors; the multiple high-dimensional vectors are concatenated to obtain high-dimensional point cloud position features, so that the point cloud position features can be subsequently input into the MLP network model 203 to capture spatial features. The position encoding function can be expressed by the following formula (1).

[0056]

[0057] in, is the point cloud position feature of the three-dimensional point cloud, is the position vector of the feature point to be encoded in the three-dimensional point cloud, is the frequency of positional encoding, Determines the dimension of the encoded feature vector, each position vector Will be converted to a length of The eigenvector of . The larger the value of , the higher the dimension of the encoded feature vector, and the model can capture more detailed location information, but it will also increase the amount of calculation and the complexity of the model.

[0058] Step 103: Input the point cloud position features into the trained multi-layer perceptron (MLP) network model for feature processing to obtain a road surface signed distance function for each feature point in the three-dimensional point cloud.

[0059] As you can understand, the signed distance function (SDF) can be understood as determining the distance from a point to the boundary of a finite region in space and defining the sign of the distance. For example, if the point is inside the region boundary, the signed distance function is positive; if the point is outside the region boundary, the signed distance function is negative; and if the point is on the region boundary, the signed distance function is zero, a scalar.

[0060] In an embodiment of the present application, the road surface signed distance function can be that, for the position of each feature point in the three-dimensional point cloud, if the position of the feature point is inside the road surface boundary, the road surface signed distance function is positive; if the feature point is outside the road surface boundary, the road surface signed distance function is negative; if the feature point is on the road surface boundary, the road surface signed distance function is 0.

[0061] In the embodiment of the present application, the MLP network model is a basic and powerful artificial neural network model that can learn the complex nonlinear relationship between point cloud position features and road surface signed distance functions. It should be noted that the relationship between point cloud position features and road surface signed distance functions is often not a simple linear relationship. For example, complex geometric shapes such as undulations and bends in the road surface can make this relationship complicated. The nonlinear mapping capability of MLP can capture these complex patterns, thereby accurately mapping point cloud position features to the corresponding road surface signed distance functions.

[0062] In one implementation, the MLP network module consists of at least two fully connected layers, each layer containing multiple neurons. In order to better learn scene information, such as Figure 3As shown, the position of each feature point in the three-dimensional point cloud (equivalent to the position of the center point of the Gaussian sphere) 301 is encoded to obtain high-dimensional point cloud position features, which are then input into the MLP network model 302 as input parameters. The hidden layer of the entire MLP network model introduces nonlinear characteristics through nonlinear activation functions to enhance the network's expressive power. Batch Normalization (BN) and Softplus activation functions are used in the MLP network model to introduce nonlinearity, improve training stability, and accelerate convergence. The output of the MLP network model is the road surface signed distance function for each feature point in the three-dimensional point cloud. .

[0063] In an embodiment of the present application, the electronic device encodes the position of each feature point in the three-dimensional point cloud to obtain the point cloud position features of the three-dimensional point cloud, and inputs the point cloud position features into the trained multi-layer perceptron MLP network model for feature processing to obtain the road surface signed distance function of each feature point in the three-dimensional point cloud. In this way, encoding the positions of the feature points in the three-dimensional point cloud can convert the spatial position information of the point cloud into a feature representation that is more suitable for model processing. Furthermore, the point cloud position features are input into the trained MLP network model, and the powerful feature learning ability of the MLP is used to extract features related to the road surface signed distance function, thereby accurately calculating the road surface signed distance function of each feature point. This helps to accurately describe the relative positional relationship between the point cloud and the road surface, which is crucial to the accuracy of road surface reconstruction.

[0064] Step 104: Generate a Gaussian sphere at the position of each feature point of the three-dimensional point cloud and initialize the attribute information of the Gaussian sphere.

[0065] In the embodiment of the present application, in the process of initializing each feature point of the three-dimensional point cloud, each feature point in the three-dimensional point cloud can be initialized as a three-dimensional Gaussian sphere (also known as a Gaussian ellipsoid), and the position of the center point of the Gaussian sphere is the position of the corresponding feature point, such as Figure 2 204 of them.

[0066] In the embodiment of the present application, each Gaussian sphere has attribute information, including the center position, covariance matrix, color, and transparency of the Gaussian sphere. Of course, the attribute information of the Gaussian sphere may also include spherical harmonic coefficients of spherical harmonic basis functions. The covariance matrix may be composed of a shape scaling matrix and a rotation matrix, and the observed color of the Gaussian sphere in each direction may be represented by a set of spherical harmonic basis functions.

[0067] In an embodiment of the present application, the electronic device generates a Gaussian sphere at each feature point position and initializes its attribute information, providing a flexible and effective representation method for road scenes. The Gaussian sphere can well approximate the local area in space, and by reasonably initializing its attribute information, it can provide a feasible starting point for subsequent optimization.

[0068] Step 105 : Iteratively optimize the attribute information of the Gaussian sphere based on the road surface image sequence and the road surface signed distance function of each feature point in the three-dimensional point cloud until the Gaussian sphere optimization stop condition is reached.

[0069] In an embodiment of the present application, the Gaussian sphere optimization stopping condition may include that the loss value of the iterative optimization of the attribute information of the Gaussian sphere is less than or equal to a first loss threshold, or that the number of optimization iterations of the attribute information of the Gaussian sphere reaches a preset maximum optimization number, or that the reduction rate of the loss value of the iterative optimization of the attribute information of the Gaussian sphere is less than or equal to a first preset reduction rate threshold.

[0070] In an embodiment of the present application, during the iterative optimization process of the attribute information of the Gaussian sphere, each time the optimization is performed, the second target loss is calculated based on the attribute information of the Gaussian sphere, the road surface image sequence, and the road surface signed distance function of each feature point in the three-dimensional point cloud. If the Gaussian sphere optimization stopping condition is reached, the subsequent step 106 is executed, and the rendered image sequence corresponding to the attribute information of the Gaussian sphere when the iterative optimization is stopped is used as the image sequence for reconstructing the road surface from a new perspective. If the Gaussian sphere optimization stopping condition is not reached, the attribute information of the Gaussian sphere is optimized, and the next optimization is entered. It should be noted that the attribute information of the Gaussian sphere can be optimized using, but not limited to, the gradient descent method and the Newton method. When using the gradient descent method to optimize the attribute information of the Gaussian sphere, the partial derivative of the loss function with respect to each attribute of the Gaussian sphere can be solved to obtain the attribute gradient, and the corresponding attribute value can be updated according to the learning rate and the attribute gradient.

[0071] Step 106: Use the rendered image sequence generated corresponding to the attribute information of the Gaussian sphere when the iterative optimization is stopped as the image sequence reconstructed from the new perspective of the road surface.

[0072] It is understandable that the rendering methods in related technologies often show problems such as loss of geometric information, blurred object contours, inaccurate depth perception, etc. under new perspectives, especially when it is necessary to accurately present surface details (such as lane lines, road surfaces). To solve this problem, the embodiment of the present application introduces a road surface signed distance function to obtain the distance and position between each feature point and the road surface boundary, and uses these rules as the true value. During iterative optimization, a road surface signed distance constraint is introduced to restrict each Gaussian sphere. This ensures that the Gaussian distribution can learn and maintain the accurate geometric information of the point cloud during the reconstruction process, so that after switching the perspective, the Gaussian sphere will not be misplaced or the direction will be confused, making the road surface clearer in the new perspective. Figure 4 This figure shows the difference between the newly reconstructed images obtained by the newly reconstructed road surface perspective method provided by the embodiment of this application and the rendered images obtained by the 3DGS algorithm. The left image is a sequence of road surface images, the middle image is a sequence of rendered images obtained by the 3DGS algorithm, and the right image is a sequence of newly reconstructed images obtained by the method of the embodiment of this application, i.e., a sequence of rendered images. It can be seen that the reconstruction accuracy of the middle image (especially the selected portion) is significantly lower than that of the right image.

[0073] As can be seen from the above, the embodiments of the present application iteratively optimize the attribute information of the Gaussian sphere based on the road surface image sequence and the road surface signed distance function, so that the Gaussian sphere representation can continuously approximate the actual road surface scene. In this way, the various attributes of the Gaussian sphere can be gradually adjusted to better fit the actual road surface texture, geometry, and other information. Until the optimization condition is reached, the resulting Gaussian sphere attribute information can be ensured to accurately reflect the road surface scene, thereby generating a high-quality rendered image sequence and accurately reconstructing the road surface from a new perspective.

[0074] In some embodiments, the training process of the MLP network model in step 103 is combined with Figure 5 Provide explanation.

[0075] Step 501: Obtain sample data.

[0076] The sample data includes sample point cloud position features of the sample three-dimensional point cloud.

[0077] In an embodiment of the present application, before obtaining sample data, the electronic device may also pre-acquire a multi-perspective sample road surface image sequence and obtain a sample three-dimensional point cloud corresponding to the sample road surface image sequence. Furthermore, the position of each feature point in the sample three-dimensional point cloud is encoded to obtain the sample point cloud position feature of the sample three-dimensional point cloud, thereby obtaining sample data.

[0078] Step 502: Input the sample point cloud position features into the MLP network model to be trained for feature processing to obtain a predicted road surface signed distance function for each feature point in the sample three-dimensional point cloud output by the MLP network model to be trained.

[0079] In an embodiment of the present application, after the electronic device obtains the sample data, the sample point cloud position features of the sample three-dimensional point cloud in the sample data are input into the MLP network model to be trained, and the sample point cloud position features are feature processed using the MLP network model to be trained to obtain the predicted road surface signed distance function for each feature point in the sample three-dimensional point cloud.

[0080] Step 503: Determine a first target loss based on the predicted road surface signed distance function of each feature point in the sample three-dimensional point cloud.

[0081] In this embodiment of the present application, the first target loss (LOSS) is used to adjust the network parameters of the MLP network model. Here, step 503 determines the first target loss based on the predicted road surface signed distance function of each feature point in the sample 3D point cloud, which can be achieved by the following steps:

[0082] Step 531: Obtain the true normal vector of each feature point in the sample 3D point cloud.

[0083] In the present embodiment, the true normal vector of each feature point in the sample 3D point cloud can be obtained using a nearest neighbor algorithm. In one embodiment, a set of neighborhood points is obtained for each feature point, an approximate tangent plane for the feature point is constructed based on the neighborhood point set, and a perpendicular vector to the approximate tangent plane is obtained, which is used as the true normal vector for the feature point.

[0084] Step 532: Determine a predicted normal vector for each feature point in the sample three-dimensional point cloud based on the predicted road surface signed distance function of each feature point in the sample three-dimensional point cloud.

[0085] In the embodiment of the present application, the predicted normal vector may be obtained by performing a differential calculation, such as a derivative calculation, on the predicted road surface signed distance function.

[0086] In an embodiment of the present application, the electronic device can perform differential calculation on the predicted road surface signed distance function of each feature point in the sample three-dimensional point cloud, thereby obtaining a predicted normal vector corresponding to the feature point in the sample three-dimensional point cloud.

[0087] Step 533: Determine a first sub-loss based on the predicted normal vector of each feature point.

[0088] Here, the first sub-loss is used to characterize the Eikonal loss for surface smoothness. Eikonal loss is a key constraint in regularized neural networks. In tasks dealing with geometry, Eikonal loss is often used to ensure the smoothness of implicitly defined surfaces and to make the geometric information learned by the network more consistent and stable. Using Eikonal loss to regularize surface smoothness ensures that the object's geometric structure is not significantly distorted at long distances or wide angles, further reducing distortion.

[0089] Here, the first sub-loss may be determined by, but is not limited to, a first difference, an average of the first differences, an average of the squares of the first differences, etc. The first difference refers to the difference between the norm of the predicted normal vector of the feature point and a first value, such as 1. For example, the average of the sum of the squares of the first differences is used as the first sub-loss value.

[0090] In some embodiments, the first sub-loss may be determined by an Eikonal loss function, where the Eikonal loss function may be represented by the following formula (2).

[0091]

[0092] in, is the first child loss, also known as Eikonal loss, represents the sample 3D point cloud, represents the feature points in the sample 3D point cloud, Representing feature points The predicted road surface signed distance function; Representing feature points Prediction road surface signed distance function The gradient of is used to represent the feature points The predicted normal vector is usually calculated by automatic differentiation in the network; Indicates taking the norm, hoping the norm of the gradient is close to 1 to satisfy the constraints of the Eikonal equation, Represents the expected value function.

[0093] Step 534: Determine a second sub-loss based on the predicted normal vector of each feature point and the corresponding true normal vector.

[0094] Here, the second sub-loss may be a normal vector loss, and the second sub-loss may be determined by, but is not limited to, a certain second difference, an average of the second differences, an average of the squares of the second differences, and the like. The second difference may be determined by calculating a first vector product between the predicted normal vector of the feature point and the corresponding true normal vector, calculating a first product between the norm of the predicted normal vector of the feature point and the norm of the corresponding true normal vector, calculating a first quotient between the first vector product and the first product, and using the difference between a second value, such as 1, and the first quotient as the second difference. For example, all second differences may be averaged, and the average value may be used as the second sub-loss value.

[0095] In some embodiments, the second sub-loss may be determined by a normal vector loss function, where the normal vector loss function may be expressed by the following formula (3).

[0096]

[0097] in, is the second sub-loss, also known as the normal vector loss, represents the sample 3D point cloud, represents the feature points in the sample 3D point cloud, Representing feature points The predicted road surface signed distance function; Representing feature points Prediction road surface signed distance function The gradient of is used to represent the feature points The predicted normal vector is usually calculated by automatic differentiation in the network; represents the norm, Representing feature points The true normal vector of Represents the expected value function.

[0098] It should be noted that the normal vector loss function is calculated by predicting the road surface sign distance function The predicted normal vector calculated by derivative and the corresponding true normal vector The loss between can enhance the geometric learning ability of the MLP network.

[0099] Step 535 : Determine the third sub-loss based on the predicted road surface signed distance function of each feature point.

[0100] Here, the third sub-loss may be an expected SDF loss. Methods for determining the third sub-loss include, but are not limited to, the absolute value of a predicted road surface signed distance function and the average of the absolute values. For example, the average of the absolute values of the predicted road surface signed distance function for all feature points may be used as the third sub-loss value.

[0101] In some embodiments, the third sub-loss may be determined by an SDF expected loss function, where the SDF expected loss function may be expressed by the following formula (4).

[0102]

[0103] in, is the third sub-loss, also known as SDF expected loss, represents the sample 3D point cloud, represents the feature points in the sample 3D point cloud, Representing feature points The predicted road surface signed distance function; Representing feature points Prediction road surface signed distance function The absolute value of Represents the expected value function.

[0104] It should be noted that the expected loss of SDF Predicted road surface signed distance function Mapped to the interval [0,1], it represents the probability function of whether the point cloud is on the surface of the scene, thereby guiding the learning of SDF.

[0105] Step 536: Determine a first target loss based on one or more of the first sub-loss, the second sub-loss, and the third sub-loss.

[0106] Here, the first target loss can be one of the first sub-loss, the second sub-loss and the third sub-loss, or it can be determined by a combination of any two of the first sub-loss, the second sub-loss and the third sub-loss. Of course, the first target loss can also be determined based on the first sub-loss, the second sub-loss and the third sub-loss.

[0107] In some embodiments, if the first target loss is determined based on multiple of the first sub-loss, the second sub-loss, and the third sub-loss, a weight coefficient may be set for each sub-loss to better balance the losses and enable the MLP to better learn geometric information. For example, if the first target loss is determined based on the first sub-loss, the second sub-loss, and the third sub-loss, the first target loss of the MLP network model can be obtained by the following formula (5).

[0108]

[0109] in, Represents the first target loss, also known as MLP loss, represents the third weight coefficient, The third son is lost, represents the first weight coefficient, represents the first child loss, represents the second weight coefficient, Indicates the loss of the second child.

[0110] Step 504: Based on the first target loss, the network parameters of the MLP network model to be trained are optimized to obtain a trained MLP network model.

[0111] Among them, the trained MLP network model meets the iterative optimization conditions of the MLP network model.

[0112] In the embodiment of the present application, the iterative optimization conditions of the MLP network model include but are not limited to: the number of iterations reaches a preset maximum MLP training number, or the MLP loss function value is less than or equal to a preset MLP loss threshold.

[0113] In the embodiment of the present application, the electronic device determines the first target loss based on the predicted road surface signed distance function of each feature point in the sample three-dimensional point cloud, such as Figure 3 The MLP loss 303 in the example is used to determine whether the MLP network model meets the iterative optimization conditions based on the first target loss. If so, the training of the MLP network model is stopped to obtain a trained MLP network model. If not, the network parameters of the trained MLP network model are optimized for the next time until the MLP network model meets the iterative optimization conditions of the MLP network model.

[0114] In this way, the MLP network is continuously iteratively trained through the position features of the sample point cloud, allowing it to learn the accurate geometric structure of the three-dimensional scene and output a road surface signed distance function that accurately describes the distance between the point cloud and the road surface, thereby improving the depth distortion and perspective distortion of the rendered image and increasing the detailed texture of the rendered image.

[0115] In some embodiments, the process of iteratively optimizing the attribute information of the Gaussian sphere based on the road surface image sequence and the road surface signed distance function of each feature point in the three-dimensional point cloud in step 105 is combined with Figure 6 Provide explanation.

[0116] Step 601: Project and rasterize the Gaussian sphere to obtain a rendered image sequence.

[0117] In the embodiment of the present application, the rendered image sequence can be a sequence of rendered images obtained by projecting and rasterizing the Gaussian sphere corresponding to each feature point in the three-dimensional point cloud. It should be noted that the rendered image sequence also has multiple perspectives, which correspond to the perspectives in the road image sequence.

[0118] In some embodiments, as Figure 2As shown, the process of projecting and rasterizing the Gaussian sphere is as follows: projecting the Gaussian sphere onto a two-dimensional image plane through a rasterizer 205 such as a differentiable tile rasterizer, and obtaining a rendered image sequence by rasterization; wherein the rasterization process may be to regard the Gaussian sphere as a "snowball" thrown at the image, leaving a diffusion trace to form the final rendered image 206. In some embodiments, the Alpha Blending technology is used in the rasterization process to fuse the projections of different Gaussian spheres to ensure that the color and transparency information of the final image are accurate. Of course, the rasterization process may also use other technologies, which will not be described here. The rendered images in the rendered image sequence correspond one-to-one to the road surface images in the road surface image sequence, such as Figure 4 shown.

[0119] Step 602 : Determine a second target loss based on the road surface image sequence, the rendered image sequence, the road surface signed distance function of each feature point in the three-dimensional point cloud, and the Gaussian sphere corresponding to each feature point in the three-dimensional point cloud.

[0120] In the embodiment of the present application, the second target loss may be the loss of a rendered image sequence obtained by projecting and rasterizing the Gaussian sphere. The second target loss may include, but is not limited to, one or more of the fourth sub-loss, the fifth sub-loss, and the sixth sub-loss. For example, the second target loss may include one of the fourth sub-loss, the fifth sub-loss, and the sixth sub-loss. The second target loss may be determined based on any combination of two of the fourth sub-loss, the fifth sub-loss, and the sixth sub-loss. The second target loss may also be determined based on the fourth sub-loss, the fifth sub-loss, and the sixth sub-loss.

[0121] Here, the fourth sub-loss may be a multi-view consistency loss during the reconstruction process. During implementation, this may be determined based on the rendered images from each viewpoint in the rendered image sequence. The fifth sub-loss may refer to the structural similarity loss and base loss during the reconstruction process. During implementation, this may be determined based on the rendered image sequence and the road surface image sequence. The sixth sub-loss may refer to the normal vector loss and scale loss during the reconstruction process. During implementation, this may be determined based on the road surface signed distance function and each Gaussian sphere based on each feature point in the 3D point cloud.

[0122] Step 603: Iteratively optimize the attribute information of the Gaussian sphere based on the second target loss.

[0123] In this embodiment of the present application, after obtaining the second target loss, the Gaussian sphere's attribute information is iteratively optimized based on the second target loss, continuously adjusting the Gaussian sphere's parameters to better fit the real road scene. Through multiple iterations, the difference between the rendered image sequence and the road image sequence can be gradually reduced, improving the accuracy of the road signed distance function and thus the Gaussian sphere's representation of the scene. This facilitates more accurate reconstruction of the road scene, providing a more reliable scene model for applications such as environmental perception in autonomous driving and scene construction in virtual reality, thereby enhancing system performance and reliability.

[0124] In some embodiments, step 602 determines the second target loss based on the road surface image sequence, the rendered image sequence, the road surface signed distance function of each feature point in the three-dimensional point cloud, and the first Gaussian sphere corresponding to each feature point in the three-dimensional point cloud. This can be achieved by the following steps:

[0125] Step 621: Determine a fourth sub-loss based on the rendered images corresponding to each perspective in the rendered image sequence.

[0126] It should be noted that the generation of new perspectives often depends on the viewing angle range of the training data. Related technologies usually have good results in generating images from specific perspectives, but when observing from unprecedented angles, the rendering quality will significantly degrade, and the geometry and details of the road surface cannot be accurately restored. This dependence limits the generalization ability of the model, resulting in poor results in practical applications.

[0127] In response to the above problem, in order to solve this perspective dependency problem, the embodiment of the present application adds a method of determining the multi-perspective consistency loss, i.e., the fourth sub-loss, based on the rendered images corresponding to each perspective in the rendered image sequence.

[0128] In some embodiments, step 621 can be implemented by determining the fourth sub-loss based on the rendered images corresponding to each perspective in the rendered image sequence through the following process:

[0129] Step 6211: Determine a seventh sub-loss based on the depth value of a common pixel point of the rendered image corresponding to any two perspectives in the multiple perspectives in the rendered image sequence.

[0130] Here, the seventh sub-loss can be a multi-perspective depth value loss, also known as a multi-perspective depth value difference. The method for determining the seventh sub-loss includes but is not limited to a certain third difference, the sum of each third difference, the square of each third difference, etc. Among them, the method for determining the third difference is to select rendering images corresponding to any two perspectives from the rendering image sequence, determine the common pixels of the rendering images corresponding to the two selected perspectives, and calculate the norm of the depth value difference of the common pixels of the rendering images corresponding to the two perspectives as the third difference. For example, for any two perspectives among all perspectives, the square of the third difference of all common pixels is used as the perspective depth value loss corresponding to any two perspectives. Since there are multiple perspective combinations, the sum of the perspective depth value losses corresponding to the multiple perspective combinations is used as the seventh sub-loss.

[0131] In some embodiments, the seventh sub-loss may be determined by a multi-view depth value difference function, where the multi-view depth value difference function may be represented by the following formula (6).

[0132]

[0133] in, Represents the seventh sub-loss, also known as the multi-view depth value difference, which represents the common pixels between any two views in the multiple views of the rendered image sequence The sum of the squared norm of the depth value differences, Represent the viewing angle index, ; Representation perspective and perspective The common point, Indicates the Common pixels in the rendered image corresponding to the perspective The depth value of the Gaussian sphere rendered, Indicates the Common pixels in the rendered image corresponding to the perspective The rendered depth value of the Gaussian sphere.

[0134] Step 6212: Determine an eighth sub-loss based on the color value of a common pixel of the rendered image corresponding to any two perspectives in the multiple perspectives in the rendered image sequence.

[0135] Here, the eighth sub-loss can be a multi-perspective color value loss, also known as a multi-perspective color value difference. The eighth sub-loss is determined by, but is not limited to, a fourth difference, the sum of the fourth differences, the square of the fourth differences, etc. The fourth difference is determined by selecting rendering images corresponding to any two perspectives from the rendering image sequence, determining the common pixels of the rendering images corresponding to the two selected perspectives, and calculating the norm of the color value difference of the common pixels of the rendering images corresponding to the two perspectives as the fourth difference. For example, for any two perspectives among all perspectives, the square of the fourth difference of all common pixels is used as the perspective color value loss corresponding to any two perspectives. Since there are multiple perspective combinations, the sum of the perspective color value losses corresponding to the multiple perspective combinations is used as the eighth sub-loss.

[0136] In some embodiments, the eighth sub-loss may be determined by a multi-view color value difference function, where the multi-view color value difference function may be expressed by the following formula (7).

[0137]

[0138] in, Represents the eighth sub-loss, also known as multi-view color value difference, which represents the common pixels between any two views in the multiple views of the rendered image sequence The sum of the squared norm of the color value differences, Represent the viewing angle index, ; Representation perspective and perspective The common point, Indicates the Common pixels in the rendered image corresponding to the perspective The color value of Indicates the Common pixels in the rendered image corresponding to the perspective Color values include but are not limited to RGB values.

[0139] Step 6213: Determine a ninth sub-loss based on the normal vectors of the feature points of the three-dimensional point cloud corresponding to the common pixels of the rendered images corresponding to any two perspectives in the multiple perspectives in the rendered image sequence.

[0140] Here, the ninth sub-loss can be a multi-perspective normal vector loss, also known as a multi-perspective normal vector difference. The method for determining the ninth sub-loss includes but is not limited to a certain fifth difference, the sum of each fifth difference, the square of each fifth difference, etc. Among them, the method for determining the fifth difference is to select rendering images corresponding to any two perspectives from the rendering image sequence, determine the common pixel points of the rendering images corresponding to the two selected perspectives, and calculate the normal vector product of the common pixel points of the rendering images corresponding to the two perspectives to obtain a second vector product, and use the difference between a third value such as 1 and the second vector product as the fifth difference. For example, for any two perspectives among all perspectives, the sum of the fifth differences of all common pixels is used as the perspective normal vector loss corresponding to any two perspectives. Since there are multiple perspective combinations, the sum of the perspective normal vector losses corresponding to the multiple perspective combinations is used as the ninth sub-loss.

[0141] In some embodiments, the ninth sub-loss may be determined by a multi-view normal vector difference function, where the multi-view normal vector difference function may be expressed by the following formula (8).

[0142]

[0143] in, Represents the ninth sub-loss, also known as the multi-view normal vector difference, which represents the common pixels between any two views in the multiple views of the rendered image sequence The normal vector is expressed as Calculates the sum of the values, Represent the viewing angle index, ; Representation perspective and perspective The common point, Indicates the Common pixels in the rendered image corresponding to the perspective The normal vector rendered by the Gaussian sphere, Indicates the Common pixels in the rendered image corresponding to the perspective The normal vector of the Gaussian sphere rendered.

[0144] Step 6214: Determine the fourth sub-loss based on the seventh sub-loss, the eighth sub-loss, and the ninth sub-loss.

[0145] Here, the fourth sub-loss can be one of the seventh sub-loss, the eighth sub-loss and the ninth sub-loss, or can be determined by a combination of any two of the seventh sub-loss, the eighth sub-loss and the ninth sub-loss. Of course, the fourth sub-loss can also be determined based on the seventh sub-loss, the eighth sub-loss and the ninth sub-loss.

[0146] In some embodiments, if the fourth sub-loss is determined based on multiple of the seventh, eighth, and ninth sub-losses, a weight coefficient may be set for each sub-loss to better balance the multi-view consistency losses. For example, if the fourth sub-loss is determined based on the seventh, eighth, and ninth sub-losses, the fourth sub-loss can be obtained using the following formula (9).

[0147]

[0148] in, It is the fourth sub-loss, also known as multi-view consistency loss. is the multi-view depth weight, For the loss of the seventh son, is the multi-view normal vector weight, For the loss of the ninth child, is the multi-view color weight, The eighth piece is lost.

[0149] Step 622: Determine a fifth sub-loss based on the road surface image sequence and the rendered image sequence.

[0150] In the embodiment of the present application, the fifth sub-loss may be the structural similarity loss and the basic loss included in the reconstruction process, and during implementation, may be determined based on the rendered image sequence and the road surface image sequence.

[0151] Here, step 622 can be implemented by determining the fifth sub-loss based on the road surface image sequence and the rendered image sequence through the following process:

[0152] Step 6221: Determine a tenth sub-loss based on the difference between the pixel value of each pixel in the road surface image sequence and the pixel value of the corresponding pixel in the rendered image sequence.

[0153] Here, the tenth sub-loss can be a basic loss. Methods for determining the tenth sub-loss include, but are not limited to, a sixth difference, an average of the sum of all sixth differences, an average of the squares of all sixth differences, and the like. The sixth difference is determined by obtaining the average pixel value of any pixel covered by a Gaussian sphere in the rendered image sequence and the average pixel value of the corresponding pixel in the road surface image sequence covered by the Gaussian sphere, and taking the difference between these two pixel value averages as the sixth difference. For example, the average of the sum of all sixth differences can be used as the tenth sub-loss.

[0154] In some embodiments, the tenth sub-loss may be determined by a basic loss function, where the basic loss function may be expressed by the following formula (10).

[0155]

[0156] in, The tenth child loss, also known as the basic loss, represents the number of Gaussian balls, Indicates the first The mean pixel value of the pixel points covered by the Gaussian balls, Indicates the first The mean pixel value of the corresponding pixel point in the road image sequence is obtained by covering the pixels of the Gaussian sphere.

[0157] Step 6222: Determine an eleventh sub-loss based on the similarity between image information of each image in the road surface image sequence and image information of the corresponding image in the rendered image sequence; wherein the image information includes at least one of the following: brightness information, contrast information, and structure information.

[0158] In some embodiments, the eleventh sub-loss may be determined by a structural similarity loss function, wherein the structural similarity loss function may be expressed by the following formula (11).

[0159]

[0160] in, is the eleventh sub-loss, also known as structural similarity loss, It represents obtaining the similarity between the image information in the rendered image sequence and the image information in the road surface image sequence.

[0161] Step 6223: Determine the fifth sub-loss based on the tenth sub-loss and the eleventh sub-loss.

[0162] In the embodiment of the present application, the fifth sub-loss may be the tenth sub-loss or the eleventh sub-loss, and the fifth sub-loss may also be determined based on the tenth sub-loss and the eleventh sub-loss.

[0163] Step 623 : Determine a sixth sub-loss based on the road surface signed distance function of each feature point in the three-dimensional point cloud and each first Gaussian sphere.

[0164] In an embodiment of the present application, the sixth sub-loss may refer to the normal vector loss and scale loss in the reconstruction process. During implementation, it may be determined based on the road surface signed distance function and each Gaussian sphere based on each feature point in the three-dimensional point cloud.

[0165] Here, step 623 determines the sixth sub-loss based on the road surface signed distance function of each feature point in the three-dimensional point cloud and each first Gaussian sphere, which can be achieved by the following process:

[0166] Step 6231: Obtain the true normal vector of each feature point in the three-dimensional point cloud.

[0167] In the embodiments of the present application, the true normal vector of each feature point in the 3D point cloud can be obtained using a nearest neighbor algorithm. In one embodiment, a set of neighborhood points is obtained for each feature point, an approximate tangent plane for the feature point is constructed based on the neighborhood point set, and a perpendicular vector to the approximate tangent plane is obtained, which is used as the true normal vector for the feature point.

[0168] Step 6232: Determine the twelfth sub-loss based on the road surface signed distance function and the corresponding true normal vector of each feature point in the three-dimensional point cloud.

[0169] In an embodiment of the present application, the electronic device can perform a differential calculation, such as a derivative calculation, on the road surface signed distance function of each feature point in the three-dimensional point cloud, thereby obtaining a predicted normal vector of the corresponding feature point in the three-dimensional point cloud. Furthermore, based on the predicted normal vector of each feature point in the three-dimensional point cloud and the corresponding true normal vector, the twelfth sub-loss is determined.

[0170] In some embodiments, the electronic device may determine the twelfth sub-loss based on the predicted normal vector and the corresponding true normal vector of each feature point in the three-dimensional point cloud using the normal vector loss function of formula (3).

[0171] Step 6233: Determine a thirteenth sub-loss based on the road surface signed distance function of each feature point in the three-dimensional point cloud and the scale information of the first Gaussian sphere corresponding to the feature point, wherein the attribute information includes the scale information.

[0172] In some embodiments, based on the road surface signed distance function of each feature point in the three-dimensional point cloud and the scale information of the first Gaussian sphere corresponding to the feature point, a thirteenth sub-loss can be determined using a scale loss function. The scale loss function can be expressed as follows:

[0173]

[0174] in, is the thirteenth child loss, also known as scale loss, represents the number of Gaussian balls; Indicates the The average value of the three axis lengths of a Gaussian sphere. The Gaussian sphere is an ellipsoid, and the three axes represent the major axis, minor axis and flattening; Indicates the The scaling function of a Gaussian sphere, , Represents the overall scale range control parameter, and its value range is , smaller scenes require smaller value to avoid excessive scale, while larger scenes can increase value; It represents the influence coefficient of SDF on the size of Gaussian ball, that is, the exponential decay rate, and its value range is ; Indicates the The characteristic points corresponding to the Gaussian sphere Road surface signed distance function The absolute value of .

[0175] It should be noted that the scale loss is set in the second objective loss. The scale loss is combined with the scale function, and the road signed distance function is introduced into the scale function, which can control the size of the Gaussian sphere in the scene, thereby retaining more details.

[0176] Step 6234: Determine the sixth sub-loss based on the twelfth sub-loss and the thirteenth sub-loss.

[0177] In the embodiment of the present application, the sixth sub-loss may be the twelfth sub-loss or the thirteenth sub-loss, and the sixth sub-loss may also be determined based on the twelfth sub-loss and the thirteenth sub-loss.

[0178] Step 624: Determine a second target loss based on the fourth sub-loss, the fifth sub-loss, and the sixth sub-loss.

[0179] In this embodiment of the present application, the fifth sub-loss includes the tenth sub-loss and the eleventh sub-loss, the sixth sub-loss includes the twelfth sub-loss and the thirteenth sub-loss, and the second target loss can be one or more of the fourth sub-loss, the tenth sub-loss, the eleventh sub-loss, the twelfth sub-loss, and the thirteenth sub-loss.

[0180] In some embodiments, to better balance the sub-losses, a weight coefficient can be set for each sub-loss. For example, if the second target loss can be determined by multiple of the fourth sub-loss, the tenth sub-loss, the eleventh sub-loss, the twelfth sub-loss, and the thirteenth sub-loss, the second target loss can be obtained by the following formula (13) for the rendering loss function.

[0181]

[0182] in, is the second target loss, also known as rendering loss, is the structural similarity weight coefficient, For the loss of the tenth child, For the loss of the eleventh child, is the normal vector weight coefficient, For the loss of the twelfth child, is the scale weight coefficient, For the loss of the thirteenth child, is the multi-view consistency weight coefficient, Loss of the fourth child.

[0183] In some embodiments, to obtain a satisfactory rendering effect, during the iterative optimization of the attribute information of the Gaussian sphere, the attribute information optimization of the Gaussian sphere and the Gaussian sphere density control are performed alternately.

[0184] In the embodiment of the present application, during the iterative Gaussian sphere optimization process, Gaussian sphere density adaptive control steps are performed alternately. For example, after performing the Gaussian sphere property optimization step a set number of times (e.g., 100 times), Gaussian sphere density control is performed once. After performing the Gaussian sphere density control, Gaussian sphere property iterative optimization is continued, and this alternating process is repeated until the Gaussian optimization stop condition is reached. In this way, by alternately performing Gaussian sphere density adaptive control, the overall quality of the reconstruction is enhanced.

[0185] In some embodiments, the Gaussian sphere density adaptive control process is combined with Figure 7 Provide explanation.

[0186] Step 701: Calculate the variance of the Gaussian sphere corresponding to each feature point in the three-dimensional point cloud.

[0187] Here, the variance of the Gaussian sphere corresponding to each feature point in the three-dimensional point cloud can be calculated using the following formula (14).

[0188]

[0189] in, For the The variance of a Gaussian ball, Indicates the The number of viewing angles that Gaussian balls participate in calculating the variance, ; Representation perspective The rendering loss function under Representation perspective Next The number of pixels covered by a Gaussian sphere; 、 Respectively represent Gaussian spheres at the viewing angle The x-axis coordinate and y-axis coordinate in the normalized device coordinates.

[0190] It should be noted that the perspective Next The number of pixels covered by a Gaussian sphere As the calculation variance The weights of , dynamically average the gradients from different perspectives, thus promoting Gaussian growth.

[0191] Step 702: Split or copy the Gaussian sphere based on the variance to obtain a processed Gaussian sphere.

[0192] In an embodiment of the present application, the electronic device can obtain a variance threshold, which is used to control the threshold for splitting and duplication. Furthermore, based on the size relationship between the variance of each Gaussian sphere and the variance threshold, the corresponding Gaussian sphere is split or duplicated to obtain a processed Gaussian sphere. In some embodiments, if the variance of the Gaussian sphere is greater than or equal to the variance threshold, the Gaussian sphere is split. If the variance of the Gaussian sphere is less than the variance threshold, a copy of the Gaussian sphere is created and the copy is moved toward the normalized device coordinate gradient of the Gaussian sphere. Normalized Device Coordinates (NDC) is a coordinate system used to represent and process graphics in computer graphics. The coordinate range of NDC is usually from [-1,1][-1,1] on the x and y axes, and from 0 to 1 on the z axis.

[0193] For example, for Gaussian balls with large position gradients in the view space (i.e., exceeding a certain threshold), it includes cloning small Gaussian balls in insufficiently reconstructed areas or splitting large Gaussian balls in over-reconstructed areas. Gaussian spheres, , is the variance threshold, split Gaussian balls, replace one large Gaussian ball with two smaller ones, and reduce their sizes by a specific factor; for clones (with too small variance), such as Gaussian spheres, , then create the A copy of the Gaussian sphere, and move the copy towards the The goal is to find the best distribution and representation of Gaussian balls in three-dimensional space, thereby enhancing the overall quality of reconstruction.

[0194] By calculating the variance of the Gaussian spheres corresponding to feature points and splitting or replicating them based on the variance, the distribution density of the Gaussian spheres can be flexibly adjusted according to the needs of the actual scene. In areas with complex point cloud feature variations, the number of Gaussian spheres can be increased by splitting, allowing for a more detailed representation of the area. In relatively simple areas, the Gaussian spheres can be replicated, ensuring a certain level of accuracy while avoiding wasted computing resources and achieving efficient scene representation.

[0195] Step 703: Eliminate Gaussian spheres that do not meet the transparency condition from the processed Gaussian spheres.

[0196] In the embodiment of the present application, the transparency condition may be that the transparency is greater than or equal to a transparency threshold, and the transparency threshold may be determined based on experience.

[0197] In the embodiment of the present application, eliminating Gaussian spheres that do not meet the transparency condition from the processed Gaussian spheres can be understood as a pruning process. The pruning process removes redundant or less influential Gaussian spheres and can be regarded as a regularization process to some extent. Generally, Gaussian spheres that are almost transparent (e.g., with a transparency below the transparency threshold) and Gaussian spheres that are too large in world space or view space are eliminated. In addition, to prevent the density of Gaussian spheres near the input camera from increasing unreasonably, these Gaussian spheres can be set to a transparency value close to 0 after a fixed number of iterations.

[0198] Eliminating Gaussian spheres that don't meet the transparency requirements can reduce computational complexity by removing those that contribute little to the final result or are unnecessary. Furthermore, properly controlling the density of Gaussian spheres and selecting those that meet the requirements can avoid over- or under-rendering during the rendering process, thereby enhancing the rendering quality of the scene. This ensures that the rendered image has appropriate detail and transparency, making the resulting image more realistic and natural, and improving the user's visual experience.

[0199] The present application provides a new perspective reconstruction device for a road surface. Figure 8 As shown, Figure 8 This is a schematic diagram of the structure of a road surface new perspective reconstruction device provided in an embodiment of the present application. The road surface new perspective reconstruction device 8 includes:

[0200] An acquisition module 801 is used to acquire a multi-view road surface image sequence and a 3D point cloud corresponding to the road surface image sequence;

[0201] The encoding module 802 is used to encode the position of each feature point in the three-dimensional point cloud to obtain the point cloud position feature of the three-dimensional point cloud;

[0202] The first processing module 803 is used to input the point cloud position features into the trained MLP network model for feature processing to obtain the road surface signed distance function of each feature point in the three-dimensional point cloud;

[0203] A generating module 804 is used to generate a Gaussian sphere at the position of each feature point of the three-dimensional point cloud and initialize the attribute information of the Gaussian sphere;

[0204] The second processing module 805 is used to iteratively optimize the attribute information of the Gaussian sphere based on the road surface image sequence and the road surface signed distance function of each feature point in the three-dimensional point cloud until the Gaussian sphere optimization stopping condition is reached; and use the rendered image sequence corresponding to the attribute information of the Gaussian sphere when the iterative optimization stops as the road surface new perspective reconstructed image sequence.

[0205] An embodiment of the present application provides an electronic device, the electronic device comprising:

[0206] at least one processor; and,

[0207] A memory communicatively connected to at least one processor; wherein the memory stores a computer program executable by at least one processor, and the computer program is executed by at least one processor to implement some or all of the steps in the above method.

[0208] The present application provides a computer-readable storage medium storing one or more computer programs, which can be executed by one or more processors to implement some or all of the steps in the above method. The storage medium can be transient or non-transient.

[0209] An embodiment of the present application provides a computer program, including computer-readable code. When the computer-readable code runs in an electronic device, a processor in the electronic device executes some or all of the steps for implementing the above method.

[0210] An embodiment of the present application provides a computer program product, comprising a non-transitory computer-readable storage medium storing a computer program. When the computer program is read and executed by a computer, the computer program implements some or all of the steps of the above-described method. The computer program product can be implemented in hardware, software, or a combination thereof. In some embodiments, the computer program product is embodied as a computer storage medium. In other embodiments, the computer program product is embodied as a software product, such as a software development kit (SDK).

[0211] It should be noted that the descriptions of the various embodiments above tend to emphasize the differences between the various embodiments, and their similarities or similarities can be referenced to each other. The descriptions of the above device, storage medium, computer program, and computer program product embodiments are similar to the descriptions of the above method embodiments and have similar beneficial effects as the method embodiments. For technical details not disclosed in the embodiments of the device, storage medium, computer program, and computer program product of this application, please refer to the description of the method embodiments of this application for understanding.

[0212] The embodiment of the present application provides a hardware entity diagram of an electronic device, which may be a vehicle or a terminal, such as Figure 9 As shown, the hardware entity of the electronic device 9 includes:

[0213] at least one processor 901; and,

[0214] A memory 902 is communicatively connected to the at least one processor 901; wherein the memory 902 stores a computer program executable by the at least one processor 901, and the computer program is executed by the at least one processor to implement the following steps:

[0215] Obtain multi-view road surface image sequences and the three-dimensional point clouds corresponding to the road surface image sequences;

[0216] Encode the position of each feature point in the three-dimensional point cloud to obtain the point cloud position feature of the three-dimensional point cloud;

[0217] The point cloud position features are input into the trained Multilayer Perceptron (MLP) network model for feature processing to obtain the road surface signed distance function for each feature point in the 3D point cloud.

[0218] Generate a Gaussian sphere at the position of each feature point in the 3D point cloud and initialize the attribute information of the Gaussian sphere;

[0219] Based on the road surface signed distance function of each feature point in the road surface image sequence and the 3D point cloud, the attribute information of the Gaussian sphere is iteratively optimized until the Gaussian sphere optimization stop condition is reached;

[0220] The rendered image sequence generated corresponding to the attribute information of the Gaussian sphere when the iterative optimization stops is used as the image sequence for reconstructing the road surface from a new perspective.

[0221] A processor 901 and a memory 902 , wherein the memory 902 stores a computer program that can be run on the processor 901 , and the processor 901 executes part or all of the steps in the above-mentioned method for reconstructing a new perspective of a road surface.

[0222] Among them, the memory 902 stores a computer program that can be run on the processor. The memory 902 is configured to store instructions and applications executable by the processor 901. It can also cache data to be processed or processed by the processor 901 and various modules in the electronic device 9 (for example, image data, audio data, voice communication data and video communication data), which can be implemented through flash memory (FLASH) or random access memory (Random Access Memory, RAM).

[0223] The processor 901 generally controls the overall operation of the electronic device 9 , and executes the program to implement any of the above steps of the method for reconstructing a road surface from a new perspective.

[0224] Continue to refer to Figure 9 The electronic device 9 may further include a communication bus 903 and a communication interface 904 .

[0225] The communication bus 903 may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus. This bus can be divided into an address bus, a data bus, a control bus, etc. The bus is configured to enable communication between the memory 902 and at least one processor 901.

[0226] The communication interface 904 is used for communication between the above-mentioned electronic device and other devices, including a network interface and a user interface. Optionally, the network interface may include a wired interface and / or a wireless interface (such as a WI-FI interface, a Bluetooth interface, etc.), which is generally used to establish a communication connection between the electronic device and other electronic devices. The user interface may be a display (Display), an input unit (such as a keyboard (Keyboard)), optionally, the user interface may also be a standard wired interface, a wireless interface. Optionally, in some embodiments, the display may be an LED display, a liquid crystal display, a touch-sensitive liquid crystal display, and an OLED (Organic Light-Emitting Diode, organic light-emitting diode) touch device, etc. Among them, the display may also be appropriately referred to as a display screen or a display unit, which is used to display information processed in the electronic device and to display a visual user interface.

[0227] It should be noted that the description of the above storage medium and device embodiments is similar to the description of the above method embodiments and has similar beneficial effects as the method embodiments. For technical details not disclosed in the storage medium and device embodiments of this application, please refer to the description of the method embodiments of this application for understanding.

[0228] The processor may be at least one of an application-specific integrated circuit (ASIC), a digital signal processor (DSP), a digital signal processing device (DSPD), a programmable logic device (PLD), a field programmable gate array (FPGA), a central processing unit (CPU), a controller, a microcontroller, and a microprocessor. It is understood that the electronic device that implements the functions of the processor may also be other electronic devices, which are not specifically limited in the embodiments of the present application.

[0229] The above-mentioned computer storage medium / memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a magnetic random access memory (FRAM), a flash memory (Flash Memory), a magnetic surface storage device, an optical disc, or a compact disc read-only memory (CD-ROM); it can also be various terminals that include one or any combination of the above-mentioned memories, such as mobile phones, computers, tablet devices, personal digital assistants, etc.

[0230] It should be understood that "one embodiment" or "an embodiment" mentioned throughout the specification means that the specific features, structures or characteristics related to the embodiment are included in at least one embodiment of the present application. Therefore, "in one embodiment" or "in an embodiment" appearing throughout the specification does not necessarily refer to the same embodiment. In addition, these specific features, structures or characteristics can be combined in one or more embodiments in any suitable manner. It should be understood that in the various embodiments of the present application, the size of the serial numbers of the above-mentioned steps / processes does not mean the order of execution, and the execution order of each step / process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application. The above-mentioned serial numbers of the embodiments of the present application are for description only and do not represent the advantages and disadvantages of the embodiments.

[0231] It should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or apparatus comprising the element.

[0232] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as: multiple units or components can be combined, or can be integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the components shown or discussed can be through some interfaces, and the indirect coupling or communication connection of the devices or units can be electrical, mechanical or other forms.

[0233] The units described above as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units; they may be located in one place or distributed across multiple convolutional network units; some or all of the units may be selected according to actual needs to achieve the purpose of this embodiment.

[0234] In addition, all functional units in the embodiments of the present application can be integrated into one processing unit, or each unit can be a separate unit, or two or more units can be integrated into one unit; the above-mentioned integrated units can be implemented in the form of hardware or in the form of hardware plus software functional units.

[0235] Those skilled in the art will understand that all or part of the steps of the above-mentioned method embodiment can be completed by hardware related to program instructions, and the aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps of the above-mentioned method embodiment; and the aforementioned storage medium includes: mobile storage devices, read-only memories (ROM), magnetic disks or optical disks, and other media that can store program codes.

[0236] Alternatively, if the above-mentioned integrated unit of the present application is implemented in the form of a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the relevant technology, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a number of instructions for enabling an on-board terminal (which can be a personal computer, server, or network device, etc.) to execute all or part of the methods described in the various embodiments of the present application. The aforementioned storage medium includes: various media that can store program codes, such as mobile storage devices, ROMs, magnetic disks, or optical disks.

[0237] The above is only an implementation method of the present application, but the scope of protection of the present application is not limited thereto. Any technician familiar with this technical field can easily think of changes or replacements within the technical scope disclosed in this application, which should be covered by the scope of protection of the present application.

Claims

1. A method for reconstructing a road surface from a new perspective, characterized in that: The method comprises: Acquire a multi-view road surface image sequence and a three-dimensional point cloud corresponding to the road surface image sequence; Encoding the position of each feature point in the three-dimensional point cloud to obtain a point cloud position feature of the three-dimensional point cloud; Inputting the point cloud position features into a trained multi-layer perceptron (MLP) network model for feature processing to obtain a road surface signed distance function for each feature point in the three-dimensional point cloud; generating a Gaussian sphere at the position of each feature point of the three-dimensional point cloud and initializing attribute information of the Gaussian sphere; Iteratively optimizing the attribute information of the Gaussian sphere based on the road surface image sequence and the road surface signed distance function of each feature point in the three-dimensional point cloud, utilizing a first correlation between the road surface signed distance function and a scale function of the Gaussian sphere, until a Gaussian sphere stopping optimization condition is reached; wherein the first correlation includes an overall scale range control parameter and an exponential decay rate of the road surface distance sign function with respect to the size of each Gaussian sphere; The rendered image sequence generated corresponding to the attribute information of the Gaussian sphere when the iterative optimization stops is used as the image sequence for reconstructing the road surface from a new perspective.

2. The method according to claim 1, characterized in that The step of encoding the position of each feature point in the three-dimensional point cloud to obtain a point cloud position feature of the three-dimensional point cloud includes: Converting the position of each feature point in the three-dimensional point cloud into a high-dimensional vector using a position encoding function to obtain a plurality of high-dimensional vectors; The multiple high-dimensional vectors are fused to obtain the point cloud position features with high dimensions.

3. The method according to claim 1, characterized in that The training process of the MLP network model includes: Obtaining sample data, wherein the sample data includes sample point cloud position features of a sample three-dimensional point cloud; Inputting the sample point cloud position features into the MLP network model to be trained for feature processing, and obtaining a predicted road surface signed distance function of each feature point in the sample three-dimensional point cloud output by the MLP network model to be trained; determining a first target loss based on a predicted road surface signed distance function for each feature point in the sample three-dimensional point cloud; Based on the first target loss, the network parameters of the MLP network model to be trained are optimized to obtain the trained MLP network model, wherein the trained MLP network model meets the iterative optimization conditions of the MLP network model.

4. The method according to claim 3, characterized in that The determining of the first target loss based on the predicted road surface signed distance function of each feature point in the sample three-dimensional point cloud includes: Obtaining a true normal vector of each feature point in the sample three-dimensional point cloud; Determining a predicted normal vector for each feature point in the sample three-dimensional point cloud based on a predicted road surface signed distance function for each feature point in the sample three-dimensional point cloud; Determining a first sub-loss based on the predicted normal vector of each of the feature points; Determining a second sub-loss based on the predicted normal vector and the corresponding true normal vector of each of the feature points; determining a third sub-loss based on the predicted road surface signed distance function of each of the feature points; The first target loss is determined based on one or more of the first sub-loss, the second sub-loss, and the third sub-loss.

5. The method according to claim 1, wherein During the iterative optimization process of the attribute information of the Gaussian sphere, the attribute information optimization of the Gaussian sphere and the density control of the Gaussian sphere are performed alternately.

6. The method according to any one of claims 1 to 5, characterized in that The iterative optimization of the attribute information of the Gaussian sphere based on the road surface signed distance function of the road surface image sequence and each feature point in the three-dimensional point cloud and utilizing a first correlation relationship between the road surface signed distance function and the scale function of the Gaussian sphere includes: Projecting and rasterizing the Gaussian sphere to obtain a rendered image sequence; determining a second target loss using the first association relationship based on the road surface image sequence, the rendered image sequence, a road surface signed distance function of each feature point in the three-dimensional point cloud, and a Gaussian sphere corresponding to each feature point in the three-dimensional point cloud; The attribute information of the Gaussian sphere is iteratively optimized based on the second target loss.

7. The method according to claim 6, characterized in that The determining of the second target loss using the first association relationship based on the road surface image sequence, the rendered image sequence, the road surface signed distance function of each feature point in the three-dimensional point cloud, and the Gaussian sphere corresponding to each feature point in the three-dimensional point cloud includes: determining a fourth sub-loss based on the rendered images corresponding to the respective perspectives in the rendered image sequence; determining a fifth sub-loss based on the road surface image sequence and the rendered image sequence; determining a sixth sub-loss based on a road surface signed distance function of each feature point in the three-dimensional point cloud and each Gaussian sphere using the first association relationship; The second target loss is determined based on the fourth sub-loss, the fifth sub-loss, and the sixth sub-loss.

8. The method according to claim 7, characterized in that The determining of the fourth sub-loss based on the rendered images corresponding to the respective perspectives in the rendered image sequence includes: determining a seventh sub-loss based on depth values of common pixels of rendered images corresponding to any two perspectives in the rendered image sequence; determining an eighth sub-loss based on color values of common pixels of rendered images corresponding to any two perspectives in the rendered image sequence; determining a ninth sub-loss based on a common pixel of the rendered image corresponding to any two perspectives in the rendered image sequence and a normal vector corresponding to a feature point of the three-dimensional point cloud; The fourth sub-loss is determined based on the seventh sub-loss, the eighth sub-loss, and the ninth sub-loss.

9. The method according to claim 7, characterized in that The determining of a fifth sub-loss based on the road surface image sequence and the rendered image sequence includes: determining a tenth sub-loss based on a difference between a pixel value of each pixel point in the road surface image sequence and a pixel value of a corresponding pixel point in the rendered image sequence; determining an eleventh sub-loss based on a similarity between image information of each image in the road surface image sequence and image information of a corresponding image in the rendered image sequence; wherein the image information includes at least one of the following: brightness information, contrast information, and structure information; The fifth sub-loss is determined based on the tenth sub-loss and the eleventh sub-loss.

10. The method according to claim 7, characterized in that The determining of the sixth sub-loss based on the road surface signed distance function of each feature point in the three-dimensional point cloud and each Gaussian sphere using the first association relationship includes: Obtaining a true normal vector for each feature point in the three-dimensional point cloud; determining a twelfth sub-loss based on a road surface signed distance function and a corresponding true normal vector of each feature point in the three-dimensional point cloud; determining a thirteenth sub-loss using the first association relationship based on a road surface signed distance function of each feature point in the three-dimensional point cloud and scale information of a Gaussian sphere corresponding to the feature point, wherein the attribute information includes the scale information; The sixth sub-loss is determined based on the twelfth sub-loss and the thirteenth sub-loss.

11. The method according to claim 5, characterized in that The method further comprises: In the Gaussian sphere density control process, the variance of the Gaussian sphere corresponding to each feature point in the three-dimensional point cloud is calculated; Splitting or replicating the Gaussian sphere based on the variance to obtain a processed Gaussian sphere; From the processed Gaussian spheres, Gaussian spheres that do not meet the transparency condition are eliminated.

12. A road surface new perspective reconstruction device, characterized in that: The device comprises: An acquisition module, configured to acquire a multi-view road surface image sequence and a three-dimensional point cloud corresponding to the road surface image sequence; an encoding module, configured to encode the position of each feature point in the three-dimensional point cloud to obtain a point cloud position feature of the three-dimensional point cloud; A first processing module is configured to input the point cloud position features into a trained MLP network model for feature processing to obtain a road surface signed distance function for each feature point in the three-dimensional point cloud; a generating module, configured to generate a Gaussian sphere at the position of each feature point of the three-dimensional point cloud and initialize attribute information of the Gaussian sphere; The second processing module is used to iteratively optimize the attribute information of the Gaussian sphere based on the road surface image sequence and the road surface signed distance function of each feature point in the three-dimensional point cloud, using a first correlation relationship between the road surface signed distance function and the scale function of the Gaussian sphere until the Gaussian sphere optimization stopping condition is reached; and use the rendered image sequence corresponding to the attribute information of the Gaussian sphere when the iterative optimization stops as the road surface new perspective reconstructed image sequence; wherein the first correlation relationship includes the overall scale range control parameter and the exponential decay rate of the road surface distance sign function with respect to the size of each Gaussian sphere.

13. An electronic device, characterized in that: The electronic device comprises: at least one processor; and, A memory communicatively connected to the at least one processor; wherein the memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor to implement the method for reconstructing a road surface from a new perspective as described in any one of claims 1 to 11.

14. A computer-readable storage medium, characterized in that The computer-readable storage medium stores one or more computer programs, and the one or more computer programs can be executed by one or more processors to implement the method for reconstructing a road surface from a new perspective as described in any one of claims 1 to 11.

Citation Information

Patent Citations

  • Multi-view three-dimensional reconstruction method and device, equipment and storage medium

    CN116310120A

  • Scene reconstruction Gaussian model generation method and scene reconstruction method

    CN118982611A