Road surface reconstruction method based on implicit neural representation, vehicle and storage medium

CN115937442BActive Publication Date: 2026-08-21安徽蔚来智驾科技有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211459639.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-17
Publication Date
2026-08-21
Estimated Expiration
2042-11-17

AI Technical Summary

Technical Problem

首先NeRF是基于体素渲染的方式来生成图像,它的几何很差甚至基本不可用,使得3D空间的标注不完整

Benefits of technology

[0017] The road surface reconstruction method based on implicit neural expression in this invention first obtains at least one two-dimensional coordinate point on the road surface and uses a first implicit neural network to obtain the height value corresponding to the two-dimensional coordinate point. Then, based on the two-dimensional coordinate point and its height value, the semantic and color information corresponding to the two-dimensional coordinate point are determined respectively. Finally, based on the height value, semantic information, and color information, a three-dimensional scene reconstruction of the road surface is performed. In this way, accurate semantic and color information is obtained, improving the accuracy of road surface reconstruction and achieving high-completeness and high-precision road surface reconstruction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115937442B_ABST
    Figure CN115937442B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of automatic driving, and particularly provides a road surface reconstruction method based on implicit neural representation, a vehicle and a storage medium, aiming at solving the technical problem of low road surface reconstruction accuracy of the existing road surface reconstruction method. To this end, the road surface reconstruction method comprises: acquiring at least one two-dimensional coordinate point on the road surface; acquiring the height value corresponding to the two-dimensional coordinate point by using a first implicit neural network; determining the semantic information and color information corresponding to the two-dimensional coordinate point based on the two-dimensional coordinate point and the height value corresponding to the two-dimensional coordinate point respectively; and performing three-dimensional scene reconstruction on the road surface based on the height value, the semantic information and the color information. In this way, the accuracy and completeness of the road surface reconstruction are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of autonomous driving technology, specifically providing a road reconstruction method, vehicle, and storage medium based on implicit neural expression. Background Technology

[0002] Currently, accurately reconstructing a road surface and automatically identifying elements on it is an important and meaningful problem in autonomous driving. The reconstructed road surface can be used for auxiliary labeling. If the reconstruction includes semantics, it can be applied to automatic labeling applications.

[0003] Existing reconstruction methods are mainly divided into two types: point cloud reconstruction and NeRF reconstruction. Point cloud reconstruction includes laser point cloud reconstruction or visual MVS reconstruction, and the output format of this type of method is a point cloud. The density of point clouds cannot reach 100%, and there are problems such as holes and gaps to varying degrees. At the same time, a point cloud is a collection of 3D points, and the points are independent of each other, lacking global constraints, and cannot effectively eliminate noise. NeRF is a method that has emerged in recent years, but NeRF methods still have significant problems. First, NeRF generates images based on voxel rendering, which has poor geometry or is even basically unusable, resulting in incomplete annotation in 3D space. At the same time, current NeRF schemes require training for tens of hours or more for large scenes, making the method impractical.

[0004] Accordingly, there is a need in the field for a new road surface reconstruction scheme based on implicit neural expression to solve the above problems. Summary of the Invention

[0005] To overcome the aforementioned deficiencies, this invention is proposed to provide a solution, or at least a partial solution, to the aforementioned technical problems. This invention provides a road surface reconstruction method based on implicit neural expression, a vehicle, and a storage medium.

[0006] In a first aspect, the present invention provides a road surface reconstruction method based on implicit neural expression, the method comprising: acquiring at least one two-dimensional coordinate point on the road surface; acquiring the height value corresponding to the two-dimensional coordinate point using a first implicit neural network; determining semantic information and color information corresponding to the two-dimensional coordinate point based on the two-dimensional coordinate point and the height value corresponding to the two-dimensional coordinate point, respectively; and reconstructing a three-dimensional scene of the road surface based on the height value, the semantic information and the color information.

[0007] In one implementation, determining the semantic information and color information corresponding to the two-dimensional coordinate point based on the two-dimensional coordinate point and the corresponding height value includes: determining the semantic probability vector and color feature vector corresponding to the two-dimensional coordinate point based on the two-dimensional coordinate point and the corresponding height value; determining the semantic information corresponding to the two-dimensional coordinate point based on the semantic probability vector; and determining the color information corresponding to the two-dimensional coordinate point based on the color feature vector.

[0008] In one implementation, determining the semantic probability vector corresponding to the two-dimensional coordinate point based on the two-dimensional coordinate point and the corresponding height value includes: acquiring a 2D image corresponding to the two-dimensional coordinate point based on the camera pose of each frame; projecting a three-dimensional coordinate point obtained by combining the two-dimensional coordinate point with the height value onto the 2D image to obtain an initial semantic probability vector; and determining the semantic probability vector corresponding to the two-dimensional coordinate point based on the semantic label value of the 2D image and the initial semantic probability vector.

[0009] In one embodiment, determining the color feature vector corresponding to the two-dimensional coordinate point based on the two-dimensional coordinate point and the height value corresponding to the two-dimensional coordinate point includes: acquiring a 2D image corresponding to the two-dimensional coordinate point at each frame's camera pose; projecting a three-dimensional coordinate point obtained by combining the two-dimensional coordinate point with the height value onto the 2D image to obtain an initial color feature vector; and determining the color feature vector corresponding to the two-dimensional coordinate point based on the color label value of the 2D image and the initial color feature vector.

[0010] In one implementation, determining the semantic information corresponding to the two-dimensional coordinate point based on the semantic probability vector includes: taking the category corresponding to the largest element in the semantic probability vector as the semantic information.

[0011] In one embodiment, determining the color information corresponding to the two-dimensional coordinate point based on the color feature vector includes: merging the color feature vector and the encoding of the camera ID, and inputting the merged feature vector into a second implicit neural network to obtain the color information corresponding to the two-dimensional coordinate point.

[0012] In one embodiment, the first implicit neural network or the second implicit neural network is a fully connected neural network.

[0013] In one implementation, the semantic information corresponding to the two-dimensional coordinate points is projected onto the image captured by the camera to obtain the semantic annotation result of the road surface.

[0014] In a second aspect, a vehicle is provided, the vehicle including a processor and a storage device adapted to store a plurality of program codes adapted to be loaded and run by the processor to perform the road surface reconstruction method based on implicit neural expression as described in any of the preceding claims.

[0015] In a third aspect, a computer-readable storage medium is provided, wherein a plurality of program codes are stored therein, the program codes being adapted to be loaded and run by a processor to perform the road surface reconstruction method based on implicit neural expression as described in any of the preceding claims.

[0016] The above-described technical solutions of the present invention have at least one or more of the following beneficial effects:

[0017] The road surface reconstruction method based on implicit neural expression in this invention first obtains at least one two-dimensional coordinate point on the road surface and uses a first implicit neural network to obtain the height value corresponding to the two-dimensional coordinate point. Then, based on the two-dimensional coordinate point and its height value, the semantic and color information corresponding to the two-dimensional coordinate point are determined respectively. Finally, based on the height value, semantic information, and color information, a three-dimensional scene reconstruction of the road surface is performed. In this way, accurate semantic and color information is obtained, improving the accuracy of road surface reconstruction and achieving high-completeness and high-precision road surface reconstruction. Attached Figure Description

[0018] The disclosure of this invention will become more readily understood with reference to the accompanying drawings. It will be readily understood by those skilled in the art that these drawings are for illustrative purposes only and are not intended to limit the scope of protection of this invention. Furthermore, similar numbers in the drawings are used to denote similar components, wherein:

[0019] Figure 1 This is a schematic flowchart of the main steps of a road surface reconstruction method based on implicit neural expression according to an embodiment of the present invention;

[0020] Figure 2(a) is a schematic diagram of obtaining road surface elements using an MPL network in one embodiment;

[0021] Figure 2(b) is a schematic diagram of a road reconstruction method based on implicit neural expression in one embodiment for obtaining road elements;

[0022] Figure 3 This is a schematic diagram of the complete process of a road surface reconstruction method based on implicit neural expression in one embodiment;

[0023] Figure 4 This is a schematic diagram of the road surface after marking in one embodiment;

[0024] Figure 5 This is a structural schematic diagram of a vehicle in one embodiment. Detailed Implementation

[0025] Some embodiments of the present invention will now be described with reference to the accompanying drawings. Those skilled in the art should understand that these embodiments are merely illustrative of the technical principles of the present invention and are not intended to limit the scope of protection of the present invention.

[0026] In the description of this invention, "module" and "processor" can include hardware, software, or a combination of both. A module can include hardware circuitry, various suitable sensors, communication ports, memory, and may also include software components, such as program code, or a combination of software and hardware. A processor can be a central processing unit, microprocessor, image processor, digital signal processor, or any other suitable processor. The processor has data and / or signal processing capabilities. The processor can be implemented in software, in hardware, or a combination of both. Non-transitory computer-readable storage media includes any suitable medium capable of storing program code, such as magnetic disks, hard disks, optical disks, flash memory, read-only memory, random access memory, etc. The term "A and / or B" means all possible combinations of A and B, such as only A, only B, or A and B. The terms "at least one A or B" or "at least one of A and B" have a similar meaning to "A and / or B" and can include only A, only B, or A and B. The singular terms "a" or "this" can also include plural forms.

[0027] Currently, existing reconstruction methods are mainly divided into two types: point cloud reconstruction and NeRF reconstruction. Point cloud reconstruction includes laser point cloud reconstruction or visual MVS reconstruction, and the output format of this type of method is a point cloud. The density of point clouds cannot reach 100%, and there are problems with varying degrees of holes and gaps. At the same time, a point cloud is a collection of 3D points, and the points are independent of each other, lacking global constraints, and cannot effectively eliminate noise. NeRF is a method that has emerged in recent years, but NeRF methods still have significant problems. First, NeRF generates images based on voxel rendering, which has poor geometry or is even basically unusable, resulting in incomplete 3D spatial annotation. At the same time, current NeRF schemes require training for tens of hours or more for large scenes, making the method impractical.

[0028] To address this, this application provides a road surface reconstruction method, vehicle, and storage medium based on implicit neural representation. First, at least one two-dimensional coordinate point on the road surface is acquired, and the height value corresponding to the two-dimensional coordinate point is obtained using a first implicit neural network. Then, semantic information and color information corresponding to the two-dimensional coordinate point are determined based on the two-dimensional coordinate point and its height value, respectively. Finally, a three-dimensional scene reconstruction of the road surface is performed based on the height value, semantic information, and color information. This method obtains accurate semantic and color information, improves the accuracy of road surface reconstruction, and achieves high-completeness and high-precision road surface reconstruction.

[0029] See appendix Figure 1 , Figure 1 This is a schematic diagram of the main steps of a road surface reconstruction method based on implicit neural expression according to an embodiment of the present invention.

[0030] like Figure 1 As shown, the road surface reconstruction method based on implicit neural expression in this embodiment of the invention mainly includes the following steps S101-S104.

[0031] Step S101: Obtain at least one two-dimensional coordinate point (x,y) on the road surface.

[0032] Step S102: Use the first implicit neural network to obtain the height value corresponding to the two-dimensional coordinate point.

[0033] Specifically, by inputting the two-dimensional coordinate point (x, y) into the first implicit neural network, the height value z corresponding to the two-dimensional coordinate point (x, y) can be obtained.

[0034] In one specific implementation, the first implicit neural network is a fully connected neural network (MLPS), but it is not limited to this. The first implicit neural network can also be other implicit neural networks, as long as it can obtain the height value corresponding to the two-dimensional coordinate point.

[0035] In one embodiment, the step is described in detail using a fully connected neural network (MLPS) as an example of a first implicit neural network.

[0036] For road surface height, the undulations are a very slow, low-frequency change, making it feasible to use fully connected neural networks (MLPs) to obtain the road surface height. Specifically, the LiDAR point cloud on the road surface is used as supervisory information. Although the supervisory information is sparse, the continuous nature of MLPs allows for the acquisition of relatively accurate height values ​​even in areas without supervisory signals. Furthermore, since the parameters of an MLP implicitly represent the entire space, noise in the supervisory information is smoothed out during training. Therefore, accurate road surface height values ​​can be obtained using fully connected neural networks (MLPs).

[0037] Unlike height values, road surface color and semantic information are high-frequency information that changes frequently and drastically with coordinate changes. For example, a zebra crossing on a road is a region of continuous alternation between black and white. Furthermore, the MLP mentioned earlier is a global function that describes the entire road scene. For an autonomous driving road several hundred meters long, a lane line only 10cm wide is a very detailed element. MLP struggles to capture these detailed areas, often resulting in over-smoothing, as shown in Figure 2(a), while these details are crucial for autonomous driving. Therefore, this application employs a road surface reconstruction method based on implicit neural representation, and the semantic information obtained is shown in Figure 2(b).

[0038] Step S103: Determine the semantic information and color information corresponding to the two-dimensional coordinate point based on the two-dimensional coordinate point and the height value corresponding to the two-dimensional coordinate point.

[0039] In one specific implementation, determining the semantic information and color information corresponding to the two-dimensional coordinate point based on the two-dimensional coordinate point and the corresponding height value includes: determining the semantic probability vector and color feature vector corresponding to the two-dimensional coordinate point based on the two-dimensional coordinate point and the corresponding height value; determining the semantic information corresponding to the two-dimensional coordinate point based on the semantic probability vector; and determining the color information corresponding to the two-dimensional coordinate point based on the color feature vector.

[0040] In one embodiment, after obtaining the semantic probability vector and color feature vector corresponding to the two-dimensional coordinate point, they can be stored in a 2D grid, where the 2D grid is a two-dimensional matrix, specifically, the real road surface is divided into grids by linear changes in scale.

[0041] In one specific implementation, determining the semantic probability vector corresponding to the two-dimensional coordinate point based on the two-dimensional coordinate point and the corresponding height value includes: acquiring a 2D image corresponding to the two-dimensional coordinate point based on the camera pose of each frame; projecting a three-dimensional coordinate point obtained by combining the two-dimensional coordinate point with the height value onto the 2D image to obtain an initial semantic probability vector; and determining the semantic probability vector corresponding to the two-dimensional coordinate point based on the semantic label value of the 2D image and the initial semantic probability vector.

[0042] For example, taking an image captured by at least one camera as an example, the image captured by the camera is pre-input into a pre-trained first neural network for semantic segmentation to obtain a semantic segmentation result, which is the semantic label value. The semantic segmentation result specifically includes, but is not limited to, sidewalks, lane lines, and curbs.

[0043] For a two-dimensional coordinate point (x, y) on the road surface, since the camera pose at that point is fixed, 2D images containing that coordinate point (x, y) captured by all cameras can be obtained based on the camera pose of each frame. Then, the three-dimensional coordinate point obtained by combining (x, y) with the height value z is projected onto the 2D image. The semantic probability vector corresponding to the three-dimensional coordinate point on the 2D image is the initial semantic probability vector. Next, a first loss function is calculated based on the semantic label values ​​of the 2D image and the initial semantic probability vector. The derivative of the first loss function yields the semantic probability vector corresponding to the two-dimensional coordinate point, which is then stored in the corresponding position in the 2D grid. Thus, each grid stores an N*1 semantic feature vector, representing the probability of N semantic categories.

[0044] In one embodiment, the cross-entropy loss function may serve as an example of the first loss function, but is not limited thereto.

[0045] In one embodiment, the size and resolution of the grid remain unchanged during training, and the semantic feature vectors stored in each grid can be updated according to the gradient descent method.

[0046] For color information, a simple approach is to directly store the color value at each coordinate in each network and update the color value using photos taken from multiple perspectives. However, common autonomous driving devices often employ surround-view camera hardware configurations, where the color temperature and exposure differ between cameras. For example, three different cameras might have different color temperatures, making direct representation using color values ​​unreasonable. Therefore, the color feature vector corresponding to a 2D coordinate point can be obtained and stored in a 2D mesh using the following method.

[0047] In one specific implementation, determining the color feature vector corresponding to the two-dimensional coordinate point based on the two-dimensional coordinate point and the height value corresponding to the two-dimensional coordinate point includes: acquiring a 2D image corresponding to the two-dimensional coordinate point based on the camera pose of each frame; projecting a three-dimensional coordinate point obtained by combining the two-dimensional coordinate point with the height value onto the 2D image to obtain an initial color feature vector; and determining the color feature vector corresponding to the two-dimensional coordinate point based on the color label value of the 2D image and the initial color feature vector.

[0048] The color label value is the R, G, and B value of the corresponding pixel point projected onto the 2D image by the three-dimensional coordinate value.

[0049] Specifically, for a two-dimensional coordinate point (x, y) on the road surface, since the camera pose at that point is fixed, all 2D images containing that coordinate point (x, y) captured by each camera can be obtained based on the camera pose of each frame. Then, the three-dimensional coordinate point obtained by combining (x, y) with the height value z is projected onto the 2D image. The color feature vector corresponding to the three-dimensional coordinate point on the 2D image is the initial color feature vector. Next, a second loss function is calculated based on the color label value of the 2D image and the initial color feature vector. Taking the derivative of the second loss function yields the color feature vector corresponding to the two-dimensional coordinate point, which is then stored in the corresponding position in the 2D grid.

[0050] In one embodiment, the cross-entropy loss function can be used as an example of the second loss function, but is not limited thereto.

[0051] In one embodiment, the size and resolution of the grid remain unchanged during training, and the color feature vector stored in each grid can be updated according to the gradient descent method.

[0052] In one specific implementation, determining the semantic information corresponding to the two-dimensional coordinate point based on the semantic probability vector includes: taking the category corresponding to the largest element in the semantic probability vector as the semantic information.

[0053] Specifically, since the semantic probability vector is an N*1 vector representing the probabilities of N semantic categories, the category corresponding to the element with the largest value in the semantic probability vector can be directly used as the semantic information corresponding to the two-dimensional coordinate point.

[0054] In one specific implementation, determining the color information corresponding to the two-dimensional coordinate point based on the color feature vector includes: merging the color feature vector and the encoding of the camera ID, and inputting the merged feature vector into a second implicit neural network to obtain the color information corresponding to the two-dimensional coordinate point.

[0055] Specifically, each camera is encoded so that each camera corresponds to a unique code for merging. In determining the color information corresponding to the two-dimensional coordinate point, an image captured by any camera is used as supervisory information. The vector resulting from merging the color feature vector with the ID code of any camera is input into a second implicit neural network to obtain the color information corresponding to the two-dimensional coordinate point. Although the supervisory information is sparse, due to the continuous nature of MLPs, relatively accurate color values ​​can be calculated even in areas without supervisory signals.

[0056] In one specific implementation, the second implicit neural network is a fully connected neural network. Specifically, a shallow fully connected neural network can be used as an example of the second implicit neural network, but it is not limited to this. The second implicit neural network can also be other implicit neural networks capable of predicting the color information corresponding to any two-dimensional coordinate point on the road surface.

[0057] Step S104: Reconstruct a three-dimensional scene of the road surface based on the height value, the semantic information, and the color information.

[0058] Based on steps S101-S104 above, at least one two-dimensional coordinate point on the road surface is first obtained, and the height value corresponding to the two-dimensional coordinate point is obtained using a first implicit neural network. Then, semantic information and color information corresponding to the two-dimensional coordinate point are determined based on the two-dimensional coordinate point and its height value, respectively. Finally, a three-dimensional scene reconstruction of the road surface is performed based on the height value, semantic information, and color information. In this way, accurate semantic and color information are obtained, improving the accuracy of road surface reconstruction and achieving high-completeness and high-precision road surface reconstruction.

[0059] In one embodiment, specifically as follows Figure 3 As shown, by inputting at least one two-dimensional coordinate point (x, y) on the road surface into a first implicit neural network, the height value of the two-dimensional coordinate point can be obtained. Then, based on the two-dimensional coordinate point and its corresponding height value, the semantic probability vector and color feature vector corresponding to the two-dimensional coordinate point are determined and stored in a 2D grid. Next, the color feature vector of the two-dimensional coordinate point stored in the grid is retrieved, merged with the camera ID encoding, and input into a second implicit neural network to obtain the color information corresponding to the two-dimensional coordinate point. Finally, the semantic feature vector of the two-dimensional coordinate point stored in the grid is retrieved, and the type corresponding to the element with the largest value in the semantic feature vector is taken as the semantic information corresponding to the two-dimensional coordinate point. Thus, by using a combination of implicit and explicit methods, the processing effect of semantic and color information is higher, while the network training time is shorter, ultimately obtaining high-precision height values, color information, and semantic information, thereby achieving a high-precision road surface reconstruction effect.

[0060] In one specific embodiment, the method further includes: projecting the semantic information corresponding to the two-dimensional coordinate points onto the image captured by the camera to obtain the semantic annotation result of the road surface.

[0061] Specifically, based on the aforementioned embodiments, the height value z corresponding to any two-dimensional coordinate point (x, y) on the roadside, as well as semantic information such as lane lines and curbs, have been obtained. For each (x, y), the first implicit neural network predicts a z, thus obtaining a 3D point. Projecting this 3D point back onto the images of each camera yields a continuous and consistent dense annotation result on the road surface. Consistency refers to the consistency of the annotation of the same object across frames in the video sequence and across images from different cameras. In one embodiment, such as... Figure 4 As shown, by projecting semantic information back to each camera, accurate lane line segmentation can be achieved.

[0062] It should be noted that although the steps in the above embodiments are described in a specific order, those skilled in the art will understand that in order to achieve the effects of the present invention, different steps do not necessarily have to be executed in such an order. They can be executed simultaneously (in parallel) or in other orders, and these variations are all within the scope of protection of the present invention.

[0063] Those skilled in the art will understand that all or part of the processes in the method of the above embodiment of the present invention can also be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable file, or some intermediate form. The computer-readable storage medium can include any entity or device capable of carrying the computer program code, a medium, a USB flash drive, a portable hard drive, a magnetic disk, an optical disk, a computer memory, a read-only memory, a random access memory, an electrical carrier signal, a telecommunication signal, and a software distribution medium, etc. It should be noted that the content included in the computer-readable storage medium can be appropriately added or removed according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable storage medium does not include electrical carrier signals and telecommunication signals.

[0064] Furthermore, the present invention also provides a vehicle. In one embodiment of the vehicle according to the present invention, such as Figure 5 As shown, the vehicle includes a processor 51 and a storage device 52. The storage device can be configured to store a program for executing the road reconstruction method based on implicit neural expression described in the above-described method embodiments. The processor can be configured to execute the program in the storage device, which includes, but is not limited to, a program for executing the road reconstruction method based on implicit neural expression described in the above-described method embodiments. For ease of explanation, only the parts related to the embodiments of the present invention are shown. For specific technical details not disclosed, please refer to the method section of the embodiments of the present invention.

[0065] Furthermore, the present invention also provides a computer-readable storage medium. In one embodiment of the computer-readable storage medium according to the present invention, the computer-readable storage medium can be configured to store a program for executing the road surface reconstruction method based on implicit neural expression described in the above-described method embodiments. This program can be loaded and run by a processor to implement the above-described road surface reconstruction method based on implicit neural expression. For ease of explanation, only the parts related to the embodiments of the present invention are shown; for specific technical details not disclosed, please refer to the method section of the embodiments of the present invention. The computer-readable storage medium can be a storage device comprising various electronic devices. Optionally, in the embodiments of the present invention, the computer-readable storage medium is a non-transitory computer-readable storage medium.

[0066] The technical solution of the present invention has been described above with reference to the preferred embodiments shown in the accompanying drawings. However, it will be readily understood by those skilled in the art that the scope of protection of the present invention is obviously not limited to these specific embodiments. Without departing from the principles of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after such changes or substitutions will all fall within the scope of protection of the present invention.

Claims

1. A road surface reconstruction method based on implicit neural expression, characterized in that, The method includes: Obtain at least one two-dimensional coordinate point on the road surface from the perspective of a BEV; The height value corresponding to the two-dimensional coordinate point is obtained using a first implicit neural network; The semantic probability vector and color feature vector corresponding to the two-dimensional coordinate point are determined based on the two-dimensional coordinate point and the height value corresponding to the two-dimensional coordinate point, respectively. The semantic information corresponding to the two-dimensional coordinate point is determined based on the semantic probability vector. The color information corresponding to the two-dimensional coordinate point is determined based on the color feature vector. A three-dimensional scene reconstruction of the road surface is performed based on the height value, the semantic information, and the color information; The determination of the semantic probability vector corresponding to the two-dimensional coordinate point based on the two-dimensional coordinate point and the corresponding height value includes: The 2D image corresponding to the two-dimensional coordinate point is obtained based on the camera pose of each frame. The three-dimensional coordinates obtained by combining the two-dimensional coordinates with the height value are projected onto the 2D image to obtain the initial semantic probability vector corresponding to the projection point of the three-dimensional coordinates in the 2D image. A first loss function is calculated based on the semantic label value of the 2D image and the initial semantic probability vector, and the semantic probability vector corresponding to the two-dimensional coordinate point is obtained by taking the derivative of the first loss function. The determination of the color feature vector corresponding to the two-dimensional coordinate point based on the two-dimensional coordinate point and the corresponding height value includes: The 2D image corresponding to the two-dimensional coordinate point is obtained based on the camera pose of each frame. The three-dimensional coordinate point obtained by combining the two-dimensional coordinate point with the height value is projected onto the 2D image to obtain the initial color feature vector corresponding to the projection point of the three-dimensional coordinate point in the 2D image. A second loss function is calculated based on the color label value of the 2D image and the initial color feature vector, and the derivative of the second loss function is used to obtain the color feature vector corresponding to the two-dimensional coordinate point.

2. The road surface reconstruction method based on implicit neural expression according to claim 1, characterized in that, Determining the semantic information corresponding to the two-dimensional coordinate point based on the semantic probability vector includes: taking the category corresponding to the largest element in the semantic probability vector as the semantic information.

3. The road surface reconstruction method based on implicit neural expression according to claim 1, characterized in that, Determining the color information corresponding to the two-dimensional coordinate point based on the color feature vector includes: merging the color feature vector and the camera ID encoding, and inputting the merged feature vector into a second implicit neural network to obtain the color information corresponding to the two-dimensional coordinate point.

4. The road surface reconstruction method based on implicit neural expression according to claim 3, characterized in that, The first implicit neural network or the second implicit neural network is a fully connected neural network.

5. The road surface reconstruction method based on implicit neural expression according to claim 1, characterized in that, The method further includes: projecting the semantic information corresponding to the two-dimensional coordinate points onto the image captured by the camera to obtain the semantic annotation result of the road surface.

6. A vehicle comprising a processor and a storage device, said storage device being adapted to store a plurality of program codes, characterized in that, The program code is adapted to be loaded and run by the processor to perform the road surface reconstruction method based on implicit neural expression as described in any one of claims 1 to 5.

7. A computer-readable storage medium storing a plurality of program codes, characterized in that, The program code is adapted to be loaded and run by a processor to perform the road surface reconstruction method based on implicit neural expression as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Image-based three-dimensional scene reconstruction method and device

    CN114742966A

  • Multi-sensor fusion vehicle-road collaborative sensing method for automatic driving

    CN114821507A