Semantic vector map construction method and device, storage medium and program product

Through knowledge of distillation rendering technology, hardware requirements are reduced, the accuracy and performance of monocular visual semantic vector map construction are improved, the problems of high hardware requirements and poor complexity improvement in the existing technology are solved, and the promotion of semantic vector map construction technology is promoted.

CN120451434APending Publication Date: 2025-08-08BEIHANG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510539745.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-27
Publication Date
2025-08-08

AI Technical Summary

Technical Problem

The existing semantic vector map construction technology has high requirements for hardware conditions, and increases model complexity and improves performance in many scenarios, resulting in difficulty in promoting and reducing efficiency.

Method used

A monocular visual semantic vector map construction method based on knowledge distillation body rendering is adopted. The circumferential image model is pre-trained as a teacher branch to construct a single-view student branch, and the three-dimensional features are projected into two-dimensional features using body rendering technology, and the loss is calculated and the model training is carried out.

Benefits of technology

It reduces hardware requirements, reduces configuration costs, improves the accuracy and performance of semantic vector map construction, and promotes the promotion and application of technology.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120451434A_ABST
    Figure CN120451434A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of automatic driving perception, in particular to a semantic vector map construction method and device, a storage medium and a program product. The semantic vector map construction method comprises the following steps: pre-training a semantic vector map construction model based on a look-around image as a knowledge distillation teacher branch; constructing a semantic vector map model based on a single view as a knowledge distillation student branch; projecting the three-dimensional features of the student branches into two-dimensional features under a plurality of visual angles by using a volume rendering technology; calculating body rendering loss and map loss of the student branch relative to the teacher branch; and fusing the body rendering loss and the map loss, and completing model training of student branches based on the fusion. According to the method, the model is constructed by depending on a semantic vector map input by a single image, so that the requirement on hardware is reduced; and fusing body rendering loss and map loss to obtain a lightweight single-mode online vector map construction model which is close to a complex multi-mode model in accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of autonomous driving perception technology, and in particular to a method, device, storage medium, and program product for constructing a monocular visual semantic vector map based on knowledge distillation rendering. Background Art

[0002] In recent years, with the rapid development of artificial intelligence technology and the continuous upgrading of autonomous driving related infrastructure, autonomous driving technology has gradually become an important development direction in the automotive field and has attracted the attention of many academic researchers.

[0003] With the continuous advancement of research, autonomous driving technology has gradually transitioned from the laboratory research stage to the commercialization and productization stage. This transition places more stringent requirements on the safety and accuracy of autonomous driving technology. To meet these requirements, autonomous driving systems must possess high-performance environmental perception capabilities in complex driving environments. Therefore, semantic understanding of driving scenarios, especially the ability to build real-time semantic maps, has become a key technical foundation for the practical application of autonomous driving technology.

[0004] In environmental perception, the task of building semantic vector maps online has attracted the attention of many researchers due to its intuitive representation and ease of subsequent application. However, due to the high hardware requirements of semantic vector map construction, the technology has been difficult to promote.

[0005] Some previous research works required the sensor to contain at least 6 cameras, which greatly limited the promotion of semantic vector map construction technology in the field of autonomous driving.

[0006] To ensure model performance, some previous research works have attempted to improve model performance by continuously increasing model complexity and improving model perception capabilities. Experiments have shown that this approach not only fails to enable the model to achieve the expected results in many scenarios, but also greatly reduces the model's perception efficiency due to its high complexity. Summary of the Invention

[0007] Based on the fact that some previous research works require that the sensor contain at least 6 cameras, which greatly limits the promotion of semantic vector map construction technology in the field of autonomous driving, and some previous research works attempt to improve model performance by continuously increasing model complexity and improving model perception capabilities. This method not only fails to enable the model to achieve the expected effect in many scenarios, but also greatly reduces the perception efficiency of the model due to its high complexity. This application proposes a monocular visual semantic vector map construction method, device, storage medium and program product based on knowledge distillation rendering.

[0008] A method for constructing a semantic vector map comprises the following steps: S1, pre-training semantic vector map construction model based on surround view images, as the knowledge distillation teacher branch; S2, builds a single-view-based semantic vector map model as a knowledge distillation student branch; S3, uses volume rendering technology to project the three-dimensional features of the student branch into two-dimensional features under multiple perspectives; S4, calculates the volume rendering loss of the student branch relative to the teacher branch; S5, calculates the map loss of the student branch relative to the teacher branch; S6 fuses the volume rendering loss and map loss, and completes the model training of the student branch based on this.

[0009] By adopting the above technical solution, this application proposes a semantic vector map construction model that relies solely on a single image input. This reduces the hardware requirements for semantic vector map construction technology, reduces configuration costs, and promotes the promotion of technology. By fusing volume rendering loss and map loss, a lightweight single-modal online vector map construction model is obtained, with accuracy approaching that of complex multimodal models. With the help of knowledge distillation and volume rendering technology, the model is more targetedly supervised, ensuring that the model has high semantic vector map construction performance even with small input sensors.

[0010] A preferred solution of the semantic vector map construction method is that step S1 specifically refers to: for the image data input of the vehicle-mounted surround view camera, extracting the features of i two-dimensional images collected by i camera sensors, and then obtaining multi-view semantic features; performing depth estimation on the two-dimensional image features of each perspective to obtain the features in the three-dimensional prism-shaped feature space under the perspective, combining the depth information of each feature point in the camera internal and external parameters, determining the three-dimensional spatial position corresponding to each feature point, and then determining the features of each voxel in the three-dimensional space, using a convolution-based BEV feature extractor to project the three-dimensional voxel features into BEV features, and finally inputting the BEV features into a Transformer-based map decoder to generate a vector map and obtain the model prediction results; pre-training the above model as a complex multi-view teacher model.

[0011] By adopting the above technical solution, the complex multi-view teacher model has good accuracy. Using a complex multi-view teacher model to supervise a simple single-view mapping model allows the intermediate quantities in the single-view mapping process to be supervised in a timely manner and the entire mapping process can be fully supervised, rather than just supervising the prediction results.

[0012] A preferred solution of the semantic vector map construction method is that in step S1, when multiple feature points are projected onto the same pixel in space, feature fusion is performed using a summation pooling method to achieve the integration of multi-view semantic features and generate three-dimensional voxel features.

[0013] By adopting the above technical solution, sum pooling retains the contribution of all projected features, avoiding the information dilution of mean pooling or the information loss of maximum pooling.

[0014] A preferred solution of the semantic vector map construction method is that step S2 specifically refers to: extracting two-dimensional semantic features from a single input image; projecting the three-dimensional voxel center point to the image coordinate system through the camera intrinsic parameter matrix, extrinsic parameter rotation matrix and translation vector to form a pairwise correspondence between voxels and pixels; using the inverse mapping function to reversely project the two-dimensional image features to the three-dimensional space to generate the three-dimensional voxel features of the student branch; projecting the three-dimensional voxel features into BEV features through a convolution-based BEV feature extractor, and finally inputting the BEV features into a Transformer-based map decoder to generate a semantic vector map and obtain the model prediction results.

[0015] By adopting the above technical solution, the process achieves efficient and lightweight three-dimensional semantic understanding through cascade processing of monocular image → 2D feature → 3D voxel → BEV → vector map.

[0016] A preferred solution of the semantic vector map construction method is that step S3 specifically refers to: in the student branch, using a classifier, taking three-dimensional voxel features as input, classifying each voxel in the 3D space to obtain a voxel classification matrix; selecting a category for each voxel, and obtaining the density information of each voxel by querying known parameters, and then obtaining a density matrix; using volume rendering technology, for each camera sensor of the teacher branch, taking the center of each pixel on its imaging plane as the starting point of the optical path, taking the direction from the camera optical center to the pixel center as the optical path direction, and accumulating the semantic vector values and densities of all visible voxels on the optical path; thereby, projecting the three-dimensional voxel features obtained in the student branch onto i imaging planes in the teacher branch to obtain two-dimensional features from the perspective of i teacher model sensors.

[0017] By adopting the above technical solution, 3D voxel classification is performed on the student branch, and then the volume rendering is projected to the teacher's perspective, achieving cross-modal (monocular → multi-sensor) knowledge distillation and 3D feature alignment.

[0018] A preferred solution of the semantic vector map construction method is that the step S4 specifically refers to: i=6, calculating the student branch body rendering two-dimensional features Relative teacher branch 2D features of losses, of which 、 and is a description of the dimension of the two-dimensional feature tensor, and the loss is ,in, .

[0019] By adopting the above technical solution, in the knowledge distillation framework, the volume rendering loss of the student branch and the teacher branch is calculated to improve the 3D perception ability of the student model under monocular input.

[0020] A preferred solution of the semantic vector map construction method is that step S5 specifically refers to: for the vector map predicted by the student branch and the vector map predicted by the teacher branch, the point-to-point loss is calculated as , the edge direction loss is , combined with hyperparameters and , the loss of the student branch prediction result relative to the teacher branch prediction result is ; The step S6 specifically refers to: using hyperparameters and , the losses are fused as follows: , and use this as the overall loss function to supervise the student branch training and obtain a unimodal online vector map construction model.

[0021] By adopting the above technical solution, in the knowledge distillation framework, map loss is used to guide the prediction results of the student branch to align with the high-quality output of the teacher branch, thereby improving the performance of the student model. After the loss is integrated, a lightweight unimodal online vector map construction model is obtained, which is close to the accuracy of complex multimodal models.

[0022] An electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the above-mentioned semantic vector map construction method is implemented.

[0023] A computer-readable storage medium stores a computer program, which implements the above-mentioned semantic vector map construction method when executed.

[0024] A computer program product, when the computer program product is run on a computer, the computer executes the above-mentioned semantic vector map construction method.

[0025] In summary, the semantic vector map construction method, device, storage medium and program product of the present application have the following beneficial effects: semantic vector map construction is realized with the help of a single image input, which will greatly reduce the configuration cost of the semantic vector map construction technology and promote technology promotion; with the help of knowledge distillation and volume rendering technology, the model can be supervised more specifically to ensure that the model has higher semantic vector map construction performance when the input sensor is small, thereby improving the accuracy of semantic vector map construction. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] In order to more clearly illustrate the embodiments of this specification or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in the embodiments of this specification. For ordinary technicians in this field, other drawings can also be obtained based on these drawings.

[0027] Figure 1 A flowchart of a method for constructing a semantic vector map provided in an embodiment of the present application. DETAILED DESCRIPTION

[0028] The technical solutions in the embodiments are described clearly and completely below. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the following embodiments, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0029] like Figure 1 As shown, a method for constructing a monocular visual semantic vector map based on knowledge distillation rendering includes the following steps: S1, pre-training semantic vector map construction model based on surround view images, as the knowledge distillation teacher branch; S2, builds a single-view-based semantic vector map model as a knowledge distillation student branch; S3, uses volume rendering technology to project the three-dimensional features of the student branch into two-dimensional features under multiple perspectives; S4, calculates the volume rendering loss of the student branch relative to the teacher branch; S5, calculates the map loss of the student branch relative to the teacher branch; S6, fuses the volume rendering loss and map loss, and completes the model training of the student branch based on this.

[0030] It should be noted that the order of steps S1 and S2 can be swapped, and the order of steps S4 and S5 can be swapped.

[0031] Step S1 specifically refers to: for the image data input of the vehicle-mounted surround view camera, using ResNet101 as the backbone network, extract the features of i two-dimensional images collected by i camera sensors Taking i=6 as an example, the features of 6 two-dimensional images collected by 6 camera sensors are extracted ,in , and It is a description of the dimension of the two-dimensional feature tensor, and then obtains multi-view semantic features f TI ={ f 1TI , f 2TI ,…, f 6TI}; Using the Transformer-based VIT network as the backbone network, the depth of the two-dimensional image features of each perspective is estimated to obtain the three-dimensional pyramidal feature space under that perspective. Features below , combined with the internal and external parameters of the camera The depth information of each feature point in the image is used to determine the 3D spatial position corresponding to each feature point, and then the features of each voxel in the 3D space are determined. In particular, when multiple feature points are projected onto the same pixel in the space, feature fusion is performed using summing pooling to achieve the integration of multi-view semantic features and generate 3D voxel features. ; Using a convolution-based BEV feature extractor, the 3D voxel features Projection as BEV features , the core part of the feature extractor consists of a convolutional network with the following convolution kernel sizes: Finally, the BEV features are input into the Transformer-based map decoder to generate a vector map and obtain the model prediction results. The above model is pre-trained as a complex multi-view teacher model.

[0032] In step S2, the student branch generates a semantic vector map using only a single image as input data. Specifically, it first uses ResNet101 as the backbone network to extract the semantic features of a single image. ; For model perception range , the coordinates of the center point of each voxel A point projected onto the image coordinate system using the camera's internal and external parameters , let the internal parameter matrix be , the extrinsic rotation matrix is , the translation vector is , obtained by the projection formula ,coordinate The pixel is located at The voxel in the input image corresponds to the pixel, thereby establishing a pairwise correspondence between the three-dimensional voxel and the two-dimensional pixel , where p is a pixel and v is a voxel. The mapping relationship h is not a single-valued function. A pixel may be mapped to multiple voxels. The subsequent process uses The inverse mapping relationship of is a single-valued function, so The non-uniqueness of does not affect the subsequent process. Back-projecting the 2D features into 3D space yields the 3D voxel features generated by a single image in the student branch. ; The 3D voxel features are extracted using a convolution-based BEV feature extractor. Projected into BEV features; finally, the BEV features are input into the Transformer-based map decoder to generate a semantic vector map and obtain the model prediction results .

[0033] Step S3 specifically refers to: in the student branch, using the MLP neural network classifier to classify the 3D voxel features As input, classify each voxel in 3D space to obtain the voxel classification matrix ; Use an Argmax layer to select the category of each voxel, and obtain the density information of each voxel by querying the known parameters, and then obtain the density matrix Using volume rendering technology, for each camera sensor of the teacher branch, the center of each pixel on its imaging plane is used as the starting point of the light path, and the direction from the camera optical center to the pixel center is used as the light path direction. The semantic vector value and density of all visible voxels on the light path are accumulated. Here, "visible" is determined by combining the density accumulation result with the volume rendering formula, that is, the density of each voxel on the light path is accumulated, and when it reaches a critical value, it is considered invisible; thus, the three-dimensional voxel features in the student branch are obtained. Projected onto the 6 imaging planes in the teacher branch, the 2D features of the 6 teacher model sensors are obtained. , can be constructed by relatively The loss is used to supervise the 3D feature generation process of the student model in different view angles.

[0034] Step S4 specifically refers to: calculating the student branch volume rendering two-dimensional features Relative teacher branch 2D features The loss of the volume rendering loss ,in .

[0035] Step S5 specifically refers to: the vector map predicted for the student branch and the vector map predicted by the teacher branch , calculate the point-to-point loss and edge direction loss , combined with hyperparameters and , get the map loss of the student branch prediction result relative to the teacher branch prediction result .

[0036] Step S6 specifically refers to: using hyperparameters and , the volume rendering loss obtained in step S4 and the map loss obtained in step S5 are fused as follows: , and use this as the overall loss function to supervise the student branch training, obtaining a lightweight unimodal online vector map construction model that is close to the accuracy of complex multimodal models.

[0037] The above implementation innovatively proposes a semantic vector map construction method based on a single image input. By introducing knowledge distillation and volume rendering techniques, this method significantly improves the accuracy of semantic vector map construction. Compared with related technologies, this method not only reduces dependence on hardware equipment but also reduces the complexity of data processing, thereby significantly reducing the configuration cost of the technology. This makes semantic vector map construction technology more accessible, providing strong support for the widespread application of autonomous driving systems and providing a sound theoretical foundation for building real-time perception systems for autonomous vehicles.

[0038] Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art can still modify the technical solutions described in the aforementioned embodiments, or make equivalent substitutions for some of the technical features therein. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present application should be included in the scope of protection of the present application.

Claims

1. A method for constructing a semantic vector map, characterized in that: The following steps are involved: S1, pre-training semantic vector map construction model based on surround view images, as the knowledge distillation teacher branch; S2, builds a single-view-based semantic vector map model as a knowledge distillation student branch; S3, uses volume rendering technology to project the three-dimensional features of the student branch into two-dimensional features under multiple perspectives; S4, calculates the volume rendering loss of the student branch relative to the teacher branch; S5, calculates the map loss of the student branch relative to the teacher branch; S6 fuses the volume rendering loss and map loss, and completes the model training of the student branch based on this.

2. The method for constructing a semantic vector map according to claim 1, wherein: The step S1 specifically includes: extracting features of i two-dimensional images captured by i camera sensors for the image data input of the vehicle-mounted surround view camera, thereby obtaining multi-view semantic features; Depth estimation is performed on the two-dimensional image features of each perspective to obtain the features in the three-dimensional prism-shaped feature space under that perspective. The depth information of each feature point in the camera's internal and external parameters is combined to determine the three-dimensional spatial position corresponding to each feature point, and then the features of each voxel in the three-dimensional space are determined. A convolution-based BEV feature extractor is used to project the three-dimensional voxel features into BEV features. Finally, the BEV features are input into a Transformer-based map decoder to generate a vector map and obtain the model prediction results. The above model is pre-trained as a complex multi-view teacher model.

3. The method for constructing a semantic vector map according to claim 2, wherein: In step S1, when multiple feature points are projected onto the same pixel in space, feature fusion is performed using a summation pooling method to integrate multi-view semantic features and generate three-dimensional voxel features.

4. The method for constructing a semantic vector map according to any one of claims 1 to 3, wherein: The step S2 specifically includes: extracting two-dimensional semantic features from a single input image; projecting the three-dimensional voxel center point to the image coordinate system through the camera intrinsic parameter matrix, extrinsic parameter rotation matrix and translation vector to form a pairwise correspondence between voxels and pixels; using the inverse mapping function to reversely project the two-dimensional image features to the three-dimensional space to generate the three-dimensional voxel features of the student branch; projecting the three-dimensional voxel features into BEV features through a convolution-based BEV feature extractor, and finally inputting the BEV features into a Transformer-based map decoder to generate a semantic vector map to obtain the model prediction results.

5. The method for constructing a semantic vector map according to claim 2, wherein: The step S3 specifically refers to: in the student branch, using a classifier, taking three-dimensional voxel features as input, classifying each voxel in the 3D space to obtain a voxel classification matrix; selecting a category for each voxel, and obtaining the density information of each voxel by querying known parameters, and then obtaining a density matrix; using volume rendering technology, for each camera sensor of the teacher branch, taking the center of each pixel on its imaging plane as the starting point of the light path, taking the direction from the camera optical center to the pixel center as the light path direction, and accumulating the semantic vector values and densities of all visible voxels on the light path; thereby, projecting the three-dimensional voxel features obtained in the student branch onto i imaging planes in the teacher branch to obtain two-dimensional features from the perspective of i teacher model sensors.

6. The method for constructing a semantic vector map according to claim 5, wherein: The step S4 specifically refers to: i=6, calculating the student branch body rendering two-dimensional features Relative teacher branch 2D features of losses, of which 、 and is a description of the dimension of the two-dimensional feature tensor, and the loss is ,in, .

7. The method for constructing a semantic vector map according to claim 6, wherein: The step S5 specifically refers to: for the vector map predicted by the student branch and the vector map predicted by the teacher branch, the point-to-point loss is calculated as , the edge direction loss is , combined with hyperparameters and , the loss of the student branch prediction result relative to the teacher branch prediction result is ; The step S6 specifically refers to: using hyperparameters and , the losses are fused as follows: , and use this as the overall loss function to supervise the student branch training and obtain a unimodal online vector map construction model.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the method for constructing a semantic vector map according to any one of claims 1 to 7 is implemented.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed, the semantic vector map construction method according to any one of claims 1 to 7 is implemented.

10. A computer program product, characterized in that When the computer program product runs on a computer, the computer executes the semantic vector map construction method according to any one of claims 1 to 7.