Semantic map construction method adopting multi-expert cross-modal distillation

The multi-expert cross-modal distillation method enhances semantic map construction by optimizing feature fusion and knowledge transfer, addressing dynamic scene challenges and resource constraints to improve map accuracy and robustness for autonomous driving.

CN120318643AActive Publication Date: 2025-07-15BEIHANG UNIV

Patent Information

Application Number
CN202510811754.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-18
Publication Date
2025-07-15
Estimated Expiration
2045-06-18

AI Technical Summary

Technical Problem

The existing high-precision map construction method has reduced the accuracy of mapping construction in dynamic scenarios, consumes a lot of computing resources, and is difficult to meet the real-time requirements of autonomous driving systems, and cannot be effectively deployed in resource-constrained environments.

Method used

The semantic map construction method of multi-expert cross-modal distillation is adopted. Through multi-view image preprocessing, geometric correction and time synchronization, combined with convolutional neural network, multi-head self-attention mechanism and cross-attention mechanism, LiDAR features, SDmap and HDmap information are integrated, and the student model is trained by using the teacher model and the Coach model to design the loss function of classification loss and regression loss to optimize the student model.

Benefits of technology

It significantly improves the construction accuracy and robustness of semantic maps, and can efficiently perform 3D object detection tasks in resource-constrained environments, adapt to changes in complex scenarios, and support the further development of autonomous driving technology.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120318643A_ABST
    Figure CN120318643A_ABST
Patent Text Reader

Abstract

The invention discloses a semantic map construction method adopting multi-expert cross-modal distillation, and belongs to the technical field of semantic map construction methods, and the method comprises the steps: carrying out the preprocessing, feature extraction and geometric correction of a multi-view image, obtaining a LiDAR feature, fusing the LiDAR feature, an SDmap and an HDmap, generating a scene feature through a cross-modal feature fusion technology, and carrying out the segmentation of the scene feature. False point cloud features and visual features are output through the Coach model, the student model begins to learn and is jointly guided by the teacher model and the Coach model, and a loss function is adopted to evaluate the student model and train the student model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for constructing a semantic map, in particular to a method for constructing a semantic map using multi-expert cross-modal distillation, belonging to the technical field of semantic map construction methods. Background Art

[0002] Although the existing high-precision map generation methods based on multi-view images can provide certain geometric and semantic expressions in static scenes, in dynamic scenes, feature fusion is insufficient, and it is easily affected by factors such as occlusion, dynamic targets, and perspective changes, resulting in a decrease in mapping accuracy and affecting the stability and reliability of the map. At the same time, although the map construction technology based on the fusion of lidar and images can provide high geometric accuracy, it consumes a large amount of computing resources and is difficult to meet the real-time requirements of the autonomous driving system. In terms of the time consistency modeling of multi-frame data, the capabilities of traditional methods are also relatively limited, and it is difficult to effectively handle the dynamic changes of target instances in the scene. Especially during the map update process, due to the lack of a strong expression of dynamic change relevance, it often leads to discontinuity and inconsistency of the results, further affecting the accuracy and reliability of the map. In addition, the existing high-precision map construction methods lack optimization for the lightweight and resource adaptability of the model, resulting in their ineffective deployment in resource-constrained scenarios. This problem limits the application of autonomous driving technology in resource-limited environments, especially in scenarios with limited computing power and storage capacity, and the potential of high-precision map construction technology cannot be fully utilized.

[0003] In summary, there are still many technical bottlenecks in the current high-precision map dynamic detection and update technology in aspects such as degraded image processing, multi-view feature fusion, dynamic change detection, and lightweight design. Therefore, there is an urgent need for an efficient solution that combines general image restoration strategies, dynamic detection, and map update technologies to achieve high-precision and real-time map construction and be able to adapt to resource-constrained application environments through lightweight optimization, providing strong support for the further development of autonomous driving technology. Summary of the Invention

[0004] The main purpose of the present invention is to provide a method for constructing a semantic map using multi-expert cross-modal distillation.

[0005] The object of the present invention can be achieved by adopting the following technical solutions: A method for constructing a semantic map using multi-expert cross-modal distillation includes the following steps: Step S1: Perform preprocessing, feature extraction, and geometric correction of multi-view images to obtain LiDAR features, fuse the LiDAR features, SDmap, and HDmap, and generate scene features through cross-modal feature fusion technology. Step S2: Output the pseudo point cloud features and visual features through the Coach model; Step S3: The student model starts to learn under the joint guidance of the teacher model and the Coach model; Step S4: Use a loss function to evaluate the student model and conduct training.

[0006] Preferably, in Step S2, when outputting the pseudo point cloud features and visual features through the Coach model, the LiDAR features extracted by the teacher model are used to guide the Coach model to generate pseudo point cloud features. Then, the Coach model combines the SDmap and HDmap for feature fusion, and the fused features output by the teacher model are further used to guide the feature extraction and learning process of the Coach model.

[0007] Preferably, the specific joint guidance in Step S3 includes that the student model outputs its visual features, and the visual features provided by the Coach model are used to guide the feature extraction process of the student model; The fused features of the teacher model and the Coach model act on the student model to optimize its learning ability when processing the fused features.

[0008] Preferably, in Step S4, the loss function includes classification loss and regression loss, which are used to evaluate the student model and verify the robustness and accuracy of the student model. If the output requirements are not met, a medium-distance perception mechanism is adopted to optimize the student model and retrain it, and the robustness and accuracy of the student model are verified according to the loss evaluation results.

[0009] Preferably, by preprocessing the data from different sensors, the geometric correction and time synchronization of multi-view images and LiDAR point cloud data are specifically as follows: Geometric correction aligns the images and point cloud data under different views so that the objects in the images can be precisely matched with the physical positions of the LiDAR point cloud; Time synchronization ensures that all sensor data are aligned in the time dimension, The preprocessed image data will enter the feature extraction module; These images are processed using a deep learning backbone network; Low-order features are extracted through a convolutional neural network. As the network deepens, the network will gradually capture higher-order abstract features, and a multi-head self-attention mechanism is introduced, which can model the semantic information in the image globally; After the image feature extraction is completed, it is fused with the LiDAR point cloud and the high-precision map; Among them, the high-precision map contains a detailed description of the road network, lane lines, and traffic sign information, while the LiDAR point cloud provides data on the three-dimensional spatial positions of the scene objects; Integrating data from different sources employs cross-modal feature fusion technology; First, the geometric information of SDmap and HDmap is processed by a geometric encoder; During the processing by the geometric encoder, the geometric features of the map are converted into high-dimensional feature representations, which can accurately describe the spatial distribution of roads and objects; The semantic information in the map is transformed into a high-dimensional feature tensor by a semantic encoder; Spatial features of LiDAR point clouds are extracted through a deep neural network. The spatial features include the shape, distance of objects, and their relative positions to other objects; Based on the spatial features, precise three-dimensional coordinates and geometric structures of the objects in the scene are obtained; Through a cross-attention mechanism, LiDAR and map features are fused. The cross-attention mechanism can calculate the correlation between modalities, enabling the model to extract key information from each modality and effectively align data from different modalities; The cross-attention mechanism calculates the correlation between the geometric and semantic features at the corresponding positions in the high-precision map for each point in the LiDAR point cloud; After feature fusion, a high-dimensional feature tensor containing rich geometric and semantic information is obtained.

[0010] Preferably, at the beginning of the training process of the Coach model in step S2, the Coach model first outputs pseudo-point cloud features and visual features; The generation process of the pseudo-point cloud features is guided by the LiDAR features extracted by the teacher model; In previous steps, the teacher model extracts features from LiDAR data through a deep learning network to capture the geometric and spatial information of the scene; Using the prior knowledge of the teacher model, that is, the LiDAR features of the teacher model, guides the Coach model to maintain a certain spatial structure and geometric information during the generation of pseudo-point clouds, ensuring that the generated pseudo-point cloud features are consistent with the actual LiDAR point cloud data; The Coach model combines SDmap and HDmap for feature fusion; Through cross-modal feature fusion technology, the Coach model can effectively integrate different features from visual and map data; SDmap and HDmap contain road information, lane markings, and semantic information of traffic signs; By encoding the geometric and semantic features in SDmap and HDmap, the Coach model can take into account both the static information in the map and the changing scene data during the generation of pseudo-point cloud features; The output features of the teacher model will serve as guidance signals to help the coach model more accurately fuse visual and map features; The teacher model fuses features from different modalities and generates high-quality fused features; During the training process, the Coach model uses the fusion features generated by the teacher model to ensure that it can better capture the geometric structure and semantic information of the scene when processing visual information and map data; In the process of feature fusion, the fused features provided by the teacher model guide the learning process of the Coach model through the supervision mechanism; Guided by cross-modal features, the Coach model can effectively integrate visual information and enhance its performance in complex environments based on the static and dynamic information of the map.

[0011] Preferably, in the knowledge distillation phase of step S3, the student model begins to enter the learning process, and gradually improves its ability in visual and map information processing through the joint guidance of the teacher model and the coach model; First, the student model outputs its preliminary visual features; The student model generates these features based on the image data, and uses the visual features learned by the coach model as a guiding signal to help the student model extract visual features more accurately; The learning process of the student model is further optimized by the fusion features of the teacher model and the coach model; The feature output of the teacher model and the fusion features of the coach model will work together on the student model, further improving the learning ability of the student model through joint guidance; The teacher model has learned how to effectively fuse visual information, LiDAR features, SDmap and HDmap features through previous training, while the Coach model builds on this and uses its experience with pseudo point clouds and visual features to help the student model better understand the spatial relationships and semantic structure of the scene.

[0012] Preferably, step S4 comprehensively evaluates and optimizes the student model by designing an effective loss function according to the student model to ensure its robustness and high-precision 3D object detection capability; The loss function includes classification loss and regression loss, which are used to evaluate the target category prediction ability and the regression accuracy of the 3D bounding box of the student model respectively; For classification accuracy, we use the quality focus loss function to measure the deviation between the target category probability distribution predicted by the student model and the true category label. The formula is:

[0013] Where S is the total number of targets, is the quality focus loss function, , and are the predicted probabilities of the teacher model, Coach model, and student model respectively, and ωi is the weight for target i, used to dynamically adjust the contributions of different targets; For the regression accuracy, the formula for evaluating the difference between the 3D bounding box parameters predicted by the student model and the ground truth using the smooth L1 loss function is:

[0014] where, Smooth L1 is the smooth L1 loss function; , and are the teacher model, ωi is the weight for target i, the bounding box regression parameters of the Coach model and the student model respectively; Finally, the classification loss and regression loss are combined into the total loss by weighted summation, and the formula for guiding the optimization of the student model is:

[0015] where: is the classification loss, is the regression loss, is the total loss; According to the evaluation index of the loss function, if the output requirement is not met, the student model is further optimized by using the medium-distance perception mechanism; This mechanism dynamically adjusts the loss weights of different targets through distance perception response distillation, and focuses on optimizing the detection accuracy of medium and short distance targets. The specific formula is:

[0016] where, is the distance between target i and the sensor, is the normalized weight of the target, is the fixed weight of the long-distance target, and A, B, C, and δ are adjustment parameters; According to the distance perception response distillation strategy, the student model is trained and optimized multiple times, and the evaluation index of the loss function is recalculated. If it meets the threshold requirement of 0.1, it is considered that the output accuracy requirement is met, ensuring that the student model can efficiently perform 3D object detection tasks in resource-constrained environments, while maintaining a high level of robustness and accuracy in different complex scenarios.

[0017] The beneficial technical effects of the present invention: A semantic map construction method using multi-expert cross-modal distillation provided by the present invention effectively integrates LiDAR features, SDmap maps, and HDmap information through multi-perspective image preprocessing, geometric correction and time synchronization, and cross-modal feature fusion technology. By using convolutional neural networks, multi-head self-attention mechanisms, and cross-attention mechanisms, it fully extracts and fuses the geometric and semantic information of multi-source data to generate scene features with rich details. Compared with traditional methods, it significantly improves the construction accuracy of semantic maps and more accurately reflects the real situation of complex scenes.

[0018] Introduce a training method of multi-expert (teacher model, Coach model) cross-modal distillation. The teacher model provides stable and reliable prior knowledge and fusion feature guidance for the Coach model and the student model. The Coach model combines map information to generate pseudo-point cloud features and visual features, enhancing the adaptability to complex environments. Under the joint guidance of the teacher model and the Coach model, the student model can better learn the association between visual and map information, improve the feature extraction and processing capabilities in different scenarios, enhance the robustness of the model, and effectively handle complex situations such as occlusion, dynamic targets, and perspective changes.

[0019] Design a loss function that includes classification loss and regression loss. Use the quality focal loss function and the smooth L1 loss function to evaluate the target category prediction ability and 3D bounding box regression accuracy of the student model respectively. Obtain the total loss through weighted summation to guide the optimization training of the student model. At the same time, use the medium-distance perception mechanism to dynamically adjust the loss weights of different targets, and focus on optimizing the detection accuracy of medium- and short-distance targets. After multiple trainings and optimizations, the student model can still efficiently perform 3D object detection tasks in resource-constrained environments, maintain high accuracy and robustness in complex scenes, and provide reliable object detection support for applications such as autonomous driving.

[0020] This method considers the lightweight and resource adaptability of the model during the design process. Through knowledge distillation, the student model learns the experience of the teacher model and the Coach model, avoiding the high computational cost brought by complex model structures. On the premise of ensuring the map construction accuracy and object detection performance, reduce the model parameters and computational amount, so that the model can be effectively deployed in scenarios with limited computing power and storage capacity, expanding the application scope of semantic map construction technology in fields such as autonomous driving. Brief Description of the Drawings

[0021] Figure 1 It is a flowchart of a preferred embodiment of a semantic map construction method using multi-expert cross-modal distillation according to the present invention. Detailed Embodiments

[0022] To make the technical solutions of the present invention clearer and more definite to those skilled in the art, the present invention will be further described in detail below in conjunction with embodiments and the accompanying drawings. However, the implementation manners of the present invention are not limited thereto.

[0023] A semantic map construction method using multi-expert cross-modal distillation. Step S1: First, perform preprocessing, feature extraction, and geometric correction on multi-view images to obtain LiDAR features. Then, fuse the LiDAR features, SDmap, and HDmap, and generate richer scene features through cross-modal feature fusion technology. At this time, freeze the parameters in the teacher model to ensure that its feature extraction and fusion modules can provide stable and efficient learning support for subsequent models.

[0024] Step S2: The Coach model first outputs pseudo-point cloud features and visual features. In this process, use the LiDAR features extracted by the teacher model to guide the Coach model to generate pseudo-point cloud features. Then, the Coach model combines the SDmap and HDmap for feature fusion, and further guides the feature extraction and learning process of the Coach model through the fusion features output by the teacher model to ensure that it can have higher accuracy and consistency when processing visual information.

[0025] Step S3: Knowledge distillation stage: The student model starts to learn under the joint guidance of the teacher model and the Coach model. First, the student model outputs its visual features, and the visual features provided by the Coach model guide the feature extraction process of the student model. Subsequently, the fusion features of the teacher model and the Coach model act on the student model together to optimize its learning ability when processing the fusion features. The goal of this stage is to help the student model more accurately understand the association between visual and map information.

[0026] Step S4: Use a loss function to evaluate and train the student model: The loss function includes classification loss and regression loss to evaluate the student model and verify the robustness and accuracy of the student model.

[0027] Step S1 includes that we first perform preprocessing on data from different sensors, and perform geometric correction and time synchronization on multi-view images and LiDAR point cloud data. This process is crucial as it ensures the consistency and coordination of information from different sources in the spatio-temporal dimension. Specifically, geometric correction aligns the images and point cloud data from different perspectives, enabling the objects in the images to be precisely matched with the physical positions of the LiDAR point cloud. Time synchronization ensures that all sensor data is aligned in the time dimension, avoiding data inconsistency problems caused by time differences. Only after this series of preprocessing operations can the subsequent feature extraction and fusion ensure the accuracy and consistency of the data.

[0028] Next, the preprocessed image data will enter the feature extraction module. We use a deep learning backbone network to process these images. First, low-level features are extracted through a convolutional neural network. These features are usually information about the shape and texture of objects in the scene. As the network deepens, it gradually captures higher-order abstract features, such as the spatial relationships between objects and structural information. To further improve the model's expressive power, we also introduce a multi-head self-attention mechanism, which can model the semantic information in the image globally. Through this mechanism, the model can better understand the context relationships between objects in the image and their interactions, generating high-dimensional feature tensors from multiple perspectives.

[0029] After the image feature extraction is completed, we fuse it with the LiDAR point cloud and high-precision maps (SDmap and HDmap). The high-precision maps contain detailed descriptions of information such as road networks, lane lines, and traffic signs, while the LiDAR point cloud provides data on the three-dimensional spatial positions of objects in the scene. To better integrate these data from different sources, we adopt cross-modal feature fusion technology.

[0030] First, the geometric information of SDmap and HDmap (such as lane boundaries, road centerlines, etc.) is processed by a dedicated geometric encoder. This process converts the geometric features of the map into high-dimensional feature representations that can accurately describe the spatial distribution of roads and objects. At the same time, the semantic information in the map (such as lane types, traffic signs, pedestrians, etc.) is also transformed into high-dimensional feature tensors through a semantic encoder. These semantic features provide important information for subsequent object recognition and environment understanding.

[0031] At the same time, the LiDAR point cloud extracts its spatial features through a deep neural network. These features include the shape, distance of objects, and their relative positions to other objects. Through these spatial features, we can obtain the precise three-dimensional coordinates and geometric structures of objects in the scene.

[0032] Then, through the Cross-attention Mechanism, we fuse the LiDAR and map features. The Cross-attention Mechanism can calculate the correlation between each modality, enabling the model to extract key information from each modality and effectively align the data of different modalities. Specifically, the Cross-attention Mechanism calculates the correlation between each point in the LiDAR point cloud and the geometric and semantic features at the corresponding position in the high-precision map. This mechanism ensures that the system not only retains the static information in the high-precision map but also dynamically fuses the real-time changing information captured by the LiDAR sensor.

[0033] After feature fusion, we obtain a high-dimensional feature tensor containing rich geometric and semantic information, which is crucial for subsequent model training. To ensure that these features can stably guide the subsequent model, we freeze the parameters in the teacher model. Freezing the parameters means that during subsequent training, the feature extraction and fusion modules of the teacher model will not change. In this way, the features generated by the teacher model will serve as fixed supervision signals, providing reliable inputs for the Coach model and the student model. This approach ensures that the teacher model can provide consistent and stable guidance signals for other models, thereby helping the subsequent models better learn the task objectives.

[0034] Step S2 includes that in step S2, the training process of the Coach model begins, focusing on how to transfer the knowledge learned by the teacher model to the Coach model and generate pseudo-point cloud features and visual features on this basis. First, the Coach model outputs pseudo-point cloud features and visual features. The generation process of the pseudo-point cloud features is guided by the LiDAR features extracted by the teacher model. Specifically, the teacher model has effectively extracted features from LiDAR data through a deep learning network in the previous step, capturing the geometric and spatial information of the scene. Therefore, the Coach model will predict pseudo-point cloud features based on the LiDAR features generated by the teacher model. These pseudo-point cloud features can simulate the structure and distribution of real LiDAR point clouds, making up for the limitations of the LiDAR sensor itself, especially in the case of insufficient LiDAR data. The generation process of the pseudo-point cloud features depends on the generation ability of the deep learning network, especially the use of technologies such as generative adversarial networks (GANs), enabling the Coach model to generate corresponding pseudo-point cloud features through visual information. The key in this process is to use the prior knowledge of the teacher model, that is, the LiDAR features of the teacher model, to guide the Coach model to maintain a certain spatial structure and geometric information when generating pseudo-point clouds, so as to ensure that the generated pseudo-point cloud features can be consistent with the actual LiDAR point cloud data.

[0035] Next, the Coach model will perform feature fusion by combining SDmap and HDmap. Through cross-modal feature fusion technology, the Coach model can effectively integrate different features from visual and map data. Specifically, SDmap and HDmap contain semantic information such as road information, lane markings, and traffic signs, which are crucial for scene understanding. By encoding the geometric and semantic features in SDmap and HDmap, the Coach model can take into account both the static information in the map and the changing scene data during the process of generating pseudo-point cloud features.

[0036] To further improve the effect of feature fusion, the output features of the teacher model will serve as guiding signals to help the Coach model more accurately fuse visual and map features. The teacher model has learned how to fuse features of different modalities together and generate high-quality fused features. Therefore, during the training process, the Coach model will rely on the fused features generated by the teacher model to ensure that it can better capture the geometric structure and semantic information of the scene when processing visual information and map data.

[0037] During the feature fusion process, the fused features provided by the teacher model guide the learning process of the Coach model through a supervision mechanism. The core of this process is to enable the Coach model to not only effectively fuse visual information but also adopt the static and dynamic information of the map through the guidance of cross-modal features, further enhancing its performance in dealing with complex environments. With the stable guidance of the teacher model, the Coach model can achieve higher accuracy and consistency in the processing of visual features, ensuring its robustness and precision in object recognition and scene understanding tasks.

[0038] Step S3 includes that in the knowledge distillation stage of step S3, the student model begins to enter the learning process and gradually improves its ability to process visual and map information under the joint guidance of the teacher model and the Coach model. First, the student model outputs its initial visual features. The student model generates these features based on image data. However, relying solely on its own feature extraction ability may be difficult to achieve ideal results. Therefore, the visual features of the Coach model play a crucial role at this stage. Specifically, the visual features learned by the Coach model serve as guiding signals to help the student model extract visual features more accurately. This process not only utilizes the experience of the Coach model but also ensures that the student model can benefit from the existing knowledge of the Coach model and quickly improve the accuracy of its visual information processing.

[0039] Next, the learning process of the student model is further jointly optimized by the fused features of the teacher model and the Coach model. At this stage, the feature output of the teacher model and the fused features of the Coach model will act on the student model together to further improve the learning ability of the student model through joint guidance. The teacher model has learned how to effectively fuse visual information, LiDAR features, SDmap, and HDmap features through previous training, while the Coach model, based on this, uses its experience of pseudo point clouds and visual features to help the student model better understand the spatial relationship and semantic structure of the scene. By inputting these two sources of features into the student model together, the student model can not only learn the deep connection between visual features and map information but also gradually improve its performance in complex environments.

[0040] Specifically, the fused features of the teacher model and the Coach model guide the feature fusion ability of the student model through a refined learning process. During this process, the student model not only relies on its own learning ability to extract features but also combines the guiding signals from the teacher and Coach models to optimize its processing strategy when fusing different modality information. Through this dual supervision mechanism, the student model can gradually strengthen its understanding of scene geometry and semantic information, enabling it to make more accurate judgments in more complex and variable environments.

[0041] The goal of this knowledge distillation stage is to enable the student model to more accurately understand the association between visual and map information. In practical applications, the student model not only needs to extract geometric features such as the shape and position of objects from visual information but also conduct global understanding by combining map information (such as lane types, traffic signs, road boundaries, etc.). The teacher model and the Coach model, through joint guidance, help the student model form a powerful fusion ability, enabling it to better process multi-modal information and lay a solid foundation for subsequent tasks (such as object detection, path planning, etc.).

[0042] Step S4 includes comprehensively evaluating and optimizing the student model according to the student model by designing an effective loss function to ensure its robustness and high-precision 3D object detection ability. The loss function includes a classification loss and a regression loss, which are used to evaluate the target category prediction ability of the student model and the regression accuracy of the 3D bounding box, respectively.

[0043] Regarding the classification accuracy, we use the Quality Focal Loss (QFL) to measure the deviation between the target category probability distribution predicted by the student model and the true category label. The formula is: ; where S is the total number of targets, is the Quality Focal Loss function, , and are the predicted probabilities of the teacher model, the Coach model, and the student model respectively, and ωi is the weight for target i, which is used to dynamically adjust the contributions of different targets.

[0044] Regarding the regression accuracy, we use the Smooth L1 Loss function to evaluate the difference between the 3D bounding box parameters (such as center coordinates, dimensions, and orientations) predicted by the student model and the true values. The formula is:

[0045] where Smooth L1 is the Smooth L1 Loss function; , and are the teacher model respectively; ωi is the weight for target i, and the bounding box regression parameters of the Coach model and the student model.

[0046] Finally, the classification loss and the regression loss are combined into the total loss by weighted summation to guide the optimization of the student model. The formula is: ; where: is the classification loss, is the regression loss, is the total loss; According to the evaluation index of the loss function, if the output requirement is not met, the medium-distance perception mechanism is further used to optimize the student model. This mechanism dynamically adjusts the loss weights of different targets through distance perception response distillation, and focuses on optimizing the detection accuracy of medium- and short-distance targets. The specific formula is:

[0047] where, is the distance between target i and the sensor, is the normalized weight of the target, is the fixed weight of the long-distance target, and A, B, C, and δ are adjustment parameters.

[0048] According to the distance perception response distillation strategy, the student model is trained and optimized multiple times, and the evaluation index of the loss function is recalculated. If it meets the threshold requirement of 0.1, it is considered that the accuracy requirement of the output is met, ensuring that the student model can efficiently perform 3D object detection tasks in resource-constrained environments, and at the same time maintaining a high level of robustness and accuracy in different complex scenarios.

[0049] The above is only a further embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the scope disclosed by the present invention, according to the technical solution and its concept of the present invention, makes equivalent substitutions or changes, all belong to the protection scope of the present invention.

Claims

1. A method for constructing a semantic map using multi-expert cross-modal distillation, characterized in that: It includes the following steps: Step S1: Preprocess, extract features, and perform geometric correction on multi-view images to obtain LiDAR features. Fusion of LiDAR features, SDmap, and HDmap, and generate scene features through cross-modal feature fusion technology; Step S2: Output pseudo-point cloud features and visual features through the Coach model; Step S3: The student model starts to learn under the joint guidance of the teacher model and the Coach model; Step S4: Use a loss function to evaluate the student model and perform training.

2. The semantic map construction method using multi-expert cross-modal distillation according to claim 1, characterized in that: In Step S2, the LiDAR features extracted by the teacher model are used in the output of pseudo-point cloud features and visual features through the Coach model to guide the Coach model to generate pseudo-point cloud features. Then, the Coach model combines SDmap and HDmap for feature fusion, and the fusion features output by the teacher model further guide the feature extraction and learning process of the Coach model.

3. A semantic map construction method using multi-expert cross-modal distillation according to claim 1, characterized in that: The specific joint guidance in Step S3 includes the student model outputting its visual features, and guiding the feature extraction process of the student model through the visual features provided by the Coach model; The fusion features of the teacher model and the Coach model act on the student model to optimize the learning ability when processing the fusion features.

4. A semantic map construction method using multi-expert cross-modal distillation according to claim 1, characterized in that: In Step S4, the loss function includes classification loss and regression loss to evaluate the student model and verify the robustness and accuracy of the student model. If the output requirements are not met, a medium-distance perception mechanism is used to optimize the student model and retrain it, and the robustness and accuracy of the student model are verified according to the loss evaluation results.

5. A semantic map construction method using multi-expert cross-modal distillation according to claim 1, characterized in that: By preprocessing data from different sensors, geometric correction and time synchronization of multi-view images and LiDAR point cloud data specifically include the following: Geometric correction aligns images and point cloud data from different perspectives, enabling objects in the images to be precisely matched with the physical positions of the LiDAR point cloud; Time synchronization ensures that all sensor data is aligned in the time dimension, The preprocessed image data will enter the feature extraction module; Use a deep learning backbone network to process these images; Extract low-level features through a convolutional neural network. As the network deepens, the network will gradually capture higher-level abstract features, and a multi-head self-attention mechanism is introduced, which can model semantic information in the image globally; After the image feature extraction is completed, it is fused with the LiDAR point cloud and the high-precision map; Among them, the high-precision map contains detailed descriptions of the road network, lane lines, and traffic sign information, while the LiDAR point cloud provides data on the three-dimensional spatial positions of scene objects; Cross-modal feature fusion technology is used to integrate data from different sources; First, the geometric information of SDmap and HDmap will be processed by a geometric encoder; The process of the geometric encoder converts the geometric features of the map into high-dimensional feature representations, which can accurately describe the spatial distribution of roads and objects; The semantic information in the map is converted into a high-dimensional feature tensor through a semantic encoder; The LiDAR point cloud uses a deep neural network to extract spatial features, including the shape, distance, and relative position of the object to other objects. Through spatial features, accurate three-dimensional coordinates and geometric structures of objects in the scene are obtained; The cross-attention mechanism is used to fuse LiDAR and map features. By calculating the correlation between the modalities, the cross-attention mechanism enables the model to extract key information from each modality and effectively align data from different modalities. The cross-attention mechanism calculates the correlation between the geometric and semantic features of each point in the LiDAR point cloud and the corresponding position in the HD map; After feature fusion, a high-dimensional feature map containing rich geometric and semantic information is obtained.

6. A semantic map construction method using multi-expert cross-modal distillation according to claim 2, characterized in that: In step S2, the training process of the Coach model begins. First, the Coach model outputs pseudo point cloud features and visual features; The pseudo point cloud feature generation process is guided by the LiDAR features extracted by the teacher model; In the previous step, the teacher model extracts features from LiDAR data through a deep learning network to capture the geometric and spatial information of the scene; The prior knowledge of the teacher model, i.e. the LiDAR features of the teacher model, is used to guide the Coach model to maintain a certain spatial structure and geometric information when generating pseudo point clouds, ensuring that the generated pseudo point cloud features are consistent with the actual LiDAR point cloud data. The Coach model combines SDmap and HDmap for feature fusion; Through cross-modal feature fusion technology, the Coach model can effectively integrate different features from visual and map data; SDmap and HDmap contain road information, lane markings, and traffic sign semantic information; By encoding the geometric and semantic features in SDmap and HDmap, the Coach model can take into account both the static information and the changing scene data in the map when generating pseudo point cloud features; The output features of the teacher model will serve as guidance signals to help the coach model more accurately fuse visual and map features; The teacher model fuses features from different modalities and generates high-quality fused features; During the training process, the Coach model uses the fusion features generated by the teacher model to ensure that it can better capture the geometric structure and semantic information of the scene when processing visual information and map data; In the process of feature fusion, the fused features provided by the teacher model guide the learning process of the Coach model through the supervision mechanism; Guided by cross-modal features, the Coach model can effectively integrate visual information and enhance its performance in complex environments based on the static and dynamic information of the map.

7. A semantic map construction method using multi-expert cross-modal distillation according to claim 3, characterized in that: In the knowledge distillation phase of step S3, the student model begins to enter the learning process and gradually improves its ability in visual and map information processing through the joint guidance of the teacher model and the coach model; First, the student model outputs its preliminary visual features; The student model generates these features based on the image data, and uses the visual features learned by the coach model as a guiding signal to help the student model extract visual features more accurately; The learning process of the student model is further jointly optimized by the combined features of the teacher model and the Coach model; The feature output of the teacher model and the combined features of the Coach model will act on the student model together, further enhancing the learning ability of the student model through joint guidance; The teacher model has learned through previous training how to effectively fuse visual information, LiDAR features, SDmap, and HDmap features. On this basis, the Coach model uses its experience of pseudo-point clouds and visual features to help the student model better understand the spatial relationship and semantic structure of the scene.

8. A semantic map construction method using multi-expert cross-modal distillation according to claim 4, characterized in that: Step S4 comprehensively evaluates and optimizes the student model according to the student model by designing an effective loss function to ensure its robustness and high-precision 3D object detection ability; The loss function includes a classification loss and a regression loss, which are used to evaluate the object category prediction ability of the student model and the regression accuracy of the 3D bounding box respectively; For the classification accuracy, we use the quality focal loss function to measure the deviation between the predicted object category probability distribution of the student model and the true category label. The formula is: where S is the total number of targets, QFL is the quality focal loss function, , are the predicted probabilities of the teacher model, Coach model, and student model respectively, and ωi is the weight for target i, which is used to dynamically adjust the contributions of different targets; For the regression accuracy, the smooth L1 loss function is used to evaluate the difference between the predicted 3D bounding box parameters of the student model and the true values. The formula is: Among them, Smooth L1 is the Smooth L1 loss function; , and are the teacher models respectively; is the weight for target i, the bounding box regression parameters of the Coach model and the student model; Finally, the classification loss and the regression loss are combined into a total loss by weighted summation to guide the optimization of the student model. The formula is: Wherein: is the classification loss, is the regression loss, is the total loss; according to the loss function evaluation index, if the output requirement is not met, a medium-distance perception mechanism is further used to optimize the student model; This mechanism dynamically adjusts the loss weights of different targets through distance-aware response distillation, focusing on optimizing the detection accuracy of medium and short-distance targets. The specific formula is: Among them, is the distance between the target and the sensor, is the normalized weight of the target, is the fixed weight of the far - distance target, and A, B, C and δ are adjustment parameters; According to the distance perception response distillation strategy, the student model is trained and optimized multiple times, and the evaluation index of the loss function is recalculated. If it meets With a threshold requirement of 0.1, it is considered that the output accuracy requirement is met, ensuring that the student model can efficiently perform 3D object detection tasks in resource-constrained environments while maintaining a high level of robustness and accuracy in different complex scenarios.

Citation Information

Patent Citations

  • Track target point prediction method based on knowledge distillation

    CN116579423A

  • Cross-modal distillation LiDAR point cloud tracking method

    CN117994694A

  • Online vector map construction method and device based on volume rendering knowledge distillation

    CN118644603A

  • Visual identification model training method and device for vehicle automatic driving

    CN119027898A

  • Vehicle-road cooperation 3D target detection method and device based on knowledge distillation

    CN119027914A

Cited By

  • Multimodal three-dimensional target detection method and device based on diffusion type knowledge distillation

    CN120953598A