High-precision map dynamic detection and updating method adopting general image recovery strategy

By adopting a general image recovery strategy and dynamic detection module, the mapping accuracy and consistency of high-precision maps in dynamic scenarios are solved, and high-quality local high-precision maps are generated, supporting the reliability and safety of the autonomous driving system.

CN120472490AActive Publication Date: 2025-08-12BEIHANG UNIV
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510550786.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-29
Publication Date
2025-08-12
Estimated Expiration
2045-04-29

AI Technical Summary

Technical Problem

The existing high-precision maps have reduced mapping accuracy in dynamic scenarios, insufficient fusion of multi-view features, large computing resources, insufficient real-time and lightweight design, resulting in discontinuous and inconsistent map generation results, making it difficult to adapt to the needs of autonomous driving.

Method used

Using a general image recovery strategy, features are extracted and enhanced layer by layer through the feature generation module and feature selection module, combined with multi-dimensional feature fusion and cross-attention mechanism, a high-precision map is generated, combined with the dynamic detection module to identify newly added, deleted or unchanged elements, and geometric adjustment and semantic correction are performed through the map update module to generate an accurate and coherent local high-precision map.

Benefits of technology

It significantly improves the details restoration and global consistency of degraded images, accurately recognizes map changes, generates high-quality local high-precision maps, and improves the reliability and safety of the autonomous driving system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120472490A_ABST
    Figure CN120472490A_ABST
Patent Text Reader

Abstract

The invention discloses a high-precision map dynamic detection and updating method adopting a general image recovery strategy, and belongs to the technical field of high-precision map dynamic detection and updating methods. Through a pre-training process, degraded images are input into a network, and features are extracted and enhanced layer by layer by adopting a feature generation module and a feature selection module; the method comprises the following steps of: S1, realizing image recovery in combination with multi-dimensional feature fusion and a cross attention mechanism, improving detail reduction and global consistency of a degraded image, identifying newly added, deleted or unchanged elements in a map by adopting a dynamic detection module according to BEV representation generated by a pre-training model in the step S1 and in combination with a historical state of a high-precision map, and outputting a difference map. S2, according to the difference map generated in the step S2, performing geometric adjustment and semantic correction on the marked newly-added and deleted areas through a map updating module, generating an accurate and coherent local high-precision map, and performing comprehensive evaluation on dynamic detection and map updating performance through a model evaluation framework.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for dynamic detection and updating of high-precision maps, and in particular to a method for dynamic detection and updating of high-precision maps using a universal image recovery strategy, belonging to the technical field of methods for dynamic detection and updating of high-precision maps. Background Art

[0002] Existing high-precision map generation methods can provide certain geometric and semantic representations in static scenes. However, when dealing with dynamic scenes, due to occlusion, dynamic objects, and changes in viewpoint, feature fusion is insufficient, resulting in reduced mapping accuracy. Furthermore, traditional methods often lack effective image restoration strategies when faced with degraded images (such as those with poor lighting, blur, or noise), resulting in insufficient input data quality and further compromising the reliability of map generation results.

[0003] On the other hand, although map construction technology based on lidar and image fusion can provide high geometric accuracy, it consumes a lot of computing resources and is difficult to adapt to the real-time requirements in actual autonomous driving scenarios. At the same time, the temporal consistency modeling capability of multi-frame data is insufficient, and the dynamic change correlation of target instances in the scene is not strongly expressed, resulting in discontinuity and inconsistency of results during the map update process. In addition, existing map construction methods do not sufficiently optimize the lightweight model and resource adaptability, making it difficult to effectively deploy in resource-constrained scenarios, further limiting the application scenarios of autonomous driving technology.

[0004] In summary, the current high-precision map dynamic detection and update technology still has many technical bottlenecks in degraded image processing, multi-view feature fusion, dynamic change detection and lightweight design. Therefore, a high-precision map dynamic detection and update method using a general image restoration strategy is designed to solve the above problems. Summary of the Invention

[0005] The main purpose of the present invention is to provide a method for dynamic detection and updating of high-precision maps using a universal image restoration strategy.

[0006] The purpose of the present invention can be achieved by adopting the following technical solutions:

[0007] A method for dynamic detection and updating of high-precision maps using a general image restoration strategy includes the following steps:

[0008] Step S1: Through the pre-training process, the degraded image is input into the network, and the feature generation module and feature selection module are used to extract and enhance features layer by layer. The image is restored by combining multi-dimensional feature fusion and cross-attention mechanism to improve the detail restoration and global consistency of the degraded image;

[0009] Step S2: Based on the BEV representation generated by the pre-trained model in step S1 and combined with the historical state of the high-precision map, a dynamic detection module is used to identify the elements that are added, deleted, or unchanged in the map, and a difference map is output;

[0010] Step S3: Based on the difference map generated in step S2, the map update module performs geometric adjustments and semantic corrections on the newly added and deleted areas to generate an accurate and coherent local high-precision map.

[0011] Step S4: Comprehensively evaluate the performance of dynamic detection and map updating through the model evaluation framework.

[0012] Preferably, in step S1, the pre-training process generates a feature map of the input image through an initial convolution operation according to the degraded image characteristics, and introduces it into the encoder stage to complete the learning of multi-level features.

[0013] Preferably, the encoder uses a cross-channel attention module to model global semantic features, and combines it with a spatial self-attention module to enhance the extraction of local detail features;

[0014] At each layer, the spatial resolution of the feature map is gradually reduced through the downsampling mechanism to ensure the layer-by-layer learning and extraction of multi-scale features. By combining global and local information, various degradation characteristics in the input image are captured.

[0015] Pre-training fuses the multi-scale features output by each layer of the encoder through a skip connection mechanism, and combines it with upsampling operations to restore the image resolution;

[0016] The decoder applies the TSAB module layer by layer to strengthen global information modeling and uses the SSAB module to optimize local detail recovery capabilities to ensure the generation of high-quality feature representations;

[0017] While features are gradually upsampled to the original input resolution, multi-scale features are jointly modeled through splicing operations to enhance the overall performance of output features;

[0018] According to the global information modeling requirements, the Multi-Dconv Transpose Attention module is used to capture long-range feature interactions in the channel dimension;

[0019] Jointly learn local and global information in the spatial dimension through the OverlappingCross-Attention module;

[0020] During the pre-training process, the alternating effects of TSAB and SSAB are introduced to constrain the consistency of features and ensure the temporal stability and structural consistency of features.

[0021] Preferably, in step S2, based on the pre-trained model, the restored image obtained by inputting and the expired high-precision map are identified as follows:

[0022] First, the geometric and semantic information of the expired high-precision map is encoded;

[0023] This process converts complex map descriptions into compact high-dimensional feature representations while preserving core information through dedicated geometric and semantic encoders;

[0024] The geometric shape and semantic information of each map element are converted into a high-dimensional feature tensor for subsequent fusion;

[0025] Process the topological structure of outdated maps to ensure their spatial consistency and fusion;

[0026] After the outdated map is encoded, the sensor image is fed into the pre-processing module, which uses geometric correction and time synchronization technology to ensure the uniformity of the data from each perspective in the spatial and temporal dimensions.

[0027] A deep learning backbone network is used to extract features from image data. Low-level features are captured through a multi-layer convolutional network, and semantic information is globally modeled through a multi-head self-attention mechanism to generate high-dimensional feature tensors from multiple perspectives.

[0028] The model integrates the real-time features of sensor data with the features of outdated HD maps through a cross-attention mechanism. The model's classification head independently predicts the state of each map element and combines these predictions into a difference map, marking all elements that have changed in the scene.

[0029] The final difference map is used as output, providing detailed change information, clearly identifying whether each element is added, deleted, or unchanged. The dynamic detection performance is quantified by the accuracy formula:

[0030]

[0031] Among them, Acc± is the dynamic detection performance;

[0032] S represents the sequence number;

[0033] Fi is the number of frames in the i-th sequence;

[0034] c and Representing the change and no change states respectively;

[0035] and are the predicted elements in the difference map and the true values of the difference map elements, respectively.

[0036] Preferably, step S3 uses a map update module to correct and improve the local high-precision map based on the output of the difference map and the sensor data to ensure that the updated map accurately reflects the current scene;

[0037] First, for the elements marked as newly added in the difference map, the system generates a geometric description of the newly added area using sensor data;

[0038] Each element in the newly added area contains detailed shape information, semantic labels, and topological relationships;

[0039] The semantic analysis module is used to further optimize the properties of the newly added elements to ensure their semantic consistency with the overall map;

[0040] For elements marked for deletion in the difference map, the system removes their geometric descriptions and related attributes from the map;

[0041] The deletion operation is completed through a precise positioning algorithm to avoid erroneous effects on unchanged elements. The element deletion accuracy is quantified by the intersection over union formula:

[0042]

[0043] After completing the addition and deletion operations, the system uses the geometric correction module to spatially align all map elements to ensure that the newly added and deleted areas are seamlessly integrated with the unchanged areas in a unified coordinate system;

[0044] The system uses spatial projection and coordinate transformation techniques to align the updated map with the global scene to eliminate potential spatial errors;

[0045] By integrating unchanged areas with newly added and updated elements, a complete local HD map is generated. The system uses the following formula to evaluate the average accuracy of the updated map to quantify its overall performance:

[0046]

[0047] in,

[0048] Where V and V′ represent the predicted and real map elements, respectively;

[0049] Chamfer distance measures the error between boundaries;

[0050] Fréchet distance assesses the similarity of center lines;

[0051] V left and V right Represents the map elements on the left and right sides of the predicted map, V' left ,V'right For the left and right map elements of the real map, V ctr ,V' ctr Represent the central elements of the predicted map and the central elements of the real map respectively;

[0052] Preferably, step S4 includes, in the final evaluation stage of the model, comprehensively verifying the model performance using multiple methods in combination with the change data in the real scene;

[0053] First, the model's ability to detect changes in each frame is evaluated using the single-frame mode, while the model's temporal stability and consistency are analyzed using the multi-frame mode. This dual-mode evaluation approach can determine the model's performance in both single and continuous scenarios.

[0054] Among them, the single-frame mode evaluation model satisfies Acc+≥0.85, Acc-≥0.90 and the multi-frame mode analysis model satisfies Acc+≥0.90, Acc-≥0.92, which can be considered to meet the requirements;

[0055] For different change types, type-dependent and type-independent indicators are calculated to evaluate the model's detection effect on specific change categories and its generalization ability to overall changes. The model's localization performance on local changes is evaluated using the localization intersection-over-union ratio, whose formula is:

[0056]

[0057] When Accloca ≥ 0.8, it is considered to meet the requirements;

[0058] The overall performance of the model updating the map is further quantified by the average precision of each map element to ensure the reliability of the model output. When the average precision AP ≥ 0.8, it is considered to meet the requirements.

[0059] Beneficial technical effects of the present invention:

[0060] This paper provides a method for dynamic detection and updating of high-precision maps using a universal image restoration strategy. Through pre-training feature generation and selection modules, as well as multidimensional feature fusion and a cross-attention mechanism, it can effectively process degraded images, significantly improving detail restoration and global consistency. Whether images are poorly lit, blurred, or noisy, the encoder and decoder, working together with cross-channel attention modules and spatial self-attention modules, can learn and fuse multi-scale features to generate high-quality feature representations. This provides an accurate and clear data foundation for subsequent map detection and updates. Compared to traditional methods, it greatly improves input data quality and reduces map construction errors caused by image quality issues.

[0061] The BEV representation generated by the pre-trained model and the historical status of the high-precision map, combined with the dynamic detection module, can accurately identify newly added, deleted or unchanged elements in the map.

[0062] By encoding the geometric and semantic information of outdated HD maps and fusing it with sensor image features, this method effectively captures real-time changes and accurately outputs difference maps even in complex scenarios such as multi-target interaction and dynamic occlusion. Its dynamic detection accuracy is Acc+ ≥ 0.85, Acc- ≥ 0.90 in single-frame mode, and Acc+ ≥ 0.90, Acc- ≥ 0.92 in multi-frame (MF) mode. This provides a reliable basis for timely map updates and enables HD maps to better adapt to dynamic environmental changes.

[0063] Based on the difference map, the map update module performs geometric adjustments and semantic corrections for added and deleted areas. It generates detailed geometric descriptions and optimizes attributes for newly added elements, while removing deleted elements precisely without affecting other parts. Through techniques such as geometric correction and spatial projection, the updated map is aligned with the global scene. The resulting local high-precision map achieves an average accuracy (AP) of ≥0.8, ensuring map accuracy and consistency. This provides high-precision map support for tasks such as autonomous driving path planning, improving the reliability and safety of autonomous driving systems.

[0064] Through the model evaluation framework, combined with single-frame and multi-frame modes, as well as type-related, type-independent indicators and positioning intersection and union multiple evaluation methods, the model's performance in change detection, positioning accuracy and map updates is comprehensively considered. This not only provides a scientific basis for model improvement, but also ensures the reliability and generalization ability of the model in different scenarios, making high-precision maps more stable and adaptable in practical applications. BRIEF DESCRIPTION OF THE DRAWINGS

[0065] Figure 1 The present invention is a flowchart of a preferred embodiment of a method for dynamic detection and updating of high-precision maps using a universal image restoration strategy. DETAILED DESCRIPTION

[0066] In order to make the technical solution of the present invention more clear and specific to those skilled in the art, the present invention is further described in detail below with reference to embodiments and drawings, but the embodiments of the present invention are not limited thereto.

[0067] A method for dynamic detection and updating of high-precision maps using a general image restoration strategy is characterized in that the method includes: step S1: through a pre-training process, a degraded image is input into a network, a feature generation module and a feature selection module are used to extract and enhance features layer by layer, and multi-dimensional feature fusion and a cross-attention mechanism are combined to achieve image restoration and improve the detail restoration and global consistency of the degraded image.

[0068] Step S2: Based on the BEV representation generated by the pre-trained model in step S1 and combined with the historical status of the high-precision map, a dynamic detection module is used to identify newly added, deleted, or unchanged elements in the map, and a difference map is output.

[0069] Step S3: Based on the difference map generated in step S2, the map update module performs geometric adjustments and semantic corrections on the newly added and deleted areas to generate an accurate and coherent local high-precision map.

[0070] Step S4: Through the model evaluation framework, the performance of dynamic detection and map update is comprehensively evaluated, including change detection accuracy, positioning accuracy, and overall consistency and accuracy after map update.

[0071] Step S1 includes, during the pre-training process, generating a feature map from the input image through an initial convolution operation according to the characteristics of the degraded image, and introducing it into the encoder stage to complete the learning of multi-level features. The encoder uses a cross-channel attention module (TSAB) to model global semantic features, and combines a spatial self-attention module (SSAB) to enhance the extraction of local detail features. At each layer, the spatial resolution of the feature map is gradually reduced through a downsampling mechanism (for example, from H×W×C to H / 2×W / 2×2C, and then expanded to 4C and 8C channel dimensions) to ensure layer-by-layer learning and extraction of multi-scale features. The goal of this stage is to capture multiple degradation characteristics in the input image by combining global and local information.

[0072] At the decoder stage, pre-training fuses the multi-scale features output by each encoder layer through skip connections, and combines this with upsampling to restore the image resolution. The decoder layer-by-layer applies the TSAB module to enhance global information modeling and the SSAB module to optimize local detail recovery, ensuring high-quality feature representation. While features are progressively upsampled to the original input resolution, multi-scale features are jointly modeled through concatenation (Concat) to enhance the overall performance of the output features.

[0073] To improve the generalization capability of the pre-trained model, the Multi-Dconv Transpose Attention (MDTA) module captures long-range feature interactions in the channel dimension, based on the need for global information modeling. To enhance local detail, the Overlapping Cross-Attention (OCA) module jointly learns local and global information in the spatial dimension. During pre-training, consistency regularization constraints (TSAB) and SSAB are introduced to ensure temporal stability and structural consistency of features.

[0074] Ultimately, the pre-trained model completes efficient learning in a variety of degraded image tasks, generates high-quality feature representations that can adapt to images of different degradation types, and provides a stable and powerful basic model for subsequent tasks.

[0075] Step S2 includes identifying the difference between the restored image and the outdated high-precision map by inputting the pre-trained model. First, the geometric and semantic information of the outdated high-precision map is encoded.

[0076] This process transforms complex map descriptions into compact high-dimensional feature representations while preserving core information through dedicated geometric and semantic encoders.

[0077] The geometry (e.g., centerline, boundary attributes) and semantic information (e.g., lane type) of each map element are converted into high-dimensional feature tensors for subsequent fusion. Furthermore, the topology of outdated maps is processed to ensure spatial consistency and fusion.

[0078] After encoding the outdated map, the sensor images are fed into a pre-processing module, which uses geometric correction and time synchronization techniques to ensure the uniformity of the data from each viewpoint in both time and space.

[0079] Subsequently, a deep learning backbone network (such as ResNet) is used to extract features from the image data.

[0080] Low-level features are captured through multi-layer convolutional networks, while semantic information is globally modeled using a multi-head self-attention mechanism, generating multi-view, high-dimensional feature tensors. These features represent the spatial positions of objects in the scene and their semantic relationships, laying the foundation for subsequent fusion steps.

[0081] The feature fusion stage is the core step of the dynamic detection process.

[0082] The real-time features of sensor data are integrated with the features of outdated high-precision maps through the cross-attention mechanism.

[0083] This approach ensures that the system retains valid information that has not changed in outdated maps while also capturing real-time changes through sensor data. The model's classification head independently predicts the state of each map element (added, deleted, or unchanged) and combines these predictions into a difference map, annotating all changed elements in the scene.

[0084] The final difference map is used as output, providing detailed change information, clearly identifying whether each element is added, deleted, or unchanged. The dynamic detection performance is quantified by the accuracy formula:

[0085]

[0086] Among them, Acc± is the dynamic detection performance;

[0087] S represents the sequence number;

[0088] Fi is the number of frames in the i-th sequence;

[0089] c and Representing the change and no change states respectively;

[0090] and are the predicted elements in the difference map and the true values of the difference map elements, respectively.

[0091] Step S3 includes using a map update module to correct and improve the local high-precision map based on the output of the difference map and sensor data to ensure that the updated map accurately reflects the current scene. First, for the elements marked as "new" in the difference map, the system generates a geometric description of the newly added area through sensor data. Each element in the newly added area contains detailed shape information, semantic labels (such as lane type or sign location), and topological relationships. In addition, the semantic analysis module further optimizes the properties of the newly added elements to ensure their semantic consistency with the overall map.

[0092] For elements marked as "deleted" in the difference map, the system removes their geometric descriptions and related attributes from the map. The deletion operation is completed using a precise positioning algorithm to avoid accidentally affecting unchanged elements. The element deletion accuracy is quantified using the intersection over union (IoU) formula:

[0093] After completing the addition and deletion operations, the system spatially aligns all map elements through the geometric correction module to ensure that the newly added and deleted areas are seamlessly integrated with the unchanged areas in a unified coordinate system. In addition, the system uses spatial projection and coordinate transformation technology to align the updated map with the global scene to eliminate potential spatial errors.

[0094] Finally, a complete local HD map is generated by integrating the unchanged areas with the newly added and updated elements. The system uses the following formula to evaluate the average precision (AP) of the updated map to quantify its overall performance:

[0095] Where V and V′ represent the predicted and real map elements, respectively;

[0096] Chamfer distance measures the error between boundaries;

[0097] Fréchet distance assesses the similarity of center lines;

[0098] Vleft and V right Represents the map elements on the left and right sides of the predicted map, V' left ,V' right For the left and right map elements of the real map, V ctr ,V' ctr They represent the central elements of the predicted map and the central elements of the true map respectively.

[0099] Through these steps, the map update module successfully generates an accurate and complete local high-precision map, while also providing reliable support for subsequent tasks such as autonomous driving path planning. This update process not only dynamically adapts to environmental changes but also significantly improves system reliability and safety through high-precision map generation capabilities.

[0100] Step S4 involves a comprehensive validation of the model's performance using multiple methods, combined with real-world scene change data, during the final evaluation phase. First, the model's ability to detect changes within each frame is evaluated using the single-frame (SF) model, while the model's temporal stability and consistency are analyzed using the multi-frame (MF) model. This dual-mode evaluation approach helps determine the model's performance in both single and continuous scenes.

[0101] The single-frame (SF) mode evaluation model satisfies Acc+≥0.85, Acc-≥0.90 and the multi-frame (MF) mode analysis model satisfies Acc+≥0.90, Acc-≥0.92, which can be considered to meet the requirements;

[0102] Secondly, type-aware and type-agnostic metrics are calculated for different change types to evaluate the model's detection performance for specific change categories (such as additions or deletions) and its generalization ability to overall changes. The model's localization performance for local changes is evaluated using the localization intersection over union (Accloca), which is formulated as:

[0103]

[0104] When Accloca ≥ 0.8, it can be considered that it meets the requirements.

[0105] In addition, the overall performance of the model in updating the map is further quantified by the average precision (AP) of each map element to ensure the reliability of the model output. When the average precision AP ≥ 0.8, it is considered to meet the requirements.

[0106] By combining multiple indicators, we comprehensively evaluate the model's change detection ability, map update accuracy, and the interpretability of the output results.

[0107] These evaluation results not only provide a scientific basis for model improvement, but also provide strong support for the application of high-precision maps in actual scenarios.

[0108] The above is only a further embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any technician familiar with the technical field can make equivalent replacements or changes based on the technical solutions and concepts of the present invention within the scope disclosed by the present invention, which fall within the scope of protection of the present invention.

Claims

1. A method for dynamic detection and updating of high-precision maps using a general image restoration strategy, characterized by: The steps include: Step S1: Through the pre-training process, the degraded image is input into the network, and the feature generation module and feature selection module are used to extract and enhance features layer by layer. The image is restored by combining multi-dimensional feature fusion and cross-attention mechanism to improve the detail restoration and global consistency of the degraded image; Step S2: Based on the BEV representation generated by the pre-trained model in step S1 and combined with the historical state of the high-precision map, a dynamic detection module is used to identify the elements that are added, deleted, or unchanged in the map, and a difference map is output; Step S3: Based on the difference map generated in step S2, the map update module performs geometric adjustments and semantic corrections on the newly added and deleted areas to generate an accurate and coherent local high-precision map. Step S4: Comprehensively evaluate the performance of dynamic detection and map updating through the model evaluation framework.

2. The method for dynamic detection and updating of high-precision maps using a universal image restoration strategy according to claim 1, characterized in that: In step S1, the pre-training process generates a feature map of the input image through an initial convolution operation according to the degraded image characteristics, and introduces it into the encoder stage to complete the learning of multi-level features.

3. The method for dynamic detection and updating of high-precision maps using a universal image restoration strategy according to claim 2, characterized in that: The encoder uses a cross-channel attention module to model global semantic features, and combines it with a spatial self-attention module to enhance the extraction of local detail features; At each layer, the spatial resolution of the feature map is gradually reduced through the downsampling mechanism to ensure the layer-by-layer learning and extraction of multi-scale features. By combining global and local information, various degradation characteristics in the input image are captured. Pre-training fuses the multi-scale features output by each layer of the encoder through a skip connection mechanism, and combines convolution operations with sampling operations to restore the image resolution; The decoder applies the TSAB module layer by layer to strengthen global information modeling and uses the SSAB module to optimize local detail recovery capabilities to ensure the generation of high-quality feature representations; While features are gradually upsampled to the original input resolution, multi-scale features are jointly modeled through splicing operations to enhance the overall performance of output features; According to the global information modeling requirements, the Multi-Dconv Transpose Attention module is used to capture long-range feature interactions in the channel dimension; Jointly learn local and global information in the spatial dimension through the OverlappingCross-Attention module; During the pre-training process, the alternating effects of TSAB and SSAB are introduced to constrain the consistency of features and ensure the temporal stability and structural consistency of features.

4. The method for dynamic detection and updating of high-precision maps using a universal image restoration strategy according to claim 1, characterized in that: In step S2, based on the pre-trained model, the restored image and the expired high-precision map are input and the differences between the two are identified as follows: First, the geometric and semantic information of the expired high-precision map is encoded; This process converts complex map descriptions into compact high-dimensional feature representations while preserving core information through dedicated geometric and semantic encoders; The geometric shape and semantic information of each map element are converted into a high-dimensional feature tensor for subsequent fusion; Process the topological structure of outdated maps to ensure their spatial consistency and fusion; After encoding the expired map, the sensor image is fed into the pre-processing module; The pre-processing module uses geometric correction and time synchronization technology to ensure the uniformity of data from each perspective in the temporal and spatial dimensions; A deep learning backbone network is used to extract features from image data. Low-level features are captured through a multi-layer convolutional network, and semantic information is globally modeled through a multi-head self-attention mechanism to generate high-dimensional feature tensors from multiple perspectives. The cross-attention mechanism integrates the real-time features of sensor data with the features of outdated high-precision maps. The classification head of the preprocessing model independently predicts the state of each map element and combines these predictions into a difference map, marking all elements that have changed in the scene. The final difference map is used as output, providing detailed change information, clearly identifying whether each element is added, deleted, or unchanged. The dynamic detection performance is quantified by the accuracy formula: Among them, Acc± is the dynamic detection performance; S represents the sequence number; Fi is the number of frames in the i-th sequence; c and Representing the change and no change states respectively; and are the predicted elements in the difference map and the true values of the difference map elements, respectively.

5. The method for dynamic detection and updating of high-precision maps using a universal image restoration strategy according to claim 1, characterized in that: Step S3: Based on the output of the difference map and the sensor data, the map update module is used to correct and improve the local high-precision map to ensure that the updated map accurately reflects the current scene; First, for the elements marked as newly added in the difference map, the system generates a geometric description of the newly added area using sensor data; Each element in the newly added area contains detailed shape information, semantic labels, and topological relationships; The semantic analysis module is used to further optimize the properties of the newly added elements to ensure their semantic consistency with the overall map; For elements marked for deletion in the difference map, the system removes their geometric descriptions and related attributes from the map; The deletion operation is completed through a precise positioning algorithm to avoid erroneous effects on unchanged elements. The element deletion accuracy is quantified by the intersection over union formula: After completing the addition and deletion operations, the system uses the geometric correction module to spatially align all map elements to ensure that the newly added and deleted areas are seamlessly integrated with the unchanged areas in a unified coordinate system; The system uses spatial projection and coordinate transformation techniques to align the updated map with the global scene to eliminate potential spatial errors; By integrating unchanged areas with newly added and updated elements, a complete local HD map is generated. The system uses the following formula to evaluate the average accuracy of the updated map to quantify its overall performance: Where V and V′ represent the predicted and real map elements, respectively; Chamfer distance measures the error between boundaries; Fréchet distance assesses the similarity of center lines; V left and V right Represents the map elements on the left and right sides of the predicted map, V' left ,V' right For the left and right map elements of the real map, V ctr ,V' ctr They represent the central elements of the predicted map and the central elements of the true map respectively.

6. The method for dynamic detection and updating of high-precision maps using a universal image restoration strategy according to claim 1, characterized in that: Said step S4 includes, in the final evaluation stage of the model, comprehensively verifying the model performance using a variety of methods in combination with the change data in the real scene; First, the model's ability to detect changes in each frame is evaluated using the single-frame mode, while the model's temporal stability and consistency are analyzed using the multi-frame mode. This dual-mode evaluation approach can determine the model's performance in both single and continuous scenarios. Among them, the single-frame mode evaluation model satisfies Acc+≥0.85, Acc-≥0.90 and the multi-frame mode analysis model satisfies Acc+≥0.90, Acc-≥0.92, which can be considered to meet the requirements; For different change types, type-dependent and type-independent indicators are calculated to evaluate the model's detection effect on specific change categories and its generalization ability to overall changes; The positioning performance of the model on local changes is evaluated using the positioning intersection-over-union ratio, and its formula is: When Accloca ≥ 0.8, it is considered to meet the requirements; The overall performance of the model updating the map is further quantified by the average precision of each map element to ensure the reliability of the model output. When the average precision AP ≥ 0.8, it is considered to meet the requirements.

Citation Information

Patent Citations

  • High-precision map increment updating method and device based on roadside camera

    CN115587106A

  • Three-dimensional scene reconstruction method and device based on neural network and multi-view consistency

    CN117523100A

  • High-precision map updating method based on environment perception in roadway in well mining scene

    CN118310507A

  • Low-light image denoising method based on local-global interactive Transform

    CN119273576A

  • Multi-stage turbulent dynamic video recovery method based on physical model

    CN119784648A