Semantic vector map construction method supporting beyond visual range and dynamic updating
Through the surround view image input network and time series fusion technology, combined with high-precision map prior data and cross-attention mechanism, a beyond-line-of-sight semantic vector map is generated, which solves the real-time and high-precision problems of map construction in dynamic environments and improves the accuracy and stability of the map.
Patent Information
- Application Number
- CN202510812660.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-18
- Publication Date
- 2025-09-16
- Estimated Expiration
- 2045-06-18
AI Technical Summary
Existing map construction methods cannot meet the real-time and high-precision requirements in dynamic environments, especially in complex scenarios where map accuracy decreases. Traditional methods also face challenges in computing resource consumption and real-time performance, making it difficult to effectively integrate real-time data with offline data to ensure spatiotemporal consistency.
The surround view image input network is used, combined with the feature generation module and mapping function to extract BEV features. Multi-frame information fusion is achieved through time series fusion and convolutional neural network. High-precision map and navigation map prior data are introduced for weighted fusion. The cross-attention mechanism is used to generate beyond-horizon BEV feature maps, and the semantic vector map is generated through the decoder. It is dynamically updated by combining the query propagation mechanism and diffusion denoising technology.
It achieves beyond-horizon expansion of map construction range, improves the spatial consistency and temporal integrity of the map, enhances environmental perception capabilities, especially the accuracy and stability of the map in static areas and long-distance scenes, and reduces computing delays and noise interference.
Smart Images

Figure CN120655776A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a semantic vector map construction method, in particular to a semantic vector map construction method supporting beyond-horizon and dynamic updating, belonging to the technical field of semantic vector map construction methods. Background Art
[0002] With the continuous development of autonomous driving technology, the importance of high-precision maps in autonomous driving systems has become increasingly prominent. High-precision maps support vehicle perception, positioning, and path planning, and can provide detailed road geometry and semantic information. However, in complex dynamic scenarios, such as multi-target interaction, dynamic occlusion, and frequent environmental changes, traditional high-precision map construction and update methods cannot meet the real-time and high-precision requirements.
[0003] Existing map construction methods can provide relatively accurate geometric and semantic expressions in static environments. However, in dynamic environments, due to perspective changes, dynamic targets, and occlusion issues, the feature fusion process is often insufficient, resulting in reduced map accuracy. Especially in scenarios with long time spans, how to compensate for the missing information in single-frame images and capture long-distance and dynamically changing information becomes a challenge for map updates.
[0004] In addition, when dealing with static areas, traditional methods have insufficient speed and scope for updating temporal information, resulting in untimely map updates. To address these issues, researchers have begun to introduce offline prior data for high-precision maps and navigation maps. These data provide stable and comprehensive road information, especially in areas without interference from dynamic objects, and play an important role in improving map accuracy and completeness. However, how to effectively integrate real-time data with offline data to ensure temporal and spatial consistency remains a technical challenge.
[0005] At the same time, existing map update technologies face challenges in terms of computing resource consumption and real-time performance. When processing large-scale data, how to efficiently implement map updates and avoid computational delays while ensuring real-time performance is one of the current technical bottlenecks. To this end, a semantic vector map construction method that supports beyond-horizon and dynamic updates is designed to solve the above problems. Summary of the Invention
[0006] The main purpose of the present invention is to provide a semantic vector map construction method that supports beyond-horizon and dynamic updating.
[0007] The purpose of the present invention can be achieved by adopting the following technical solutions:
[0008] A method for constructing a semantic vector map that supports beyond-horizon and dynamic updating includes the following steps:
[0009] Step S1: Input the surround view image into the network;
[0010] Step S2: Fusion features based on the above time series.
[0011] Step S3: Dynamic map updates are achieved through the query propagation mechanism to ensure the consistency of map elements in the time dimension.
[0012] Preferably, in step 1, a feature generation module and a mapping function are used to extract and enhance BEV features layer by layer, and time series fusion and convolutional neural network are combined to achieve multi-frame information fusion and improve the spatial and temporal integrity of environmental perception;
[0013] In step 2, the high-precision map and navigation map prior data are input into the network, and weighted fusion is performed using the cross-attention mechanism to generate a beyond-visual-range (BEV) feature map. Subsequently, the fused features are mapped into a semantic vector map with the help of a decoder, and a deep neural network is used to convert the features into semantic elements of actual roads, obstacles, and lane lines to realize beyond-visual-range map construction.
[0014] Preferably, in step three, historical memory latent variables are combined with current frame features to improve map update accuracy. At the same time, diffusion denoising technology is used to remove noise. Finally, high-resolution feature maps are refined and fused to dynamically generate accurate semantic vector maps to adapt to environmental changes in real time.
[0015] Preferably, the step 1 includes using a surround view image, wherein the surround view image is a 360-degree panoramic view generated by stitching images captured by multiple cameras around the vehicle;
[0016] The feature of each frame image is set to I. On this basis, the bird's-eye view feature can be expressed as follows:
[0017]
[0018] Among them, I t It represents the surround image of the current frame;
[0019] n refers to the number of history frames;
[0020] f BEV It is a function used to map the surround view image to a bird's-eye view by synthesizing images from multiple perspectives;
[0021] By fusing the BEV features of the current frame with the BEV features of the previous frames, we can obtain richer spatial and temporal information. The temporal fusion module will fuse the BEV features of the current frame with the BEV features of the previous n frames, and finally obtain the features after temporal fusion.
[0022]
[0023] Among them, f temporal It refers to the time series fusion function, which is achieved by convolutional layers and recurrent neural networks, also known as RNNs;
[0024] Time series fusion integrates historical information with current information to make up for the missing information in distant areas and expand the scope of map construction.
[0025] Preferably, the step 2 further includes introducing offline prior data, that is, using the navigation map and the high-precision map as prior information to further enrich the map construction process;
[0026] High-precision maps typically include more detailed road information than traditional maps;
[0027] After introducing offline prior data, a cross-attention mechanism is adopted;
[0028] Assume that the features after time series fusion are The offline map prior data is Then the fusion process of the cross attention mechanism can be expressed as:
[0029]
[0030] Among them, f prior Represents the cross attention mechanism function, which gradually completes the weighted fusion of features by calculating the correlation between time series features and offline map prior data;
[0031] The cross-attention mechanism enables the system to selectively integrate data from offline maps with temporal fusion information.
[0032] Preferably, after introducing time series fusion technology and offline prior data, the autonomous driving system can generate a BEV feature map beyond visual range;
[0033] The feature map contains the environmental information of the current frame, and also integrates the spatial and temporal information of historical frames, as well as the prior data in the high-precision map and navigation map;
[0034] During the BVR map construction process, the system decodes the fused BEV features through a deep neural network to generate a semantic vector map containing rich semantic information. The decoding process uses a fully convolutional network to map image features into semantic elements of actual roads, obstacles, lane lines, and traffic signs.
[0035] The goal of the decoding process is to generate a beyond-line-of-sight semantic vector map, thereby providing the autonomous driving system with more accurate environmental perception, especially in extended-line-of-sight scenarios.
[0036] Specifically, the decoding process can be expressed by the following formula:
[0037]
[0038] Among them, g() represents the decoding function, which maps the fused features into a vector map with semantic annotations
[0039] Preferably, step S3 includes assuming that a certain map element has been successfully identified and tracked in the map of the previous frame, and the location information of these elements will be transmitted to the current frame by means of the query propagation mechanism. We rely on calculating the spatial correlation between the current frame and the previous frame to obtain the updated position of the corresponding element in the current frame. This process can be expressed by the following formula:
[0040] E t =f query (E t-1 ,I t );
[0041] Among them, E t Is the map element of the current frame, E t-1 It is the map element of the previous frame, I t is the image of the current frame, f query is a function of the query propagation mechanism;
[0042] The concept of historical memory latent variables was introduced;
[0043] Historical memory latent variables are potential feature information accumulated in historical frames. They can help the model better understand the map elements of the current frame, integrate historical information with the map elements of the current frame, make up for key information that may be missing in the current frame, and improve the accuracy of map updates.
[0044] Specifically, the historical information and the characteristics of the current map elements are integrated with the help of the historical memory latent variable. The latent variable of the historical frame is set to H t-1 , the map element of the current frame is E t , then the fusion process of historical information and current map elements can be expressed as:
[0045] E fusion =f fusion (E t ,H t-1 );
[0046] where f fusion As a fusion function, the map elements of the current frame and the historical latent variables are combined to obtain the fused map element features E fusion ;
[0047] This fusion approach allows for a deeper understanding of map elements, improving the accuracy of map updates in a dynamically changing environment.
[0048] In the map update process, the semantic vector map must first be initialized with noise, and then the noise is removed step by step through the diffusion process.
[0049] Specifically, the diffusion process can be defined as the following formula:
[0050]
[0051] Among them, ε noisy is a preliminary noise map;
[0052] τ is the time step of the diffusion process;
[0053] It is a diffusion operation, which removes noise by means of step-by-step iteration and finally obtains the denoised semantic map E denosied ;
[0054] After noise removal, feature information at different levels will be refined and fused to improve the spatial resolution and detail of the map. Refinement is achieved by increasing the resolution of the feature map, making the details in the map clearer, while fusion combines multiple levels of feature information to improve map accuracy and more effectively capture road and obstacle details in complex environments.
[0055] Specifically, the refinement and fusion process can be expressed by the following formula:
[0056] E refined =f refine (E denosied ,F high-res );
[0057] Among them, E denosied Refers to the map obtained after denoising;
[0058] F high-res It is a feature map with high resolution characteristics;
[0059] f refine is the refinement function.
[0060] Beneficial technical effects of the present invention:
[0061] The present invention provides a semantic vector map construction method that supports beyond-horizon and dynamic updating. It integrates the bird's-eye view features of the current frame and historical frames through temporal fusion technology, and uses convolutional layers and recurrent neural networks (RNN) to model time dimension information, effectively compensating for the problem of missing long-distance information in single-frame images, and expanding the map construction range from the close-range area of traditional methods to the beyond-horizon range.
[0062] By introducing high-precision maps and navigation map prior data, and using a cross-attention mechanism to weightedly fuse temporal features with offline data, we further enhance map integrity in static areas. For example, within a 120m x 60m mapping area, the spatial consistency of map features improved by 12.27% after incorporating prior data. BRIEF DESCRIPTION OF THE DRAWINGS
[0063] Figure 1 This is a flow chart of a method for constructing a semantic vector map that supports beyond-horizon and dynamic updating, provided by one embodiment of the present invention;
[0064] Figure 2 This is a schematic diagram showing the model effect of a semantic vector map construction method that supports beyond-horizon and dynamic updating, provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0065] In order to make the technical solution of the present invention more clear and specific to those skilled in the art, the present invention is further described in detail below with reference to embodiments and drawings, but the embodiments of the present invention are not limited thereto.
[0066] The method comprises:
[0067] Step S1: The surround view image is input into the network, and the BEV features are extracted and enhanced layer by layer using the feature generation module and mapping function. Combined with time series fusion and convolutional neural network (RNN), multi-frame information fusion is achieved to improve the spatial and temporal integrity of environmental perception.
[0068] Step S2: Based on the aforementioned temporal fusion features, the high-precision map and navigation map prior data are fed into the network and weightedly fused using a cross-attention mechanism to generate a Beyond Visual Range (BEV) feature map. Subsequently, a decoder is used to map the fused features into a semantic vector map. A deep neural network is then used to transform the features into semantic elements of actual roads, obstacles, and lane markings, enabling BEV map construction.
[0069] Step S3: Dynamic map updates are achieved through a query propagation mechanism, ensuring the consistency of map elements across time. Historical memory latent variables are combined with current frame features to improve map update accuracy. Diffusion denoising techniques are also used to remove noise. Finally, high-resolution feature maps are refined and fused to dynamically generate an accurate semantic vector map that adapts to environmental changes in real time.
[0070] The step S1 includes:
[0071] In traditional autonomous driving systems, maps are generally constructed using forward-view images or surround-view images. However, single-perspective images often have limitations in capturing long-distance road information. In complex urban environments, surround-view images are first used to generate bird's-eye view features.
[0072] The surround view image is a 360-degree panoramic view generated by stitching together images taken by multiple cameras around the vehicle. The BEV feature is a projection method viewed from above, which can effectively present the road topology and lane information. By fusing images from multiple perspectives, high-quality BEV features can be obtained.
[0073] The feature of each frame image is set to I. On this basis, the bird's-eye view feature can be expressed as follows:
[0074]
[0075] Here, I t represents the surround image of the current frame, and n refers to the number of historical frames, f BEV is a function used to map the surround view image to a bird's-eye view. By synthesizing images from multiple perspectives, a more complete environmental information can be constructed from the BEV's perspective.
[0076] The introduction of temporal fusion technology aims to make up for the missing information in single-frame images and the lost road information over a long time span. By modeling temporal information, this technology allows historical frame images to supplement long-distance information that is difficult to capture in the current frame, effectively expanding the viewing range of map construction.
[0077] The core idea of the temporal fusion module is to obtain richer spatial and temporal information by fusing the BEV features of the current frame with the BEV features of the previous frames. The temporal fusion module fuses the BEV features of the current frame with the BEV features of the previous n frames, and finally obtains the features after temporal fusion.
[0078]
[0079] Among them, f temporal This refers to the time series fusion function, which is achieved by using convolutional layers and recurrent neural networks (RNNs). Time series fusion integrates historical information with current information, filling in missing information in distant areas and expanding the scope of map construction.
[0080] The step S2 includes:
[0081] While temporal fusion technology can compensate for the shortcomings of the current frame by leveraging historical frame information and extending the viewing range, in some scenarios, particularly in static environments, the speed and range of temporal information updates may not be sufficient to meet practical needs. For example, in areas free of dynamic objects, where vehicle trajectory changes little, temporal fusion may lag or even lose critical information.
[0082] To overcome the limitations of time-series fusion technology, offline prior data can be introduced, using navigation maps and high-precision maps as prior information to further enrich the map-building process. This data can provide more stable and comprehensive road information, compensating for the shortcomings of time-series information in static areas.
[0083] High-definition maps typically include more detailed road information than traditional maps, such as lane widths, lane markings, intersection shapes, terrain information, traffic signs, and traffic lights. HD maps are updated less frequently, but their information is highly stable and accurate, making them ideal for navigation and positioning in static environments. Navigation maps, on the other hand, typically include routes, traffic information, navigation instructions, and dynamic route changes. They can be updated more frequently and at a lower cost.
[0084] After introducing offline prior data, how to effectively combine it with temporal fusion information becomes a key issue. To achieve this goal, a cross-attention mechanism can be used. This mechanism can calculate the correlation between temporal features and offline map prior data and selectively retain valuable information through weighted fusion.
[0085] The core idea of the cross-attention mechanism is to determine which information to retain based on the similarity and correlation between temporal features and prior data, and then incorporate this information into the system's current view. This mechanism enables autonomous driving systems to better integrate real-time perception information with static information in high-precision maps and navigation maps, providing stronger support for map updates and environmental perception.
[0086] Assume that the features after time series fusion are The offline map prior data is Then the fusion process of the cross attention mechanism can be expressed as:
[0087]
[0088] Among them, f prior The cross-attention mechanism function gradually performs weighted feature fusion by calculating the correlation between time series features and offline map prior data. The cross-attention mechanism enables the system to selectively integrate data from the offline map with the time series fusion information, thereby improving the system's environmental perception capabilities.
[0089] Through this process, the system not only effectively compensates for deficiencies in time-series fusion but also leverages static data provided by high-precision maps and navigation maps to further enhance the accuracy and completeness of map construction. Particularly in static environments, the introduction of offline prior data significantly enriches the system's understanding of the surrounding environment, avoiding the errors and lags that can arise from relying solely on time-series information.
[0090] By incorporating time series fusion technology and offline prior data, the autonomous driving system can generate a BEV feature map beyond visual range. This feature map not only incorporates the current frame's environmental information but also incorporates spatial and temporal information from historical frames, as well as prior data from high-precision maps and navigation maps, encompassing a wider visual range.
[0091] During the BVR map construction process, the system decodes the fused BEV features through a deep neural network (DNN) to generate a semantic vector map containing rich semantic information. The decoding process uses a fully convolutional network (FCN) to map image features into semantic elements of actual roads, obstacles, lane markings, and traffic signs.
[0092] The goal of the decoding process is to generate a beyond-line-of-sight semantic vector map, thereby providing the autonomous driving system with more accurate environmental perception, especially in extended-line-of-sight scenarios.
[0093] Specifically, the decoding process can be expressed by the following formula:
[0094]
[0095] Among them, g() represents the decoding function, which maps the fused features into a vector map with semantic annotations
[0096] Building beyond-horizon maps offers significant advantages in autonomous driving. By combining time-series fusion technology with offline prior data, the system can achieve a wider viewing range and enhance its awareness of the surrounding environment. Especially when the vehicle enters static areas or moves away from dynamic objects, offline map data can effectively compensate for the lack of time-series information, ensuring map accuracy and stability.
[0097] 4. The method according to claim 1, wherein step S3 comprises:
[0098] As road conditions continue to change, autonomous driving systems need to be able to update semantic vector maps in real time to adapt to emerging traffic conditions. However, traditional map update methods often suffer from poor temporal consistency and significant noise interference. To improve map update capabilities, a map update technology based on map element tracking and diffusion denoising is proposed to address these issues.
[0099] Traditional map update methods generally ignore the issue of temporal consistency. This can lead to inaccurate positional changes of map elements at different points in time, and even duplicate or missing map elements. To improve temporal consistency, a query propagation mechanism has been proposed. Its function is to explicitly associate tracked road elements from the previous frame with the current frame. This mechanism ensures that map elements remain consistent across time, improving the accuracy of the map update process.
[0100] To elaborate, assume that a map element (such as a lane line or traffic sign) has been successfully identified and tracked in the previous frame's map. Using the query propagation mechanism, the location information of these elements is transmitted to the current frame. We calculate the spatial correlation between the current and previous frames to determine the updated position of the corresponding element in the current frame. This process can be expressed as follows:
[0101] E t =f query (E t-1 ,I t );
[0102] Among them, E t Is the map element of the current frame, E t-1 It is the map element of the previous frame, I t is the image of the current frame, f query is a function of the query propagation mechanism;
[0103] To further improve the accuracy of map updates, we introduced the concept of historical memory latent variables. Historical memory latent variables are potential feature information accumulated in historical frames. They help the model better understand the map elements in the current frame. By integrating historical information with the map elements in the current frame, they can compensate for key information that may be missing in the current frame, improving the accuracy of map updates.
[0104] To expand on this, we use the historical memory latent variable to integrate the historical information and the characteristics of the current map elements. The latent variable of the historical frame is set to H t-1 , the map element of the current frame is E t , then the fusion process of historical information and current map elements can be expressed as:
[0105] E fusion =ffusion (E t ,H t-1 );
[0106] where f fusion As a fusion function, the map elements of the current frame and the historical latent variables are combined to obtain the fused map element features E fusion This fusion approach allows for a deeper understanding of map elements, improving the accuracy of map updates in dynamically changing environments.
[0107] In actual applications, sensor noise and environmental interference factors will cause a large amount of noise in the generated semantic vector map. This noise will affect the map accuracy, causing map inaccuracy and affecting the decision-making of the autonomous driving system. In order to remove noise and improve map update accuracy, diffusion denoising technology is introduced.
[0108] Diffusion denoising is an image processing method based on the diffusion process. It simulates the diffusion of information in space to gradually remove noise. During the map update process, the semantic vector map must first be initialized with noise. Then, the diffusion process is used to gradually remove noise, ultimately resulting in a clearer and more accurate map.
[0109] Specifically, the diffusion process can be defined as the following formula:
[0110]
[0111] Among them, ε noisy is the preliminary noise map, τ is the time step of the diffusion process, It is a diffusion operation, which removes noise by means of step-by-step iteration and finally obtains the denoised semantic map E denosied With this diffusion denoising technology, the accuracy of map updates can be improved and unnecessary noise interference can be eliminated.
[0112] After the noise removal operation is completed, the feature information at different levels will be refined and fused to improve the spatial resolution and detail of the map. Refinement is achieved by increasing the resolution of the feature map, which can make the details in the map clearer. The fusion process combines multiple levels of feature information to improve the accuracy of the map. In complex environments, it can more effectively capture detailed information about roads and obstacles.
[0113] Specifically, the refinement and fusion process can be expressed by the following formula:
[0114] E refined =f refine (Edenosied ,F high-res );
[0115] Here, E denosied Refers to the map obtained after denoising, F high-res is a feature map with high resolution, f refine To refine the function, it fuses high-resolution feature information to improve the map's detail and spatial resolution. This allows for the generation of a refined semantic vector map, improving both map accuracy and practicality.
[0116] Example: The experimental comparison table of the IOU (an indicator that measures the degree of overlap between the predicted box and the true box) and AP (prediction accuracy) values of the map construction system obtained by this patent and other existing examples shows that under different mapping ranges, the IOU and AP values of this patent are higher than those of existing examples, verifying the effectiveness of this patent solution.
[0117] The IOU calculation formula is:
[0118]
[0119] Where A: prediction area.
[0120] B: The actual marked area.
[0121] |A∩B|: the number of pixels in the intersection area of the predicted area and the true annotated area, |A∪B|: the number of pixels in the union area of the predicted area and the true annotated area.
[0122] The formula for calculating the AP value is:
[0123]
[0124] r i It is the recall value of different map elements, usually taking a series of predetermined recall rates, such as [0, 0.1, 0.2, ..., 1.0].
[0125] Precision i ) is the recall rate r i time accuracy.
[0126] N is the number of map elements.
[0127] Based on the above indicators, the baseline method and the patented method were experimentally compared, and the final experimental data were as follows:
[0128]
[0129]
[0130] The above is only a further embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any technician familiar with the technical field can replace or change the technical solution and concept of the present invention within the scope disclosed by the present invention, which falls within the scope of protection of the present invention.
Claims
1. A semantic vector map construction method that supports beyond-horizon and dynamic updating, characterized by: The steps include: Step S1: Input the surround view image into the network; Step S2: Fusion features based on the above temporal sequence; Step S3: Dynamic map updates are achieved through the query propagation mechanism to ensure the consistency of map elements in the time dimension.
2. The method for constructing a semantic vector map that supports beyond-horizon and dynamic updating according to claim 1, characterized in that: In step 1, the feature generation module and mapping function are used to extract and enhance BEV features layer by layer. Combined with time series fusion and convolutional neural network, multi-frame information fusion is achieved to improve the spatial and temporal integrity of environmental perception. In step 2, the high-precision map and navigation map prior data are input into the network, and weighted fusion is performed using the cross-attention mechanism to generate a beyond-visual-range (BEV) feature map. Subsequently, the fused features are mapped into a semantic vector map with the help of a decoder, and a deep neural network is used to convert the features into semantic elements of actual roads, obstacles, and lane lines to realize beyond-visual-range map construction.
3. The method for constructing a semantic vector map that supports beyond-horizon and dynamic updating according to claim 1, characterized in that: In step three, historical memory latent variables are combined with current frame features to improve map update accuracy. At the same time, diffusion denoising technology is used to remove noise. Finally, high-resolution feature maps are refined and fused to dynamically generate accurate semantic vector maps that adapt to environmental changes in real time.
4. The method for constructing a semantic vector map that supports beyond-horizon and dynamic updating according to claim 1, characterized in that: The first step includes using a surround view image, wherein the surround view image is generated by stitching images captured by multiple cameras around the vehicle to form a 360-degree panoramic view; The feature of each frame image is set to I. On this basis, the bird's-eye view feature can be expressed as follows: Among them, I t It represents the surround image of the current frame; n refers to the number of history frames; f BEV It is a function used to map the surround view image to a bird's-eye view by synthesizing images from multiple perspectives; By fusing the BEV features of the current frame with the BEV features of the previous frames, we can obtain richer spatial and temporal information. The temporal fusion module will fuse the BEV features of the current frame with the BEV features of the previous n frames, and finally obtain the features after temporal fusion. Among them, f temporal It refers to the time series fusion function, which is achieved by convolutional layers and recurrent neural networks, also known as RNNs; Time series fusion integrates historical information with current information to make up for the missing information in distant areas and expand the scope of map construction.
5. The method for constructing a semantic vector map that supports beyond-horizon and dynamic updating according to claim 1, characterized in that: The second step also includes introducing offline prior data, that is, using navigation maps and high-precision maps as prior information to further enrich the map construction process; High-precision maps typically include more detailed road information than traditional maps; After introducing offline prior data, a cross-attention mechanism is adopted; Assume that the features after time series fusion are The offline map prior data is Then the fusion process of the cross attention mechanism can be expressed as: Among them, f prior Represents the cross attention mechanism function, which gradually completes the weighted fusion of features by calculating the correlation between time series features and offline map prior data; The cross-attention mechanism enables the system to selectively integrate data from offline maps with temporal fusion information.
6. The method for constructing a semantic vector map that supports beyond-horizon and dynamic updating according to claim 5, characterized in that: After introducing time series fusion technology and offline prior data, the autonomous driving system can generate a BEV feature map beyond visual range; The feature map contains the environmental information of the current frame, and also integrates the spatial and temporal information of historical frames, as well as the prior data in the high-precision map and navigation map; During the BVR map construction process, the system decodes the fused BEV features through a deep neural network to generate a semantic vector map containing rich semantic information. The decoding process uses a fully convolutional network to map image features into semantic elements of actual roads, obstacles, lane lines, and traffic signs. The goal of the decoding process is to generate a beyond-line-of-sight semantic vector map, thereby providing the autonomous driving system with more accurate environmental perception, especially in extended-line-of-sight scenarios. Specifically, the decoding process can be expressed by the following formula: Among them, g() represents the decoding function, Represents map prior information and maps the fused features into a vector map with semantic annotations 7. The method for constructing a semantic vector map supporting beyond-horizon and dynamic updating according to claim 1, characterized in that: Step S3 includes assuming that a certain map element has been successfully identified and tracked in the map of the previous frame. With the help of the query propagation mechanism, the location information of these elements will be transmitted to the current frame. We calculate the spatial correlation between the current frame and the previous frame to obtain the updated position of the corresponding element in the current frame. This process can be expressed by the following formula: HAVE BEEN t =f query (HAVE BEEN t-1 ,I t ); Among them, E t Is the map element of the current frame, E t-1 It is the map element of the previous frame, I t is the image of the current frame, f query is a function of the query propagation mechanism; The concept of historical memory latent variables was introduced; Historical memory latent variables are potential feature information accumulated in historical frames. They can help the model better understand the map elements of the current frame, integrate historical information with the map elements of the current frame, make up for key information that may be missing in the current frame, and improve the accuracy of map updates. Specifically, the historical information and the characteristics of the current map elements are integrated with the help of the historical memory latent variable. The latent variable of the historical frame is set to H t-1 , the map element of the current frame is E t , then the fusion process of historical information and current map elements can be expressed as: E fusion =f fusion (E t ,H t-1 ); where f fusion As a fusion function, the map elements of the current frame and the historical latent variables are combined to obtain the fused map element features E fusion ; This fusion approach allows for a deeper understanding of map elements, improving the accuracy of map updates in a dynamically changing environment. In the map update process, the semantic vector map must first be initialized with noise, and then the noise is removed step by step through the diffusion process. Specifically, the diffusion process can be defined as the following formula: Among them, ε noisy is a preliminary noise map; τ is the time step of the diffusion process; It is a diffusion operation, which removes noise by means of step-by-step iteration and finally obtains the denoised semantic map E denosied ; After noise removal, feature information at different levels will be refined and fused to improve the spatial resolution and detail of the map. Refinement is achieved by increasing the resolution of the feature map, making the details in the map clearer, while fusion combines multiple levels of feature information to improve map accuracy and more effectively capture road and obstacle details in complex environments. Specifically, the refinement and fusion process can be expressed by the following formula: E refined =f refine (E denosied ,F high-res ); Among them, E denosied Refers to the map obtained after denoising; F high-res It is a feature map with high resolution characteristics; f refine is the refinement function.
Citation Information
Patent Citations
Map construction method and device, vehicle, storage medium and computer program product
CN118443005A
Online vectorization high-precision map construction method based on multi-modal instance fusion
CN118864651A