Self-adaptive high-precision map construction method and system based on visual language model
The adaptive high-precision map building system based on visual language models solves the problems of high sensor dependence and low multimodal fusion efficiency in existing technologies. It achieves efficient and high-precision map building in complex scenarios, with real-time performance and robustness, and reduces hardware costs and manual annotation requirements.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-16
- Publication Date
- 2026-03-31
AI Technical Summary
Existing autonomous driving map building methods rely on high-precision sensors, resulting in high hardware costs and difficulty in achieving efficient and accurate map building in complex scenarios. The efficiency of multimodal data fusion is low, making it difficult to meet real-time requirements, especially in complex traffic scenarios where inference efficiency is low.
An adaptive high-precision map building system based on a visual language model is adopted. Multimodal prior information is generated in the cloud and combined with real-time data from the vehicle. A BEV feature builder and a query modeler are used to build the map. The collaborative reasoning module performs real-time mapping and deep reasoning in both normal and complex scenarios, forming a closed-loop optimization system.
It enables high-precision map construction in complex scenarios with relatively low hardware costs, improves the real-time performance and accuracy of map construction, reduces inference latency and false detection rate, has adaptive capabilities and interpretability, and reduces manual intervention and annotation costs.
Smart Images

Figure CN121767581A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of high-precision map construction technology, and in particular to an adaptive high-precision map construction method and system based on a visual language model. Background Technology
[0002] Existing autonomous driving map building methods primarily rely on massive amounts of data collected by high-precision sensors (such as LiDAR and high-definition cameras) to generate high-precision maps. This typically involves dividing the road network topology of the high-precision map into multiple road topology structures and determining road centerlines based on lane lines and road junction information. However, existing solutions require a large amount of high-resolution sensor data, resulting in high system hardware costs and limiting their widespread application in resource-constrained scenarios. Multimodal data fusion efficiency (such as images and lane data) is low, and topological reasoning capabilities in complex traffic scenarios are insufficient, making it difficult to meet the demands of real-time map updates. Furthermore, when dealing with complex road environments (such as intersections or roundabouts), the problem of inaccurate data segmentation persists, causing the generated maps to be unable to effectively adapt to changes in the dynamic traffic environment. Especially in large road network environments, processing speed is slow, making it difficult to meet real-time requirements. Summary of the Invention
[0003] In view of the above problems, the present invention provides an adaptive high-precision map construction method and system based on a visual language model to solve the technical problem that the existing technology has a strong dependence on pre-built high-precision maps and is difficult to achieve stable operation in areas not covered by the map (such as newly built roads or temporary construction sections) or in scenarios with limited sensor perception (such as occlusion or severe weather).
[0004] This invention provides an adaptive high-precision map construction system based on a visual language model. The system includes: a cloud-based prior information generation subsystem, located in the cloud, used to perform multimodal processing on a standard-defined map and multi-view environmental images using a visual language model to generate point prior information, topological prior information, and mask prior information, and then send this prior information to the vehicle; and a vehicle-side real-time map construction subsystem, located on the vehicle and connected to the cloud-based prior information generation subsystem, used to fuse the prior information through a BEV feature builder and a query modeler, and then combine it with the standard-defined map and multi-view environmental images. The system generates lane and traffic element models based on a bird's-eye view. The collaborative reasoning subsystem, located on the vehicle, is connected to the vehicle-side real-time map building subsystem. It is used to build maps in real time in normal scenarios based on the lane and traffic element models, and to perform deep visual language reasoning in complex scenarios to complete topological and semantic relationships and form high-precision maps. The system feedback and optimization subsystem, located in the cloud, is connected to the cloud-based prior information generation subsystem and the collaborative reasoning subsystem. It is used to retrain the visual language model and dynamically update the prior information based on the real-time mapping results, thereby forming a continuously evolving closed-loop system.
[0005] Furthermore, the system also includes an input module, which is connected to the cloud-based prior information generation subsystem and the vehicle-side real-time map construction subsystem, respectively, for acquiring standard-defined maps, multi-view environmental images and radar perception data, and sending them to the cloud-based prior information generation subsystem and the vehicle-side real-time map construction subsystem.
[0006] Furthermore, the cloud-based prior information generation subsystem includes: an input alignment module, used to encode and align the input standard definition map, multi-view environmental images and text, unify the input format, generate feature space aligned embedding representations, and provide structured input for the visual language model; a multimodal fusion module, connected to the input alignment module, used to perform feature fusion on image, semantic text and navigation map information using the visual language model; and a prior information generation module, connected to the multimodal fusion module, used to extract point prior information, mask prior information and topological prior information from the features output by the visual language model, and send them to the vehicle end, while caching them in the cloud.
[0007] Furthermore, the point prior information represents road nodes and key geometric points; the mask prior information represents lane regions and semantic segmentation structures; and the topological prior information represents the topological connectivity between roads.
[0008] Furthermore, the vehicle-side real-time map construction subsystem includes: a feature fusion module, which uses a feature pyramid network to extract multi-scale features and projects image features onto a bird's-eye view plane to form bird's-eye view features; a fusion input module, which aligns the network and fuses prior information with bird's-eye view features through a BEV feature builder and a query modeler to form a bird's-eye view map; and an inference structure module, which combines a standard definition map and multi-view environmental images, uses a neural network framework to capture geometric relationships, updates bird's-eye view features, performs lane line prediction, generates a lane and traffic element model based on the bird's-eye view, and completes the initial construction of the map.
[0009] Furthermore, the collaborative reasoning subsystem includes: a fast reasoning module, used to output real-time lane lines and topological information based on a pre-constructed map in normal scenarios to form a high-precision map; and a slow reasoning module, used to perform deep reasoning based on a pre-constructed map in complex scenarios to complete topological and semantic relationships and form a high-precision map.
[0010] Furthermore, the system feedback and optimization subsystem includes: an error calculation module, used to calculate the error between the vehicle-side inference result and the actual result; and a parameter update module, which updates the visual language model parameters based on the error calculation result, optimizes the inference process, forms a closed-loop learning system, and continuously optimizes the map construction process.
[0011] This invention provides an adaptive high-precision map construction method based on a visual language model. The method includes: Step 1, performing multimodal processing on a standard-defined map and multi-view environmental images using a visual language model to generate point prior information, topological prior information, and mask prior information, and sending this prior information to the vehicle; Step 2, fusing the prior information using a BEV feature builder and a query modeler, and then combining it with the standard-defined map and multi-view environmental images to generate a lane and traffic element model based on a bird's-eye view; Step 3, performing real-time mapping in conventional scenarios and deep visual language inference in complex scenarios based on the lane and traffic element model to complete the topological and semantic relationships and form a high-precision map; Step 4, retraining and dynamically updating the visual language model based on the real-time mapping results to form a continuously evolving closed-loop system.
[0012] Furthermore, step 1 includes: Step 11, performing feature encoding and alignment on the input standard definition map, multi-view environment image, and text according to the following formula, unifying the input format, generating a feature space-aligned embedding representation, and providing structured input for the visual language model. ,in, , , These are map, image, and text encoders, respectively. , , For the corresponding embedding representation; Step 12, the visual language model fuses features from the image, semantic text, and navigation map information using the following formula: , ,in, , , For query, key, value mapping matrix, For feature dimension, For the fused multimodal feature representation; Step 13, based on the point prior generative formula Extract key geometric points and generate prior information for those points. It is a linear transformation matrix. The activation function; Step 14, based on the topological prior generative formula Based on the connection probabilities between lanes, topological prior information between lanes is generated, where... lane With lane The connectivity probability, This represents vector concatenation. Represents the topological relationship prediction matrix; Step 15, extract visual features using convolutional layers and upsample them, based on the mask prior generative formula. Generate semantic mask prior information, where, This indicates that the convolutional layer is used to extract local visual features. Indicates an upsampling operation. Indicate the feature fusion weights; Step 16, based on the output compression formula All prior information is compressed and quantized for easy transmission to the vehicle. It is the output after quantization. Prior information for points, This is topological prior information. This is the prior information for the mask.
[0013] Furthermore, step 2 includes: step 21, extracting multi-scale features using a feature pyramid network and projecting the image features onto a bird's-eye view plane, according to a functional... Feature fusion is performed to form a bird's-eye view feature, among which... The feature pyramid network represents the extraction of multi-scale features. This indicates that image features are projected onto the BEV plane. Indicate the vehicle's current pose; Step 22, align the network, and through the BEV feature builder and query modeler, according to the functional... By fusing prior information with bird's-eye view features, a bird's-eye view map is formed, in which... To align the network, The initial input features include bird's-eye view features and prior information; step 23, combining the standard defined map and multi-view environmental images, according to the functional formula... The system uses a neural network framework to capture geometric relationships, update bird's-eye view features, predict lane lines, and generate a model of lanes and traffic elements based on a bird's-eye view, thus completing the initial map construction. This indicates that the internal geometry of the query is captured. This indicates the fusion of external prior information. This indicates that the feedforward layer performs feature updates; step 24, based on the following function, implements the graph construction output. , , ,in, Indicates lane line prediction, This represents semantic mask prediction. Indicates topological relationship prediction. This represents the learning weight matrix related to lane prediction. This represents the learning weight matrix associated with mask prediction. This represents the input feature matrix, sigmoid is the activation function, and || is the vector concatenation operator.
[0014] This invention provides an adaptive high-precision map construction method and system based on a visual language model, mainly addressing the following problems in existing technologies: 1) Traditional map generation methods are overly reliant on high-precision sensors: Existing autonomous driving map construction methods often rely on large amounts of sensor data (such as LiDAR and high-definition cameras), resulting in high hardware costs and limitations imposed by hardware performance. Especially in complex scenarios (such as intersections or occluded areas), existing technologies struggle to provide efficient and accurate map construction solutions. 2) Lack of multimodal fusion of prior information: Existing methods are insufficient in multimodal data fusion, failing to effectively combine SD maps (standard defined maps) and visual information, thus failing to fully utilize prior information obtained from different modalities (images, text, etc.). 3) Low inference efficiency and inability to handle complex scenarios in real time: Existing symbolic reasoning-based methods are inefficient in handling complex scenarios, especially in complex scenarios with multiple traffic elements and intersections, making it difficult to achieve fast and efficient reasoning. Attached Figure Description
[0015] Figure 1 A schematic diagram of an adaptive high-precision map construction system based on a visual language model provided by the present invention; Figure 2 A flowchart of an adaptive high-precision map construction method based on a visual language model provided by the present invention; Figure 3This is a flowchart of a method for generating prior information provided by the present invention; Figure 4 This is a flowchart of a preliminary map construction method provided by the present invention. Detailed Implementation
[0016] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.
[0017] Device Example: This invention provides an adaptive high-precision map construction system based on a visual language model, such as... Figure 1 As shown, the system includes a cloud-based prior information generation subsystem, a vehicle-side real-time map construction subsystem, a collaborative reasoning subsystem, and a system feedback and optimization subsystem.
[0018] The cloud-based prior information generation subsystem, located in the cloud, is used to perform multimodal processing on standard-defined maps and multi-view environmental images through a visual language model to generate point prior information, topological prior information, and mask prior information, and then send this prior information to the vehicle. The cloud-based system utilizes a Vision-Language Model (VLM, also known as a visual-language model, which possesses cross-modal understanding capabilities and can simultaneously process image and text information; deployed in the cloud in this solution, it generates prior information related to road geometry, semantics, and topology) to perform multimodal processing on standard-defined maps (SD Maps) and multi-view environmental images. This generates three types of high-value prior information: point-priors, topology-priors, and mask-priors. This prior information is transmitted to the vehicle via the network. The point-priors represent road nodes and key geometric points; the mask-priors represent lane regions and semantic segmentation structures; and the topology-priors represent the topological connectivity relationships between roads. The cloud-based prior information generation subsystem includes: an input alignment module, used to encode and align the input standard definition map, multi-view environmental images and text, unify the input format, generate feature space aligned embedding representations, and provide structured input for the visual language model; a multimodal fusion module, connected to the input alignment module, used to perform feature fusion on images, semantic text and navigation map information using the visual language model; and a prior information generation module, connected to the multimodal fusion module, used to extract point prior information, mask prior information and topological prior information from the features output by the visual language model, and send them to the vehicle end, while caching them in the cloud.
[0019] The vehicle-side real-time map building subsystem is set up on the vehicle side and connected to the cloud-based prior information generation subsystem. It is used to fuse prior information through the BEV feature builder and query modeler, and then combine it with standard definition maps and multi-view environmental images to generate a lane and traffic element model based on a bird's-eye view. The prior information is fused by the vehicle-side BEV feature builder (Bird's Eye View Constructor) and query modeler (SDQuery) to generate a lane and traffic element model based on BEV (Bird's Eye View, which converts multi-view images into a top-down perspective to facilitate unified modeling of the spatial location and structural relationships of roads). The vehicle-side real-time map construction subsystem includes: a feature fusion module, which uses a feature pyramid network to extract multi-scale features and projects image features onto the bird's-eye view plane to form bird's-eye view features; a fusion input module, which aligns the network and fuses the prior information with the bird's-eye view features through the BEV feature builder and query modeler to form a bird's-eye view image; and an inference structure module, which combines a standard-defined map and multi-view environmental images, uses a neural network framework to capture geometric relationships, updates bird's-eye view features, performs lane line prediction, and generates a lane and traffic element model based on the bird's-eye view, completing the initial map construction.
[0020] The collaborative reasoning subsystem, located on the vehicle side and connected to the vehicle-side real-time map building subsystem, is used to build maps in real time in normal scenarios based on lane and traffic element models, and to perform deep visual language reasoning in complex scenarios to complete topological and semantic relationships and form high-precision maps. The collaborative reasoning subsystem employs a fast-slow dual-reasoning structure to achieve a balance between efficiency and high accuracy: the fast system handles real-time mapping for standard scenarios, while the slow system invokes the Visual-Language Model (VLM) to perform deep visual-linguistic reasoning in complex or occluded scenarios, completing topological and semantic relationships. The collaborative reasoning subsystem includes: a fast reasoning module, used in standard scenarios to output real-time lane lines and topological information based on the initially constructed map, forming a high-precision map; and a slow reasoning module, used in complex scenarios to perform deep reasoning based on the initially constructed map, completing topological and semantic relationships, and forming a high-precision map.
[0021] The system feedback and optimization model subsystem uploads the latest perception and mapping results from the vehicle to the cloud for VLM retraining and prior dynamic updates, forming a continuously evolving "generation-application-feedback-optimization" closed-loop system. The system feedback and optimization subsystem, located in the cloud, is connected to the cloud-based prior information generation subsystem and collaborative reasoning subsystem, respectively. It is used to retrain the visual language model and dynamically update the prior information based on real-time mapping results, thus forming a continuously evolving closed-loop system.
[0022] The system feedback and optimization subsystem includes: an error calculation module, used to calculate the error between the vehicle-side inference result and the actual result; and a parameter update module, which updates the visual language model parameters based on the error calculation result, optimizes the inference process, forms a closed-loop learning system, and continuously optimizes the map construction process.
[0023] The system also includes an input module, which is connected to the cloud-based prior information generation subsystem and the vehicle-side real-time map construction subsystem, respectively. The input module is used to acquire standard-defined maps, multi-view environmental images and radar perception data, and send them to the cloud-based prior information generation subsystem and the vehicle-side real-time map construction subsystem.
[0024] This invention provides an adaptive high-precision map construction method and system based on a visual language model to solve the technical problem that existing technologies rely heavily on pre-built high-precision maps and are difficult to operate stably in areas not covered by the map (such as newly built roads or temporary construction sections) or in scenarios where sensor perception is limited (such as occlusion or severe weather).
[0025] Method Example: This invention provides an adaptive high-precision map construction method based on a visual language model, such as... Figure 2 As shown, the method includes the following steps.
[0026] Step 1: Perform multimodal processing on the standard definition map and multi-view environmental images using a visual language model to generate point prior information, topological prior information and mask prior information, and send this prior information to the vehicle end. This step is the core upstream component of this invention, responsible for performing multimodal fusion analysis on the input standard definition map, multi-view images, and text prompts (Prompt). It generates three types of key prior information through a Vision-Language Model (VLM): point prior, topological prior, and mask prior. This prior information will be input as knowledge constraints to the vehicle-side system, supporting BEV feature construction and high-precision map generation. For example... Figure 3 As shown, step 1 includes: Step 11: Based on the following formula, perform feature encoding and alignment on the input standard definition map, multi-view environment images, and text to unify the input format and generate feature space-aligned embedding representations, providing structured input for the visual language model. ,in, , , These are map, image, and text encoders, respectively. , , This is the corresponding embedded representation; Step 12: The visual language model fuses features from images, semantic text, and navigation map information using the following formula: , , in, , , For query, key, value mapping matrix, For feature dimension, This represents the multimodal features after fusion. Step 13, based on the point prior generative formula Extract key geometric points and generate prior information for those points. It is a linear transformation matrix. For activation functions; Step 14, based on the topological prior generative formula Based on the connection probabilities between lanes, topological prior information between lanes is generated, where... lane With lane The connectivity probability, This represents vector concatenation. Represents the topological relationship prediction matrix; Step 15: Extract visual features using convolutional layers and upsample them, based on the mask prior generative formula. Generate semantic mask prior information, where, This indicates that the convolutional layer is used to extract local visual features. Indicates an upsampling operation. Indicates the feature fusion weights; Step 16, according to the output compression format All prior information is compressed and quantized for easy transmission to the vehicle. It is the output after quantization. Prior information for points, This is topological prior information. This is the prior information for the mask.
[0027] Step 2: The prior information is fused by the BEV feature builder and query modeler, and then combined with the standard definition map and multi-view environmental images to generate a lane and traffic element model based on a bird's-eye view. This step is the core execution stage of the invention, responsible for fusing multimodal prior information (point prior, topological prior, mask prior) transmitted from the cloud with vehicle-side perception data (camera, radar, etc.) to generate a structured map representation based on a bird's-eye view (BEV). The invention employs a lightweight Transformer architecture and a hierarchical feature fusion mechanism, ensuring both real-time performance and maintaining topological accuracy and semantic consistency. The Transformer architecture is a deep learning model used for various natural language processing tasks, possessing advantages such as long-distance dependency modeling, parallel computing capabilities, and general performance, and has been widely applied to the processing of serial data. Figure 4 As shown, step 2 includes: Step 21: Use a feature pyramid network to extract multi-scale features and project the image features onto a bird's-eye view plane, according to the functional... Feature fusion is performed to form a bird's-eye view feature, among which... The feature pyramid network represents the extraction of multi-scale features. This indicates that image features are projected onto the BEV plane. Indicates the current position of the vehicle; Step 22, align the network, using the BEV feature builder and query modeler, according to the functional... By fusing prior information with bird's-eye view features, a bird's-eye view map is formed, in which... To align the network, The initial input features include bird's-eye view features and prior information.
[0028] Step 23, combining the standard definition map and multi-view environmental images, according to the functional... The system uses a neural network framework to capture geometric relationships, update bird's-eye view features, predict lane lines, and generate a model of lanes and traffic elements based on a bird's-eye view, thus completing the initial map construction. This indicates that the internal geometry of the query is captured. This indicates the fusion of external prior information. This indicates that the feedforward layer implements feature updates; Step 24: Based on the following function, generate the graph output. , , , in, Indicates lane line prediction, This represents semantic mask prediction. Indicates topological relationship prediction. This represents the learning weight matrix related to lane prediction. This represents the learning weight matrix associated with mask prediction. This represents the input feature matrix, sigmoid is the activation function, and || is the vector concatenation operator.
[0029] Step 3: Based on the lane and traffic element model, perform real-time mapping in normal scenarios and deep visual language reasoning in complex scenarios to complete the topological and semantic relationships and form a high-precision map. The collaborative reasoning subsystem is used to adaptively select the reasoning path under different scene complexities, achieving a balance between real-time performance and accuracy. This invention adopts a collaborative structure of a fast system and a slow system: the former is responsible for fast processing of normal scenes, while the latter performs deep symbolic reasoning in complex or uncertain scenes. The process of selecting the reasoning path is as follows: (1) According to the formula Then, the reasoning path is selected, where... Indicates the complexity of the scenario; Indicates the confidence level of the inference; , Indicates the selection threshold. (2) Calculate the complexity function: ,in, Lane geometry complexity; Visual uncertainty; Average topological connectivity , , : Weighting coefficients, used to balance the contributions of different terms in the overall cost function. (3) Fast system output: The Transformer quickly generates prediction results. This is the output of the fast system, representing the prediction results generated by the fast inference system in typical scenarios. This is a fast inference function, representing the inference result quickly generated by the Transformer model. This function utilizes BEV feature maps and prior information to perform fast prediction. A bird's-eye view (BEV) feature map, typically obtained through image projection, is used to represent the vehicle's position and environmental information in three-dimensional space. It is multimodal input data, which usually includes point priors, topological priors and mask priors, etc., to assist the reasoning process. (4) Slow system reasoning: , : Language model recursive inference function; Output the reasoning result of the last logical chain. This represents the updated state after reasoning in the slow system. This state is derived through recursive reasoning and is used to represent the result of the next step of reasoning. It is a language model adaptation inference function used for inference via a visual-language model (VLM). This function receives the current state. and input features And generate inference results. This refers to the input feature data, typically acquired from sensors or other data sources, which is provided to the VLM for inference. This is the output of the slow system, representing the reasoning results generated by the slow system in complex or uncertain scenarios. The slow system utilizes deep vision-language reasoning to complete topological relationships and semantic information. It is an output function used to output the state of the previous step. It processes and outputs the final inference result. It is typically used to map states to the final inference decision or the last layer of the inference process.
[0030] Step 4: Based on the real-time mapping results, the visual language model is retrained and dynamically updated with priors, thus forming a continuously evolving closed-loop system.
[0031] This step constitutes the core mechanism of the "cloud-edge closed-loop learning" of this invention, which transmits the map data and inference results generated by the vehicle during actual operation back to the cloud to realize the continuous optimization and dynamic evolution of the model. Specific process: (1) Error calculation: , : Vehicle-side inference results; : Accurate labeling; : Comprehensive loss function. (2) Parameter update: , Model parameters; Learning rate; Gradient update. (3) Definition of feedback loss: The optimization is achieved by weighting and combining geometric errors, semantic errors, and topological errors. This is the feedback loss function, representing the loss calculated by the system during the feedback process. This loss is calculated using a multinomial weighted loss function and is mainly used for model optimization. , , These are weighting coefficients used to balance the importance of different types of loss in the feedback loss, corresponding to geometric loss ( semantic loss ) and topological loss ( ), Geometric loss represents the loss associated with the map's geometry (such as the prediction accuracy of lane lines). This loss measures the difference between the predicted geometric information and the actual geometric information. This is semantic loss, representing the loss associated with semantic information (such as road type, obstacle classification, etc.). This loss measures the accuracy of semantic information prediction. It is the topological loss, which represents the loss associated with the road topology. This loss measures the connectivity and accuracy of the topological relationships between road elements. (4) Closed-loop process: This represents the cyclical process of the system from perception and mapping to optimization, enabling continuous model evolution. Indicates the input of the sensing module ( ) and perceived data ( These data include environmental information collected from sensors or other sources, used for subsequent map generation and inference. This represents the output after processing sensor data (such as BEV feature maps), which is used as input to the subsequent inference system. This is a bird's-eye view (BEV) feature map, representing a top-down perspective converted from an image, which is very useful for road environment modeling. The output of the inference module is based on sensor data and generated prior information, producing environment-related predictions. Here, represents the feedback loss function, which is calculated through the feedback process. This loss is used to optimize the model and update it by comparing the differences between the predictions and the actual data. This indicates the updated model parameters. After the feedback optimization process, the model parameters are improved.
[0032] This invention provides an adaptive high-precision map construction method and system based on a visual language model to solve the technical problem that existing technologies rely heavily on pre-built high-precision maps and are difficult to operate stably in areas not covered by the map (such as newly built roads or temporary construction sections) or in scenarios where sensor perception is limited (such as occlusion or severe weather).
[0033] In summary, this invention provides an adaptive high-precision map construction method and system based on a visual language model. This invention improves the accuracy and generalization ability of prior generation: In the cloud, VLM is used to perform cross-modal fusion of standard-defined maps and images, generating three types of high-confidence prior information: points, topology, and masks. This mechanism maintains high geometric consistency and semantic stability even under complex road conditions, lighting changes, and occlusion scenarios; it achieves real-time mapping and enhanced structural consistency on the edge: On the vehicle side, a BEV feature builder fuses prior information with multi-sensor data, using a lightweight Transformer structure to achieve real-time inference of ≥30FPS on an embedded platform, while ensuring lane connectivity and semantic consistency at the topology level, improving map construction efficiency and accuracy; it possesses adaptive fast / slow inference capabilities: The collaborative inference module introduces dual criteria of scene complexity and confidence, automatically switching between Fast and Slow based on the actual scenario. The system achieves efficient operation in conventional scenarios and semantic completion in complex scenarios, significantly reducing inference latency and false detection rate. It establishes a closed-loop optimization mechanism for data and models: the system feedback module automatically calculates error gradients based on vehicle-side operational data, performs incremental model updates and prior regeneration, constructing a "generation-inference-feedback-optimization" loop structure, enabling continuous evolution and optimization of the model during long-term operation across multiple scenarios. It reduces manual intervention and annotation costs: through automatic sampling on the edge and weakly supervised retraining in the cloud, it reduces the need for manual annotation, achieving a semi-automated map update process and improving the economy and scalability of map construction. It enhances the interpretability and robustness of the system: by adopting explicit topology and mask priors, the inference process has multi-layered structural constraints of geometry, semantics, and logic, making the results traceable and easily tuned, significantly improving robustness in occlusion, dynamic obstacles, and unstructured road environments.
[0034] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.
Claims
1. An adaptive high-precision map construction system based on a visual language model, characterized by, The system comprises: A cloud prior information generation subsystem is arranged in the cloud, and is configured to perform multimodal processing on the standard definition map and the multi-view environment image by using a visual language model, to generate point prior information, topological prior information, and mask prior information, and to send the prior information to the vehicle end; A vehicle end real-time map construction subsystem is arranged in the vehicle end, and is connected to the cloud prior information generation subsystem, and is configured to fuse the prior information by using a BEV feature constructor and a query modeling device, and to combine the standard definition map and the multi-view environment image to generate a lane and traffic element model based on a bird's eye view; A collaborative reasoning subsystem is arranged in the vehicle end, and is connected to the vehicle end real-time map construction subsystem, and is configured to perform real-time mapping in a conventional scene according to the lane and traffic element model, and to perform deep visual language reasoning in a complex scene to complete the topological and semantic relationship, thereby forming a high-precision map; A system feedback and optimization subsystem is arranged in the cloud, and is connected to the cloud prior information generation subsystem and the collaborative reasoning subsystem, and is configured to retrain the visual language model and dynamically update the prior information according to the real-time mapping result, thereby forming a closed-loop system that continuously evolves. 2.The adaptive high-precision map construction system based on visual language model according to claim 1, wherein, The system further comprises an input module connected to the cloud prior information generation subsystem and the vehicle end real-time map construction subsystem, and configured to acquire the standard definition map, the multi-view environment image, and radar sensing data, and to send the data to the cloud prior information generation subsystem and the vehicle end real-time map construction subsystem. 3.The adaptive high-precision map construction system based on visual language model of claim 1, wherein, The cloud prior information generation subsystem comprises: An input alignment module configured to encode and align the input standard definition map, multi-view environment image, and text, to unify the input format, to generate embedded representations of feature space alignment, and to provide structured input for the visual language model; A multimodal fusion module connected to the input alignment module, and configured to use the visual language model to fuse features of the image, semantic text, and navigation map information; A prior information generation module connected to the multimodal fusion module, and configured to extract point prior information, mask prior information, and topological prior information from the features output by the visual language model, and to send the information to the vehicle end and to cache the information in the cloud. 4.The adaptive high-precision map construction system based on visual language model according to claim 3, wherein, The point prior information represents road nodes and key geometric points; the mask prior information represents lane regions and semantic segmentation structures; and the topological prior information represents topological connectivity between roads. 5.The adaptive high-precision map construction system based on visual language model according to claim 2, wherein, The vehicle end real-time map construction subsystem comprises: A feature fusion module configured to use a feature pyramid network to extract multi-scale features, and to project image features to a bird's eye plane to form bird's eye view features; A fusion input module configured to align the network, and to fuse the prior information and the bird's eye view features by using the BEV feature constructor and the query modeling device to form a bird's eye view map; A reasoning structure module configured to combine the standard definition map and the multi-view environment image, to use a neural network framework to capture geometric relationships, to update the bird's eye view map features, to perform lane line prediction, to generate a lane and traffic element model based on a bird's eye view, and to complete preliminary construction of the map. 6.The adaptive high-precision map construction system based on visual language model according to claim 1, wherein, The collaborative reasoning subsystem comprises: A fast inference module is configured to output real-time lane line and topology information based on the preliminary constructed map in a regular scenario, thereby forming a high-precision map. A slow inference module is configured to perform deep inference based on the preliminary constructed map in a complex scenario, thereby completing the topology and semantic relationship and forming a high-precision map. 7.The adaptive high-precision map construction system based on visual language model according to claim 6, wherein, The system feedback and optimization subsystem comprises: An error calculation module is configured to calculate the error between the inference result and the actual result. A parameter updating module is configured to update the visual language model parameters based on the error calculation result, optimize the inference process, form a closed-loop learning system, and continuously optimize the map construction process.
8. A method of adaptive high-precision map construction using the visual language model-based system according to any one of claims 1-7, characterized in that, The method comprises: Step 1: The visual language model is used to perform multi-modal processing on the standard definition map and multi-view environment images, to generate point prior information, topology prior information and mask prior information, and to send the prior information to the vehicle end. Step 2: The BEV feature constructor and the query modeler are used to fuse the prior information, and the standard definition map and multi-view environment images are combined to generate a bird's eye view-based lane and traffic element model. Step 3: According to the lane and traffic element model, real-time mapping is performed in a regular scenario, and deep visual language inference is performed in a complex scenario to complete the topology and semantic relationship, thereby forming a high-precision map. Step 4: The visual language model is retrained and prior dynamic updated according to the real-time mapping result, thereby forming a continuously evolving closed-loop system.
9. The adaptive high-precision map construction method based on the visual language model according to claim 8, wherein step 1 comprises: Step 11, encode and align the input standard definition map, multi-view environment image and text according to the following formula, unify the input format, generate the embedding representation aligned in the feature space, and provide structured input for the visual language model, wherein, , , are map, image and text encoders respectively, , , are the corresponding embedding representations; Step 12: The visual language model is used to perform feature fusion on the image, semantic text and navigation map information according to the following formula: , , wherein, , , is a query, key, value mapping matrix, is a feature dimension, is a fused multi-modal feature representation; Step 13, generating formula according to point prior , extracting key geometric points to generate point prior information, wherein, is a linear transformation matrix, is an activation function; Step 14, generating formula according to topology prior , generating topology prior information between lanes based on connection probability between lanes, wherein, represents the lane and the communication probability of the lane , represents vector splicing, represents the topology relationship prediction matrix; Step 15, visual features are extracted by using a convolutional layer and up-sampling, and a mask prior generation formula is used according to a mask prior , semantic mask prior information is generated, wherein, indicates that the convolutional layer is used for extracting local visual features, indicates an up-sampling operation, indicates a feature fusion weight; Step 16, according to the output compression formula P All prior information is compressed and quantitatively coded to facilitate transmission to the vehicle end, wherein, is the output after quantization processing, is the point prior information, is the topological prior information, is the mask prior information.
10. The adaptive high-precision map construction method based on the visual language model according to claim 8, wherein step 2 comprises: Step 21, using a feature pyramid network to extract multi-scale features, and projecting the image features to a bird's eye view plane according to a function perform feature fusion to form bird's eye view features, wherein, indicates that the feature pyramid network extracts multi-scale features, indicates that the image features are projected to the BEV plane, indicates the current pose of the vehicle; Step 22, aligning the network, through the BEV feature builder and the query modeler, according to the function Fusing the prior information with the bird's eye view features to form a bird's eye view, wherein, for the aligning network, for the initial input features, including the bird's eye view features and the prior information; Step 23, combining the standard definition map and the multi-view environment image, according to the function , using the neural network framework to capture the geometric relationship, update the bird's eye view feature, perform lane line prediction, generate the lane and traffic element model based on the bird's eye view, and complete the preliminary construction of the map, wherein, represents capturing the geometric relationship inside Query, represents fusing external prior information, represents that the feedforward layer realizes feature updating; Step 24: The mapping output is realized according to the following function, , , , wherein, represents a lane line prediction, represents a semantic mask prediction, represents a topological relationship prediction, represents a learning weight matrix related to the lane prediction, represents a learning weight matrix related to the mask prediction, represents an input feature matrix, sigmoid is an activation function, and || is a vector concatenation operator.