Lightweight online high-precision map construction method, device and system and storage medium

By pruning and optimizing the Self-Attention and Cross-Attention modules, the high computational complexity of the online high-precision map construction model was solved, achieving lightweight and real-time performance improvements.

CN120685107APending Publication Date: 2025-09-23LANZHOU JIAOTONG UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202511035015.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-25
Publication Date
2025-09-23

AI Technical Summary

Technical Problem

The existing online high-precision map construction model has large parameter scale and high computational complexity, which makes it difficult to adapt to application scenarios with limited computing resources and complex and dynamic traffic environments, affecting real-time performance.

Method used

A two-stage pruning optimization framework is adopted to prune and optimize the Self-Attention and Cross-Attention modules. Redundant structures are identified through the Fisher information matrix and mask variables are introduced for fine-tuning to reduce computational complexity and maintain model accuracy.

Benefits of technology

Significantly reduce computational complexity and FLOPs, improve inference efficiency, and enhance the real-time performance and deployment feasibility of lightweight online high-precision map generation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120685107A_ABST
    Figure CN120685107A_ABST
Patent Text Reader

Abstract

The invention discloses a lightweight on-line high-precision map construction method, device and system, and a storage medium. The method comprises the following steps: S1, a map decoder carries out refining processing on vehicle surrounding environment data by using SA and CA modules to construct a lightweight on-line high-precision map; s2, redundant structures existing in the SA module and the CA module are searched, and pruning is carried out by introducing mask variables; and meanwhile, refining and fine tuning are carried out on the pruned map decoder. By adopting the technical scheme of the invention, the model performance can be effectively maintained, and meanwhile, the calculation complexity and FLOPs are remarkably reduced, so that the reasoning efficiency of online high-precision map construction is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of data processing technology, and in particular relates to a lightweight online high-precision map construction method and device, system, and storage medium. Background Art

[0002] High-precision maps play a crucial role in autonomous driving systems, providing precise support for environmental perception and path planning. However, existing online high-precision map construction models suffer from large parameter sizes and high computational complexity, severely limiting their inference efficiency and real-time performance. They are particularly difficult to adapt to application scenarios with limited computing resources and complex and dynamic traffic environments. Summary of the Invention

[0003] The technical problem to be solved by the present invention is to provide a lightweight online high-precision map construction method and device, system, and storage medium.

[0004] To achieve the above object, the present invention adopts the following technical solutions: A lightweight online high-precision map construction method, comprising: Step S1: The map decoder uses the SA and CA modules to refine the vehicle's surrounding environment data to construct a lightweight online high-precision map; Step S2: Perform Fisher mask search on the redundant structures in the SA and CA modules, and implement pruning by introducing mask variables; at the same time, refine and fine-tune the pruned map decoder through continuously adjustable mask variables.

[0005] The present invention also provides a lightweight online high-precision map construction device, comprising: The first processing module is used for the map decoder to refine the vehicle's surrounding environment data using the SA and CA modules to build a lightweight online high-precision map; The second processing module is used to perform Fisher mask search on the redundant structures existing in the SA and CA modules, and implement pruning by introducing mask variables; at the same time, the pruned map decoder is refined and fine-tuned through continuously adjustable mask variables.

[0006] The present invention also provides a lightweight online high-precision map construction system, comprising: a memory and a processor, wherein the memory stores a computer program run by the processor, and the computer program executes a lightweight online high-precision map construction method when run by the processor.

[0007] The present invention also provides a storage medium, on which a computer program is stored, and the computer program executes a lightweight online high-precision map construction method when running.

[0008] This paper designs a two-stage pruning optimization framework for the Self-Attention and Cross-Attention modules in lightweight online high-precision map construction models. The first stage uses the Fisher information matrix to identify redundant structures in the network and perform preliminary pruning. The second stage introduces mask variables to fine-tune the pruned modules, restoring the accuracy of the pruned models in a stable and rapid manner. Compared to traditional pruning methods that require extensive training, this method effectively maintains model performance while significantly reducing computational complexity and FLOPs, thereby improving inference efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are merely embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying any creative work.

[0010] Figure 1 This is a flow chart of a lightweight online high-precision map construction method according to an embodiment of the present invention; Figure 2 This is a detailed diagram of the map encoder; Figure 3 This is a detailed working diagram of the map decoder; Figure 4 Schematic diagram of multi-head attention (SA and CA) pruning; Figure 5 Schematic diagram of mask fine-tuning lightweight online high-precision map accuracy recovery; where (a) is the initial mask, (b) is the mask search, and (c) is the initial tuning. DETAILED DESCRIPTION

[0011] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0012] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.

[0013] Example 1:

[0014] like Figure 1 As shown, an embodiment of the present invention provides a lightweight online high-precision map construction method, including: Step S1: The map decoder uses the SA and CA modules to refine the vehicle's surrounding environment data to construct a lightweight online high-precision map; Step S2: Perform Fisher mask search on the redundant structures in the SA and CA modules, and implement pruning by introducing mask variables; at the same time, refine and fine-tune the pruned map decoder through continuously adjustable mask variables.

[0015] As one implementation of an embodiment of the present invention, in step S1, multi-view images of the vehicle's surroundings, captured by multiple cameras onboard the vehicle, are input as data sources into a map encoder. The map encoder's backbone network then extracts features from the multi-view images. A 2D-to-BEV conversion module efficiently maps the two-dimensional image features into bird's-eye-view spatial features. The resulting BEV features are then fed into a map decoder for further analysis, processing, and progressive refinement of the output map element features. After this processing, a map prediction result is ultimately output.

[0016] In the aforementioned end-to-end lightweight online HD map construction process, the map decoder, after receiving BEV features, refines them using self-attention (SA) and cross-attention (CA) techniques, organically integrating the interrelated elements within the BEV features and ultimately outputting a lightweight online HD map. The map decoder plays a critical role in the lightweight online HD map construction process. Its core task is to accurately reconstruct complex map elements, such as lane dividers and road boundaries, from BEV features. To achieve this, the decoder utilizes SA and CA modules to further process and optimize the BEV feature representation, resulting in high-quality HD map element prediction and output.

[0017] Figure 2 The detailed architecture of the map encoder is as follows: First, the original multi-view image input is fed into a shared ResNet-50 backbone network to extract preliminary image semantics. Multi-scale feature fusion is then performed through the Feature Pyramid Network (FPN) to enhance the perception of objects of varying sizes. The output is 2D image features at multiple scales, one for each camera. Finally, the resulting 2D image features are converted into BEV features.

[0018] Figure 3The detailed architecture of the map decoder is presented, which combines scatter-aggregate query and location embedding mechanisms. The core process of scatter-aggregate query first copies each instance query into multiple scatter queries through a scattering operation. These scatter queries perform feature sampling at different positions in the BEV feature space, aiming to capture the local geometric features of map elements. Then, all scatter queries belonging to the same instance are connected and aggregated through a multi-layer perceptron (MLP) to finally generate an enhanced instance query representation. In this architecture, instance queries serve as the input of the decoder, and each instance query represents an instance of a map element, such as a lane divider or a crosswalk. Location embedding introduces location information for each query through a reference point ref during the query initialization phase. Each scatter-aggregate query generates different location information based on its corresponding reference point. These reference points help the model accurately locate the query in the BEV feature space, so that the query can capture spatial relationships more accurately. Specifically, each instance consists of four reference points, namely 、 、 、 .in Indicates the The location of the points; Indicates the The x-coordinate of a point at the initial moment (or initial state); Indicates the The y coordinate of a point at the initial moment (or initial state); the superscript (0) indicates the "initial value". Position offset 、 、 、 represents the spatial offset between each corresponding reference point and the scatter aggregation query, and Respectively expressed as The offset of each point in the x-coordinate and y-coordinate. By adding the reference point coordinates and the corresponding offsets, the precise location information of each scatter aggregation query is finally calculated, which are 、 、 、 . Indicates the position after the offset The new coordinates of the points, and Represent the x- and y-coordinates of the i-th point after adding the offsets to the x- and y-coordinates, respectively. This process, by combining the position information of the reference point with the position offset, ensures that the spatial relationships of map elements can be accurately modeled using the queried position information, thereby improving the accuracy and efficiency of map construction.

[0019] On this basis, the decoder further enhances the feature representation of scatter-aggregate queries through SA and CA modules, such as Figure 3 As shown in the figure, the SA module acts between instance queries. As shown in formula (1), Q is the query, which indicates the information required to find the "currently processed position". K is the key, which indicates the "identification of all positions" and determines which positions the current query should focus on. V is the value, which indicates the "actual content information of all positions" and is the feature that will be weighted and added to generate the output. 、 and , From the same input sequence, 、 、 is the learned weight matrix. By calculating the similarity between instances, it is possible to capture the content consistency between different points in the same map instance, while allowing information interaction between different map instances, thereby extracting the global structural features of the scene.

[0020]

[0021] Different from the self-attention mechanism, in the CA module 、 、 ,in, is the feature representation of the reference point, Represents BEV characteristics, 、 、 Represents the learned weight matrix. is the operation of scaling the dot product attention, This design allows query information to interact with external features, thereby achieving local feature perception based on location guidance, as shown in formula (2):

[0022] The CA module achieves precise location-guided feature extraction by fusing scattered queries embedded with reference point locations with BEV features. By explicitly decoupling the query's semantic and location information—semantic information shared by instance queries, while location information is independently generated by reference points—the model can independently optimize spatial localization and semantic understanding. This decoupling not only improves the model's prediction accuracy but also enhances its perception of scene structure.

[0023] like Figure 4Figure 2 shows a schematic diagram of the attention-head-based structured pruning method used. In this method, each attention head in the multi-head attention is treated as a functionally coupled unit, including its corresponding query, key, value linear projection, and output mapping structure. To ensure the consistency of the pruned model structure and computational correctness, when pruning an attention head, all its dependent modules must be simultaneously pruned, thus forming a complete structural pruning path. The pruning method follows the principle of finding the optimal mask under sparse constraints described below.

[0024] As an implementation method of an embodiment of the present invention, in step S2, although the SA module and CA module demonstrate good performance in the lightweight online HD map construction process, their high computational complexity, especially when the spatial dimension is large, seriously affects the real-time reasoning efficiency of the model, thereby limiting the timeliness of lightweight online HD map generation. In addition, excessive global attention calculations easily introduce redundant information, resulting in higher computational overhead in a real-time environment. To address the above issues, the embodiment of the present invention proposes pruning optimization for the SA and CA modules to effectively reduce the computational burden of the model while ensuring mapping accuracy, improve reasoning efficiency, and thus enhance the real-time performance and deployment feasibility of lightweight online HD map generation.

[0025] Furthermore, as shown in the definition of SA and CA modules, the pruning problem can be regarded as finding the optimal mask under the sparsity constraint. However, it is very difficult to solve this problem directly without a lot of complex training. To this end, the lightweight online high-precision map construction task is defined as pruning optimization of the SA and CA modules in the model decoder. The specific optimization process is as follows Figure 5 As shown, it is mainly divided into two stages, including mask search and mask fine-tuning, where For the Layer multi-head attention mask. First, redundant structures in the SA and CA modules of the network are searched and identified, and preliminary pruning is performed by introducing mask variables. Second, based on the preliminary pruning, the pruned modules are refined and fine-tuned to achieve rapid recovery of model accuracy and improved stability.

[0026] In the task of lightweight online high-precision map construction, the significant advantage of the pruning method is that its optimization process mainly focuses on the adjustment of mask variables rather than updating the parameters of the entire model. Since the number of pruned mask variables is much smaller than the overall parameter amount of the model, this method allows mask optimization to be performed using only a small number of samples without causing serious overfitting problems. Compared with traditional pruning methods that require retraining based on the complete data set, this method greatly improves computational efficiency by simplifying the pruning optimization process to the selection of mask variables. In addition, since the pruning process only involves the selection of mask variables and does not adjust the core parameters of the model, the embodiment of the present invention regards the model weights as fixed values ​​and uses the mask variables as the only optimization parameters in the pruning process, making it the only optimization variable for the pruning problem, ensuring that the pruning optimization is efficient and feasible, while still maintaining a high level of lightweight online high-precision map construction accuracy in scenarios with limited computing budgets.

[0027] To this end, the embodiment of the present invention formulates the SA and CA pruning problems in lightweight online high-precision map construction as an optimization problem of selecting the optimal mask under given constraints:

[0028] in, for , indicating that the following constraints are satisfied; represents the model loss function during the pruning process, represents the inference latency overhead of the pruned model, This is the preset upper limit for inference latency, which is used to constrain the inference latency of the pruned model to not exceed this threshold.

[0029] To solve the above formula (1) This is usually a tricky thing. Since the loss function Within the domain of definition, Whether it is continuous within the interval is unknown, and it is difficult to optimize directly. The embodiment of the present invention needs to use a suitable approximation method to Solve to obtain the optimal lightweight pruning process .

[0030] Therefore, the embodiment of the present invention adopts the most common Taylor expansion formula for approximation Taylor expansion is often used to approximate complex functions. Given a loss function When m=1, it indicates that the model and the SA and CA modules of the embodiment of the present invention are in the initial state. Therefore, the second-order Taylor expansion is used here to approximate the loss function:

[0031] in, Represents the loss value when not pruned; is the gradient, which measures the loss with respect to the pruning mask The first-order rate of change of is the Hessian matrix, which measures the second-order curvature information of the loss; represents all terms higher than the second order, whose magnitude is in Approaching 1 The velocities of powers of 1 and 2 tend to zero and are usually ignored.

[0032] During the Taylor expansion process, the 0th order uses only the original loss value, ignoring gradient and second-order information, resulting in a crude approximation. The 1st order uses gradient information, but the gradient term (g) can only reflect local linear changes and cannot describe nonlinear trends. The 2nd order Hessian term provides more reliable information and can be used to determine which parameters have the least impact on the loss. Therefore, the embodiment of the present invention uses a second-order Taylor expansion to approximate the loss function. Formula (6) is derived based on the fact that the gradient term approaches 0 at the local optimum (assuming the model converges to the local optimum).

[0033]

[0034] Among them, m is the pruning mask vector; due to is a constant and does not affect the optimization objective. It can be removed and the final objective becomes Formula (7). The optimal mask is determined by the Hessian matrix H. Since the Hessian matrix H is difficult to calculate directly (the computational complexity is high and may lead to numerical instability), the embodiment of the present invention uses the Fisher information matrix Approximate. The Fisher information matrix is ​​the covariance matrix of the first-order gradient, which can approximate the Hessian under maximum likelihood estimation. Ultimately, the optimization objective is converted to:

[0035] in, It means "defined as", Represents the training dataset, the size of ; Represents a sample pair in the dataset, where is the input, is the target label; Represents the input sample when not pruned The loss function of It is the first-order derivative of the loss function with respect to the pruning mask variable m, which measures the impact of pruning on the loss.

[0036] After completing the definition of the redundant modules of SA and CA, the Fisher information matrix needs to be fully calculated when solving Equation 7. , but this is computationally infeasible. Therefore, the embodiment of the present invention adopts a diagonal approximation method, assuming that the Fisher information matrix is ​​a diagonal matrix, thereby simplifying the calculation. Based on this assumption, the optimization objective can be rewritten as:

[0037] in, Indicates the The importance score of the Fisher information corresponding to each mask variable. This approximation greatly reduces the computational complexity while maintaining the effectiveness of pruning optimization. Taking only values ​​0 or 1 can further simplify the optimization objective: 、 in, Represents the index set of pruned model units. Therefore, the pruning process is to preferentially remove the parts that contribute less to the loss function, thereby minimizing the amount of computation while reducing model performance loss.

[0038] In this optimization framework, the diagonal elements of the Fisher information matrix It can be considered as a metric to measure the importance of the model units. The pruning process is to select a set of units with the lowest importance to remove, so as to ensure that the reduction in computational cost does not significantly affect the model performance.

[0039] Under a given delay budget C, the embodiment of the present invention models the pruning process as a delay-constrained optimization problem, whose goal is to find an optimal set of mask variables , while ensuring that the model inference delay does not exceed the budget C, minimize the performance loss caused by pruning. This problem can be formally expressed as:

[0040] in, for , indicating that the following constraints are satisfied; Indicates the first To efficiently solve this constrained optimization problem, the present invention draws on the idea of ​​solving the knapsack problem: on the one hand, increasing the number of SA and CA heads is always beneficial, which helps to make better use of the available computing budget; on the other hand, in order to meet the inference delay constraint and minimize the model performance loss, the importance score should be pruned first. The head with smaller value in the middle is retained, thereby retaining the part of the structure that has a greater impact on the overall performance.

[0041] Finally, the algorithm searches for possible mask values ​​and finds the possible optimal solution by minimizing the importance loss within the computational budget in the corresponding solution space, so that the optimal mask variable The following conditions are met:

[0042] in, represents the optimal pruning mask. Formula (12) states that the total importance score of the pruned model cannot be higher than that of other pruning strategies. This condition ensures that the importance loss after pruning will not be higher than that of other feasible solutions, thus ensuring the optimality of pruning.

[0043] In the above stage, the mask variable is limited to binary values ​​of 0 or 1. This setting may cause the model structure after SA and CA pruning to be too sparse. For tasks such as lightweight online high-precision map construction that require extremely high output accuracy, if certain structural units are directly completely shielded, that is, masked to 0, the expression ability and prediction performance of the model may be significantly weakened. Therefore, the embodiment of the present invention further introduces a continuously adjustable mask variable, allowing it to take continuous values ​​within a given real number interval, thereby improving the accuracy recovery ability and stability of the model after pruning. And the value of the continuous value satisfies a fixed value interval.

[0044] In order to obtain the corresponding optimal mask value from the above continuous interval and minimize the error in the output results of SA and CA, that is, in the SA and CA layers, the embodiment of the present invention aims to optimize the pruned mask variable To minimize the reconstruction error, that is, to make the pruned model as close as possible to the output of the unpruned model. The embodiment of the present invention models this as a layer-by-layer reconstruction error minimization problem based on the linear least squares method. In order to maximize the performance of the pruned model, the embodiment of the present invention optimizes the mask variables layer by layer to restore the activation output of the unpruned model as much as possible. This process can be expressed as:

[0045] in, Indicates SA and CA modules; is the input of the pruned model, is the input of the original model; Indicates application mask of The output of the layer; Indicates the case of no pruning To ensure that the pruning structure in (13) remains consistent, the embodiment of the present invention introduces the following constraints:

[0046] Required optimized mask With the given mask There are exactly the same 0 value positions, which fix which headers cannot be used, that is, the parts with zero masks remain unchanged.

[0047] Since the above optimization objective involves nonlinear transformation, it is difficult to solve it directly. To improve computational efficiency, the embodiment of the present invention converts it into a linear least squares problem:

[0048] In formula 15, is the output activation matrix of the multi-head attention mechanism, which represents the output contribution of all pruned attention heads. is the difference between the pruned output and the original model output. The effectiveness of this method relies on retaining some structure after pruning and optimizing it to restore the original activation as much as possible.

[0049] In formula (16) represents the number of layers in the network, Indicates the first The mask of the layer. represents the output of the attention head after pruning. In formula (17) Represents the output of SA and CA from the first to the Nth layer. is the input of the original model, Represents the input of the pruned multi-head attention layer to ensure that the error calculation is consistent with the residual connection.

[0050] It is difficult to quickly obtain a stable optimal solution for the above formulas (16) and (17) within the corresponding solution interval. The reason is that Figure 3 The output activation matrix A of the multi-head attention mechanism corresponding to is too large. Directly solving the least squares problem may lead to numerical instability. Therefore, the embodiment of the present invention adopts the LSMR solver and introduces a damping parameter to enhance stability. The final form of the optimization problem is:

[0051] The embodiment of the present invention will Reparameterized to That is, fine-tuning is performed based on the original mask 1. Indicates the The continuous real-valued mask vector used for fine-tuning in the layer, which is used to adjust the heads retained in the pruned model, Represents the perturbation vector to be learned. This value is typically small and is used for fine-tuning the mask. The damping parameter is fixed to 1 to prevent over-adjustment of the mask variable during the optimization process. Furthermore, to prevent over-adjustment of the mask variable and thus impacting model stability, this embodiment of the present invention limits its range to a continuous interval. If the adjusted mask value exceeds this range, this embodiment of the present invention terminates the optimization and retains the original value.

[0052] This optimization strategy can improve the accuracy of the pruned model by layer-by-layer reconstruction without affecting the computational complexity, providing a more robust post-pruning accuracy recovery method for the lightweight online high-precision map construction model.

[0053] Example 2:

[0054] The embodiment of the present invention further provides a lightweight online high-precision map construction device, comprising: The first processing module is used for the map decoder to refine the vehicle's surrounding environment data using the SA and CA modules to build a lightweight online high-precision map; The second processing module is used to perform Fisher mask search on the redundant structures existing in the SA and CA modules, and implement pruning by introducing mask variables; at the same time, the pruned map decoder is refined and fine-tuned through continuously adjustable mask variables.

[0055] Example 3:

[0056] An embodiment of the present invention also provides a lightweight online high-precision map construction system, including: a memory and a processor, wherein the memory stores a computer program run by the processor, and the computer program executes a lightweight online high-precision map construction method when run by the processor.

[0057] Example 4:

[0058] An embodiment of the present invention also provides a storage medium, on which a computer program is stored, and the computer program executes a lightweight online high-precision map construction method when running.

[0059] The embodiments described above are merely descriptions of preferred embodiments of the present invention and are not intended to limit the scope of the present invention. Without departing from the spirit of the present invention, various modifications and improvements made to the technical solutions of the present invention by persons skilled in the art should fall within the scope of protection defined by the claims of the present invention.

Claims

1. A lightweight online high-precision map construction method, characterized in that: include: Step S1: The map decoder uses the SA and CA modules to refine the vehicle's surrounding environment data to construct a lightweight online high-precision map; Step S2: Perform Fisher mask search on the redundant structures in the SA and CA modules, and implement pruning by introducing mask variables; at the same time, refine and fine-tune the pruned map decoder through continuously adjustable mask variables.

2. A lightweight online high-precision map construction device, characterized in that: include: The first processing module is used for the map decoder to refine the vehicle's surrounding environment data using the SA and CA modules to build a lightweight online high-precision map; The second processing module is used to perform Fisher mask search on the redundant structures existing in the SA and CA modules, and implement pruning by introducing mask variables; at the same time, the pruned map decoder is refined and fine-tuned through continuously adjustable mask variables.

3. A lightweight online high-precision map construction system, characterized by: include: A memory and a processor, wherein the memory stores a computer program executed by the processor, and when the computer program is executed by the processor, the lightweight online high-precision map construction method as claimed in claim 1 is executed.

4. A storage medium, characterized in that The storage medium stores a computer program, which, when running, executes the lightweight online high-precision map construction method according to claim 1.

Citation Information

Patent Citations

  • Online high-precision vector map generation method based on mask guidance

    CN118247382A

  • Progressive end-to-end trajectory planning method and system based on BEV characteristics

    CN120143673A

  • Transform-based video behavior recognition algorithm

    CN120318724A