Vector Map Construction Method and Related Device Based on Autoregressive Depth Estimation

Through the autoregressive depth estimation method, the pre-trained depth estimation model is used to estimate depth information based on camera images, and the problem of being unable to construct accurate vector maps in autonomous driving is solved, and high-precision environment perception and path planning are realized without a depth camera.

CN120047640BActive Publication Date: 2025-07-08HANGZHOU INNOVATION RES INST OF BEIJING UNIV OF AERONAUTICS & ASTRONAUTICS +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510518181.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-24
Publication Date
2025-07-08
Estimated Expiration
2045-04-24

AI Technical Summary

Technical Problem

In the field of autonomous driving, it is difficult for the prior art to directly acquire depth images through cameras, resulting in the inability to construct accurate vector maps, affecting the accuracy of vehicle path planning and environmental perception.

Method used

The autoregressive depth estimation method is used to estimate the depth information based on the scene image taken by the camera through the pre-trained depth estimation model, and a vector map is constructed in combination with the depth image. The depth estimation model of the autoregressive equation uses a loss function to optimize the residual characteristics of the depth prediction map and the real depth map during the training process to improve the accuracy of depth estimation.

Benefits of technology

In the absence of a depth camera, accurate vector maps are built based on camera images, and environmental perception and path planning accuracy of autonomous driving vehicles are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120047640B_ABST
    Figure CN120047640B_ABST
Patent Text Reader

Abstract

An embodiment of the present application provides a method and related device for constructing a vector map based on autoregressive depth estimation, which relates to the field of computer technology. The method includes: obtaining a depth image corresponding to a scene image according to the scene image obtained by a camera and a pre-trained autoregressive depth estimation model, where the loss used in the training process of the depth estimation model is calculated based on depth prediction maps of different scales and corresponding second target depth maps, the second target depth map is a map obtained by enhancing residual features for a first target depth map, the first target depth map is a map with a fixed size obtained by processing a real depth image, and the residual features are calculated based on the first target depth map and the obtained depth prediction maps; constructing a vector map according to the scene image and the depth image. In this way, the corresponding depth information can be accurately estimated based on the scene image captured by the camera, and then the vector map can be constructed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular, to a method and related device for constructing a vector map based on autoregressive depth estimation. Background Art

[0002] In the field of autonomous driving, vector maps can provide precise road information, lane lines, traffic signs and other details, providing a basic guarantee for the path planning, decision-making and control of autonomous driving vehicles. At the same time, vehicles need to accurately perceive the surrounding environment through vector maps under various weather conditions to ensure safe driving.

[0003] Currently, it is generally necessary to construct a vector map based on the images captured by a camera and the corresponding depth images. However, in actual scenarios, due to the influence of various factors, it is impossible to directly collect depth images. For example, it is not convenient to install a light field camera for obtaining depth images. Therefore, how to obtain a depth image corresponding to the image captured by the camera to construct a vector map has become a technical problem that those skilled in the art urgently need to solve. Summary of the Invention

[0004] Embodiments of the present application provide a method and related device for constructing a vector map based on autoregressive depth estimation, which can accurately estimate the corresponding depth information based on the scene image captured by the camera, and then construct a vector map by combining the scene image and the depth information.

[0005] Embodiments of the present application can be implemented as follows:

[0006] In a first aspect, embodiments of the present application provide a method for constructing a vector map based on autoregressive depth estimation, the method including:

[0007] Obtaining a scene image obtained by image acquisition through a camera;

[0008] According to the scene image and a pre-trained autoregressive depth estimation model, obtaining a depth image corresponding to the scene image, where the loss used in the training process based on the image pair is calculated based on depth prediction maps of different scales and corresponding second target depth maps, the image pair includes a corresponding sample scene image and a true depth image, the second target depth map is a map obtained by enhancing the residual features of the first target depth map, the first target depth map is a fixed-size map obtained by processing the true depth image, and the residual features are calculated based on the first target depth map and the obtained depth prediction map;

[0009] Constructing a vector map according to the scene image and the depth image.

[0010] Second aspect, an embodiment of the present application provides a vector map construction device based on autoregressive depth estimation, the device comprising:

[0011] An image acquisition module, configured to acquire a scene image obtained by image acquisition through a camera;

[0012] A depth estimation module, configured to obtain a depth image corresponding to the scene image according to the scene image and a pre-trained autoregressive depth estimation model, wherein the loss used in the training process based on image pairs of the depth estimation model is calculated based on depth prediction maps of different scales and corresponding second target depth maps, the image pairs include corresponding sample scene images and real depth images, the second target depth map is a map obtained by enhancing residual features for a first target depth map, the first target depth map is a fixed-size map obtained by processing the real depth image, and the residual features are calculated based on the first target depth map and the obtained depth prediction map;

[0013] A construction module, configured to construct a vector map according to the scene image and the depth image.

[0014] Third aspect, an embodiment of the present application provides an electronic device, including a processor and a memory, the memory stores machine-executable instructions that can be executed by the processor, and the processor can execute the machine-executable instructions to implement the vector map construction method based on autoregressive depth estimation described in the foregoing embodiments.

[0015] Fourth aspect, an embodiment of the present application provides a readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the vector map construction method based on autoregressive depth estimation described in the foregoing embodiments.

[0016] The vector map construction method and related device based on autoregressive depth estimation provided by the embodiments of the present application first obtain a scene image acquired by image capture using a camera; then, according to the scene image and a pre-trained autoregressive depth estimation model, obtain a depth image corresponding to the scene image. The loss used in the training process based on image pairs of this depth estimation model is calculated based on depth prediction maps at different scales and corresponding second target depth maps. An image pair includes a corresponding sample scene image and a true depth image. The second target depth map is a map obtained by enhancing the residual features of a first target depth map. The first target depth map is a fixed-size map obtained by processing the true depth image, and the residual features are calculated based on the first target depth map and the obtained depth prediction map; finally, a vector map is constructed according to the scene image and the depth image. In this way, the corresponding depth information can be accurately estimated based on the scene image captured by the camera, and then a vector map can be constructed by combining the scene image and the depth information. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] To more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings required to be used in the embodiments. It should be understood that the following drawings only show some embodiments of the present application, and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.

[0018] Figure 1 It is a block diagram of an electronic device provided by an embodiment of the present application;

[0019] Figure 2 It is one of the flow diagrams of the vector map construction method based on autoregressive depth estimation provided by an embodiment of the present application;

[0020] Figure 3 It is another flow diagram of the vector map construction method based on autoregressive depth estimation provided by an embodiment of the present application;

[0021] Figure 4 For Figure 3 It is a flow diagram of the sub-steps included in step S100 in

[0022] Figure 5 It is a schematic diagram of a vector map construction process provided by an embodiment of the present application;

[0023] Figure 6 For Figure 2 It is a flow diagram of the sub-steps included in step S400 in

[0024] Figure 7 For Figure 6One of the schematic flowcharts of the sub-steps included in sub-step S420;

[0025] Figure 8 is Figure 6 Another schematic flowchart of the sub-steps included in sub-step S420;

[0026] Figure 9 is Figure 8 The schematic flowchart of the sub-steps included in sub-step S425;

[0027] Figure 10 is Figure 8 The schematic flowchart of the sub-steps included in sub-step S426;

[0028] Figure 11 One of the block diagrams of the vector map construction device based on autoregressive depth estimation provided by the embodiments of the present application;

[0029] Figure 12 Another block diagram of the vector map construction device based on autoregressive depth estimation provided by the embodiments of the present application.

[0030] Icons: 100 - Electronic device; 110 - Memory; 120 - Processor; 130 - Communication unit; 200 - Vector map construction device based on autoregressive depth estimation; 210 - Training module; 220 - Image acquisition module; 230 - Depth estimation module; 240 - Construction module. Detailed implementation manners

[0031] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. Usually, the components of the embodiments of the present application described and illustrated in the accompanying drawings here can be arranged and designed in various different configurations.

[0032] Therefore, the following detailed description of the embodiments of the present application provided in the accompanying drawings is not intended to limit the scope of the present application claimed, but merely represents the selected embodiments of the present application. All other embodiments obtained by those skilled in the art based on the embodiments of the present application without creative efforts fall within the scope of protection of the present application.

[0033] It should be noted that relational terms such as "first" and "second" are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the existence of other identical elements in the process, method, article or device including the elements.

[0034] In conjunction with the accompanying drawings, some embodiments of the present application are described in detail below. In the absence of conflict, the following embodiments and features in the embodiments can be combined with each other.

[0035] Please refer to Figure 1 , Figure 1 A block diagram of an electronic device 100 provided in an embodiment of the present application. The electronic device 100 may be, but is not limited to, a vehicle-mounted terminal device, a cloud device, a computer, a server, etc. The electronic device 100 may include a memory 110, a processor 120, and a communication unit 130. The memory 110, the processor 120, and the communication unit 130 are electrically connected to each other directly or indirectly to achieve data transmission or interaction. For example, these components can be electrically connected to each other via one or more communication buses or signal lines.

[0036] The memory 110 is used to store programs or data. The memory 110 may be, but is not limited to, a random access memory (RAM), a read only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), etc.

[0037] The processor 120 is configured to read / write data or programs stored in the memory 110 and perform corresponding functions. For example, a vector map construction device 200 based on autoregressive depth estimation is stored in the memory 110. The vector map construction device 200 based on autoregressive depth estimation includes at least one software function module that can be stored in the memory 110 in the form of software or firmware. The processor 120 executes various functional applications and data processing by running software programs and modules stored in the memory 110, such as the vector map construction device 200 in the embodiments of the present application, thereby implementing the vector map construction method based on autoregressive depth estimation in the embodiments of the present application.

[0038] The communication unit 130 is configured to establish a communication connection between the electronic device 100 and other communication terminals through a network and is configured to transmit and receive data through the network.

[0039] It should be understood that Figure 1 The structure shown is only a schematic diagram of the structure of the electronic device 100. The electronic device 100 may further include more or fewer components than those shown Figure 1 in the figure, or have a different configuration from that shown Figure 1 in the figure. Figure 1 Each component shown in the figure may be implemented by hardware, software, or a combination thereof.

[0040] Please refer to Figure 2 , Figure 2 which is one of the flow diagrams of the vector map construction method based on autoregressive depth estimation provided by the embodiments of the present application. The method can be applied to the above-mentioned electronic device. The specific process of the vector map construction method based on autoregressive depth estimation will be elaborated in detail below. In this embodiment, the method may include step S200 to step S400.

[0041] Step S200: Obtain a scene image obtained by image acquisition through a camera.

[0042] Step S300: Obtain a depth image corresponding to the scene image according to the scene image and a pre-trained autoregressive depth estimation model.

[0043] Step S400: Construct a vector map according to the scene image and the depth image.

[0044] In this embodiment, a scene image captured by a camera for constructing a map can be obtained first. The scene image can be an RGB image, that is, a color image. It can be understood that the above scene image can be an image obtained in real time to construct an online vector map; or multiple images can be obtained in advance and saved in an image library, and when a map needs to be constructed, an image is obtained from the image library as the scene to construct a vector map. Among them, the number of obtained scene images can be one or more, and can be specifically determined in combination with actual requirements. When there are multiple scene images, the multiple scene images can be images of the same scene obtained at different shooting angles.

[0045] For example, if there are 6 cameras installed on a vehicle, the images obtained by each of the 6 cameras can be used as the obtained scene images, and then an online vector map can be constructed based on the 6 scene images, so that the vehicle can perform environmental perception based on the constructed online vector map.

[0046] An autoregressive depth estimation model that has been pre-trained can be stored in the electronic device. The depth estimation model can be pre-trained by the electronic device itself, or can be trained by other devices and sent to the electronic device. The depth estimation model is trained based on multiple image pairs, and the parameters are adjusted based on the loss during the training process. Each image pair includes a corresponding sample scene image and a ground truth depth image. During the training based on an image pair, the loss for this time can be calculated in the following way: the loss is calculated based on depth prediction maps at different scales and the second target depth maps corresponding to each scale of the depth prediction maps. The scale of a depth prediction map and the corresponding second target depth map is the same. Among them, the second target depth map is a map obtained by enhancing the residual features of the first target depth map, and the first target depth map is a fixed-size map obtained by processing the ground truth depth image. The residual features are calculated based on the first target depth map and the obtained depth prediction map. During the training process of the depth estimation model, in order to make the prediction result at the current scale closer to the real value, the model is made to focus on learning the difference between the prediction result and the real value to better compensate for the prediction error. In this way, it is convenient to improve the prediction accuracy of the obtained depth estimation model.

[0047] After obtaining the scene image, the depth estimation model can be used to obtain the depth image corresponding to the scene image based on the obtained scene image. Among them, if there are multiple obtained scene images, there will also be multiple obtained depth images, and the number of both is the same. Moreover, the scales of the scene image and the corresponding depth image are the same.

[0048] After obtaining the scene image and the depth image, a vector map can be constructed by the set map construction method. The vector map can well describe the geometric shape and semantic information of map elements. Among them, the specific method for constructing the map based on the scene image and the depth image can be set according to actual needs, and no specific limitation is made here. In this way, in the case where the depth image is not directly obtained through a camera or the like, the corresponding depth image can be accurately predicted by using the above depth estimation model, and then a vector map can be constructed in combination with the scene image and the depth information.

[0049] Please refer to Figure 3 , Figure 3 FIG. 2 is a second flowchart of the method for constructing a vector map based on autoregressive depth estimation provided by an embodiment of the present application. In this embodiment, before step S300, the method may further include step S100.

[0050] Step S100, pre-train the depth estimation model.

[0051] The initial depth estimation model parameters can be set first, and a plurality of image pairs are collected, and then trained based on the plurality of image pairs to adjust the initial depth estimation model parameters until it is determined that the training is completed, that is, a trained depth estimation model is obtained.

[0052] As a possible implementation manner, the depth estimation model can be obtained by the Figure 4 shown manner. Please refer to Figure 4 , Figure 4 is Figure 3 a flowchart of the sub-steps included in step S100 in

[0053] Sub-step S110, for each of the plurality of obtained image pairs, obtain a first image sequence corresponding to the sample scene image in the image pair.

[0054] Sub-step S120, for each sample mapped image, perform depth prediction using the initial depth estimation model to obtain a depth prediction map of the corresponding scale.

[0055] Sub-step S130, perform image coding on the real depth image to obtain the first target depth map.

[0056] Sub-step S140, for each scale, according to the depth prediction maps smaller than the scale corresponding to the sample scene image, perform upsampling processing and weighted summation processing to obtain a depth map to be used with the same scale as the first target depth map.

[0057] Sub-step S150: Obtain residual features based on the first target depth map and the depth map to be used.

[0058] Sub-step S160: Obtain the second target depth map at this scale according to the first target depth map and the residual features.

[0059] Sub-step S170: Calculate the loss corresponding to this image pair according to the depth prediction maps at each scale and the corresponding second target depth maps.

[0060] Sub-step S180: Optimize the initial depth estimation model according to the calculated loss and continue training until the depth estimation model is obtained.

[0061] In this embodiment, for an image pair, the sample scene image in the image pair is used as the image for which the depth map is to be predicted. The sample scene image in the image pair can be processed to obtain an image sequence, and the image sequence is used as the first image sequence corresponding to the sample scene image. The first image sequence includes multiple sample mapping images with different scales, and the scale of one sample mapping image can be the same as the scale of the sample scene image. The sample mapping images in the first image sequence are multi-scale discretized representations, respectively representing image information at different scales (resolutions). Among them, the encoder can be used to process the above sample scene image to obtain the first image sequence. The specific encoder used can be determined according to actual needs. For example, a Vector Quantized Variational Auto Encoder (VQ-VAE) is used, and the encoder used is not specifically limited here.

[0062] The above first image sequence can be input into the initial depth estimation model, and multiple depth prediction maps are obtained. Among them, the scales corresponding to the multiple depth prediction maps are the same as the scale corresponding to the first image sequence.

[0063] Optionally, taking one scale as the current scale, the depth prediction map at the current scale can be predicted in the following way: According to the depth prediction maps smaller than the current scale obtained based on the first image sequence and the sample mapping images smaller than or equal to the current scale in the first image sequence, the depth prediction map at the current scale is predicted, that is: , where represents the depth prediction map at the current scale, represents an autoregressive depth prediction model, represents the depth prediction maps smaller than the current scale obtained based on the first image sequence, Represent the sample mapped images in the first image sequence that are less than or equal to the current scale. In this way, accurate prediction is facilitated.

[0064] Alternatively, the depth prediction map at the current scale can also be predicted in the following way: Based on the depth prediction map corresponding to the largest scale less than the current scale obtained from the first image sequence, and the sample mapped image in the first image sequence that is equal to the current scale, predict the depth prediction map at the current scale, that is: . In this way, rapid completion of prediction is facilitated.

[0065] It is also possible to perform image encoding on the ground truth depth image in the image pair to obtain a first target depth map with a fixed scale. The scale of the first target depth map can be greater than the largest scale in the scales corresponding to the first image sequence. The scale of the first target depth map can be specifically determined in combination with actual requirements. Among them, the encoder used to obtain the first image sequence can be utilized for processing, but discretization is not performed in this encoding, and it is directly encoded into continuous features.

[0066] For each scale, the second target depth map at this scale can be constructed according to the above-mentioned first target depth map and the depth prediction maps at scales lower than this scale that have been predicted. This process can be represented by the following formula:

[0067]

[0068] Among them, Represents the depth prediction maps of all scales (i.e., scales less than the current scale) of the initial depth estimation model before the current scale Through the upsampling operation (This operation includes upsampling processing and weighting processing) after mapping to the resolution of scale k of the first target image depth, and the cumulative sum is obtained by calculating them to obtain the depth map to be used; Represents the first target depth map; Represents the ground truth depth feature And the cumulative prediction The residual (i.e., the residual feature); Represents the target feature after enhancing the residual feature, Represents the weighting coefficient; Represents the current scale (resolution); Represents the scaling operation, adjusting the target feature To the current scale ; Represents the quantization operation, discretizing the continuous feature into a label map; Represents the second target depth map at the current scale n.

[0069] In this embodiment, in order to make the prediction result at the current scale closer to the true value, the initial depth estimation model is made to focus on learning the difference between the prediction result and the true value, so as to better compensate for the prediction error. Therefore, the weighted residual feature is summed with the true value, aiming to increase the training weight of the difference part and make the difference part be more emphasized in training at the current scale.

[0070] The above autoregressive depth estimation method can generate a high-precision depth map by modeling the dependency relationship between pixels. Depth estimation can provide the three-dimensional geometric information of the scene, helping the model for generating the map to better understand the shape and structure of the road environment. By fusing the depth feature with the image feature, a richer feature representation can be generated, thereby improving the detection and classification performance of map elements.

[0071] Next, calculate the losses of the depth prediction maps at each scale and the corresponding second target depth maps, so as to obtain the loss corresponding to the image pair. Furthermore, based on the loss corresponding to the image pair, the parameters of the initial depth estimation model can be optimized. After that, repeat the above process until it is determined that the training is completed, thus obtaining a trained autoregressive depth estimation model. Among them, the loss corresponding to an image pair can be obtained by calculating the cross-entropy loss, and this process can be expressed by the following formula: , where the sample mapped image in the first image sequence is N.

[0072] To facilitate the understanding of the above training process, the following gives a simple example of the above training process based on an image pair.

[0073] First, take the image to be predicted (i.e., a sample scene image) as the input, and obtain a series of token maps through the vector quantization variational autoencoder (VQ-VAE) , where is a multi-scale discretized representation, representing the image information at different scales (resolutions) respectively, and the image details included increase in turn. Specifically, is the token map at the lowest scale (resolution), containing the coarsest information in the sample scene image. As increases, the resolution of the token map gradually increases ( is used to represent any token map), and thus more detailed information is included. Among them, among the series of token maps, the scale of one token map is the same as the scale of the sample scene image, so as to facilitate predicting a depth prediction map with the same scale as the sample scene image.

[0074] The image token map As the starting sequence, this starting sequence is used as the input of the initial depth estimation model to serve as the condition for depth estimation. Then, the already predicted depth label map at a low scale (i.e., the depth prediction map) is also used as the input to autoregressively predict the depth label map at the current scale n. 。

[0075] Subsequently, the ground truth depth image corresponding to the aforementioned sample scene image is used as the input. Through a vector quantization variational autoencoder (VQ-VAE) that is exactly the same as the image encoding, but this time the encoding is not discretized and is directly encoded into a continuous feature of a fixed size. 。Next, based on the continuous feature and the already predicted depth label map at a low scale, a second target depth map at the current scale n is constructed. 。

[0076] Finally, the depth label maps at each scale corresponding to the image pair and the corresponding second target depth maps are calculated, and the cross-entropy loss corresponding to the image pair is calculated. The formula can be recorded as: 。

[0077] This model construction method uses dynamic targets (i.e., the prediction results of the model itself) during the training process and does not rely on predefined static targets (the quantization label maps provided by the VQ-VAE, i.e., the aforementioned continuous features ). This enables the model to self-correct during the training phase and there are multimodal solutions, thereby improving the accuracy of depth estimation.

[0078] As a possible implementation, as Figure 5 shown, the obtained scene images are multiple (i.e., multi-view images of the same scene). In this case, the depth images corresponding to each scene image can be obtained using the depth estimation model. Among them, for each scene image, the scene image can be processed first to obtain a second image sequence corresponding to the scene image, and then the second image sequence is input into the depth estimation model to obtain a depth image with the same scale as the scene image.

[0079] As a possible implementation, a vector map can be constructed as Figure 6 shown. Please refer to Figure 6 , Figure 6 For Figure 2 is a schematic flowchart of the sub-steps included in step S400. In this embodiment, step S400 may include sub-steps S410 to sub-step S430.

[0080] Sub-step S410, obtaining global bird's-eye view features according to the scene image and the depth image.

[0081] In this embodiment, the scene image can be projected into the BEV (Bird's Eye View) coordinate system in combination with the depth image, and then the global bird's-eye view feature can be obtained through processing. Alternatively, based on the above method, the scene image can be further processed by other methods, and the processing results can be fused with the processing results of the above method to obtain the global bird's-eye view feature. The following will Figure 5 illustrate by way of example how to obtain the global bird's-eye view feature through the latter method.

[0082] Suppose that at the t-th moment, a scene image is obtained from each of multiple cameras, and thus multiple scene images are obtained. The depth information (i.e., depth image) corresponding to the multiple scene images can be obtained by using the depth estimation model. Assume that there are 6 cameras on the vehicle, then the obtained depth information is where represents different cameras on the vehicle. The optical center of the internal parameters of the camera (i.e., the camera) can be obtained. focal length external parameter rotation matrix and translation vector .

[0083] Subsequently, coordinate transformation is performed to convert each pixel point in the image coordinate system of each camera (i.e., the pixel point in the scene image) into a three-dimensional point in the camera coordinate system . The specific transformation formula is as follows: , , .

[0084] After that, the three-dimensional point is transformed into the world coordinate system, the obtained world coordinates are projected into the BEV coordinate system, and finally the BEV feature of the depth projection is obtained through feature extraction and feature fusion, etc. . The BEV feature can be regarded as a feature map.

[0085] The panoramic multi-view images (i.e., the 6 scene images obtained by the above 6 cameras in total) are subjected to feature extraction through the backbone network Resnet-101 to obtain the features of different camera views , where is the feature of the i-th view; without using depth information, each view feature is projected into the BEV space, and the image BEV feature at the current moment is obtained through the spatial cross-attention mechanism and the feed-forward neural network . Among them, the view feature can be projected into the BEV space by BEVformer. The image BEV feature ​Can be regarded as a feature map.

[0086] Fuse the depth-projected BEV features through the attention mechanism and the image BEV features , and perform weighted fusion on the feature map according to the attention weights. Then, through max pooling and upsampling layers, the final global bird's-eye view feature is obtained .

[0087] Sub-step S420, according to the target bird's-eye view feature obtained from the global bird's-eye view feature, obtain the target decoding result through the map decoder.

[0088] In this embodiment, the target bird's-eye view feature can be first determined based on the global bird's-eye view feature , and then the pre-trained map decoder is used to perform multi-granularity queries based on the target bird's-eye view feature, so as to obtain the target decoding result. Among them, the target decoding result includes the corresponding target instance-level query result and the target point-level query result. The target instance-level query result is used to indicate the category of the predicted map element, and the target point-level query result is used to indicate the position of the final predicted point of interest. The position of the point of interest is the position in the target bird's-eye view feature. The multi-granularity query method can capture both the local geometric details and the overall category information of the map element, which is convenient for improving the accuracy of the generated vector map.

[0089] Optionally, as a possible implementation, the above global bird's-eye view feature can be directly used as the target bird's-eye view feature , which is convenient for quickly determining the target bird's-eye view feature.

[0090] Since the lengths of map elements are different, relying only on the BEV features of a single scale cannot meet the requirements for detecting elements of different lengths. Therefore, as another possible implementation, the target bird's-eye view feature can be obtained through the Figure 7 shown method. Please refer to Figure 7 . Figure 7 is Figure 6 One of the flow diagrams of the sub-steps included in sub-step S420. In this embodiment, sub-step S420 may include sub-steps S4211 to S4212.

[0091] Sub-step S4211, downsample the global bird's-eye view feature to obtain the downsampled global bird's-eye view feature.

[0092] Sub-step S4212, concatenate the global bird's-eye view feature and the downsampled global bird's-eye view feature using a flattened tensor to obtain the target bird's-eye view feature.

[0093] In this embodiment, a downsampling module can be used to downsample the global bird's-eye view feature The spatial resolution is reduced by half to obtain the downsampled global bird's-eye view feature Then, flat tensors are concatenated in the encoder and to obtain the multi-scale BEV feature as the target bird's-eye view feature Such a dual scale can capture both global features and local details simultaneously.

[0094] In this embodiment, the map decoder includes a plurality of decoding layers connected in sequence. The target bird's-eye view feature and the initial instance-level query result can be input into the map decoder, and the plurality of decoding layers are used to perform iterative queries to obtain the target decoding result. Among them, the initial instance-level query result includes the instances determined initially, specifically, it can be the type, etc. For two adjacent decoding layers, the output result of the previous decoding layer (i.e., the decoding result of the previous decoding layer) is optimized by the next decoding layer, and the output result of the last decoding layer is the target decoding result. Each decoding layer includes a multi-granularity attention layer.

[0095] As a possible implementation, as Figure 5 shown, the map decoder has L layers, and each decoding layer includes a self-attention layer, a multi-granularity attention layer, and a feed-forward neural network connected in sequence. That is, each decoding layer can be composed of a self-attention mechanism, a multi-granularity attention mechanism, and a feed-forward neural network. The self-attention layer, based on the self-attention mechanism, enables different instance-level queries to interact and captures the relationships between different instances, which helps the map construction model (which can include modules for obtaining the aforementioned global bird's-eye view feature, the map decoder, modules for generating maps based on the target decoding result, etc.) to better understand the overall structure of map elements. The multi-granularity attention layer is used to perform multi-granularity queries through the multi-granularity attention mechanism. The feed-forward neural network: performs further non-linear transformation on the current instance-level query result and the point-level query result (i.e., the output result of the multi-granularity attention layer in the current decoding layer) to obtain a more complex representation; at the same time, it further fuses and optimizes the internal connection between the instance-level query and the point-level query.

[0096] Instance-level queries can effectively capture the overall category information of road elements, but lack accurate geometric position representation; point-level queries can provide accurate geometric position information, but multiple queries need to be aggregated to represent one instance. In summary, in order to capture comprehensive and accurate instance features, in this embodiment, a multi-granularity attention mechanism is used. The multi-granularity attention mechanism consists of two components: a multi-granularity aggregator and point-instance interaction. That is, the multi-granularity attention layer includes a multi-granularity aggregator and a point-instance interaction component.

[0097] In the multi-granularity aggregator, the instance-level query result and the target bird's-eye view feature Interact to generate point-level query results. Specifically, introduce multiple reference points for each instance-level query to aggregate features at a long distance from the target bird's-eye view feature to aggregate information more widely. from the target bird's-eye view feature, enabling the map construction model to aggregate more extensive information from the target bird's-eye view feature

[0098] As a possible implementation, the output result of a decoding layer can be obtained in the manner shown Figure 8 . Please refer to Figure 8 , Figure 8 is Figure 6 a second flow chart of the sub-steps included in sub-step S420 in

[0099] Sub-step S422: Obtain the first reference point set of this decoding layer, and obtain the position encoding according to the first reference point set and the target bird's-eye view feature

[0100] In this embodiment, the multi-granularity aggregator in the first decoding layer (i.e., the first-layer decoding layer) takes the obtained initial instance-level query result as input. The instances in the initial instance-level query result can be pre-specified manually or obtained by other means. The multi-granularity aggregator in other decoding layers takes the point-level query result output by the previous decoding layer (subsequently referred to as the first point-level query result) and the reference point set used in the point-level query in the previous decoding layer (subsequently referred to as the second reference point set) as input. Among them, represents the total number of instance-level queries, is the total number of points belonging to an instance (the points of the instance can subsequently be referred to as points of interest). The reference points of the first decoding layer are predicted by , and the reference points of subsequent layers are updated by the reference points of the previous layer

[0101] That is, when this decoding layer is the first decoding layer, the first reference point set is determined according to the initial instance-level query result. When this decoding layer is not the first decoding layer, the first reference point set is determined according to the second reference point set used in the point-level query of the previous decoding layer and the first point-level query result output by the previous decoding layer

[0102] The determination method of the above first reference point set can be expressed by the following formula

[0103]

[0104] Among them, represents the current layer, represents the sigmoid activation function, represents the inverse sigmoid activation function, represents the set of reference points for the layer,

[0105] Since an instance is represented as a sequence of points, positional encoding is added to the instance-level query. In the case of obtaining the first set of reference points, i.e., given the position of the reference point positional encoding is generated using :

[0106]

[0107] where is a projection layer for generating positional embedding features from the reference points. The above positional encoding is also obtained by combining the target bird's-eye view features

[0108] Sub-step S423: Calculate the position offset corresponding to the first set of reference points according to the positional encoding and the first instance-level query result output by the previous decoding layer.

[0109] Sub-step S424: Calculate the sampling points corresponding to each first reference point in the first set of reference points according to the first set of reference points and the position offset.

[0110] In this embodiment, sampling points are assigned to each reference point, and these sampling points are used to aggregate features to enhance the features of the reference points. The position offset of each first reference point in the first set of reference points can be calculated according to the obtained positional encoding and the first instance-level query result output by the previous decoding layer (i.e., the instance-level query result input to this decoding layer) . The calculation process of the position offset of each first reference point in the multi-granularity attention layer of the layer decoding layer can be shown by the following formula: where is correspondingly expanded to match the

[0111]

[0112] shape. By using the position offset and the first reference point, i.e., through and , the sampling position and is obtained (i.e., the sampling points of the first reference point are obtained). ​​​

[0113] Sub-step S425: Obtain a second instance-level query result and a second point-level query result according to the sampling points of each first reference point and the target bird's-eye view feature.

[0114] After that, feature extraction and analysis can be performed based on the sampling points of each first reference point and the target bird's-eye view feature, so as to obtain a second instance-level query result and a second point-level query result. Optionally, the second instance-level query result and the second point-level query result can be obtained in the Figure 9 shown manner. Please refer to Figure 9 , Figure 9 is Figure 8 a schematic flowchart of the sub-steps included in sub-step S425 in

[0115] Sub-step S4251: Obtain the first weight corresponding to each sampling point according to the position encoding and the first instance-level query result output by the previous decoding layer.

[0116] Sub-step S4252: Calculate the second weight corresponding to each instance of the sampling point and the third weight corresponding to each point of interest of the sampling point according to the first weight corresponding to each sampling point.

[0117] Sub-step S4253: Obtain the second instance-level query result according to the target bird's-eye view feature, each sampling point corresponding to each instance, and the second weight corresponding to each instance.

[0118] Sub-step S4254: Obtain the third point-level query result according to the target bird's-eye view feature, each sampling point corresponding to each point of interest, and the third weight corresponding to each point of interest.

[0119] In this embodiment, after obtaining the position encoding , it is also possible to analyze and obtain the weight corresponding to each sampling point as the first weight according to the position encoding and the first instance-level query result output by the previous decoding layer . This process can be expressed by the following formula: .

[0120] Then, the second instance-level query result and the third point-level query result can be generated through the weighted sum of the sampling features, that is, generate the instance-level query result and the point-level query result first generated at this layer. This process can be expressed by the following formula:

[0121]

[0122] where is on the instance Index of a point, is the index between the sampling points assigned to the reference point, indicating on the normalized weight (i.e., the second weight corresponding to the instance), indicating on the normalized weight (i.e., the third weight corresponding to the point of interest), is the bilinear sampling operator.

[0123] Sub-step S426, according to the second instance-level query result and the second point-level query result, through point-instance interaction, obtain the third instance-level query result and the third point-level query result.

[0124] In this embodiment, the second instance-level query result and the second point-level query result can be input into the point-instance interaction component to process and obtain the third instance-level query result and the third point-level query result. The purpose of point-instance interaction is to enhance the interaction of position information and category information between two different granularity queries. Point-instance interaction includes two different attention operators: P2P (point-to-point) attention and P2I (point-to-instance) attention. Through the above two operators, the third instance-level query result and the third point-level query result are processed. The output result of this decoding layer is determined based on the third instance-level query result and the third point-level query result. For example, in Figure 5 In the shown structure, a decoding layer further includes a feed-forward neural network after the multi-granularity attention layer in this decoding layer. Then, the multi-granularity attention layer outputs the third instance-level query result and the third point-level query result to the feed-forward neural network, and the output of the feed-forward neural network based on the third instance-level query result and the third point-level query result is the output result of this decoding layer.

[0125] Optionally, the third instance-level query result and the third point-level query result can be obtained in the Figure 10 shown manner. Please refer to Figure 10 , Figure 10 is Figure 8 the schematic flow diagram of the sub-steps included in sub-step S426 in

[0126] Sub-step S4261, according to the obtained sampling points, the second weights corresponding to each instance, and the target bird's-eye view feature, obtain the first position information feature corresponding to the instance.

[0127] Sub-step S4262: Obtain the second position information feature corresponding to the interest point according to the obtained sampling points, the third weight corresponding to each interest point, and the target bird's-eye view feature.

[0128] Sub-step S4263: According to the second point-level query result and the first position information feature, calculate the initial third point-level query result through the P2P attention mechanism. When this decoding layer is not the first decoding layer, the initial third point-level query result is obtained based on the second point-level query result, the first position information feature, the first point-level query result, and the position information feature corresponding to the first point-level query result.

[0129] Sub-step S4264: According to the initial third point-level query result, the second position information feature, the second instance query result, and the first position information feature, obtain the third point-level query result through the P2I attention mechanism.

[0130] Sub-step S4265: According to the third point-level query result, aggregate the interest points corresponding to the same instance in the third point-level query result to obtain the third instance-level query result.

[0131] In this embodiment, the sampling positions (i.e., sampling points) obtained by aggregating multiple granularities in the th decoding layer and the attention weights (i.e., the second weight corresponding to the instance), (i.e., the third weight corresponding to the interest point) are tiled and cascaded to encode the position information features in the P2P attention mechanism and the P2I attention mechanism:

[0132]

[0133] Among them, is the multi-layer perceptron for instance-level query, is the multi-layer perceptron for point-level query. and are the corresponding generated position embeddings. represents the first position information feature corresponding to the instance, represents the second position information feature corresponding to the interest point, and the above position information features are also obtained by combining the target bird's-eye view feature

[0134] In the P2P attention mechanism, since the coordinates of the map elements are refined based on the point-level queries in the multi-granularity attention layer of the previous layer (the th layer), these point-level queries play a key role in predicting the coordinates of the current layer (the th layer). The P2P attention module is designed to combine the information from the current The point-level queries of the current layer and the previous layer are used as the input of the P2P attention module. The processing based on the P2P attention mechanism is as follows:

[0135]

[0136] Since there is no previous decoding layer in the multi-granularity attention layer of the first decoding layer, self-attention operation is performed in the first multi-granularity attention layer. In subsequent multi-granularity attention layers, the previous point-level query results are mixed with the currently generated point-level query results to perform cross-attention operation. represents the initial third point-level query result.

[0137] That is, when this decoding layer is the first decoding layer, the initial third point-level query result is obtained through self-attention operation based on the second point-level query result and the first position information feature. When this decoding layer is not the first decoding layer, the initial third point-level query result is obtained through cross-attention operation based on the second point-level query result, the first position information feature, the first point-level query result and the first point-level query result.

[0138] In the P2I attention mechanism, following the P2P attention mechanism, the P2I attention operation realizes information interaction between different granularities. The point-level query exchanges geometric information with the instance-level query using cross-attention: , represents the third point-level query result.

[0139] Finally, the third point-level query results belonging to the same instance-level query are aggregated to update the corresponding instance-level query, so as to obtain the third instance-level query result of this decoding layer. This process can be as follows:

[0140]

[0141] where, represents the index of.

[0142] The point-granularity query is used to predict the point position using the MLP as the regression head, while the instance-granularity query is used to predict the category of the map element using another MLP. By utilizing the multi-granularity aggregator and the point-instance interaction component, multi-granularity queries can be generated and updated. At the same time, the geometric shapes and categories of each map element can be effectively perceived.

[0143] The above processing is performed for each decoding layer. When the output of the last decoding layer is obtained, the target decoding result is obtained.

[0144] Sub-step S430, obtaining the vector map according to the target decoding result.

[0145] After obtaining the target decoding result, a vector map can be generated based on the target decoding result in any manner.

[0146] The above method provided by the embodiments of the present application is a multi-granularity vector map construction method based on autoregressive depth estimation. By constructing an autoregressive depth model, this method can effectively improve the accuracy of depth estimation information; moreover, by using multi-granularity queries, it can capture global features and local information simultaneously, which is beneficial for more comprehensive identification of map information; it also improves the geometric position information of the global bird's-eye view feature by fusing depth information, and then constructs a more accurate vector map.

[0147] To execute the corresponding steps in the above embodiments and various possible manners, the following presents an implementation manner of a vector map construction device 200 based on autoregressive depth estimation. Optionally, the vector map construction device 200 based on autoregressive depth estimation may adopt the Figure 1 device structure of the electronic device 100 shown above. Further, please refer to Figure 11 , Figure 11 which is one of the block diagrams of the vector map construction device 200 based on autoregressive depth estimation provided by the embodiments of the present application. It should be noted that for the vector map construction device 200 based on autoregressive depth estimation provided in this embodiment, its basic principle and the technical effects produced are the same as those in the above embodiments. For the sake of brief description, for the parts not mentioned in this embodiment, reference can be made to the corresponding content in the above embodiments. In this embodiment, the vector map construction device 200 based on autoregressive depth estimation may include: an image acquisition module 220, a depth estimation module 230, and a construction module 240.

[0148] The image acquisition module 220 is configured to acquire a scene image obtained by image acquisition through a camera.

[0149] The depth estimation module 230 is configured to obtain a depth image corresponding to the scene image according to the scene image and a pre-trained autoregressive depth estimation model. Wherein, the loss used in the training process based on image pairs of the depth estimation model is calculated based on depth prediction maps of different scales and corresponding second target depth maps. The image pairs include corresponding sample scene images and real depth images. The second target depth map is a map obtained by enhancing the residual features for the first target depth map. The first target depth map is a fixed-size map obtained by processing the real depth image, and the residual features are calculated based on the first target depth map and the obtained depth prediction map.

[0150] The construction module 240 is configured to construct a vector map according to the scene image and the depth image.

[0151] Please refer to Figure 12 , Figure 12 The second block diagram of the vector map construction device 200 based on autoregressive depth estimation provided in the embodiment of the present application. In this embodiment, the vector map construction device 200 based on autoregressive depth estimation may further include a training module 210. The training module 210 is used to pre-train the depth estimation model.

[0152] Optionally, the above modules can be stored in the form of software or firmware. Figure 1 The memory 110 shown in the figure may be fixed in the operating system (OS) of the electronic device 100 and may be Figure 1 Meanwhile, the data and program codes required for executing the above modules may be stored in the memory 110.

[0153] An embodiment of the present application also provides a readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the vector map construction method based on autoregressive depth estimation is implemented.

[0154] In summary, the embodiment of the present application provides a method and related device for constructing a vector map based on autoregressive depth estimation. First, a scene image is obtained by acquiring an image through a camera; then, based on the scene image and a pre-trained autoregressive depth estimation model, a depth image corresponding to the scene image is obtained. The loss used in the depth estimation model in the image pair training process is calculated based on depth prediction maps of different scales and corresponding second target depth maps. The image pair includes a corresponding sample scene image and a true depth image. The second target depth map is a map obtained by enhancing the residual features of the first target depth map. The first target depth map is a fixed-size map obtained by processing the true depth image. The residual features are calculated based on the first target depth map and the obtained depth prediction map; finally, a vector map is constructed based on the scene image and the depth image. In this way, the corresponding depth information can be accurately estimated based on the scene image taken by the camera, and then a vector map can be constructed in combination with the scene image and the depth information.

[0155] In several embodiments provided by this application, it should be understood that the disclosed devices and methods can also be implemented in other ways. The device embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the accompanying drawings show the possible architectures, functions, and operations of devices, methods, and computer program products according to multiple embodiments of this application. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code, and the module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks can actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, as well as the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.

[0156] In addition, in each embodiment of this application, the various functional modules can be integrated together to form an independent part, or each module can exist alone, or two or more modules can be integrated to form an independent part.

[0157] If the above functions are implemented in the form of software functional modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of this application. The aforementioned storage medium includes various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs.

[0158] The above are only optional embodiments of this application and are not used to limit this application. For those skilled in the art, this application can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of this application shall be included in the protection scope of this application.

Claims

1. A method for constructing a vector map based on autoregressive depth estimation, characterized in that The method includes: Obtaining a scene image acquired by image capture through a camera; According to the scene image and a pre-trained autoregressive depth estimation model, obtaining a depth image corresponding to the scene image, wherein the loss used in the training process based on image pairs of the depth estimation model is calculated based on depth prediction maps of different scales and corresponding second target depth maps. An image pair includes a corresponding sample scene image and a ground truth depth image. The second target depth map is a map obtained by enhancing residual features for a first target depth map. The first target depth map is a fixed-size map obtained by processing the ground truth depth image. The residual features are calculated based on the first target depth map and the obtained depth prediction maps. Wherein, the depth estimation model is trained in the following manner: for each of the multiple obtained image pairs, obtaining a first image sequence corresponding to the sample scene image in the image pair, wherein the first image sequence includes multiple sample mapped images with different scales; for each sample mapped image, using an initial depth estimation model to perform depth prediction to obtain a depth prediction map corresponding to the scale; performing image encoding on the ground truth depth image to obtain the first target depth map; for each scale, according to each depth prediction map corresponding to the sample scene image that is smaller than the scale, through upsampling processing and weighted summation processing, obtaining a depth map to be used with the same scale as the first target depth map; according to the first target depth map and the depth map to be used, obtaining residual features; according to the first target depth map and the residual features, obtaining a second target depth map at this scale; according to the depth prediction maps of each scale and the corresponding second target depth maps, calculating the loss corresponding to the image pair; optimizing the initial depth estimation model according to the calculated loss and continuing the training until the depth estimation model is obtained; According to the scene image and the depth image, constructing a vector map.

2. The method according to claim 1, characterized in that, The method further includes: Pre-training to obtain the depth estimation model.

3. The method according to claim 1 or 2, characterized in that, The constructing a vector map according to the scene image and the depth image includes: Obtaining a global bird's-eye view feature according to the scene image and the depth image; According to a target bird's-eye view feature obtained from the global bird's-eye view feature, obtaining a target decoding result through a map decoder, wherein the target decoding result includes a corresponding target instance-level query result and a target point-level query result; Obtaining the vector map according to the target decoding result.

4. The method according to claim 3, wherein The map decoder includes a plurality of decoding layers connected in sequence. The obtaining a target decoding result through the map decoder according to a target bird's-eye view feature obtained from the global bird's-eye view feature includes: According to the target bird's-eye view feature and an initial instance-level query result, using the plurality of decoding layers to perform iterative queries to obtain the target decoding result, wherein for two adjacent decoding layers, the output result of the previous decoding layer is optimized by the next decoding layer, and the output result of the last decoding layer is the target decoding result. Each decoding layer includes a multi-granularity attention layer.

5. The method according to claim 4, characterized in that Performing iterative queries using the multiple decoding layers according to the target bird's-eye view feature and the initial instance-level query result to obtain the target decoding result, including: Obtaining a first set of reference points for the current decoding layer, and obtaining position encoding according to the first set of reference points and the target bird's-eye view feature, where when the current decoding layer is the first decoding layer, the first set of reference points is determined according to the initial instance-level query result; when the current decoding layer is not the first decoding layer, the first set of reference points is determined according to the second set of reference points used in the point-level query of the previous decoding layer and the first point-level query result output by the previous decoding layer; Calculating a position offset corresponding to the first set of reference points according to the position encoding and the first instance-level query result output by the previous decoding layer; Calculating sampling points corresponding to each first reference point in the first set of reference points according to the first set of reference points and the position offset; Obtaining a second instance-level query result and a second point-level query result according to the sampling points of each first reference point and the target bird's-eye view feature; Obtaining a third instance-level query result and a third point-level query result through point-instance interaction according to the second instance-level query result and the second point-level query result, where the output result of the current decoding layer is determined based on the third instance-level query result and the third point-level query result.

6. The method according to claim 5, wherein The obtaining the second instance-level query result and the second point-level query result according to the sampling points of each first reference point and the target bird's-eye view feature includes: Obtaining a first weight corresponding to each sampling point according to the position encoding and the first instance-level query result output by the previous decoding layer; Calculating a second weight for each instance corresponding to the sampling points and a third weight for each interest point corresponding to the sampling points according to the first weights corresponding to the sampling points; Obtaining the second instance-level query result according to the target bird's-eye view feature, each sampling point corresponding to each instance, and the second weight corresponding to each instance; Obtaining the third point-level query result according to the target bird's-eye view feature, each sampling point corresponding to each interest point, and the third weight corresponding to each interest point.

7. The method according to claim 6, characterized in that, The obtaining the third instance-level query result and the third point-level query result through point-instance interaction according to the second instance-level query result and the second point-level query result includes: Obtaining a first position information feature corresponding to an instance according to the obtained sampling points, the second weight corresponding to each instance, and the target bird's-eye view feature; Obtaining a second position information feature corresponding to an interest point according to the obtained sampling points, the third weight corresponding to each interest point, and the target bird's-eye view feature; Calculating an initial third point-level query result through a P2P attention mechanism according to the second point-level query result and the first position information feature, where when the current decoding layer is not the first decoding layer, the initial third point-level query result is obtained based on the second point-level query result, the first position information feature, the first point-level query result, and the position information feature corresponding to the first point-level query result; Based on the initial third-level query result, the second position information feature, the second instance query result, and the first position information feature, the third-level query result is obtained through the P2I attention mechanism; Based on the third-level query result, by aggregating the points of interest corresponding to the same instance in the third-level query result, the third instance-level query result is obtained.

8. The method according to claim 3, wherein The target bird's-eye view feature is obtained in the following manner: Downsample the global bird's-eye view feature to obtain the downsampled global bird's-eye view feature; Use a flattened tensor to concatenate the global bird's-eye view feature and the downsampled global bird's-eye view feature to obtain the target bird's-eye view feature.

9. A vector map construction device based on autoregressive depth estimation, characterized in that, The device includes: An image acquisition module for acquiring a scene image obtained by image acquisition through a camera; A depth estimation module for obtaining a depth image corresponding to the scene image according to the scene image and a pre-trained autoregressive depth estimation model, wherein the loss used in the training process based on image pairs is calculated based on depth prediction maps of different scales and corresponding second target depth maps. An image pair includes a corresponding sample scene image and a real depth image. The second target depth map is a map obtained by enhancing the residual feature for the first target depth map. The first target depth map is a fixed-size map obtained by processing the real depth image. The residual feature is calculated based on the first target depth map and the obtained depth prediction map; wherein, the depth estimation model is trained in the following manner: for each of the obtained multiple image pairs, obtain a first image sequence corresponding to the sample scene image in the image pair, wherein the first image sequence includes multiple sample mapping images with different scales; for each sample mapping image, use an initial depth estimation model to perform depth prediction to obtain a depth prediction map corresponding to the scale; perform image encoding on the real depth image to obtain the first target depth map; for each scale, according to the depth prediction maps corresponding to the sample scene image smaller than this scale, through upsampling processing and weighted summation processing, obtain a depth map to be used with the same scale as the first target depth map; according to the first target depth map and the depth map to be used, obtain the residual feature; according to the first target depth map and the residual feature, obtain the second target depth map at this scale; according to the depth prediction maps of each scale and the corresponding second target depth maps, calculate the loss corresponding to the image pair; optimize the initial depth estimation model according to the calculated loss, and continue training until the depth estimation model is obtained; A construction module for constructing a vector map according to the scene image and the depth image.

10. An electronic device, characterized in that, It includes a processor and a memory. The memory stores machine-executable instructions that can be executed by the processor. The processor can execute the machine-executable instructions to implement the vector map construction method based on autoregressive depth estimation according to any one of claims 1-8.

Citation Information

Patent Citations

  • Aerial image geographic positioning method based on spatial scale attention mechanism and vector map

    CN113239952A

  • Image-laser radar data fusion method based on mixed attention mechanism

    CN114398937A