Multi-view bev detector enhanced by using ground real target information and construction method
By introducing TT-bev and TT-Q modules into bevformer, integrating ground real information, the encoder dependence on decoder and lack of interpretability in bev model is solved, and a more accurate and robust bev detection effect is achieved.
Patent Information
- Application Number
- CN202510009784.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-03
- Publication Date
- 2025-05-06
AI Technical Summary
In the existing bev model, the encoding effect of the encoder depends on the output of the decoder, lacks interpretability, and the decoding process is similar to a black box.
In bevformer's algorithm framework, the TT-bev module and the TT-Q module are added to generate ground real information TT-bev and query TT-Q, which is used to enhance the capabilities of bev encoder and bev decoder, so that the ground real information interacts with the bev feature, and increase the interpretability of decoding.
By integrating ground real information, the detection accuracy and robustness of the bev detector are enhanced and the interpretability of the model is improved.
Smart Images

Figure CN119942495A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present invention relate to the field of intelligent driving technology, and in particular to a multi-view BEV detector enhanced by ground-truth target information and a construction method thereof. Background Art
[0002] Automated driving systems (ADS) have witnessed rapid advancements, revolutionizing the transportation landscape. Current state-of-the-art vision-based ADS approaches heavily rely on bird’s-eye view (BEV) representations extracted from multi-view images to perceive and understand the surrounding environment. BEV detectors, serving as the backbone of ADS, convert multi-view images into top-down view representations, called BEV feature maps. The effectiveness of BEV detectors significantly impacts the success of perception tasks in automated driving, emphasizing their critical role in enhancing the system’s situational awareness.
[0003] However, in current BEV models such as BevFormer, the encoding effect of the encoder can only be judged by the output of the decoder, which makes the encoder overly dependent on the decoder's capabilities. In addition, the decoder uses the initialized query to decode, and the decoding result is aligned with the ground truth information. The decoding process is similar to a black box and lacks interpretability. Summary of the invention
[0004] In view of this, an embodiment of the present invention provides a multi-view bev detector enhanced with ground truth target information. On the basis of the original bevformer algorithm framework (bev encoder and bev decoder), a TT-bev module and a TT-Q module are added. The former generates ground truth information TT-bev to align the bev feature generated by the bev encoder, and the latter generates a query TT-Q of the ground truth information and adds it to the query of the bev decoder, so that the ground truth information in the bev feature interacts with the ground truth information in the TT-Q, thereby increasing the interpretability of the bev decoder.
[0005] According to a first aspect of an embodiment of the present invention, a method for constructing a multi-view bev detector enhanced with real target information is provided, including: S1, building a bevformer architecture of the bev detector, using the backbone network backbone of the architecture to extract features of six panoramic images, and inputting the features of the six panoramic images into the bev encoder bev encoder of the architecture to generate bev features bev feature; S2, inputting the ground-truth target information into the target information encoder TTencoder for encoding to obtain target information enhancement features TT-bev and target information query features TT-Q, the target information enhancement features TT-bev and the target information query features TT-Q are respectively used to enhance the encoding and decoding capabilities of the bevformer architecture; S3, using the bev feature bev feature and the target information enhancement features TT-bev for comparative learning to obtain the ground-truth target supervision loss loss TT-bev ; S4, input the randomly initialized bev query bev query, bev feature bev feature and target information query feature TT-Q into the bev decoder bev decoder of the architecture to obtain the baseline bev detector loss loss normal and the perceptual loss of ground truth object query TT-Q ; S5. Use ground truth target supervision loss loss TT-bev Update the network parameters of the bev encoder bev encoder, using the baseline bev detector loss loss normal Update the network parameters of the bev encoder bevencoder and bev decoder bev decoder, using the perceptual loss loss of the ground truth target query TT-Q Update the network parameters of the bev decoder; S6, optimize the bev detector through the updated network parameters.
[0006] In one implementation, in step S1, the features of the six panoramic images are input into the bev encoder of the architecture, and the bev encoder is used to merge the image features of the six panoramic images into a unified top-down view bev feature map, that is, The expression of bev encoder is:
[0007] z bev =BEVEnc(X views )
[0008] Among them, X views is a multi-view camera image.
[0009] In another implementation, the target information encoder TT encoder in step S2 is composed of a simple multi-layer perceptron MLP, which aligns the generated target information enhancement feature TT-bev with the bev feature bev feature, ensures that the bev elements are clearly arranged according to their class labels, positions and boundaries, and takes the class label of the i-th instance in the bev map as t i , the ground truth bounding box is b i , the expression of the target information encoder TT encoder is:
[0010] α i =TTEnc(t i ,b i )
[0011] Among them, α i It has the same feature dimension C as the bev feature and represents the encoded true information feature of the i-th instance on the bev map.
[0012] In another implementation, the step S3 uses the generated z bev and the generated α i Contrastive learning is performed. Prior to this, for each instance, in order to enhance its clarity on the bev map, the tensor within its true information bounding box is cropped from the bev feature map, and the cropped tensor is pooled as an object γ i The expression is:
[0013] γ i =Pool(Crop(z bev ,b i ))
[0014] Among them, b i represents the ground truth bounding box.
[0015] In another implementation, in order to make the target information enhancement feature TT bev and the bev feature bev feature embedding closer in step S3, contrastive learning is used to optimize the element relationship and distance in the bev feature space, expression:
[0016]
[0017]
[0018] where N and I represent the generated and target similarity matrices between the object bev features and the object real information features, respectively, μ is the logit scale learned during contrastive learning, represents matrix multiplication, L CEis the cross entropy loss for similarity matrix optimization and is the final loss produced from the TT bev module.
[0019] In another implementation, the step S4 introduces the ground truth information query α i To the query pool of bev decoder bevdecoder for decoding, after decoding, the same head Head and perceptual loss L per Applied to the processed ground truth information query, the above process can be expressed as:
[0020] Q tt =α,Q t ' t = Dec(Q tt ),
[0021]
[0022]
[0023] Among them, Q tt represents the initial ground truth query, Q t ' t and denote the initial ground truth information query after processing by the perceptual decoder and the perceptual head, respectively, and loss TT-Q is the final loss for ground truth information query.
[0024] In another implementation, the training formula in step S5 includes three key components: baseline bev detector loss loss normal , the perceptual loss for ground truth object query TT-Q and ground truth target supervision loss loss TT -bev , the expression is as follows:
[0025] L = loss normal +loss TT-Q +loss TT-bev
[0026] Among them, loss normal is the baseline bev detector loss, loss TT-Q is the perceptual loss for ground truth target query, loss TT-bev Supervised loss for ground truth targets.
[0027] According to a second aspect of an embodiment of the present invention, a multi-view bev detector enhanced with real target information is provided, including: a bev encoder module, which merges image features into the same top-down bev feature map, fuses historical bev features into a bev query bev query through temporal self-attention, and then uses the features of the bev query bev query and six panoramic images for cross-attention, performs residual connection, regularization and full connection layer on the obtained results, and finally obtains a bev feature bev feature; a TT-bev module, which obtains a target information enhanced feature TT-bev from the ground truth target information through an encoder composed of a target information encoder TTEncoder layer, the bev encoder bev encoder generates a bev feature bev feature, regards the target information enhanced feature TT-bev as the result of a text compiler, regards the bev feature bevfeature as the result of an image compiler, uses a contrastive learning method to enhance the perception result of the target information enhanced feature TT-bev and the bev feature bev feature, and finally obtains the ground truth target supervision loss loss TT-bev ; TT-Q module, which passes the ground truth target information through an encoder composed of target information encoder TT Encoder layers to obtain the target information query feature TT-Q; bev decoder module, which performs attention during decoding in a parallel manner to prevent information leakage, and finally obtains the baseline bev detector loss loss normal and the perceptual loss of ground truth object query TT-Q .
[0028] According to a third aspect of an embodiment of the present invention, there is provided a computer storage medium on which a computer program is stored. When the program is executed by a processor, the method of the first aspect described above is implemented.
[0029] Compared with the prior art, the present invention has the following beneficial effects:
[0030] (1) The present invention integrates ground truth information into the BEV encoding and decoding process, enhances the BEV detector, and improves the accuracy and robustness of the BEV detector.
[0031] (2) The additional code layer structure TT encoder is a simple MLP layer that is only used during training and does not incur additional time during inference. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in the embodiments of the present invention. For ordinary technicians in this field, other drawings can also be obtained based on these drawings.
[0033] Figure 1 A flowchart of the steps of a method for constructing a multi-view BEV detector enhanced by real target information according to the present invention;
[0034] Figure 2 A code layer structure diagram of a multi-view BEV detector enhanced with real target information according to the present invention;
[0035] Figure 3 A logic diagram for comparative learning of the bev feature TT-bev directly encoded by the ground truth information in the present invention and the bev feature bev feature generated by the bev encoder;
[0036] Figure 4 This is a schematic diagram of the relationship between the main code modules of a multi-view BEV detector enhanced using real target information of the present invention. DETAILED DESCRIPTION
[0037] In order to have a clearer understanding of the technical features, purposes and effects of the embodiments of the present invention, the specific implementation of the embodiments of the present invention is now described with reference to the accompanying drawings.
[0038] In this document, “exemplary” means “serving as an example, instance or illustration”, and any illustration or implementation described in this document as “exemplary” should not be construed as a more preferred or more advantageous technical solution.
[0039] In order to enable those skilled in the art to better understand the technical solutions in the embodiments of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in the field based on the embodiments in the embodiments of the present invention should fall within the scope of protection of the embodiments of the present invention.
[0040] The specific implementation of the embodiment of the present invention is further described below in conjunction with the accompanying drawings of the embodiment of the present invention.
[0041] See also Figure 1 , Figure 2 The present invention provides a method for constructing a multi-view BEV detector using real target information enhancement, comprising:
[0042] S1. Build the bevformer architecture of the bev detector, use the backbone network backbone of the architecture to extract the features of six panoramic pictures, and input the features of the six panoramic pictures into the bev encoder of the architecture to generate bev features;
[0043] S2. Input the ground truth target information into the target information encoder TT encoder for encoding, and obtain the target information enhancement feature TT-bev and the target information query feature TT-Q. The target information enhancement feature TT-bev and the target information query feature TT-Q are used to enhance the encoding and decoding capabilities of the bevformer architecture respectively;
[0044] S3. Use bev feature bev feature and target information enhancement feature TT-bev for comparative learning to obtain the ground truth target supervision loss loss TT-bev ;
[0045] S4. Use bev decoder to input the randomly initialized bev query, bev feature and target information query feature TT-Q into the bev decoder of this architecture to obtain the baseline bev detector loss loss normal and the perceptual loss of ground truth object query TT-Q ;
[0046] S5, back propagation, update network parameters. Use ground truth target supervision loss loss TT-bev Update the network parameters of the bev encoder bev encoder, using the baseline bev detector loss loss normal Update the network parameters of the bev encoder and bev decoder, using the perceptual loss loss of the ground truth target query TT-Q Update the network parameters of bev decoder;
[0047] S6. Optimize the bev detector using the updated network parameters.
[0048] Optionally, in step S1, the features of the six panoramic pictures are input into the bev encoder bev encoder of the architecture, and the bev encoder bev encoder is used to merge the image features of the six panoramic pictures into a unified top-down view bev feature map, that is, The expression of bev encoder is:
[0049] z bev =BEVEnc(X views )
[0050] Among them, X views is a multi-view camera image.
[0051] Furthermore, the target information encoder TT encoder in step S2 is composed of a simple multi-layer perceptron MLP, which aligns the generated target information enhancement feature TT-bev with the bev feature bev feature, ensures that the bev elements are clearly arranged according to their class labels, positions and boundaries, and takes the class label of the i-th instance in the bev map as t i , the ground truth bounding box is b i , the expression of the target information encoder TT encoder is:
[0052] α i =TTEnc(t i ,b i )
[0053] Among them, α i It has the same feature dimension C as the bev feature and represents the encoded true information feature of the i-th instance on the bev map.
[0054] like Figure 3 The logic diagram of contrastive learning shown in the figure is based on the z generated in step S3. bev and the generated α i Contrastive learning is based on the bev feature (z bev ) and TT-bev(α i ) for contrastive learning. Prior to this, for each instance, in order to enhance its clarity on the bev map, the tensor within its true information bounding box is cropped from the bev feature map, and the cropped tensor is pooled as an object γ i The expression is:
[0055] γ i =Pool(Crop(z bev ,b i ))
[0056] Among them, b i represents the ground truth bounding box.
[0057] Specifically, in order to make the target information enhancement feature TT bev and bev feature bev feature embedding closer, contrastive learning is used to optimize the element relationship and distance in the bev feature space. The expression is:
[0058]
[0059]
[0060] where N and I represent the generated and target similarity matrices between the object bev features and the object real information features, respectively, μ is the logit scale learned during contrastive learning, represents matrix multiplication, L CE is the cross entropy loss for similarity matrix optimization and is the final loss produced from the TT bev module.
[0061] Furthermore, in order to make up for the problem that the query bev query of bev decoder is randomly initialized and lacks real information, the ground truth information query α is introduced in step S4 i To the query pool of the bev decoder, it undergoes the same process and modules as a normal query. This simulates the real information communication in the actual scene through the self-attention mechanism in the ground truth information query, and communicates the global environment through the cross attention with the bev map. The ground truth query information is further processed using a feedforward neural network. Although the processes and modules of TT-Q and Q query bev feature are shared, the attention during decoding is performed in parallel to prevent information leakage. And TT-Q is a ground truth information query for a specific instance, so the query of TT-Q does not need to be matched.
[0062] Among them, after decoding, the same head Head and perceptual loss L per Applied to the processed ground truth information query, the above process can be expressed as:
[0063] Q tt =α,Q t ' t = Dec(Q tt ),
[0064]
[0065]
[0066] Among them, Q tt represents the initial ground truth query, Q t ' t and denote the initial ground truth information query after processing by the perceptual decoder and the perceptual head, respectively, and loss TT-Q is the final loss for ground truth information query.
[0067] Specifically, by injecting ground truth information in the decoding stage of training, the ground truth information in the bev feature interacts with the ground truth information in TT-Q, which not only enhances the robustness of the model but also enhances the ability to detect various objects in the bev graph.
[0068] Furthermore, the training formula of the present invention contains three key components, each of which is designed to optimize a specific aspect of model performance: baseline bev detector loss loss normal , the perceptual loss for ground truth object query TT-Q , ground truth target supervision loss loss TT-bev , the expression is as follows:
[0069] L = loss normal +loss TT-Q +loss TT-bev
[0070] Among them, loss normal is the baseline bev detector loss, loss TT-Q is the perceptual loss for ground truth target query, loss TT-bev Supervised loss for ground truth targets.
[0071] The ground truth target information is only used to update the network parameters through TT-Q and TT-bev during the training phase, and no additional parameters or calculations are introduced during the inference process. This ensures that the efficiency of the original model is maintained during the inference phase.
[0072] See also Figure 4 The embodiment of the present invention further provides a multi-view BEV detector enhanced by real target information, including:
[0073] The bev encoder module merges the image features into the same top-down bev feature map, fuses the historical bev features into the bev query through temporal self-attention, and then uses the bev query and the features of the six panoramic images for cross-attention. The result is subjected to residual connection, regularization and full connection layer to finally obtain the bev feature;
[0074] The TT-bev module passes the ground truth target information through an encoder composed of TT Encoder layers to obtain TT-bev features, and the bev encoder generates the original bev features. The TT-bev features are regarded as the results of the text compiler, and the original bev features are regarded as the results of the image compiler. The contrastive learning method is used to enhance the perception results of TT-bev and original bev features, and finally the loss loss is obtained. TT-bev .
[0075] The TT-Q module is generated by the TT Encoder part, and then goes through the same process and modules as the normal query, and finally obtains the query TT-Q;
[0076] bev decoder module, TT-Q and Q query bev feature process and module are shared, but the attention during decoding is performed in parallel to prevent information leakage, and the final loss is normal and loss normal .
[0077] Based on the classic multi-view BEV detector BevFormer, the present invention enhances the BEV detector by integrating the ground truth target information into the BEV encoding and perceptual decoding process.
[0078] Compared with the prior art, the present invention has the following beneficial effects:
[0079] (1) The present invention integrates ground truth information into the BEV encoding and decoding process, enhances the BEV detector, and improves the accuracy and robustness of the BEV detector.
[0080] (2) The additional code layer structure TT encoder is a simple MLP layer that is only used during training and does not incur additional time during inference.
[0081] The exemplary embodiments of the present invention further provide a computer storage medium storing a computer program, wherein the computer program, when executed by a processor, implements the methods of the various embodiments of the present invention. The corresponding process descriptions in the aforementioned method embodiments may be referred to and will not be repeated here.
[0082] The above-described method according to an embodiment of the present invention may be implemented in hardware, firmware, or as software or computer code that may be stored in a recording medium (such as a CD ROM, RAM, floppy disk, hard disk, or magneto-optical disk), or as computer code that is originally stored in a remote recording medium or a non-temporary machine-readable medium downloaded over a network and will be stored in a local recording medium, so that the method described herein may be stored in such software processing on a recording medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware (such as an ASIC or FPGA). It is understood that a computer, processor, microprocessor controller, or programmable hardware includes a storage component (e.g., RAM, ROM, flash memory, etc.) that can store or receive software or computer code, and when the software or computer code is accessed and executed by a computer, processor, or hardware, the method described herein is implemented. In addition, when a general-purpose computer accesses the code for implementing the method shown herein, the execution of the code converts the general-purpose computer into a dedicated computer for executing the method shown herein.
[0083] Thus far, specific embodiments of the present invention have been described. Other embodiments are within the scope of the appended claims. In some cases, the actions recited in the claims may be performed in a different order and still achieve the desired results. Additionally, the processes depicted in the accompanying drawings do not necessarily require the specific order or sequential order shown to achieve the desired results. In some embodiments, multitasking and parallel processing may be advantageous.
[0084] It should be understood that although this specification is described according to various embodiments, not every embodiment contains only one independent technical solution. This narrative method of the specification is only for the sake of clarity. Those skilled in the art should regard the specification as a whole. The technical solutions in each embodiment may also be appropriately combined to form other implementation methods that can be understood by those skilled in the art.
[0085] Finally, it should be noted that the above implementation methods are only used to illustrate the embodiments of the present invention, and are not limitations of the embodiments of the present invention. Ordinary technicians in the relevant technical field can make various changes and modifications without departing from the spirit and scope of the embodiments of the present invention. Therefore, all equivalent technical solutions also belong to the scope of the embodiments of the present invention. The patent protection scope of the embodiments of the present invention should be defined by the claims.
Claims
1. A method for constructing a multi-view BEV detector enhanced by real target information, characterized in that: include: S1. Build the bevformer architecture of the bev detector, use the backbone network backbone of the architecture to extract the features of six panoramic pictures, and input the features of the six panoramic pictures into the bev encoder of the architecture to generate bev features; S2. Input the ground truth target information into the target information encoder TT encoder for encoding, and obtain the target information enhancement feature TT-bev and the target information query feature TT-Q. The target information enhancement feature TT-bev and the target information query feature TT-Q are used to enhance the encoding and decoding capabilities of the bevformer architecture respectively; S3. Use bev feature bev feature and target information enhancement feature TT-bev for comparative learning to obtain the ground truth target supervision loss loss TT-bev ; S4. Input the randomly initialized bev query bev query, bev feature bev feature and target information query feature TT-Q into the bev decoder bev decoder of the architecture to obtain the baseline bev detector loss loss normal and the perceptual loss of ground truth object query TT-Q ; S5. Use ground truth target supervision loss TT-bev Update the network parameters of the bev encoder bev encoder, using the baseline bev detector loss loss normal Update the network parameters of the bev encoder and bev decoder, using the perceptual loss loss of the ground truth target query TT-Q Update the network parameters of bev decoder; S6. Optimize the bev detector using the updated network parameters.
2. The method according to claim 1, characterized in that In step S1, the features of the six panoramic images are input into the bev encoder of the architecture, and the bev encoder is used to merge the image features of the six panoramic images into a unified top-down view bev feature map, that is, The expression of bev encoder is: z bev =BEVEnc(X views ) Among them, X views is a multi-view camera image.
3. The method according to claim 2, characterized in that The target information encoder TTencoder in step S2 is composed of a simple multi-layer perceptron MLP, which aligns the generated target information enhancement feature TT-bev with the bev feature bevfeature, ensures that the bev elements are clearly arranged according to their class labels, positions and boundaries, and takes the class label of the i-th instance in the bev map as t i , the ground truth bounding box is b i , the expression of the target information encoder TT encoder is: α i =TTEnc(t i ,b i ) Among them, α i It has the same feature dimension C as the bev feature and represents the encoded true information feature of the i-th instance on the bev map.
4. The method according to claim 3, characterized in that The generated z is used in step S3 bev and the generated α i Contrastive learning is performed. Prior to this, for each instance, in order to enhance its clarity on the bev map, the tensor within its true information bounding box is cropped from the bev feature map, and the cropped tensor is pooled as an object γ i The expression is: γ i =Pool(Crop(z bev ,b i )) Among them, b i represents the ground truth bounding box.
5. The method according to claim 4, characterized in that In step S3, in order to make the target information enhancement feature TTbev and the bev feature bev feature embedding closer, contrastive learning is used to optimize the element relationship and distance in the bev feature space, the expression is: where N and I represent the generated and target similarity matrices between the object bev features and the object real information features, respectively, μ is the logit scale learned during contrastive learning, represents matrix multiplication, L CE is the cross entropy loss for similarity matrix optimization and is the final loss produced from the TT bev module.
6. The method according to claim 5, characterized in that The step S4 introduces the ground truth information query α i To the query pool of bev decoder bev decoder for decoding, after decoding, the same head Head and perceptual loss L per Applied to the processed ground truth information query, the above process can be expressed as: Q tt =α,Q t ′ t =Dec(Q tt ), Among them, Q tt represents the initial ground truth query, Q t ' t and denote the initial ground truth information query after processing by the perceptual decoder and the perceptual head, respectively, and loss TT-Q is the final loss for ground truth information query.
7. The method according to claim 1, characterized in that The training formula in step S5 contains three key components: baseline bev detector loss loss normal , the perceptual loss for ground truth object query TT-Q and ground truth target supervision loss loss TT-bev , the expression is as follows: L=loss normal +loss TT-Q +loss TT-bev Among them, loss normal is the baseline bev detector loss, loss TT-Q is the perceptual loss for ground truth target query, loss TT-bev Supervised loss for ground truth targets.
8. A multi-view BEV detector enhanced with real target information, characterized in that: The method for constructing a multi-view BEV detector using real target information enhancement as claimed in any one of claims 1 to 7 is adopted, comprising: The bev encoder module merges the image features into the same top-down bev feature map, fuses the historical bev features into the bev query through temporal self-attention, and then uses the features of the bev query and the six panoramic images for cross-attention. The obtained results are subjected to residual connection, regularization and full connection layer to finally obtain the bev feature. TT-bev module, the ground truth target information is passed through an encoder composed of target information encoder TT Encoder layer to obtain target information enhanced feature TT-bev, bev encoder bev encoder generates bev feature bev feature, target information enhanced feature TT-bev is regarded as the result of text compiler, bev feature bev feature is regarded as the result of image compiler, target information enhanced feature TT-bev and bev feature bev feature are enhanced by contrastive learning method, and finally the ground truth target supervision loss loss is obtained TT-bev ; The TT-Q module passes the ground truth target information through an encoder composed of a target information encoder TT Encoder layer to obtain the target information query feature TT-Q; The bev decoder module performs attention during decoding in a parallel manner to prevent information leakage, and finally obtains the baseline bev detector loss loss normal and the perceptual loss of ground truth object query TT-Q .
9. A computer storage medium, characterized in that A computer program is stored thereon, and when the program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.