BEV feature decoding method and apparatus, electronic device, and storage medium
By adopting the multi-dimensional attention decoding model and layer-by-layer loss optimization method in BEV feature decoding, the problem of slow training convergence speed and easy to fall into local optimality in the prior art is solved, and faster training convergence and higher model effects are achieved.
Patent Information
- Application Number
- PCT/CN2024/080438
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-04
- Filing Date
- 2024-03-07
- Publication Date
- 2025-06-12
AI Technical Summary
The existing BEV feature decoding methods converge slowly during training, easily deviate, and the model is easily trapped in local optimality.
The multi-dimensional attention decoding model is adopted, and the network layer is decoding through the N-layer attention mechanism network layer, and the backpropagation process is optimized in combination with hierarchical reference and layer-by-layer loss, so that the overall training converges faster and avoids local optimization.
A faster training convergence process is achieved, reducing the huge deviation of the results and improving the training effect of the model.
Smart Images

Figure CN2024080438_12062025_PF_FP_ABST
Abstract
Description
BEV feature decoding method, device, electronic device and storage medium Technical Field
[0001] The present application mainly relates to the field of autonomous driving technology, and in particular to a BEV feature de-coding method, device, electronic device and storage medium. Background Art
[0002] A bird's eye view (BEV) is a perspective of viewing an object or scene from above, much like a bird looking down from the sky. In the fields of autonomous driving and robotics, data acquired by sensors (such as lidar and cameras) is often converted into a BEV representation to facilitate tasks such as object detection and path trajectory prediction.
[0003] Summary of the Invention
[0004] The present application provides a BEV feature de-coding method, device, electronic device and storage medium.
[0005] In a first aspect, the present application provides a BEV feature de-coding method, comprising:
[0006] Get the BEV feature map corresponding to the image at the current moment;
[0007] The BEV feature map is input into the preset multidimensional attention decoding model for decoding to obtain a multi-task decoding result, wherein the preset multidimensional attention decoding model includes N layers of attention mechanism network layers, each attention mechanism network layer is trained according to the feature sampling point set of the layer and the preset loss function, and N is an integer greater than or equal to 1.
[0008] In an optional embodiment, the BEV feature map is input into a preset multi-dimensional attention de-encoding model for de-encoding to obtain a multi-task de-encoding result, including:
[0009] For each layer in the N-layer attention mechanism network layer:
[0010] Obtain the query parameter, the query position code corresponding to the query parameter, and the reference point; wherein the reference point is the absolute position on the BEV feature map;
[0011] Perform multi-head self-attention mechanism processing according to the Query parameter to obtain the target Query parameter corresponding to the Query parameter;
[0012] Performing cross-attention mechanism processing according to the target Query parameter, the reference point, and the BEV feature map to obtain an updated Query parameter;
[0013] Obtaining the relative position offset of the reference point on the BEV grid according to the updated Query parameter through a fully connected layer and a normalization layer;
[0014] The relative position offset and the reference point are summed to obtain an updated reference point.
[0015] In the above method, the multi-head self-attention mechanism processing process can establish connections between the targets and prevent the output targets from being repeated.
[0016] In a possible implementation, obtaining the Query parameter, the Query position code corresponding to the Query parameter, and the reference point includes:
[0017] If the attention mechanism network layer is the first layer in N layers, the query parameter, the query position encoding corresponding to the query parameter, and the reference point are obtained by random initialization;
[0018] If the attention mechanism network layer is a layer other than the first layer in the N layers, the Query parameter is the updated Query parameter of the previous layer, the Query position encoding corresponding to the Query parameter is the same as the Query position encoding of the previous layer, and the reference point is the updated reference point of the previous layer.
[0019] In one possible implementation, if the attention mechanism network layer is the last layer, after summing the relative position offset and the reference point to obtain the updated reference point, the method further includes: outputting the updated reference point as the final absolute position of the BEV grid and outputting a multi-task inverse encoding result.
[0020] In one possible implementation, performing a cross-attention mechanism processing according to the target Query parameter, the reference point, and the BEV feature to obtain an updated Query parameter includes:
[0021] Obtaining the sampling offset and attention weight of the reference point according to the target Query parameter;
[0022] According to the BEV characteristic graph, obtaining a Value parameter corresponding to the BEV characteristic graph;
[0023] Summing the reference point and the sampling offset to obtain a sampling position of the reference point on the BEV characteristic map, and combining the Value parameter to obtain a sampling characteristic value of the BEV characteristic map;
[0024] The sampled feature values are weighted and summed according to the attention weights to obtain the updated Query parameters.
[0025] Through the above method, the absolute position of the target Query parameter and the multi-task inverse encoding result are obtained.
[0026] In an optional embodiment, the fully connected layer can be a multi-layer or a single layer.
[0027] In an optional embodiment, the multidimensional attention de-encoding model is trained in the following manner:
[0028] The final absolute position output of the BEV grid is input into the N-layer attention mechanism network layer in the multi-dimensional attention inverse encoding model;
[0029] In the N-layer attention mechanism network layer, the final absolute position output of the BEV grid is decoded through the inverse encoder to obtain the multi-task prediction results of each of the N-layer attention mechanism network layers;
[0030] Based on the multi-task prediction results and the true value, the loss of each layer of the N-layer attention mechanism network is calculated;
[0031] The sum of the losses of the N attention mechanism network layers is taken as the total loss of the preset multi-dimensional attention inverse encoding model.
[0032] Through the above method, the multi-task prediction results of each layer of the attention mechanism network layer are compared with the true value to calculate the loss, and the loss of each layer of the attention mechanism network layer is obtained. The total loss of the preset multi-dimensional attention inverse encoding model is the sum of the losses of each layer of the attention mechanism network layer. In this way, the backpropagation process is optimized, and the model is not easy to fall into the local optimum. Therefore, the results are unlikely to have huge deviations.
[0033] In a second aspect, the present application provides a BEV de-coding device, the device comprising:
[0034] An acquisition module is used to obtain a BEV feature map corresponding to the image at the current moment;
[0035] The decoding module is used to input the BEV feature map into the preset multi-dimensional attention decoding model for decoding to obtain a multi-task decoding result, wherein the preset multi-dimensional attention decoding model includes N layers of attention mechanism network layers, each attention mechanism network layer is trained according to the feature sampling point set of the layer and the preset loss function, and N is an integer greater than or equal to 1.
[0036] In an optional embodiment, when the BEV feature map is input into a preset multidimensional attention de-coding model for de-coding to obtain a multi-task de-coding result, the de-coding module is specifically used to:
[0037] For each layer in the N-layer attention mechanism network layer:
[0038] Obtaining a query parameter, a query position code corresponding to the query parameter, and a reference point; wherein the reference point is an absolute position on the BEV characteristic map;
[0039] Perform multi-head self-attention mechanism processing according to the Query parameter to obtain the target Query parameter corresponding to the Query parameter;
[0040] Performing cross-attention mechanism processing according to the target Query parameter, the reference point, and the BEV feature map to obtain an updated Query parameter;
[0041] Obtaining the relative position offset of the reference point on the BEV grid according to the updated Query parameter through a fully connected layer and a normalization layer;
[0042] The relative position offset and the reference point are summed to obtain an updated reference point.
[0043] In an optional implementation, when obtaining the Query parameter, the Query position code corresponding to the Query parameter, and the reference point, the decoding module is specifically configured to:
[0044] If the attention mechanism network layer is the first layer in N layers, the query parameter, the query position encoding corresponding to the query parameter, and the reference point are obtained by random initialization;
[0045] If the attention mechanism network layer is a layer other than the first layer in the N layers, the Query parameter is the updated Query parameter of the previous layer, the Query position encoding corresponding to the Query parameter is the same as the Query position encoding of the previous layer, and the reference point is the updated reference point of the previous layer.
[0046] In an optional embodiment, if the attention mechanism network layer is the last layer, after summing the relative position offset and the reference point to obtain the updated reference point, the decoding module is further used to: output the updated reference point as the final absolute position of the BEV grid, and output the multi-task decoding result.
[0047] In an optional embodiment, when performing a cross-attention mechanism processing according to the target Query parameter, the reference point, and the BEV feature to obtain an updated Query parameter, the de-encoding module is specifically configured to:
[0048] Obtaining the sampling offset and attention weight of the reference point according to the target Query parameter;
[0049] According to the BEV characteristic graph, obtaining a Value parameter corresponding to the BEV characteristic graph;
[0050] Summing the reference point and the sampling offset to obtain a sampling position of the reference point on the BEV characteristic map, and combining the Value parameter to obtain a sampling characteristic value of the BEV characteristic map;
[0051] The sampled feature values are weighted and summed according to the attention weights to obtain the updated Query parameters.
[0052] In an optional embodiment, the fully connected layer is a multi-layer or a single layer.
[0053] In an optional embodiment, the multidimensional attention de-encoding model is trained in the following manner:
[0054] The final absolute position output of the BEV grid is input into the N-layer attention mechanism network layer in the multi-dimensional attention inverse encoding model;
[0055] Decoding the final absolute position output of the BEV grid through the inverse encoder in the N-layer attention mechanism network layer to obtain the multi-task prediction results of each of the N-layer attention mechanism network layers;
[0056] Calculate the loss of each of the N attention mechanism network layers based on the multi-task prediction results and the true value;
[0057] The sum of the losses of the N attention mechanism network layers is taken as the total loss of the preset multi-dimensional attention inverse encoding model.
[0058] In a third aspect, the present application provides an electronic device, comprising:
[0059] Memory for storing computer programs;
[0060] The processor is configured to implement the steps of the above-mentioned BEV feature de-coding method when executing the computer program stored in the memory.
[0061] In a fourth aspect, the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above-mentioned BEV feature decoding method are implemented.
[0062] Through the technical solutions in one or more of the above embodiments of the present application, the embodiments of the present application have at least the following beneficial effects:
[0063] In an embodiment of the present application, the BEV feature map corresponding to the image at that moment is first obtained; then the BEV feature map is input into the preset multi-dimensional attention de-coding model for de-coding to obtain a multi-task de-coding result. The preset multi-dimensional attention de-coding model is set to an N-layer attention mechanism network layer, and each layer of the attention mechanism network layer is trained according to the feature sampling point set of the layer and the preset loss function. The loss of each layer of the attention mechanism network layer is used as an auxiliary loss, and the sum of the losses of each layer of the attention mechanism network layer is used as the total loss of the preset multi-dimensional attention de-coding model. During the training process, the back propagation process of the model is optimized by combining hierarchical reference and layer-by-layer loss, so that the overall training convergence process is faster. At the same time, the model training is not easy to fall into the local optimum, so that the results are unlikely to have huge deviations.
[0064] For each of the above-mentioned aspects from the second to the fourth aspects and the technical effects that may be achieved by each of the aspects, please refer to the above-mentioned description of the technical effects that can be achieved by the first aspect and the various possible solutions in the first aspect, and no further details will be given here. BRIEF DESCRIPTION OF THE DRAWINGS
[0065] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for describing the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without inventive work. In the drawings:
[0066] FIG1 is a schematic diagram of an implementation flow of a BEV feature de-coding method provided in an embodiment of the present application;
[0067] FIG2 is a schematic diagram of the structure of a multidimensional attention de-coding model provided in an embodiment of the present application;
[0068] FIG3 is a schematic diagram of an absolute position output of a reference point provided in an embodiment of the present application;
[0069] FIG4 is a schematic structural diagram of a BEV feature de-coding device provided in an embodiment of the present application;
[0070] FIG5 is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0071] To make the purpose, technical solutions, and advantages of the embodiments of this application more clear, the technical solutions of this application will be clearly and completely described below in conjunction with the drawings in the embodiments of this application. Obviously, the described embodiments are part of the embodiments of the technical solutions of this application, but not all of them. Based on the embodiments described in this application document, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the technical solutions of this application.
[0072] It should be noted that in the description of this application, "multiple" is understood to mean "at least two." "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. A and B are connected, which can mean: A and B are directly connected, and A and B are connected through C. In addition, in the description of this application, words such as "first" and "second" are only used for the purpose of distinguishing descriptions, and cannot be understood as indicating or implying relative importance, nor can they be understood as indicating or implying an order.
[0073] The embodiments of the present application are described in detail below with reference to the accompanying drawings.
[0074] A bird's-eye view (BEV) is a perspective of viewing an object or scene from above, much like a bird looking down at the ground from the air. In the fields of autonomous driving and robotics, data acquired by sensors (such as lidar and cameras) is often converted into a BEV representation to facilitate tasks such as object detection and path trajectory prediction.
[0075] In autonomous driving technology, BEV is often used to output multiple tasks, such as 3D target detection, local high-precision map vector fragment output, trajectory prediction, etc., all of which require BEV feature processing. The existing inverse encoding method uses a CNN detection head, which uses a convolutional layer to detect or regress the BEV feature map to obtain multi-task results. However, this method has a slow training process convergence speed and is prone to deviations. This is because when the CNN detection head processes the BEV features, it first obtains a total loss and then optimizes this total loss. Therefore, a local optimum may appear when training the model, resulting in insufficient convergence of the model training loss, which in turn leads to large deviations.
[0076] In view of this, an embodiment of the present application provides a BEV feature de-coding method, which includes: first, obtaining the BEV feature map corresponding to the image at the current moment; then, inputting the BEV feature map into a preset multi-dimensional attention de-coding model for de-coding to obtain a multi-task de-coding result. The preset multi-dimensional attention de-coding model includes N layers of attention mechanism network layers, each of which is trained according to the corresponding feature sampling point set and the preset loss function, and the model is trained in a flexible training manner. First, each layer of encoder is regarded as a de-coding of the target, and the loss of each layer output is calculated separately. Except for the loss calculated for the de-coding output of the last layer, the losses of the remaining layers are auxiliary losses, which are used to assist model training. Therefore, the total loss of the model is the sum of the losses of multiple layers. In this way, the back propagation process of each layer of the attention mechanism network layer is optimized, so that the overall training converges faster. At the same time, the model is not easy to fall into the local optimum, so the results are unlikely to have huge deviations.
[0077] It should be noted that the preferred embodiments of the present application are described below in conjunction with the drawings in the specification. The preferred embodiments described here are only used to illustrate and explain the present application and are not used to limit the present application. In addition, the embodiments of the present application and the features in their embodiments may be combined with each other if there is no conflict.
[0078] 1 , which is a schematic diagram of an implementation flow of a BEV feature de-coding method provided in an embodiment of the present application, the specific implementation flow of the method is as follows:
[0079] S1, obtain the BEV feature map corresponding to the image at the current moment.
[0080] S2: Input the BEV feature map into the preset multi-dimensional attention de-encoding model for de-encoding to obtain the multi-task de-encoding result.
[0081] First of all, the method provided in the embodiment of the present application can be applied to the multidimensional attention decoding model as shown in Figure 2. The input of the multidimensional attention decoding model can be a BEV feature map. First, the corresponding sensor data is collected by various types of vehicle-mounted sensors. Then, the corresponding features are extracted through the feature extraction network and the BEV feature map is obtained after corresponding preprocessing. The method for obtaining the BEV feature map is not described in detail here. The output of the multidimensional attention decoding model can be a multi-task decoding result.
[0082] In the embodiments of the present application, the multidimensional attention de-encoding model can be deployed on a designated computing chip, such as a designated on-chip computing unit on a vehicle. Furthermore, the multidimensional attention model can be trained on a cloud server platform. The embodiments of the present application do not impose any specific restrictions on the deployment and training locations of the multidimensional attention de-encoding model.
[0083] In the embodiments of the present application, the BEV features in the BEV feature map are features obtained from a bird's-eye view. That is, the BEV features are global features and a cross-view feature representation that fuses information from images of different viewpoints and can effectively represent information near the origin of the reference point coordinates. The fused image information is a multi-dimensional feature, for example, 1000+ dimensions, including color, gradient, edge, luminosity, grayscale, etc.
[0084] First, multiple images of the vehicle at the current moment are acquired through onboard sensors. Depending on the location of the onboard sensors on the vehicle, images from multiple different perspectives can be acquired accordingly. These multiple perspectives may include, but are not limited to, the perspective directly in front of the vehicle, the perspective from the right front of the vehicle, the perspective from the left front of the vehicle, the perspective from the rear of the vehicle, the perspective from the right rear of the vehicle, and the perspective from the left rear of the vehicle. The images are then processed to obtain a BEV feature map for each of the vehicle's multiple perspectives. After acquiring the BEV feature map corresponding to the image at the current moment, the BEV feature map is input into a pre-set multi-dimensional attention de-encoding model for de-encoding, resulting in a multi-task de-encoding result.
[0085] Among them, the preset multi-dimensional attention decoding model includes N layers of attention mechanism network layers, and each attention mechanism network layer is trained according to the feature sampling point set of the layer and the preset loss function.
[0086] Specifically, when de-encoding BEV features, connections need to be established between targets to avoid duplication. For each of the N attention layers, the following steps are performed: first, the query parameters, the corresponding query position encoding, and the reference point are obtained; the reference point is the absolute position on the BEV feature map; then, a multi-head self-attention mechanism is applied to the query parameters to obtain the target query parameters corresponding to the query parameters; then, a cross-attention mechanism is applied to the target query parameters, the reference point, and the BEV feature map to obtain the updated query parameters.
[0087] Among them, the Query parameter consists of n elements, that is, the sequence length of the Query parameter is n. The multi-head self-attention mechanism can make the n internal elements of the Query parameter focus on different areas.
[0088] Furthermore, a relative position offset of the reference point on the BEV grid is obtained according to the updated Query parameter through a fully connected layer and a normalization layer; and the relative position offset and the reference point are summed to obtain an updated reference point.
[0089] In some embodiments, the fully connected layer can be multiple layers or a single layer.
[0090] Specifically, when the attention mechanism network layer is the first layer in N layers, the Query parameter, the Query position encoding corresponding to the Query parameter, and the reference point are obtained through random initialization.
[0091] Because the self-attention mechanism is used, the initial Query parameter, initial Value parameter, and initial Key parameter have the same value. It should be noted that the Query parameter represents the query in the attention mechanism, the Key parameter represents the key in the attention mechanism, and the Value parameter represents the value in the attention mechanism.
[0092] When the attention mechanism network layer is a layer other than the first layer in N layers, the Query parameter is the updated Query parameter of the previous layer, the Query position encoding corresponding to the Query parameter is the same as the Query position encoding of the previous layer, and the reference point is the updated reference point of the previous layer.
[0093] Furthermore, when the attention mechanism network layer is the last layer, the relative position offset of the reference point on the BEV grid is obtained according to the query parameters output by the last attention mechanism network layer through the fully connected layer and the normalization layer, and then the relative position offset and the reference point are summed to obtain an updated reference point and the updated reference point is output as the final absolute position of the BEV grid, and the multi-task inverse encoding result is output at the same time.
[0094] For example, in the embodiment of the present application, the task decoding result may include outputting the absolute position of the BEV grid, generating a 3D target, estimating the 3D motion trajectory of the 3D target, and determining the 3D occupied area of the 3D target based on the 3D motion trajectory estimation result of the 3D target. The 3D target may be other participants on the road, such as pedestrians, other vehicles, or other moving objects.
[0095] Referring to FIG3 , in some embodiments, the sampling offset and attention weight of the reference point are obtained according to the target Query parameter. According to the BEV feature map, the Value parameter corresponding to the BEV feature map is obtained. The reference point and the sampling offset are summed to obtain the sampling position of the reference point on the BEV feature map, and the sampling characteristic value of the BEV feature map is obtained in combination with the Value parameter. According to the attention weight, the sampling characteristic value is weightedly summed to obtain the updated Query parameter. It should be noted that this process adopts the cross-attention mechanism, and the Value value in the cross-attention mechanism is different from the Value value in the aforementioned self-attention focus.
[0096] In an optional embodiment, a preset multi-dimensional attention inverse encoding model is trained by flexible training, specifically including: the final absolute position output of the BEV grid obtained in the above process is input into the N-layer attention mechanism network layer in the multi-dimensional attention inverse encoding model; then, the final absolute position output of the BEV grid is decoded by the inverse encoder in the N-layer attention mechanism network layer to obtain the multi-task prediction results of each of the N-layer attention mechanism network layers; then, a preset loss function is used to independently calculate the loss of the multi-task prediction results and the true value obtained by each layer of the attention mechanism network layer to obtain the loss of each of the N-layer attention mechanism network layers; further, the sum of the losses of the N-layer attention mechanism network layers is used as the total loss of the preset multi-dimensional attention inverse encoding model. In an embodiment of the present application, the multi-dimensional attention inverse encoding model can be set to a 6-layer attention mechanism network layer. Of course, it can also be set to other numbers of layers according to actual needs, and the target loss finally obtained is the sum of the losses of each layer of the N-layer attention mechanism network layer.
[0097] It's understandable that each attention network layer is considered a de-encoding of the target. Therefore, the loss of each attention network layer's output can be calculated separately. Except for the loss of the final de-encoding output, the losses of each subsequent layer are auxiliary losses used to assist model training. The final total loss of the multi-dimensional attention de-encoding model is the sum of the losses of each attention network layer.
[0098] In an embodiment of the present application, the multi-dimensional attention de-coding model is based on the attention mechanism under natural language as the core. By superimposing N layers of attention mechanism network layers, all image coding information (i.e., feature coding information) is fused according to weight exchange (i.e., feature sampling point weights), and the multi-task de-coding result is output end-to-end. Moreover, during training, each attention mechanism network layer in the multi-dimensional attention de-coding network optimizes the back propagation process between the N layers of attention mechanism network layers by combining the sampling points of the features at the reference points of each attention mechanism network layer (i.e., the set of feature sampling points) and the loss function, so that the results will not have huge deviations, thereby ensuring the accuracy of the multi-dimensional attention de-coding model.
[0099] Based on the same inventive concept, the embodiment of the present application also includes a BEV feature de-coding device, as shown in FIG4 , the device includes: an acquisition module 401 and a de-coding module 402, wherein:
[0100] An acquisition module 401 is used to acquire a BEV feature map corresponding to the image at the current moment;
[0101] The decoding module 402 is used to input the BEV feature map into a preset multi-dimensional attention decoding model for decoding to obtain a multi-task decoding result, wherein the preset multi-dimensional attention decoding model includes N layers of attention mechanism network layers, each attention mechanism network layer is trained according to the feature sampling point set of the layer and the preset loss function, and N is an integer greater than or equal to 1.
[0102] In an optional embodiment, when the BEV feature is input into a preset multi-dimensional attention de-encoding model for de-encoding to obtain a multi-task de-encoding result, the de-encoding module 402 is specifically configured to:
[0103] For each layer of the N-layer attention network: obtain a query parameter, a BEV position encoding corresponding to the query parameter, and a reference point; wherein the reference point is an absolute position on the BEV feature map; perform a multi-head self-attention mechanism on the query parameter to obtain a target query parameter corresponding to the query parameter; perform a cross-attention mechanism on the target query parameter, the reference point, and the BEV feature map to obtain an updated query parameter;
[0104] Obtaining the relative position offset of the reference point on the BEV grid according to the updated Query parameter through a fully connected layer and a normalization layer;
[0105] The relative position offset and the reference point are summed to obtain an updated reference point.
[0106] In an optional embodiment, when obtaining the Query parameter, the Query position code corresponding to the Query parameter, and the reference point, the decoding module is specifically used to: if the attention mechanism network layer is the first layer in N layers, then the Query parameter, the Query position code corresponding to the Query parameter, and the reference point are obtained through random initialization; if the attention mechanism network layer is other layers in N layers except the first layer, then the Query parameter is the updated Query parameter of the previous layer, the Query position code corresponding to the Query parameter is the same as the Query position code of the previous layer, and the reference point is the updated reference point of the previous layer.
[0107] In an optional embodiment, if the attention mechanism network layer is the last layer, after summing the relative position offset and the reference point to obtain the updated reference point, the decoding module is further used to: output the updated reference point as the final absolute position of the BEV grid, and output the multi-task decoding result.
[0108] In an optional embodiment, when a cross-attention mechanism is performed according to the target Query parameter, the reference point and the BEV feature to obtain an updated Query parameter, the decoding module is specifically used to: obtain the sampling offset and attention weight of the reference point according to the target Query parameter; obtain the Value parameter corresponding to the BEV feature map according to the BEV feature map; sum the reference point and the sampling offset to obtain the sampling position of the reference point on the BEV feature map, and obtain the sampling feature value of the BEV feature map in combination with the Value parameter; and perform weighted summation of the sampling feature value according to the attention weight to obtain the updated Query parameter.
[0109] In an optional embodiment, the fully connected layer can be a multi-layer or a single layer.
[0110] In an optional embodiment, the multidimensional attention de-encoding model is trained in the following manner:
[0111] The final absolute position output of the BEV grid is input into the N-layer attention mechanism network layer in the multi-dimensional attention inverse encoding model;
[0112] In the N-layer attention mechanism network layer, the final absolute position output of the BEV grid is decoded through the inverse encoder to obtain the multi-task prediction results of each of the N-layer attention mechanism network layers;
[0113] Based on the multi-task prediction results and the true value, the loss of each layer of the N-layer attention mechanism network is calculated;
[0114] The sum of the losses of each of the N-layer attention mechanism network layers is used as the total loss of the preset multi-dimensional attention de-encoding model.
[0115] It should be noted here that the above-mentioned device provided in the embodiment of the present application can implement all the method steps in the above-mentioned BEV feature decoding method embodiment, and can achieve the same technical effect. The parts and beneficial effects of this embodiment that are the same as the method embodiment will not be described in detail here.
[0116] Based on the same inventive concept, an electronic device is further provided in an embodiment of the present application. The electronic device can implement the functions of the aforementioned BEV feature de-coding method. As shown in FIG5 , the electronic device includes:
[0117] At least one processor 501, and a memory 502 connected to at least one processor 501. The specific connection medium between the processor 501 and the memory 502 is not limited in the embodiments of the present application. FIG5 takes the connection between the processor 501 and the memory 502 via the bus 500 as an example. The bus 500 is represented by a bold line in FIG5. The connection method between other components is only for schematic illustration and is not intended to be limiting. The bus 500 can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, FIG5 only uses a bold line to represent it, but this does not mean that there is only one bus or one type of bus. Alternatively, the processor 501 can also be called a controller, and there is no limitation on the name.
[0118] In this embodiment of the present application, memory 502 stores instructions executable by at least one processor 501. At least one processor 501 can execute the BEV feature decoding method discussed above by executing the instructions stored in memory 502. Processor 501 can implement the functions of each module in the apparatus shown in FIG4 .
[0119] Among them, the processor 501 is the control center of the device, which can use various interfaces and lines to connect the various parts of the entire control device, and monitor the device as a whole by running or executing instructions stored in the memory 502 and calling data stored in the memory 502, the various functions of the device and processing data.
[0120] In one possible design, processor 501 may include one or more processing units. Processor 501 may integrate an application processor and a modem processor. The application processor primarily processes the operating system, user interface, and application programs, while the modem processor primarily processes wireless communications. It is understood that the modem processor may not be integrated into processor 501. In some embodiments, processor 501 and memory 502 may be implemented on the same chip. In some embodiments, they may also be implemented on separate chips.
[0121] The processor 501 can be a general-purpose processor, such as a central processing unit (CPU), a digital signal processor, an application-specific integrated circuit, a field-programmable gate array or other programmable logic device, a discrete gate or transistor logic device, or a discrete hardware component, and can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of this application. A general-purpose processor can be a microprocessor or any conventional processor. The steps of the BEV feature decoding method disclosed in the embodiments of this application can be directly implemented and executed by a hardware processor, or by a combination of hardware and software modules in the processor.
[0122] The memory 502 is a non-volatile computer-readable storage medium that can be used to store non-volatile software programs, non-volatile computer executable programs and modules. The memory 502 may include at least one type of storage medium, such as a flash memory, a hard disk, a multimedia card, a card-type memory, a random access memory (Random Access Memory, RAM), a static random access memory (Static Random Access Memory, SRAM), a programmable read-only memory (Programmable Read Only Memory, PROM), a read-only memory (Read Only Memory, ROM), an electrically erasable programmable read-only memory (Electrically Erasable Programmable Read-Only Memory, EEPROM), a magnetic memory, a disk, an optical disk, etc. The memory 502 is any other medium that can be used to carry or store a desired program code in the form of an instruction or data structure and can be accessed by a computer, but is not limited thereto. The memory 502 in the embodiment of the present application can also be a circuit or any other device that can realize a storage function, for storing program instructions and / or data.
[0123] By programming the processor 501, the code corresponding to the BEV feature decoding method described in the aforementioned embodiment can be embedded in the chip, enabling the chip to execute the steps of the BEV feature decoding method of the embodiment shown in FIG1 during operation. Designing and programming the processor 501 is well known to those skilled in the art and will not be further described here.
[0124] Based on the same inventive concept, an embodiment of the present application further provides a storage medium storing computer instructions. When the computer instructions are executed on a computer, the computer executes the BEV feature decoding method discussed above.
[0125] In some possible implementations, various aspects of the BEV feature decoding method provided in the present application can also be implemented in the form of a program product, which includes program code. When the program product is run on the device, the program code is used to enable the control device to execute the steps of the BEV feature decoding method according to various exemplary embodiments of the present application described above in this specification.
[0126] Those skilled in the art will appreciate that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0127] The present application is described with reference to the flow chart and / or block diagram of the method, device (system), and computer program product according to the embodiment of the present application. It should be understood that each flow process and / or box in the flow chart and / or block diagram and the combination of the flow process and / or box in the flow chart and / or block diagram can be realized by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processing machine or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device produce a device for realizing the function specified in one flow chart flow or multiple flows and / or one box or multiple boxes of the block diagram.
[0128] These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce a product including an instruction device that implements the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.
[0129] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.
[0130] Obviously, those skilled in the art may make various changes and modifications to this application without departing from the spirit and scope of this application. Thus, if these modifications and variations of this application fall within the scope of the claims of this application and their equivalents, this application is intended to include these modifications and variations.
Claims
1. A BEV feature de-coding method, the method comprising: Get the BEV feature map corresponding to the image at the current moment; The BEV feature map is input into a preset multidimensional attention de-coding model for de-coding to obtain a multi-task de-coding result, wherein the preset multidimensional attention de-coding model includes N layers of attention mechanism network layers, each of which is trained according to the feature sampling point set of the layer and a preset loss function, and N is an integer greater than or equal to 1.
2. The method according to claim 1, characterized in that The step of inputting the BEV feature map into a preset multi-dimensional attention de-coding model for de-coding to obtain a multi-task de-coding result includes: For each layer in the N-layer attention mechanism network layer: Obtaining a Query parameter, a Query position code corresponding to the Query parameter, and a reference point; wherein the reference point is an absolute position on the BEV characteristic map; Perform multi-head self-attention mechanism processing according to the Query parameter to obtain a target Query parameter corresponding to the Query parameter; Perform cross-attention mechanism processing according to the target Query parameter, the reference point and the BEV feature map to obtain updated Query parameters; Obtaining the relative position offset of the reference point on the BEV grid according to the updated Query parameter through a fully connected layer and a normalized layer; The relative position offset and the reference point are summed to obtain an updated reference point.
3. The method according to claim 2, characterized in that Acquiring the Query parameter, the Query position code corresponding to the Query parameter, and the reference point includes: If the attention mechanism network layer is the first layer among N layers, the Query parameter, the Query position encoding corresponding to the Query parameter, and the reference point are obtained by random initialization; If the attention mechanism network layer is a layer other than the first layer in the N layers, the Query parameter is the updated Query parameter of the previous layer, the Query position encoding corresponding to the Query parameter is the same as the Query position encoding of the previous layer, and the reference point is the updated reference point of the previous layer.
4. The method according to claim 3, characterized in that If the attention mechanism network layer is the last layer, after summing the relative position offset and the reference point to obtain the updated reference point, the method further includes: The updated reference point is output as the final absolute position of the BEV grid, and the multi-task inverse encoding result is output.
5. The method according to claim 2, characterized in that A cross attention mechanism is performed according to the target Query parameter, the reference point and the BEV feature to obtain an updated Query parameter, including: According to the target Query parameter, obtain the sampling offset and attention weight of the reference point; According to the BEV characteristic graph, obtaining a Value parameter corresponding to the BEV characteristic graph; The reference point and the sampling offset are summed to obtain the sampling position of the reference point on the BEV characteristic map, and combined with the Value parameter to obtain the sampling characteristic value of the BEV characteristic map; According to the attention weights, the sampled feature values are weighted summed to obtain the updated Query parameters.
6. The method according to claim 2, characterized in that The fully connected layer is a multi-layer or a single layer.
7. The method according to any one of claims 1 to 6, characterized in that The multi-dimensional attention anti-coding model is trained in the following way: Input the final absolute position output of the BEV grid into the N-layer attention mechanism network layer in the multi-dimensional attention inverse encoding model; In the N-layer attention mechanism network layer, the final absolute position output of the BEV grid is decoded by the inverse encoder to obtain the multi-task prediction results of each of the N-layer attention mechanism network layers; According to the multi-task prediction results and the true value, the loss of each of the N-layer attention mechanism network layers is calculated; The sum of the losses of each of the N attention mechanism network layers is taken as the total loss of the preset multi-dimensional attention de-encoding model.
8. A BEV feature de-coding device, the device comprising: An acquisition module is used to acquire a BEV feature map corresponding to the image at the current moment; A decoding module is used to input the BEV feature map into a preset multidimensional attention decoding model for decoding to obtain a multi-task decoding result, wherein the preset multidimensional attention decoding model includes N layers of attention mechanism network layers, each of which is trained according to the feature sampling point set of the layer and a preset loss function, and N is an integer greater than or equal to 1.
9. The device according to claim 8, characterized in that When the BEV feature map is input into the preset multi-dimensional attention de-coding model for de-coding to obtain a multi-task de-coding result, the de-coding module is specifically used for: For each layer of the N-layer attention mechanism network layer: obtain a Query parameter and a Query position code corresponding to the Query parameter, and a reference point; wherein the reference point is an absolute position on the BEV feature map; Perform multi-head self-attention mechanism processing according to the Query parameter to obtain a target Query parameter corresponding to the Query parameter; Perform cross-attention mechanism processing according to the target Query parameter, the reference point and the BEV feature map to obtain updated Query parameters; Obtaining the relative position offset of the reference point on the BEV grid according to the updated Query parameter through a fully connected layer and a normalized layer; The relative position offset and the reference point are summed to obtain an updated reference point.
10. The device according to claim 9, characterized in that When acquiring the Query parameter, the Query position code corresponding to the Query parameter, and the reference point, the decoding module is specifically used to: If the attention mechanism network layer is the first layer among N layers, the Query parameter, the Query position encoding corresponding to the Query parameter, and the reference point are obtained by random initialization; If the attention mechanism network layer is a layer other than the first layer in the N layers, the Query parameter is the updated Query parameter of the previous layer, the Query position encoding corresponding to the Query parameter is the same as the Query position encoding of the previous layer, and the reference point is the updated reference point of the previous layer.
11. The device according to claim 10, characterized in that If the attention mechanism network layer is the last layer, after summing the relative position offset and the reference point to obtain the updated reference point, the de-coding module is further used to: The updated reference point is output as the final absolute position of the BEV grid, and the multi-task inverse encoding result is output.
12. The device according to claim 9, characterized in that When the cross attention mechanism is processed according to the target Query parameter, the reference point and the BEV feature to obtain the updated Query parameter, the decoding module is specifically used to: According to the target Query parameter, obtain the sampling offset and attention weight of the reference point; According to the BEV characteristic graph, obtaining a Value parameter corresponding to the BEV characteristic graph; The reference point and the sampling offset are summed to obtain the sampling position of the reference point on the BEV characteristic map, and combined with the Value parameter to obtain the sampling characteristic value of the BEV characteristic map; According to the attention weights, the sampled feature values are weighted summed to obtain the updated Query parameters.
13. The device according to claim 9, characterized in that The fully connected layer is a multi-layer or a single layer.
14. The device according to any one of claims 8 to 13, characterized in that The multi-dimensional attention anti-coding model is trained in the following way: Input the final absolute position output of the BEV grid into the N-layer attention mechanism network layer in the multi-dimensional attention inverse encoding model; In the N-layer attention mechanism network layer, the final absolute position output of the BEV grid is decoded by the inverse encoder to obtain the multi-task prediction results of each of the N-layer attention mechanism network layers; According to the multi-task prediction results and the true value, the loss of each of the N-layer attention mechanism network layers is calculated; The sum of the losses of each of the N attention mechanism network layers is taken as the total loss of the preset multi-dimensional attention de-encoding model.
15. An electronic device, comprising: Memory, used to store computer programs; A processor, for implementing the method according to any one of claims 1 to 7 when executing a computer program stored in the memory.
16. A computer-readable storage medium, wherein a computer program is stored in the computer-readable storage medium, and when the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Multi-modal fusion three-dimensional target robust detection method based on automatic driving scene
CN116486368A
Bird-eye view feature coding method, system and device of image and storage medium
CN116543059A
Systems and Methods for Generating Motion Forecast Data for a Plurality of Actors with Respect to an Autonomous Vehicle
US20210009166A1
Cited By
BEV space construction method, automatic driving system, equipment and medium
CN121074319A
Air-ground collaborative topology understanding method and device based on potential world model
CN121482313A
Vectorized high-precision map construction method and device, electronic equipment and medium
CN122066820A