Bird's-eye view feature generation method, device and equipment based on spatiotemporal attention
By constructing a bird's-eye view feature generation method based on space-time attention, using the relative position matrix between cameras and the timing relationship between video frames, the cross attention results of the features and corresponding candidate feature points in the BEV feature map are calculated, and feature fusion is carried out, which solves the shortcomings of the existing BEV feature learning methods in view angle conversion, spatial distribution information utilization and video timing information utilization, and achieves high accuracy and rich semantic expression of BEV features.
Patent Information
- Application Number
- CN202311463080.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-06
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2043-11-06
AI Technical Summary
The existing BEV feature learning methods have shortcomings in perspective conversion, spatial distribution information utilization and video timing information utilization, resulting in limited accuracy and expression ability of BEV features.
By constructing a bird's-eye and spatial attention feature generation method based on space-time attention, using the relative position matrix between cameras and the timing relationship between video frames, the cross attention results of the features and corresponding candidate feature points in the BEV feature map are calculated, and feature fusion is performed to improve the expression accuracy of BEV features.
This method not only accurately expresses and fuses BEV features from the two dimensions of space and time domain, but also makes full use of the spatial position relationship between multiple cameras and the relative timing relationship of video frames, improving the expression accuracy and semantic richness of BEV features.
Smart Images

Figure CN117671623B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image data processing, and specifically to a method, device and equipment for generating bird's-eye view features based on spatiotemporal attention. Background Art
[0002] The combination and application of multiple on-board sensors has become the development trend of intelligent driving. The application of multiple sensors, especially multi-vision sensors, has greatly expanded the intelligent driving system's ability to perceive the surrounding environment. However, due to the differences in internal and external parameters of different vision sensors and the complex spatial topological relationship of visual information, the perception results of these sensors are difficult to accurately unify into the same coordinate system, which greatly limits the perception performance of the visual system. Therefore, the bird's-eye view (BEV) perception algorithm for multi-view image fusion has become a research hotspot and research difficulty in industry and academia. This method maps multi-view images to the BEV feature space, and then uses the BEV features to generate a bird's-eye view covering the vehicle's 360-degree visual information. Therefore, how to learn / generate BEV features with high accuracy and strong expressiveness has become one of the most important and challenging issues in this research field.
[0003] The existing BEV feature learning methods have the following three problems: (1) There is a deviation in the perspective conversion. The existing methods mostly perform spatial transformation and fusion on the image features captured by different cameras based on the internal and external parameters of the camera. Due to the deviation of the parameters themselves and the visual misalignment caused by the movement of the vehicle, there is a certain error in the perspective conversion; (2) The spatial distribution information is not fully utilized. In the process of multi-perspective feature extraction and fusion, most methods ignore the deep spatial correlation between different perspectives, resulting in the limited expression ability of the final learned BEV features; (3) The video time sequence information is not fully utilized. When learning time domain features, the existing methods ignore the influence of the frame sequence information of adjacent frames on feature learning, resulting in the lack of spatiotemporal continuity in the learned time domain features. Summary of the invention
[0004] The present application provides a method, device and equipment for generating bird's-eye view features based on spatiotemporal attention, which can improve the accuracy of BEV feature expression.
[0005] In a first aspect, an embodiment of the present application provides a method for generating bird's-eye view features based on spatiotemporal attention, and the method for generating bird's-eye view features based on spatiotemporal attention includes:
[0006] The video images captured by multiple cameras on the vehicle are characterized to obtain a FOV feature map;
[0007] According to the relative position matrix between cameras and the coordinates in the relative position matrix, a relative position index table is obtained to construct the camera space relative position code;
[0008] Based on the positional relationship between the FOV feature map and the BEV feature map, and the spatial relative position encoding, the cross-attention results of the features in the BEV feature map and the corresponding candidate feature points are calculated;
[0009] Segment the video image video frame to obtain multiple time windows, generate the relative position encoding of the time sequence in the time window, and calculate the cross attention result between the features in the BEV feature map of the current frame and the previous frame;
[0010] Based on the calculated two cross attention results, BEV feature fusion is realized.
[0011] In combination with the first aspect, in one implementation, the step of performing feature expression on video images captured by multiple cameras on a vehicle to obtain a FOV feature map specifically includes:
[0012] Acquire video images captured by multiple cameras on the vehicle;
[0013] Based on the convolutional neural network, the video images captured by multiple cameras are respectively expressed with features, and the FOV feature map of the image is extracted.
[0014] In combination with the first aspect, in one implementation, the relative position index table is obtained according to the relative position matrix between cameras and the coordinate processing in the relative position matrix, wherein the coordinate processing in the relative position matrix is specifically as follows:
[0015] According to the shooting position parameters of each camera, the distance between each camera is calculated;
[0016] Obtain a set number of cameras, and based on the distances between the cameras, construct a relative position matrix for representing the distances between the cameras;
[0017] Obtain the elements with positive horizontal coordinates in the relative position matrix and sort the obtained elements, and replace the values of the elements with the arrangement numbers corresponding to the elements;
[0018] The elements with negative horizontal coordinates in the relative position matrix are obtained and the obtained elements are sorted, the values of the elements are replaced with the arrangement numbers corresponding to the elements, and the inverse number calculation is performed on the replaced element values.
[0019] In combination with the first aspect, in one implementation, the constructing of the camera space relative position code is specifically:
[0020] Get the relative position matrix after element replacement, and integerize the coordinates of each element to get an integer matrix;
[0021] Adjust the horizontal and vertical coordinates of each element of the integer matrix to obtain a relative position bias matrix;
[0022] A vector whose dimension is related to the number of selected cameras is generated as a relative position index table to realize the construction of the spatial relative position encoding of the cameras.
[0023] In combination with the first aspect, in one implementation, based on the positional relationship between the FOV feature map and the BEV feature map, and the spatial relative position encoding, the cross-attention result of the feature in the BEV feature map and the corresponding candidate feature point is calculated, specifically:
[0024] Based on the spatial correspondence between the elements in the BEV feature map and the elements in the FOV feature map, the corresponding candidate sampling points of the BEV feature points in the FOV feature map are obtained;
[0025] The constructed spatial relative position encoding is used as the relative position encoding of the corresponding candidate sampling point to perform cross attention calculation:
[0026]
[0027] in, Indicates that in Z j High Q p The cross attention calculation result of the corresponding candidate feature point, Q p represents the BEV feature at position p in the BEV feature map, DeformAttn represents the DeformableDETR algorithm, P(p,N p (i),Z j ) indicates Q p The corresponding candidate sampling point in the FOV feature map, i.e., Q p At height Z j On the i-th FOV feature map, the corresponding candidate feature point set, N p (i) represents the corresponding point set of position p on the i-th FOV feature map, P(p,Z j ) indicates that in Z j High Q p The corresponding feature point set in all FOV feature maps, N p Represents the corresponding point set of position p on the FOV feature map, R k Indicates Q p Corresponding to the relative position encoding of k views, Represents the BEV feature at position p in the BEV feature map of the current frame in the video image.
[0028] In combination with the first aspect, in one implementation, the video image video frame is segmented to obtain multiple time windows, and the relative position coding of the time sequence in the time window is generated, specifically:
[0029] In the continuous video frames of the video image, a set number of adjacent frames are selected as time sequence sliding windows, and the continuous video frames are segmented with a set step size to divide the video into a plurality of time sequence windows;
[0030] In the divided time series window, the time position of the BEV feature graph is assigned according to the time sequence, and according to the difference between the sequence numbers, a matrix for representing the difference between the time series is constructed;
[0031] Generate a temporal relative position code according to a method for generating a spatial relative position code;
[0032] According to the vehicle speed corresponding to the current video frame, the position offset between the BEV feature map of the next video frame and the BEV feature map of the current video frame is calculated, and the position offset is used to align the adjacent video frame images.
[0033] In combination with the first aspect, in one implementation, the calculation obtains the cross attention result between the features in the BEV feature map of the current frame and the previous frame, and the specific calculation method is:
[0034]
[0035] in, represents the BEV feature at position p in the BEV feature map of the previous frame, T k Relative position encoding that represents time series.
[0036] In combination with the first aspect, in one implementation, the BEV feature fusion is realized based on the calculated cross attention result, specifically:
[0037] BEV feature fusion is achieved based on the calculated cross-attention results of the features in the BEV feature map and the corresponding candidate feature points, as well as the calculated cross-attention results of the features in the BEV feature map of the current frame and the features in the BEV feature map of the previous frame.
[0038] In a second aspect, an embodiment of the present application provides a bird's-eye view feature generation device based on spatiotemporal attention, comprising:
[0039] An extraction module is used to express the features of the video images captured by multiple cameras on the vehicle to obtain an image FOV feature map;
[0040] A construction module is used to obtain a relative position index table according to the relative position matrix between cameras and the coordinates in the relative position matrix, and to construct a relative position code of the camera space;
[0041] A first calculation module, which is used to calculate the cross-attention results of the features in the BEV feature map and the corresponding candidate feature points based on the position relationship between the feature map FOV and the BEV feature map, and the spatial relative position encoding;
[0042] The second calculation module is used to segment the video image video frame to obtain multiple time windows, generate relative position coding of the time sequence in the time window, and calculate the cross attention result between the features in the BEV feature map of the current frame and the previous frame;
[0043] The fusion module is used to realize BEV feature fusion based on the two cross-attention results calculated.
[0044] In the third aspect, an embodiment of the present application provides a bird's-eye view feature generation device based on spatiotemporal attention, wherein the bird's-eye view feature generation device based on spatiotemporal attention comprises a processor, a memory, and a bird's-eye view feature generation program based on spatiotemporal attention stored in the memory and executable by the processor, wherein when the bird's-eye view feature generation program based on spatiotemporal attention is executed by the processor, the steps of the above-mentioned bird's-eye view feature generation method based on spatiotemporal attention are implemented.
[0045] The beneficial effects brought by the technical solution provided by the embodiments of the present application include:
[0046] (1) The calculation method proposed in this application can not only accurately express and fuse BEV features from the two dimensions of space and time, but also integrate the spatial position relationship between multiple cameras and the relative temporal relationship of video frames into BEV feature learning in the form of encoding, further improving the accuracy and semantic richness of BEV feature expression;
[0047] (2) The relative position encoding method proposed in this application can be extended to a variety of cutting-edge algorithm models and can effectively improve model performance at the cost of increasing minimal computational overhead. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] Figure 1 A flowchart of a method for generating bird's-eye view features based on spatiotemporal attention for this application;
[0049] Figure 2 This is a structural schematic diagram of a bird's-eye view feature generation device based on spatiotemporal attention in this application;
[0050] Figure 3 This is a schematic diagram of the hardware structure of the bird's-eye view feature generation device based on spatiotemporal attention in this application. DETAILED DESCRIPTION
[0051] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0052] In order to make the objectives, technical solutions and advantages of the present application clearer, the implementation methods of the present application will be further described in detail below in conjunction with the accompanying drawings.
[0053] On the first aspect, an embodiment of the present application provides a bird's-eye view feature generation method based on spatiotemporal attention to solve the problem of unified learning / expression of bird's-eye view features in multi-perspective image input scenarios. First, a pre-trained feature learning network is used to extract features of images from different perspectives to obtain feature maps of the image at a conventional perspective. The feature maps are then put into BEV spatial feature learning and BEV temporal feature learning respectively to obtain two types of features with the same dimensions. Finally, the above features are added in corresponding dimensions to obtain the final feature map.
[0054] Reference Figure 1 , Figure 1 Schematic diagram of the process of generating bird's-eye view features based on spatiotemporal attention in this application. Figure 1 As shown, the bird's-eye view feature generation method based on spatiotemporal attention includes:
[0055] S1: Perform feature expression on the video images captured by multiple cameras on the vehicle to obtain a FOV feature map; that is, obtain video images captured by multiple cameras on the same vehicle, and perform feature expression on the acquired images through the FOV viewing angle feature extraction network to map the images, thereby realizing the extraction of the image FOV feature map.
[0056] S2: A relative position index table is obtained according to the relative position matrix between cameras and the coordinates in the relative position matrix, and the camera space relative position encoding is constructed; specifically, the relative position index table is obtained by calculating the relative position matrix between cameras and processing the non-negative coordinates in the relative position matrix and the negative coordinates in the relative position matrix, thereby realizing the construction of the camera space relative position encoding.
[0057] S3: Based on the position relationship between the FOV feature map and the BEV feature map, and the spatial relative position encoding, the cross-attention result of the feature in the BEV feature map and the corresponding candidate feature point is calculated;
[0058] That is, based on the positional relationship between the BEV feature map and the extracted FOV feature map, and the constructed spatial relative position encoding, the cross-attention results of the features in the BEV feature map and the corresponding candidate feature points are obtained through cross-attention calculation.
[0059] S4: Segment the video image video frame to obtain multiple time windows, generate the relative position encoding of the time sequence in the time window, and calculate the cross attention result between the features in the BEV feature map of the current frame and the previous frame;
[0060] Specifically, the video frames of the video image are segmented to obtain multiple time windows, and the relative position encoding of the time sequence in the time window is generated, and through cross-attention calculation, the cross-attention results of the features in the BEV feature map of the current frame and the features in the BEV feature map of the previous frame are obtained.
[0061] S5: Based on the two calculated cross-attention results, BEV feature fusion is realized. Specifically, BEV feature fusion is realized based on the calculated cross-attention results of the features in the BEV feature map and the corresponding candidate feature points, and the calculated cross-attention results of the features in the BEV feature map of the current frame and the features in the BEV feature map of the previous frame.
[0062] Furthermore, in one embodiment, the video images captured by multiple cameras on the vehicle are characterized to obtain a FOV feature map, and the specific steps include:
[0063] S101: Obtain video images captured by multiple cameras on a vehicle; specifically, obtain video images captured by n cameras on the same vehicle, where the image captured by the i-th camera at the j-th frame is represented by a tensor as Right now Among them, H represents the height of the image, and W represents the width of the image.
[0064] S102: Based on a convolutional neural network, feature expressions are performed on the video images captured by multiple cameras, and FOV feature maps of the images are extracted. Specifically, a Resnet50 network (a convolutional neural network with a depth of 50 layers) pre-trained on the ImageNet dataset (an image classification dataset) is used as a FOV (Field of view) perspective feature extraction network, and feature expressions are performed on the images in the video images captured by n cameras, and the images are mapped to Right now Among them, h and w represent the spatial scale of the feature map, and C represents the number of channels of the feature map.
[0065] Furthermore, in one embodiment, a relative position index table is obtained according to the relative position matrix between cameras and the coordinate processing in the relative position matrix, wherein the coordinate processing in the relative position matrix is specifically as follows:
[0066] S201: Calculate the distance between each camera according to the shooting position parameters of each camera;
[0067] S202: Acquire a set number of cameras, and construct a relative position matrix for representing the distance between the cameras based on the distance between the cameras;
[0068] That is, the construction of the relative position matrix in this application is as follows: according to the shooting position parameters of each camera, the distance d between each camera is calculated. ij , where d ij The horizontal axis represents the horizontal difference vector between the i-th camera and the j-th camera, d ij The vertical coordinate represents the vertical difference vector between the i-th camera and the j-th camera, and then according to d ij , randomly select k cameras from n cameras to form a k×k relative position matrix.
[0069] S203: Obtain elements with positive horizontal coordinates in the relative position matrix and sort the obtained elements, and replace the values of the elements with the arrangement numbers corresponding to the elements;
[0070] That is, the processing of the non-negative coordinates of the relative position matrix in the present invention is specifically as follows: for all elements in the k×k relative position matrix, obtain the elements whose horizontal coordinates are greater than 0 and sort them in ascending order, and replace the values of the elements with the arrangement numbers corresponding to the elements. In any k×k relative position matrix, the diagonal elements are (0,0). For example, for {2.3, 1.2, 4.6, 1.2}, after the above-mentioned processing of the non-negative coordinates of the relative position matrix, it becomes {2, 1, 3, 1}.
[0071] S204: Obtain elements with negative horizontal coordinates in the relative position matrix and sort the obtained elements, replace the values of the elements with the arrangement numbers corresponding to the elements, and perform inverse number calculation on the replaced element values.
[0072] That is, the processing of negative coordinates of the relative position matrix in the present invention is specifically as follows: for all elements in the k×k relative position matrix, obtain the elements whose horizontal coordinates are less than 0 and sort them in ascending order, replace the values of the elements with the arrangement numbers corresponding to the elements, and multiply the replaced element values by -1. For example, for {-2.3, -1.2, -4.6, -1.2}, after the above processing of negative coordinates of the relative position matrix, it becomes {-2, -1, -3, -1}.
[0073] Furthermore, in one embodiment, the camera space relative position encoding is constructed, specifically:
[0074] S211: Obtain the relative position matrix after element replacement, and perform integer processing on the coordinates of each element to obtain an integer matrix;
[0075] Specifically, for the k×k relative position matrix after element value replacement, the horizontal and vertical coordinates of each element are converted into integers to obtain the integer matrix K; that is, according to the above steps of processing the non-negative coordinates of the relative position matrix and processing the negative coordinates of the relative position matrix, the horizontal and vertical coordinates of the k×k relative position matrix are converted into integers to obtain the integer matrix K.
[0076] S212: adjusting the horizontal and vertical coordinates of each element of the integer matrix to obtain a relative position bias matrix;
[0077] Specifically, add the horizontal and vertical coordinates of each element in the integer matrix K to Then multiply the horizontal coordinate of each element by k-1, and finally add the horizontal and vertical coordinates of each element to get the relative position bias matrix K b .
[0078] S213: Generate a vector whose dimension is related to the number of selected cameras as a relative position index table to realize the construction of spatial relative position coding of the cameras.
[0079] Specifically, generate a 2n-1 dimensional vector V with an initial value of 0 table , as the relative position index table, and the relative position index table and the relative position bias matrix K b The values in V correspond to each other, realizing the construction of the spatial relative position encoding of the camera. table is a learnable vector whose vector value is automatically adjusted during the training process.
[0080] The above steps realize the relative position encoding of the camera space. For the generated k×k relative position matrix, there is a corresponding k×k matrix R k , R k Element R ij Represents the bias of the image in the i-th camera relative to the image captured by the j-th camera. It is used in the subsequent calculation of the cross-attention mechanism, and the usage method is consistent with the relative position encoding method in Cross-Attention (In-depth Understanding of Cross-Attention Mechanism).
[0081] Furthermore, in one embodiment, based on the positional relationship between the FOV feature map and the BEV feature map, and the spatial relative position encoding, the cross-attention result of the feature in the BEV feature map and the corresponding candidate feature point is calculated, specifically:
[0082] S301: Based on the spatial correspondence between the elements in the BEV feature map and the elements in the FOV feature map, obtaining the corresponding candidate sampling points of the BEV feature points in the FOV feature map;
[0083] S302: Using the constructed spatial relative position code as the relative position code of the corresponding candidate sampling point, and performing cross attention calculation:
[0084]
[0085] in, Indicates that in Z j High Q p The cross attention calculation result of the corresponding candidate feature point, Q p represents the BEV feature at position p in the BEV feature map, DeformAttn represents the DeformableDETR algorithm, P(p,N p (i),Z j ) indicates Q p The corresponding candidate sampling point in the FOV feature map, i.e., Q p At height Z j On the i-th FOV feature map, the corresponding candidate feature point set, N p (i) represents the corresponding point set of position p on the i-th FOV feature map, P(p,Z j ) indicates that in Z j High Q p The corresponding feature point set in all FOV feature maps, N p Represents the corresponding point set of position p on the FOV feature map, R k Indicates Q p Corresponding to the relative position encoding of k views, Represents the BEV feature at position p in the BEV feature map of the current frame in the video image.
[0086] The following is a detailed description of the cross-attention calculation of the features and the corresponding candidate feature points in the BEV feature map:
[0087] a: Based on the spatial correspondence between the elements in the BEV feature map and the elements in the FOV feature map, obtain the BEV feature point Q p The corresponding candidate sampling point P(p,N p (i),Z j ), where Qp is the BEV feature at position p in the BEV feature map, P is Q p The corresponding set of candidate feature points, N p (i) represents the corresponding point set of position p on the i-th FOV feature map. Specifically, P(p,N p (i),Z j ) indicates Q p At height Z j On the i-th FOV feature map, the corresponding candidate feature point set is P(p,Z j ) indicates that in Z j High Q p The corresponding feature point set in all FOV feature maps; the present invention uses the spatial point sampling method of spatial cross-attention in the BEVFormer algorithm;
[0088] b: For any video frame l in the video image, a There are candidate feature points corresponding to one or more FOV feature maps, set There are corresponding candidate feature points in the k FOV feature maps, and the constructed spatial relative position encoding is used as P(p,N p (i),Z j ) to perform cross-attention calculation:
[0089]
[0090] c:Q p Final calculation results for:
[0091]
[0092] Among them, L represents the total number of selected heights and is a hyperparameter.
[0093] Furthermore, in one embodiment, the video image and video frame are segmented to obtain a plurality of time windows, and the relative position codes of the time in the time windows are generated, specifically:
[0094] S401: selecting a set number of adjacent frames from continuous video frames of a video image as time sequence sliding windows, segmenting the continuous video frames with a set step length, and dividing the video into a plurality of time sequence windows;
[0095] Specifically, in the video frame of the 0:T frame of the video image, t adjacent frames are selected as the time sequence sliding window, and the video frame of the 0:T frame is segmented with a step size of t, and the video is divided into A timing window;
[0096] S402: in the divided time series window, assigning a time position to the BEV feature graph according to the time sequence, and constructing a matrix for representing the difference between the time series according to the difference between the sequence numbers;
[0097] Specifically, in any time series window, the BEV feature graph is assigned a time position {1, 2, L, t} in chronological order, and a t×t matrix t is constructed according to the difference between the sequence numbers. ij , t ij Represents the difference between the i-th time series and the j-th time series.
[0098] S403: Generate a temporal relative position code according to a method for generating a spatial relative position code;
[0099] Specifically, based on the generation method of spatial relative position coding, the relative position coding of the time series is generated, which is recorded as T k ; That is, the relative position code of the timing is generated according to steps S212 and S213.
[0100] S404: Calculate the position offset between the BEV feature map of the next video frame and the BEV feature map of the current video frame according to the vehicle speed corresponding to the current video frame, and use the position offset to align adjacent video frame images to align t frame video frame images.
[0101] Furthermore, in one embodiment, the cross attention result between the features in the BEV feature map of the current frame and the previous frame is calculated, and the specific calculation method is:
[0102]
[0103] in, It represents the BEV feature at position p in the BEV feature map of the current frame, that is, the BEV feature at position p in the BEV feature map of video frame l is recorded as represents the BEV feature at position p in the BEV feature map of the previous frame, that is, the BEV feature at position p in the BEV feature map of the previous t-1 frame is recorded as Where, i = 0, 1, L t-1, T k Relative position encoding that represents time series.
[0104] It should be noted that the BEV feature fusion is implemented in this application, specifically by and The result of adding is the BEV feature of the lth video frame.
[0105] because and have the same data dimension, and according to the timing information and spatial position relationship, the corresponding elements of the two can be added at the same position in the same frame. When p traverses all positions, and The sum is the BEV feature of the lth video frame.
[0106] This application makes full use of the position distribution information of the camera and proposes a spatial attention calculation method based on camera extrinsics to assist in image spatial feature extraction. At the same time, it fully considers the feature distribution characteristics of image features in adjacent time series and proposes a temporal attention calculation method based on local time series distribution to achieve accurate expression of image time domain features. Compared with existing technologies and methods, the present invention has the following advantages:
[0107] (1) Strong generalization ability, suitable for a variety of BEV feature learning models. The spatial attention calculation method and temporal attention calculation method proposed in the present invention are applicable to most current BEV feature learning frameworks, especially the spatial attention calculation method, which can be directly embedded in most existing BEV feature learning modules.
[0108] (2) Make full use of relative position / time distribution information. Most existing BEV feature learning models ignore the relative shooting position information of images under different viewing angles and the relative distribution characteristics of visual semantic features under fixed time sequence, which results in such models being unable to make full use of the two types of attention mechanisms to improve the expressive power of BEV features. The two types of attention mechanisms proposed in this invention just use the above two important types of information, which can effectively improve the expressive power of BEV features.
[0109] (3) Low computing power consumption. The two types of calculation methods proposed in the present invention can be embedded in multiple BEV feature learning frameworks in the form of additional modules. However, both types of calculation methods draw on the idea of relative position encoding and can improve feature expression capabilities at a very low computing cost.
[0110] In a second aspect, an embodiment of the present application also provides a bird's-eye view feature generation device based on spatiotemporal attention.
[0111] In one embodiment, referring to Figure 2 , Figure 2 This is a functional module diagram of the device for generating bird's-eye view features based on spatiotemporal attention in this application. As shown in the figure, the device for generating bird's-eye view features based on spatiotemporal attention includes an extraction module, a construction module, a first calculation module, a second calculation module and a fusion module.
[0112] The extraction module is used to express the features of the video images taken by multiple cameras on the vehicle to obtain the image FOV feature map; the construction module is used to obtain the relative position index table according to the relative position matrix between cameras and the coordinate processing in the relative position matrix, and construct the camera space relative position coding; the first calculation module is used to calculate the cross-attention results of the features and the corresponding candidate feature points in the BEV feature map based on the position relationship between the feature map FOV and the BEV feature map, and the spatial relative position coding; the second calculation module is used to segment the video image video frame to obtain multiple time windows, and generate the relative position coding of the time sequence in the time window, and calculate the cross-attention results between the features in the BEV feature map of the current frame and the previous frame; the fusion module is used to realize BEV feature fusion based on the two cross-attention results calculated.
[0113] On the third aspect, an embodiment of the present application provides a bird's-eye view feature generation device based on spatiotemporal attention. The bird's-eye view feature generation device based on spatiotemporal attention can be a personal computer (PC), a laptop computer, a server, or other device with data processing capabilities.
[0114] Reference Figure 3 , Figure 3 The hardware structure diagram of the device for generating bird's-eye view features based on spatiotemporal attention involved in the embodiment of the present application is shown in FIG. In the embodiment of the present application, the device for generating bird's-eye view features based on spatiotemporal attention may include a processor, a memory, a communication interface, and a communication bus.
[0115] The communication bus may be of any type and is used to interconnect the processor, the memory, and the communication interface.
[0116] The communication interface includes an input / output (I / O) interface, a physical interface, and a logical interface, etc., which are used to realize the interconnection of devices inside the spatiotemporal attention-based bird's-eye view feature generation device, and an interface for realizing the interconnection of the spatiotemporal attention-based bird's-eye view feature generation device with other devices (such as other computing devices or user devices). The physical interface can be an Ethernet interface, an optical fiber interface, an ATM interface, etc.; the user device can be a display (Display), a keyboard (Keyboard), etc.
[0117] The memory can be various types of storage media, such as random access memory (RAM), read-only memory (ROM), non-volatile RAM (NVRAM), flash memory, optical storage, hard disk, programmable ROM (PROM), erasable PROM (EPROM), electrically erasable PROM (EEPROM), etc.
[0118] The processor may be a general-purpose processor, which may call a bird's-eye view feature generation program based on spatiotemporal attention stored in a memory, and execute the bird's-eye view feature generation method based on spatiotemporal attention provided in an embodiment of the present application. For example, the general-purpose processor may be a central processing unit (CPU). The method executed when the bird's-eye view feature generation program based on spatiotemporal attention is called may refer to the various embodiments of the bird's-eye view feature generation method based on spatiotemporal attention of the present application, which will not be repeated here.
[0119] Those skilled in the art will understand that Figure 3 The hardware structure shown in the figure does not constitute a limitation on the present application, and may include more or less components than shown in the figure, or combine certain components, or arrange the components differently.
[0120] The terms "including" and "having" and any variations thereof in the specification and claims of this application and the above-mentioned drawings are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but optionally includes steps or units that are not listed, or optionally includes other steps or units inherent to these processes, methods, products or devices. The terms "first", "second" and "third" are used to distinguish different objects, etc., and do not represent a sequence, nor do they limit "first", "second" and "third" to different types.
[0121] In the description of the embodiments of the present application, "exemplary", "for example" or "for example" are used to indicate examples, illustrations or descriptions. Any embodiment or design described as "exemplary", "for example" or "for example" in the embodiments of the present application should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Specifically, the use of words such as "exemplary", "for example" or "for example" is intended to present related concepts in a specific way.
[0122] In the description of the embodiments of the present application, unless otherwise specified, “ / ” means or, for example, A / B can mean A or B; the “and / or” in the text is merely a description of the association relationship of associated objects, indicating that three relationships may exist, for example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. In addition, in the description of the embodiments of the present application, “multiple” refers to two or more than two.
[0123] In some processes described in the embodiments of the present application, multiple operations or steps that appear in a specific order are included, but it should be understood that these operations or steps may not be executed in the order in which they appear in the embodiments of the present application or in parallel, and the sequence number of the operation is only used to distinguish the different operations, and the sequence number itself does not represent any execution order. In addition, these processes may include more or fewer operations, and these operations or steps may be executed in sequence or in parallel, and these operations or steps may be combined.
[0124] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present application is essentially or the part that contributes to the prior art can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, disk, CD) as described above, and includes a number of instructions for a terminal device to execute the methods described in each embodiment of the present application.
[0125] The above are only preferred embodiments of the present application, and are not intended to limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made using the contents of the present application specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present application.
Claims
1. A method for generating bird's-eye view features based on spatiotemporal attention, characterized in that: The method for generating bird's-eye view features based on spatiotemporal attention includes: The video images captured by multiple cameras on the vehicle are characterized to obtain a FOV feature map; According to the relative position matrix between cameras and the coordinates in the relative position matrix, a relative position index table is obtained to construct the camera space relative position code; Based on the positional relationship between the FOV feature map and the BEV feature map, and the spatial relative position encoding, the cross-attention results of the features in the BEV feature map and the corresponding candidate feature points are calculated; Segment the video image video frame to obtain multiple time windows, generate the relative position encoding of the time sequence in the time window, and calculate the cross attention result between the features in the BEV feature map of the current frame and the previous frame; Based on the calculated two cross attention results, BEV feature fusion is realized; Among them, based on the position relationship between the FOV feature map and the BEV feature map, and the spatial relative position encoding, the cross-attention result of the feature in the BEV feature map and the corresponding candidate feature point is calculated, specifically: Based on the spatial correspondence between the elements in the BEV feature map and the elements in the FOV feature map, the corresponding candidate sampling points of the BEV feature points in the FOV feature map are obtained; The constructed spatial relative position encoding is used as the relative position encoding of the corresponding candidate sampling point to perform cross attention calculation: in, Indicates that in Z j High Q p The cross attention calculation result of the corresponding candidate feature point, Q p represents the BEV feature at position p in the BEV feature map, DeformAttn represents the DeformableDETR algorithm, P(p,N p (i),Z j ) indicates Q p The corresponding candidate sampling point in the FOV feature map, i.e., Q p At height Z j On the i-th FOV feature map, the corresponding candidate feature point set, N p (i) represents the corresponding point set of position p on the i-th FOV feature map, P(p,Z j ) indicates that in Z j High Q p The corresponding feature point set in all FOV feature maps, N p Represents the corresponding point set of position p on the FOV feature map, R k Indicates Q p Corresponding to the relative position encoding of k views, Represents the BEV feature at position p in the BEV feature map of the current frame in the video image.
2. A method for generating bird's-eye view features based on spatiotemporal attention as claimed in claim 1, characterized in that: The step of performing feature expression on the video images captured by multiple cameras on the vehicle to obtain the FOV feature map specifically includes: Acquire video images captured by multiple cameras on the vehicle; Based on the convolutional neural network, the video images captured by multiple cameras are respectively expressed with features, and the FOV feature map of the image is extracted.
3. The method for generating bird's-eye view features based on spatiotemporal attention according to claim 1, characterized in that: The relative position index table is obtained according to the relative position matrix between cameras and the coordinate processing in the relative position matrix, wherein the coordinate processing in the relative position matrix is specifically as follows: According to the shooting position parameters of each camera, the distance between each camera is calculated; Obtain a set number of cameras, and based on the distances between the cameras, construct a relative position matrix for representing the distances between the cameras; Obtain the elements with positive horizontal coordinates in the relative position matrix and sort the obtained elements, and replace the values of the elements with the arrangement numbers corresponding to the elements; The elements with negative horizontal coordinates in the relative position matrix are obtained and the obtained elements are sorted, the values of the elements are replaced with the arrangement numbers corresponding to the elements, and the inverse number calculation is performed on the replaced element values.
4. A method for generating bird's-eye view features based on spatiotemporal attention as claimed in claim 3, characterized in that: The construction of the camera space relative position encoding is specifically as follows: Get the relative position matrix after element replacement, and integerize the coordinates of each element to get an integer matrix; Adjust the horizontal and vertical coordinates of each element of the integer matrix to obtain a relative position bias matrix; A vector whose dimension is related to the number of selected cameras is generated as a relative position index table to realize the construction of the spatial relative position encoding of the cameras.
5. A method for generating bird's-eye view features based on spatiotemporal attention as claimed in claim 4, characterized in that: The video image video frame is segmented to obtain multiple time windows, and the relative position coding of the time sequence in the time window is generated, specifically: In the continuous video frames of the video image, a set number of adjacent frames are selected as time sequence sliding windows, and the continuous video frames are segmented with a set step size to divide the video into a plurality of time sequence windows; In the divided time series window, the time position of the BEV feature graph is assigned according to the time sequence, and according to the difference between the sequence numbers, a matrix for representing the difference between the time series is constructed; Generate a temporal relative position code according to a method for generating a spatial relative position code; According to the vehicle speed corresponding to the current video frame, the position offset between the BEV feature map of the next video frame and the BEV feature map of the current video frame is calculated, and the position offset is used to align the adjacent video frame images.
6. A method for generating bird's-eye view features based on spatiotemporal attention as claimed in claim 5, characterized in that: The calculation obtains the cross attention result between the features in the BEV feature map of the current frame and the previous frame. The specific calculation method is: in, represents the BEV feature at position p in the BEV feature map of the previous frame, T k Relative position encoding that represents time series.
7. The method for generating bird's-eye view features based on spatiotemporal attention according to claim 1, characterized in that: The cross attention result obtained based on the calculation is used to realize BEV feature fusion, specifically: BEV feature fusion is achieved based on the calculated cross-attention results of the features in the BEV feature map and the corresponding candidate feature points, as well as the calculated cross-attention results of the features in the BEV feature map of the current frame and the features in the BEV feature map of the previous frame.
8. A bird's-eye view feature generation device based on spatiotemporal attention, characterized in that: include: An extraction module is used to express the features of the video images captured by multiple cameras on the vehicle to obtain an image FOV feature map; A construction module is used to obtain a relative position index table according to the relative position matrix between cameras and the coordinates in the relative position matrix, and to construct a relative position code of the camera space; A first calculation module, which is used to calculate the cross-attention results of the features in the BEV feature map and the corresponding candidate feature points based on the position relationship between the feature map FOV and the BEV feature map, and the spatial relative position encoding; The second calculation module is used to segment the video image video frame to obtain multiple time windows, generate relative position coding of the time sequence in the time window, and calculate the cross attention result between the features in the BEV feature map of the current frame and the previous frame; A fusion module is used to realize BEV feature fusion based on the two cross-attention results calculated; Among them, based on the position relationship between the FOV feature map and the BEV feature map, and the spatial relative position encoding, the cross-attention result of the feature in the BEV feature map and the corresponding candidate feature point is calculated, specifically: Based on the spatial correspondence between the elements in the BEV feature map and the elements in the FOV feature map, the corresponding candidate sampling points of the BEV feature points in the FOV feature map are obtained; The constructed spatial relative position encoding is used as the relative position encoding of the corresponding candidate sampling point to perform cross attention calculation: in, Indicates that in Z j High Q p The cross attention calculation result of the corresponding candidate feature point, Q p represents the BEV feature at position p in the BEV feature map, DeformAttn represents the DeformableDETR algorithm, P(p,N p (i),Z j ) indicates Q p The corresponding candidate sampling point in the FOV feature map, i.e., Q p At height Z j On the i-th FOV feature map, the corresponding candidate feature point set, N p (i) represents the corresponding point set of position p on the i-th FOV feature map, P(p,Z j ) indicates that in Z j High Q p The corresponding feature point set in all FOV feature maps, N p Represents the corresponding point set of position p on the FOV feature map, R k Indicates Q p Corresponding to the relative position encoding of k views, Represents the BEV feature at position p in the BEV feature map of the current frame in the video image.
9. A bird's-eye view feature generation device based on spatiotemporal attention, characterized in that: The spatiotemporal attention-based bird's-eye view feature generation device comprises a processor, a memory, and a spatiotemporal attention-based bird's-eye view feature generation program stored in the memory and executable by the processor, wherein the spatiotemporal attention-based bird's-eye view feature generation program, when executed by the processor, implements the steps of the spatiotemporal attention-based bird's-eye view feature generation method as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Automatic driving visual perception feature extraction method and device
CN116259025A
3D target detection algorithm based on combination of time attention and deformable cross attention
CN116977623A