Image Processing Method, Readable Storage Medium, Program Product, and Vehicle-mounted Device
By sampling and feature extraction of multiple sampling rates for multi-view images in vehicle-mounted equipment and fusing feature data of feature maps of different sizes, the problem of insufficient BEV space in the prior art is solved, and more accurate detection of the vehicle's surrounding environment is achieved.
Patent Information
- Application Number
- CN202510501624.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-21
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-04-21
AI Technical Summary
The prior art fails to fully utilize the features of the multi-view image when projecting a multi-view image onto a BEV space, resulting in the obtained BEV space being inaccurate enough.
By sampling multiple images at multiple sampling rates in the on-board device, multiple feature maps of different sizes are generated. Then, the first feature maps of these feature maps are projected into the setting plane, feature extraction is performed, and the first feature data is obtained. At the same time, feature extraction is performed on the feature map other than the first feature map to obtain the second feature data. The first characteristic data and the second characteristic data are fused to generate the third characteristic data for detecting the surrounding environment of the vehicle.
By fusing the feature information of multi-view images at different sampling rates, the generated third feature data can retain more information, thereby improving the accuracy of vehicle environment detection and enhancing the safety of vehicle driving.
Smart Images

Figure CN120014605B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing technologies, and in particular, to an image processing method, a readable storage medium, a program product, and a vehicle-mounted device. Background Art
[0002] Multi-view image three-dimensional object detection is currently widely used in many fields and application scenarios, such as autonomous driving scenarios, intelligent transportation scenarios, industrial automation, and virtual reality and augmented reality scenarios.
[0003] For example, in assisted driving technology, through the bird's eye view (BEV) technology, an image of the multi-view image features of the vehicle's surrounding environment projected into the BEV space can be provided, and then the global view of the positions, sizes, and attributes of dynamic obstacles can be estimated, thus facilitating users to perform path planning and decision-making.
[0004] However, currently, in the process of projecting multi-view images into the BEV space, the features of the multi-view images are not fully utilized, resulting in an inaccurate BEV space obtained based on the multi-view images. Summary of the Invention
[0005] Embodiments of this application provide an image processing method, a readable storage medium, a program product, and a vehicle-mounted device.
[0006] In a first aspect, embodiments of this application provide an image processing method, which is applied to a vehicle-mounted device on a vehicle. The method includes: the vehicle-mounted device obtains multiple images in response to a request for detecting the surrounding environment of the vehicle, where the multiple images are images of different perspectives collected by multiple sensors of the vehicle. The vehicle-mounted device samples the multiple images based on multiple sampling rates to obtain multiple feature maps of each image, and the sizes of the multiple feature maps are different. The vehicle-mounted device projects a first feature map among the multiple feature maps into a set plane to obtain a first plane image. The vehicle-mounted device extracts features from the first plane image to obtain first feature data. The vehicle-mounted device extracts features from the feature maps other than the first feature map among the multiple feature maps to obtain second feature data. The vehicle-mounted device fuses the first feature data and the second feature data to obtain third feature data. Then, the vehicle-mounted device detects the surrounding environment of the vehicle based on the third feature data.
[0007] In some embodiments of the present application, after the in-vehicle device obtains multiple images of different perspectives of the vehicle, each image can be sampled based on multiple sampling rates, so that each image has feature maps of multiple sizes. For example, the original image can be downsampled by 4 times, downsampled by 32 times, and downsampled by 64 times, etc. In the process of generating the first environmental data, a first planar image can be generated using a first feature map of one size. The first feature map is, for example, the feature map corresponding to a 4-fold downsampling, and the first planar image is, for example, the initial BEV space image. Then, the in-vehicle device can extract the feature data of the first planar image and the second feature data of the feature maps other than the first feature map among the feature maps of multiple sizes, and fuse the first feature data and the second feature data to obtain the third feature data. It can be understood that the third feature data includes information sampled from multiple images at different sampling rates. Therefore, the third feature data can retain more information in multiple images, so that the in-vehicle device can detect the surrounding environment of the vehicle more accurately based on the third feature data. For example, the in-vehicle device can perform more accurate detection of pedestrians, other vehicles, lane lines, etc. around the vehicle based on the third feature data, thereby improving the safety of vehicle driving.
[0008] In a possible implementation of the first aspect above, projecting the first feature map among the multiple feature maps into a set plane to obtain a first planar image includes: The in-vehicle device obtains the acquisition parameters of multiple sensors, and based on the acquisition parameters, performs an inverse perspective transformation on the first feature map to generate a first planar image.
[0009] In some embodiments of the present application, the process of projecting the first feature map is, for example, to project the first feature map into the BEV space through an inverse perspective transformation to obtain the feature map of the initial BEV space.
[0010] In a possible implementation of the first aspect above, extracting the feature maps other than the first feature map among the multiple feature maps to obtain the second feature data includes: The in-vehicle device aligns the feature maps of different sampling rates among the feature maps other than the first feature map among the multiple feature maps to obtain a second feature map. The in-vehicle device obtains the acquisition parameters of multiple sensors, and encodes the acquisition parameters based on the size of the second feature map to obtain sensor features. The in-vehicle device extracts the feature data of the second feature map to obtain fourth feature data. The in-vehicle device fuses the fourth feature data and the sensor features to determine the second feature data.
[0011] In some embodiments of the present application, the in-vehicle device can align feature maps with different sampling rates in the feature maps other than the first feature map through a pyramid feature network to generate a second feature map. It can be understood that the sizes of the respective second feature maps are the same. During the generation of the second feature data, the in-vehicle device can also fuse the encoding corresponding to the acquisition parameters in the vehicle sensors with the fourth feature data extracted from the second feature map to obtain the second feature data. It can be understood that the second feature data also includes the acquisition data of the sensors, so as to obtain more information about the image acquired by the sensor.
[0012] In a possible implementation of the above first aspect, the above-mentioned multiple sensors include a camera, and the acquisition parameters include the internal parameters and external parameters of the camera.
[0013] It can be understood that in some embodiments of the present application, the sensor can be a camera, and the acquisition parameters of the sensor include the internal parameters and external parameters of the camera. The internal parameters can include, for example, the focal length, distortion coefficient, etc. The external parameters can include the position and orientation of the camera, etc. The internal parameters and external parameters of the camera are used to describe the internal characteristics of the camera and the relative spatial information with the vehicle origin coordinates externally.
[0014] In a possible implementation of the above first aspect, the above-mentioned fusion of the first feature data and the second feature data to obtain the third feature data includes: the in-vehicle device fuses the first feature data and the second feature data based on the self-attention mechanism to obtain the third feature data.
[0015] In some embodiments of the present application, the fusion of the first feature data and the second feature data through the self-attention mechanism can enable the self-attention mechanism to have a global view, capture the connections between different positions in the input features, and thus consider global information to improve the accuracy of the third feature data.
[0016] In a possible implementation of the above first aspect, the fusion of the first feature data and the second feature data based on the self-attention mechanism includes: The first feature data includes N partial regions, each region includes N partial first sub-features, and the positions of the N partial first sub-features in the same region are different. The second feature data includes a first key feature, where the first key feature includes N partial regions, each region includes N partial second sub-features, and the positions of the N partial second sub-features in the same region are different. The second feature data includes a first value feature, where the first value feature includes N partial regions, each region includes N partial third sub-features, and the positions of the N partial third sub-features in the same region are different. The vehicle-mounted device fuses the N partial first sub-features, the N partial second sub-features, and the N partial second sub-features in the same region based on the self-attention mechanism to obtain a first query feature, and determines the third feature data based on the first query feature, where N is greater than or equal to 2.
[0017] In some embodiments of the present application, in order to improve the speed of obtaining the third feature data, during the fusion of the first feature data and the second feature data, the first feature data and the second feature data can be fused in regions, so as to improve the speed of the self-attention mechanism fusion.
[0018] In a possible implementation of the above first aspect, the above first query feature includes N partial regions, each region includes N partial fourth sub-features, and the positions of the N partial fourth sub-features in the same region are different. Determining the third feature data based on the second query feature includes: The vehicle-mounted device fuses the N partial fourth sub-features, the N partial second sub-features, and the N partial third sub-features at the same positions in N different regions based on the self-attention mechanism to obtain the third feature data.
[0019] In some embodiments of the present application, in order to ensure the effect of the self-attention mechanism in fusing the first feature data and the second feature data, the result of the first fusion of the first feature data and the second feature data based on the self-attention mechanism can be used as new data, that is, the first query data, and then cross-fused with the second feature data based on the self-attention mechanism again, and the corresponding partitions are different during the cross-fusion process. In this way, the accuracy of obtaining the third feature data through the self-attention mechanism can be further improved, so that the third feature data retains more information.
[0020] In a second aspect, the present application provides a vehicle-mounted device, including: a memory for storing instructions; at least one processor for executing the instructions to enable the device to implement the method provided by the above first aspect and any possible implementation of the above first aspect. The beneficial effects that can be achieved by the second aspect can refer to the beneficial effects of the method provided by any implementation manner of the first aspect, which will not be elaborated here.
[0021] In a third aspect, the present application provides a computer-readable storage medium storing instructions which, when executed by a device, cause a computer to implement the method provided in the above first aspect and any possible implementation of the first aspect. The beneficial effects achievable in the third aspect can be referred to the beneficial effects of the method provided in any embodiment of the first aspect, which will not be elaborated here.
[0022] In a fourth aspect, the present application provides a computer program product which, when running on a device, causes the device to implement the method provided in the above first aspect and any possible implementation of the first aspect. The beneficial effects achievable in the fourth aspect can be referred to the beneficial effects of the method provided in any embodiment of the first aspect, which will not be elaborated here. Description of the Drawings
[0023] Figure 1A Shows a flowchart for generating a BEV space;
[0024] Figure 1B Shows a schematic diagram during vehicle driving;
[0025] Figure 1C Shows a flowchart for generating a BEV space based on IPM;
[0026] Figure 2 Shows an implementation flowchart of an image processing method according to an embodiment of the present application;
[0027] Figure 3 According to some embodiments of the present application, shows a schematic diagram for generating third feature data;
[0028] Figure 4 According to some embodiments of the present application, shows a schematic diagram of a cross-attention mechanism;
[0029] Figure 5A According to some embodiments of the present application, shows a schematic diagram of local window partitioning;
[0030] Figure 5B According to an embodiment of the present application, shows a schematic diagram of global window partitioning;
[0031] Figure 6 According to some embodiments of the present application, shows a schematic diagram of the structure of an in-vehicle device. Detailed Embodiments
[0032] The illustrative embodiments of the present application include but are not limited to image processing methods, readable storage media, program products, and in-vehicle devices.
[0033] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions in the embodiments of this application will be described in detail below in conjunction with the accompanying drawings of the specification and specific implementation manners.
[0034] For ease of understanding, some terms and related technologies involved in this application are explained below.
[0035] 1. BEV technology:
[0036] BEV technology is a technology that converts image information from the image space to the BEV space by a neural network. Through BEV technology, a complex three-dimensional environment can be simplified into a two-dimensional image. For example, in the field of autonomous driving, through BEV technology, a panoramic view looking down from above the vehicle can be generated based on the spatial information around the vehicle, thereby comprehensively displaying the environment around the vehicle, including the front, rear, left, right, and top situations. This enables the autonomous driving system to better understand the surrounding environment and improve the accuracy of perception and decision-making.
[0037] Next, the process of the vehicle generating the BEV space is introduced.
[0038] For example, Figure 1A shows a flowchart for generating the BEV space.
[0039] It can be understood that the following processes can be executed by in-vehicle devices. The in-vehicle devices in the embodiments of this application can also be referred to as in-vehicle terminals. The in-vehicle terminal can be a mobile phone, a car computer, a terminal in self-driving, a wireless terminal in transportation safety, a terminal in a smart city, etc. The car computer will be used as an example for illustration below. However, it can be understood that the technical solutions described in this application are applicable to various in-vehicle devices for the above-mentioned three-dimensional object detection, not limited to the car computer.
[0040] As Figure 1A shown, this process includes:
[0041] S101, obtain images of multiple perspectives of the vehicle.
[0042] Exemplarily, Figure 1B shows a schematic diagram during the driving process of a vehicle.
[0043] It can be understood that vehicle 10 is usually equipped with multiple cameras (such as including front-view, side-view, and rear-view cameras, etc.). As Figure 1A shown, during the driving process of vehicle 10, through a timestamp synchronization mechanism or a hardware synchronization mechanism, images of multiple perspectives of vehicle 10 can be captured by multiple cameras at the same moment.
[0044] For example, vehicle 10 can capture images near the vehicle 10 through multiple cameras. Refer to Figure 1B , the images captured by the cameras can include, for example, the first sedan 01, pedestrian 02, first building 03, second sedan 04, and second building 05, etc.
[0045] The in-vehicle computer deployed on vehicle 10 can obtain the pictures taken by multiple cameras on vehicle 10 at the same moment, and use the pictures as images from multiple perspectives of vehicle 10.
[0046] S102. Extract feature maps from the multi-perspective images based on the feature extraction module.
[0047] Exemplarily, an image extraction module can be deployed on the in-vehicle computer. Through the image extraction module, feature extraction can be performed on the multi-perspective images (i.e., the original images) to generate feature maps of the images from each perspective. For example, the feature map obtained through the feature extraction module can be a feature map that is downsampled by n times (e.g., 4 times) compared to the original image.
[0048] Among them, downsampling (such as pooling operation) is a method used in convolutional neural networks to reduce the data dimension. When the image is downsampled by 4 times, it means that in a certain dimension (usually the width and height) of the feature map, the data resolution is reduced by 4 times. This operation is usually used to reduce the computational amount, extract more abstract features, and increase the robustness of the model. The feature map after downsampling by 4 times retains the important information in the original image, but the details are ignored or simplified. In the embodiments of the present application, the image downsampled by 4 times can be called a stride4 feature map. It can be understood that through the feature extraction module, image features of downsampling by 8 times, 16 times, 32 times, 64 times, etc. of the multi-perspective images can also be extracted. The present application does not limit the process of image feature extraction.
[0049] S103. Perform BEV space feature fusion on the feature maps of the multi-perspective images based on the BEV space feature fusion module to obtain the BEV space.
[0050] Exemplarily, after the in-vehicle computer extracts the feature maps of the multi-perspective images, it can perform inverse perspective mapping (IPM) on the image features according to the internal and external parameters of each camera, so as to fuse the BEV space features to obtain the BEV space features.
[0051] Among them, IPM is based on the geometric principle of perspective projection and converts a two-dimensional image containing three-dimensional information into a top view or a bird's-eye view through a mathematical algorithm. This process relies on a transformation matrix, which is usually determined by the internal parameters of the camera (such as focal length, principal point, etc.) and external parameters (such as the position and orientation of the camera).
[0052] For example, Figure 1C shows a flowchart for generating a BEV space based on IPM.
[0053] As Figure 1C shown, exemplarily, after the vehicle-mounted computer extracts the feature maps of images from multiple perspectives of the vehicle through the feature extraction module, it can perform IPM transformation on the feature maps. For example, in some embodiments of the present application, the image features extracted by the vehicle-mounted computer are feature maps of stride4. The vehicle-mounted computer can input the feature maps of stride4 into the spatial feature fusion module for IPM transformation. Then, the vehicle-mounted computer can convert the two-dimensional image containing three-dimensional information into a top view or a bird's-eye view, thereby generating a BEV space.
[0054] S104, fuse the BEV space based on the temporal fusion module.
[0055] Exemplarily, the temporal fusion module fuses spatio-temporal information, can make up for the limitations of single-frame BEV space perception, thereby increasing the receptive field, improving the problem of frame-to-frame jump and target occlusion in object detection, more accurately judging the target movement speed, and also playing an important role in target prediction and tracking. By introducing temporal information, the temporal fusion module makes the perception result more stable, improves the accuracy of the vehicle's judgment of the road conditions, and thus enhances the safety of autonomous driving.
[0056] In some embodiments, by integrating the pose and position information of the vehicle itself into the temporal fusion module, objects and obstacles in the surrounding environment can be perceived more accurately. These information helps the model better understand the relative relationship between the vehicle and the surrounding environment, thereby improving the accuracy of perception.
[0057] S105, optimize the BEV space based on the BEV deep feature extraction module to obtain deeper feature information of the image.
[0058] Exemplarily, in some embodiments, in the BEV space, the deep feature extraction module can use deep learning algorithms (such as convolutional neural network CNN, etc.) to extract deep feature information from the converted image. These feature information include the outline, shape, texture, etc. of the object, which are the basis for subsequent recognition and tracking tasks. For example, semantic information such as lane lines, obstacles, pedestrians, vehicles, etc. can be marked in the BEV map. And dynamic update, according to the real-time update of the sensor data, dynamically adjust the BEV space. Exemplarily, the BEV space generated by the vehicle-mounted computer can also refer toFigure 1B View.
[0059] S106, perform multi-task detection based on the BEV space.
[0060] Exemplarily, after the vehicle head unit obtains the BEV space, it can perform multi-task detection based on the BEV space. For example, it can perform 3D vehicle detection, 3D pedestrian / cyclist detection, lane line detection, ground discrete element detection, ground passable area detection, or other detections, etc.
[0061] It can be understood that in the above process of generating the BEV space, mainly in S103, BEV space feature fusion is performed based on the BEV space feature fusion module to obtain the BEV space. However, during the IPM transformation, only one downsampled feature map of the multi-view images (such as the feature map with stride 4) is used, resulting in the loss of some information in the image during the sampling process, and the BEV space obtained by the IPM transformation is not accurate enough.
[0062] As mentioned above, during the process of processing multi-view images (such as performing IPM transformation to generate the BEV space), the information of the multi-view images is not fully utilized, resulting in inaccurate spatial information generated according to the multi-view images.
[0063] To solve the problem that the spatial information generated based on a single image feature is not precise enough, the present application proposes an image processing method. In response to a request to detect the surrounding environment of the vehicle, multiple images are acquired, and the multiple images are images of different views collected by multiple sensors of the vehicle. The multiple images are sampled based on multiple sampling rates to obtain multiple feature maps of each image, and the sizes of the multiple feature maps are different. The first feature map (such as the feature map with stride 4) in the multiple feature maps is projected onto a set plane to obtain a first plane image (such as generating a BEV bird's-eye view).
[0064] Feature extraction is performed on the first plane image to obtain first feature data (features corresponding to the BEV bird's-eye view, such as the BEV query feature in the following text). Feature extraction is performed on the feature maps other than the first feature map in the multiple feature maps (such as the feature map with stride 16 and the feature map with stride 32) to obtain second feature data (such as the multi-view key feature and value feature in the following text). The first feature data and the second feature data are fused to obtain third feature data (such as the BEV spatial fusion feature in the following text); based on the third feature data, the surrounding environment of the vehicle is detected.
[0065] Through the above solution, the in-vehicle device can extract feature maps of multiple different scales from images of multiple perspectives respectively, and then fuse the features corresponding to the feature maps of different sizes to obtain third feature data, so that the third feature data retains more information of the multi-perspective images, thereby improving the accuracy of the in-vehicle device to monitor the vehicle surrounding environment based on the third feature data.
[0066] In some embodiments of the present application, the first feature data and the second feature data can be fused through a self-attention mechanism. The self-attention mechanism can capture long-distance dependencies in the image, thereby improving the accuracy of generating the first environment image.
[0067] Next, the image processing method in the embodiments of the present application will be introduced.
[0068] For example, Figure 2 According to an embodiment of the present application, a flowchart of an implementation of an image processing method is shown.
[0069] Exemplarily, the execution subject of each of the following processes can be an in-vehicle device.
[0070] As Figure 2 shown, the method includes:
[0071] S201, in response to a request to detect the vehicle surrounding environment, obtain multiple images, where the multiple images are images of different perspectives collected by multiple sensors of the vehicle.
[0072] Exemplarily, in some embodiments of the present application, the vehicle can collect images of different perspectives of the vehicle surrounding environment through multiple sensors during driving. For example, the vehicle can be configured with front-view, side-view, rear-view cameras, etc. At the same moment, the cameras on the vehicle can collect images of different perspectives.
[0073] After the in-vehicle device on the vehicle starts the assisted driving or autonomous driving function, it can detect a request to detect the vehicle surrounding environment. After the in-vehicle device responds to the request to detect the vehicle surrounding environment, it can obtain multiple images of different perspectives collected by each camera on the vehicle at the same moment.
[0074] S202, sample the multiple images based on multiple sampling rates to obtain multiple feature maps of each image, where the sizes of the multiple feature maps are different.
[0075] Exemplarily, in some embodiments of the present application, after the in-vehicle device obtains multiple images (hereinafter referred to as multi-perspective images), it can downsample the multi-perspective images through multiple different sampling rates, thereby obtaining feature maps of different sizes.
[0076] For example, the in-vehicle device can perform downsampling through a neural network (backbone network). The downsampling rate can include, for example, downsampling by a factor of 4, downsampling by a factor of 32, downsampling by a factor of 64, etc. The embodiments of the present application do not limit the downsampling rate.
[0077] S203. Project the first feature map among multiple feature maps onto a set plane to obtain a first planar image.
[0078] In some embodiments of the present application, the in-vehicle device can project the first feature map among multiple feature maps onto a set plane to obtain a first planar image. For example, the in-vehicle device can obtain the acquisition parameters of multiple sensors; based on the acquisition parameters, perform an inverse perspective transformation on the first feature map to generate a first planar image.
[0079] Exemplarily, the in-vehicle device can select a first feature map of a certain size from feature maps of different sizes and perform an inverse perspective transformation on the first feature map, thereby generating a first planar map centered on the vehicle. The first planar map can be a BEV initial spatial feature map.
[0080] For example, the first feature map can be a feature map downsampled by a factor of 4, that is, a feature map with a stride of 4. The in-vehicle device can obtain the internal parameters and external parameters of each camera on the vehicle and perform an inverse perspective transformation on the stride-4 feature map based on the internal parameters and external parameters of each camera, thereby generating a BEV initial spatial feature map.
[0081] In other embodiments, the first feature map can also be a feature map downsampled by a factor of 16, a feature map downsampled by a factor of 32, or a feature map downsampled by a factor of 64, etc. The embodiments of the present application do not limit the specific downsampling rate of the first feature map.
[0082] S204. Perform feature extraction on the first planar image to obtain first feature data.
[0083] In some embodiments of the present application, the in-vehicle device can perform feature extraction on the first planar image to obtain first feature data.
[0084] For example, the in-vehicle device downsamples the BEV initial spatial feature map. The downsampling rate can be, for example, 32 times. It can be understood that the downsampling rate of the BEV initial spatial feature map is the target sampling rate preset by the in-vehicle device. In other embodiments, the sampling rate can also be 16 times, 8 times, 64 times, etc. After downsampling the BEV initial spatial feature map, perform feature extraction on the BEV initial spatial feature map to generate first feature data.
[0085] In some other embodiments, the first feature data further includes the BEV spatial position encoding of the BEV initial spatial feature map. For example, in some embodiments, the downsampled BEV initial spatial feature map can be encoded by a BEV spatial position encoding module to obtain the BEV spatial position encoding, and the BEV spatial position encoding is added and fused with the features corresponding to the downsampled BEV initial spatial feature map, so as to obtain the first feature data of the BEV initial spatial feature map. In some embodiments, the first spatial data can be used as the query feature of the BEV for self-attention fusion with other features, and the self-attention fusion process is described in detail below.
[0086] S205. Feature extraction is performed on the feature maps other than the first feature map among multiple feature maps to obtain second feature data.
[0087] In some embodiments of the present application, since the first feature map is used to be converted into the BEV initial spatial feature map, the vehicle-mounted device can perform feature extraction on the remaining feature maps with different sizes among multiple feature maps to obtain second feature data.
[0088] In some embodiments, the following method can be used to perform feature extraction on the feature maps other than the first feature map among multiple feature maps to obtain second feature data:
[0089] Align the feature maps with different sampling rates among the feature maps other than the first feature map among multiple feature maps to obtain a second feature map.
[0090] For example, in some embodiments of the present application, since multiple feature maps include a feature map with a stride of 4, a feature map with a stride of 32, and a feature map with a stride of 64. The feature map with a stride of 4 is used to be converted into the BEV initial spatial feature map, then the remaining feature map with a stride of 64 and the feature map with a stride of 32 can be aligned, and then the aligned feature map is determined. For example, the feature map with a stride of 64 can be upsampled by a pyramid feature module and then fused with the feature map with a stride of 32 to obtain a multi-view channel aligned feature map with a stride of 32, and this multi-view channel aligned feature map is the second feature map.
[0091] In some embodiments of the present application, the vehicle-mounted device can also fuse the features of the acquisition parameters of multiple sensors with the features corresponding to the second feature map, so as to generate second feature data.
[0092] For example, the vehicle-mounted device can obtain the acquisition parameters of multiple sensors and encode the acquisition parameters based on the size of the second feature map to obtain sensor features.
[0093] In some embodiments, the sensors of the vehicle-mounted device may include, for example, a camera, and the acquisition parameters of the camera may be, for example, the internal parameters and external parameters of the camera. The vehicle-mounted device may encode the internal parameters and external parameters of the camera to obtain sensor features.
[0094] The vehicle-mounted device may perform feature extraction on the second feature map to obtain fourth feature data. The vehicle-mounted device fuses the fourth feature data and the sensor features to determine the second feature data.
[0095] It can be understood that the second feature map is the multi-view channel alignment feature map described above. The vehicle-mounted device may extract features from the multi-view channel alignment feature map to obtain fourth feature data. It can be understood that the size of the multi-channel alignment feature map is the same as the size of the feature map corresponding to the target sampling rate preset by the vehicle-mounted device, that is, the feature map with stride 32. The vehicle-mounted device may also fuse the sensor features and the fourth feature data to determine the second feature data. It can be understood that the second feature data includes information in the feature map with stride 32 and the feature map with stride 64.
[0096] It can be understood that the second feature data can also be generated by feature maps with more sampling rates or fewer sampling rates. The embodiments of the present application do not limit the types and numbers of sampling rates of the feature maps for generating the second feature data.
[0097] In some embodiments of the present application, the second feature data may include a key feature (as an example of the first key feature) and a value feature (as an example of the first value feature), where the key feature is obtained by fusing the feature obtained by the second feature map based on one sampling method and the sensor features. The value feature is the feature obtained by the second feature map based on another sampling method. The acquisition process of the second feature data is described in detail below.
[0098] S206, fuse the first feature data and the second feature data to obtain third feature data.
[0099] In some embodiments of the present application, the vehicle-mounted device may fuse the first feature data and the second feature data through a self-attention mechanism to obtain third feature data.
[0100] For example, linearly transform the first feature data to obtain a query feature, and perform different linear transformations on the second feature data respectively to obtain corresponding key features and value features.
[0101] Then, the vehicle-mounted device performs an operation based on the query feature, the key feature, and the value feature through equation (1):
[0102] (1);
[0103] Where Q is the query feature, K is the key feature, and V is the value feature, represents performing self-attention operation on Q, K, and V, represents that after multiplying the Q and K matrices and passing through the softmax operator, a normalized attention weight is obtained. It can be understood that The output feature can be the feature fused with the query feature of BEV. By upsampling the feature fused with the query feature of BEV by 32 times, the third feature data can be obtained. It can be understood that since the first feature map is obtained by downsampling in S201, and in S202, the first feature data for generating the query is obtained by the BEV space encoding module extracting features from the downsampled BEV initial space feature map, it is necessary to restore the feature fused with the query feature of BEV through upsampling to obtain the third feature data. That is to say, the third feature data includes the information in the feature maps with stride 4, stride 32, and stride 64 of the multi-view image, so as to retain the information in the multi-view image more completely.
[0104] S207, based on the third feature data, detect the surrounding environment of the vehicle.
[0105] Exemplarily, in some embodiments of the present application, after the in-vehicle device obtains the third feature data, it can detect the surrounding environment of the vehicle according to the third feature data.
[0106] It can be understood that the third feature data includes the information retained by sampling the multi-view image at different sampling rates. Therefore, the third feature data can retain more data in the multi-view image. Therefore, the accuracy of detecting the surrounding environment of the vehicle through the third feature data is higher.
[0107] For example, the in-vehicle device can more accurately detect other vehicles, pedestrians, road lines, traffic signs, etc. around the vehicle based on the third feature, and then provide various services according to the detection results, such as autonomous driving, assisted driving, road planning, etc., so as to improve the safety of vehicle driving.
[0108] Next, introduce the process of obtaining the third feature data in the embodiments of the present application.
[0109] For example, Figure 3 According to some embodiments of the present application, a schematic diagram of generating the third feature data is shown.
[0110] As Figure 3As shown, an IPM inverse perspective transformation module 305, an internal and external parameter encoding module 306, a pyramid feature module, a BEV feature downsampling module 309, a BEV spatial position encoding module 310, a multi-view key feature and value feature generation module 314, and a transformer hierarchical self-attention module 316 can be configured in the vehicle-mounted device. The process of generating the third feature data is introduced below.
[0111] Multi-view stride 4 image feature map 301.
[0112] In some embodiments of the present application, the multi-view stride 4 image feature map (as an example of the first feature map) 301 is obtained by downsampling each perspective image by 4 times through an image feature extraction module. Exemplarily, the size of each multi-view image is (3, img_h, img_w), where 3 represents the number of channels, corresponding to the three color channels of RGB (red, green, blue). img_h represents the height of the image, that is, the number of pixels in the vertical direction of the image, and img_w represents the width of the image, that is, the number of pixels in the horizontal direction of the image.
[0113] The size of the feature map after downsampling the multi-view image by 4 times is (c4, img_h / 4, img_w / 4), where c4 represents the number of channels of the multi-view stride 4 image feature map 301. It can be understood that a convolutional layer is applied during the downsampling process, and the value of c4 will depend on the output channels of the convolutional layer. For example, if a convolutional layer with 16 output channels is applied, then the value of c4 will be 16.
[0114] Camera internal and external parameters 302.
[0115] The camera internal and external parameters 302 include the internal parameters and external parameters of the camera. Among them, the internal parameters are related to the hardware structure and imaging principle of the camera and are used to represent the mapping relationship between light from three-dimensional space passing through the camera and mapping to a two-dimensional plane. The internal parameters are represented by Intri_mat, a 3×3 matrix. In some embodiments, the internal parameters also include camera distortion coefficients. The camera external parameters are represented by Extri_mat, a 4×4 matrix. One perspective image corresponds to one set of camera internal and external parameters 302, which are used to describe the internal characteristics of the camera and the relative spatial information of the external coordinates with respect to the vehicle origin.
[0116] Multi-view stride 32 image feature map 303.
[0117] Similar to the multi-view stride4 image feature map 301, the size of the feature map obtained after the multi-view image is downsampled 32 times by the image feature extraction module is (c32, img_h / 32, img_w / 32), that is, the multi-view stride32 image feature map 303 is generated. It can be understood that the multi-view stride32 image feature map 303 is one of the feature maps other than the first feature map among multiple feature maps.
[0118] Multi-view stride64 image feature map 304.
[0119] Similar to the multi-view stride4 image feature map 301, the multi-view stride64 image feature map 304 is obtained after the multi-view image is downsampled 64 times by the image feature extraction module, and its feature map size is (c64, img_h / 64, img_w / 64). It can be understood that the multi-view stride64 image feature map 304 is also one of the feature maps other than the first feature map among multiple feature maps.
[0120] IPM inverse perspective transformation module 305.
[0121] The IPM inverse perspective transformation module 305 can obtain the internal and external camera parameters 302, so as to perform IPM inverse perspective transformation on the multi-view stride4 image feature map 301.
[0122] Referring to the process of S203, the vehicle-mounted device can perform IPM transformation through the multi-view stride4 image feature map 301 and the internal and external camera parameters 302, so as to generate the BEV initial spatial feature map (as an example of the first plane image) 308. Based on this initialization method, it is ensured that the attention mechanism of the subsequent conversion model (transformer) has better convergence and the model training is more stable. After IPM, the BEV initial spatial feature map 308 is obtained, and its size is (bev_c, bev_h, bev_w), where bev_c, bev_h, and bev_w are the number of channels, height, and width of the BEV initial spatial feature map 308 respectively. It can be understood that the first feature data can be obtained by performing feature extraction on the BEV initial spatial feature map 308.
[0123] Internal and external parameter encoding module 306.
[0124] In some embodiments of the present application, for each perspective, the internal and external parameter encoding module 306 inputs the corresponding internal parameter matrix intri_mat and external parameter matrix extri_mat, and generates multi-perspective camera encoding features 312 for each stride32 image feature. It can be understood that since the target sampling rate of the vehicle-mounted device is stride32, that is, downsampled by 32 times, therefore, when the internal and external parameter encoding module 306 encodes the internal and external parameters of the camera, it can be aligned with the stride32 image feature.
[0125] For example, the internal and external parameter encoding module 306 can generate a coordinate point matrix img_plane at the pixel positions corresponding to the multi-perspective stride32 image feature map 303. Its size is (3, img_h / 32, img_w / 32), where 3 represents three channels. Among them, channels 1 and 2 respectively represent the x (e.g., abscissa) and y (e.g., ordinate) points of each pixel in the image coordinate system of the multi-perspective stride32 image feature map 303, and channel 3 is a fixed filling value of 1.
[0126] The internal and external parameter encoding module 306 obtains intri_mat_inv and extri_mat_inv by performing an inverse transformation on the internal parameter matrix intri_mat and the external parameter matrix extri_mat. Then, through matrix multiplication and fixed filling value operation, img_pos is generated, where img_pos = extri_mat_inv × pad(intri_mat_inv × img_plane). Here, × represents matrix multiplication, and pad represents making the dimensions aligned by filling with fixed values. img_pos is used as the image space position matrix, and its size is (4, img_h / 32, img_w / 32), indicating the position correspondence of each image feature pixel in the vehicle coordinate system (VCS) of the ego vehicle.
[0127] Finally, img_pos passes through a fully connected network for feature dimension transformation to obtain multi-perspective camera encoding features 312, whose dimension is (d, img_h / 32, img_w / 32). Performing the same process operation on the cameras of each perspective can obtain multi-perspective camera encoding features 312.
[0128] Feature Pyramid Network (FPN) 307.
[0129] The FPN 307 can fuse feature maps of different scales to achieve effective detection of objects of different sizes. For example, in some embodiments of the present application, image features with two downsampling scales, including stride32 and stride64, can be fused through the FPN 307 to obtain channel-aligned image features after scale fusion, with a size of (d, img_h / 32, img_w / 32), whose length and width scales are consistent with the stride32 image feature map 301, and the channel dimension is consistent with the camera-encoded features. Fusing the feature maps of different scales for each perspective can obtain multi-perspective channel-aligned feature maps (as an example of the second feature map) 313. It can be understood that second feature data can be obtained by performing feature extraction on the multi-perspective channel-aligned feature maps 313.
[0130] BEV feature downsampling module 309.
[0131] In some embodiments of the present application, the BEV initial spatial feature map 308 is a dense BEV spatial representation (a BEV space that includes information of each pixel point of the feature map), and the computational time complexity of performing transformer attention operations in this dense space is high. The BEV feature downsampling module 309 can change the BEV initial spatial feature map 308 with a size of (bev_c, bev_h, bev_w) into a downsampled BEV feature with a size of (d, bev_h / 4, bev_w / 4) through the stacking of a convolutional network and PixelUnshuffle operations, reducing the pixel resolution in the BEV space, while the channel dimension is consistent with the multi-perspective camera-encoded features 312.
[0132] BEV spatial position encoding module 310.
[0133] The encoding process of the BEV spatial position encoding module 310 is similar to the img_pos encoding in the internal and external parameter encoding module 306. Each pixel point of the downsampled initial BEV spatial feature map corresponds to the coordinates of a specific position point in the vehicle's VCS coordinate system, represented by the bev_pos matrix, with a size of (3, bev_h / 4, bev_w / 4). Among them, channels 1 and 2 represent the x and y points of each pixel point in the VCS left coordinate system, and channel 3 is the value of sqrt(x×x + y×y), representing the true distance of the pixel point from the origin coordinate in the VCS coordinate system. The bev_pos matrix passes through a fully connected network to align the channel dimensions and obtain the BEV spatial position encoding, with a size of (d, bev_h / 4, bev_w / 4). Finally, the BEV spatial position encoding is added and fused with the corresponding features of the downsampled initial BEV spatial feature map to obtain the BEV query feature (as an example of the first feature data) 311. The BEV query feature 311 can be used as one of the key inputs of the transformer hierarchical self-attention module 316.
[0134] Multi-view key feature and value feature generation module 314.
[0135] In some embodiments of the present application, the multi-view channel-aligned feature map 313 can be processed in two ways. For example, in one processing method, the multi-view channel-aligned feature map 313 first passes through a fully connected network for feature transformation and is added and fused with the multi-view camera encoding feature 312 to generate multi-view key features, with the size remaining unchanged, still (d, img_h / 32, img_w / 32). In another processing method, the multi-view channel-aligned feature map 313 passes through another fully connected network for feature transformation to generate multi-view value features. It can be understood that the key features (as an embodiment of the first key feature) and value features (as an embodiment of the first value feature) can be used as an embodiment of the second feature data in the embodiments of the present application, that is, the second feature data includes multi-view key features and value features 315.
[0136] Transformer hierarchical self-attention module 316.
[0137] In some embodiments of the present application, the BEV query feature 311 and the multi-view key features and value features 315 are used as inputs for an attention operation to obtain the BEV query fusion feature 317. It can be understood that better BEV feature expressions can be obtained by fusing multi-view image features through the self-attention mechanism. Exemplarily, the process of self-attention mechanism fusion is described in detail below.
[0138] BEV Feature Upsampling Module 318.
[0139] Since the BEV initial spatial feature map 308 has been downsampled before the transformer hierarchical self-attention module 316, it is necessary to restore the resolution size of the query fusion feature 317 of BEV in the VCS space through a convolutional network and Upsample operation. After passing through the upsampling module, the spatial fusion feature 319 of BEV with the same size as the BEV initial spatial feature map 308 is obtained. The spatial fusion feature 319 of BEV can be used as the final output of the transformer hierarchical self-attention module 316. It can be understood that the spatial fusion feature 319 of BEV can be used as an embodiment of the third feature data.
[0140] Exemplarily, the self-attention operation can usually be processed by Equation (1). In some embodiments of the present application, the number of query features of BEV is BEV_H × BEV_W (where BEV_H is bev_h / 4 and BEV_W is bev_w / 4), the number of cameras is M, the size of a single image feature map is D × H × W, the time complexity is O(BEV_H × BEV_W × M × H × W × D), and the spatial complexity is O(BEV_H × BEV_W × N × H × W + (BEV_H × BEV_W + M × H × W) × D). The computational complexity of the transformer hierarchical self-attention module 316 is strongly related to the size of the query features of BEV, the size of the key features and value features of the multi-view images. Therefore, the way to reduce its computational complexity is to reduce the size of the query features of BEV for attention calculation and the size of the key features and value features of the multi-view images.
[0141] For example, in some embodiments of the present application, through the transformer hierarchical self-attention module 316, the query features 311 of BEV, the key features, and the value features of the multi-view images are simultaneously divided into different regions, and cross-attention operations are performed between the regions, greatly reducing the computational complexity of the attention operation, so that the transformer hierarchical self-attention module 316 can be deployed on the chip of the vehicle-mounted device (such as the J6 chip) with a higher frame rate to meet the real-time detection efficiency.
[0142] Therefore, in some embodiments of the present application, the process of fusing the first feature data and the second feature data by the vehicle-mounted device based on the self-attention mechanism can be, for example, that the vehicle-mounted device can divide the first feature data into N partial regions, and each region includes N partial first sub-features, and the positions of the N partial first sub-features in the same region are different.
[0143] It can be understood that the first feature data is the feature data obtained by extracting the BEV initial spatial feature map 308, that is, the query feature 311 of BEV. In some embodiments of the present application, in order to implement the cross-attention operation, the query feature 311 of BEV can be divided into N partial regions, each region includes N partial first sub-features, and the positions of the N partial first sub-features in the same region are different.
[0144] In some embodiments of the present application, the key feature in the second feature data also includes N partial regions, each region includes N partial second sub-features, and the positions of the N partial second sub-features in the same region are different.
[0145] The value feature in the second feature data also includes N partial regions, each region includes N partial third sub-features, and the positions of the N partial third sub-features in the same region are different. It can be understood that the process of extracting the key feature and the value feature can refer to the part of the multi-view key feature and value feature generation module 314.
[0146] Then, the N partial first sub-features, N partial second sub-features, and N partial second sub-features in the same region are fused based on the self-attention mechanism to obtain a second query feature, and the third feature data is determined based on the second query feature, where N is greater than or equal to 2.
[0147] It can be understood that by dividing the query feature, key feature, and value feature of BEV into multiple regions, and performing self-attention optimization on the multi-part sub-features (including the first sub-feature, second sub-feature, and third sub-feature) on each region, the size of the self-attention optimization can be reduced, thereby improving the efficiency of the self-attention optimization.
[0148] For example, Figure 4 Some embodiments according to the present application show a flowchart of the implementation of a cross-attention mechanism.
[0149] Exemplarily, the execution subject of each of the following processes is the transformer hierarchical self-attention module.
[0150] As Figure 4 shown, the process includes:
[0151] S401A, perform local window partitioning on the query feature of BEV.
[0152] In some embodiments of the present application, the transformer hierarchical self-attention module evenly divides the query feature of BEV (as an embodiment of the first feature data) into N windows in a local area (i.e., divides it into N partial regions), and promotes the windows to the Batch dimension, that is, each window includes N parts of the first sub-features, and each part of the first sub-feature in the same window is at different positions in the window.
[0153] For example, Figure 5A According to some embodiments of the present application, a schematic diagram of local window division is shown.
[0154] As Figure 5A shown, in some embodiments of the present application, taking the query feature of BEV as the original feature as an example, where N is 4, that is to say, the query feature of BEV can be divided into 4 regions, for example, region 1, region 2, region 3, and region 4. Each region includes 4 parts of sub-features. In the case where the original feature is the query feature of BEV, the sub-feature corresponding to the original feature is the first sub-feature. Taking region 1 as an example, the 4 parts of the first sub-features are respectively at positions 1, 2, 3, and 4.
[0155] S401B, perform local window division on the key feature and the value feature.
[0156] Exemplarily, the local window division of the key feature and the value feature is similar to the local window division method of the query feature of BEV.
[0157] For example, referring to Figure 5A , taking the key feature as the original feature as an example, where the key feature is divided into 4 regions, for example, region 1, region 2, region 3, and region 4. Each region includes 4 parts of sub-features. In the case where the original feature is the key feature, the sub-feature corresponding to the original feature is the second sub-feature. Taking region 1 as an example, the 4 parts of the second sub-features are respectively at positions 1, 2, 3, and 4.
[0158] Referring to Figure 5A , taking the value feature as the original feature as an example, where the value feature is divided into 4 regions, for example, region 1, region 2, region 3, and region 4. Each region includes 4 parts of sub-features. In the case where the original feature is the value feature, the sub-feature corresponding to the original feature is the third sub-feature. Taking region 1 as an example, the 4 parts of the third sub-features are respectively at positions 1, 2, 3, and 4.
[0159] S402, perform cross-attention fusion on the query feature, key feature, and value feature of BEV to obtain the updated query feature of BEV.
[0160] After the local window division of each feature is completed, the four sub-features of each feature in the same area can perform cross-attention operations, and the sub-features in different areas will not perform cross-attention operations. In this way, the size of self-attention fusion can be reduced, thereby improving the efficiency of self-attention fusion. For example, when N = 4, the size of the query feature of BEV will change from (1, D, BEV_H, BEV_W) to (4, D, BEV_H / 2, BEV_W / 2), and its time complexity will be reduced to 1 / 4 of the original global attention operation, thereby improving the efficiency of self-attention fusion.
[0161] S403A, perform local window restoration on the BEV updated query feature.
[0162] Exemplarily, after the process of S402 is executed, the transformer hierarchical self-attention module restores the output of self-attention optimization, and the output of self-attention optimization can be used as the BEV updated query feature. It can be understood that the BEV updated query feature output by self-attention optimization is also windowed, and the BEV updated query feature can be restored from (4, D, BEV_H / 2, BEV_W / 2) to (1, D, BEV_H, BEV_W). It can be understood that the query feature of BEV is the first query feature after restoration.
[0163] S403B, perform local window restoration on the key feature and the value feature.
[0164] Exemplarily, the transformer hierarchical self-attention module can restore the key feature and the value feature in the same way as restoring the BEV updated query feature.
[0165] In some embodiments of the present application, after the vehicle-mounted device obtains the first query feature through the transformer hierarchical self-attention module, cross-attention optimization can continue to be performed.
[0166] For example, the first query feature includes N partial regions, each region includes N partial fourth sub-features, and the positions of the N partial fourth sub-features in the same region are different.
[0167] In some embodiments of the present application, the first query feature is the BEV updated query feature, and the BEV updated query feature can also include N partial regions, each region includes N partial fourth sub-features of the BEV updated query feature. The positions of the N partial fourth sub-features in the same region are different.
[0168] In some embodiments, the third feature data can be determined based on the second query feature in the following manner:
[0169] The vehicle-mounted device can fuse N parts of the fourth sub-features, N parts of the second sub-features, and N parts of the third sub-features at the same position in N different regions based on the self-attention mechanism to obtain the third feature data.
[0170] For example, the following introduces the fusion process of the fourth sub-features, the corresponding second sub-features, and the third sub-features in different regions based on the self-attention mechanism.
[0171] S404A, perform global window partitioning on the BEV updated query feature.
[0172] Exemplarily, the transformer hierarchical self-attention module divides the BEV updated query feature into N windows. Similar to 401A, the BEV updated query feature is evenly divided into N windows (i.e., divided into N partial regions), and the windows are promoted to the Batch dimension, that is, each window includes N parts of the fourth sub-features, and each part of the fourth sub-feature is at a different position in the window.
[0173] For example, Figure 5B According to an embodiment of the present application, a schematic diagram of global window partitioning is shown.
[0174] As Figure 5B shown, in some embodiments of the present application, taking the original feature as the BEV updated query feature as an example, where N is 4, that is, the BEV updated query feature is divided into 4 regions, for example, region 1, region 2, region 3, and region 4. Each region includes 4 parts of sub-features. In the case where the original feature is the BEV updated query feature, the sub-feature corresponding to the original feature is the fourth sub-feature. Taking region 1 as an example, the 4 parts of the fourth sub-features are respectively at positions 1, 2, 3, and 4.
[0175] S404B, perform global window partitioning on the key feature and the value feature.
[0176] Exemplarily, the local window partitioning of the key feature and the value feature is similar to the local window partitioning method of the query feature of the BEV.
[0177] For example, referring to Figure 5B , taking the original feature as the key feature, the key feature is divided into 4 regions, for example, region 1, region 2, region 3, and region 4. Each region includes 4 parts of sub-features. In the case where the original feature is the key feature, the sub-feature corresponding to the original feature is the second sub-feature. Taking region 1 as an example, the 4 parts of the second sub-features are respectively at positions 1, 2, 3, and 4.
[0178] Referring to Figure 5B, the original feature is the value feature, and the value feature is divided into 4 regions. For example, region 1, region 2, region 3, and region 4. Each region includes 4 parts of sub-features. When the original feature is the value feature, the sub-feature corresponding to the original feature is the third sub-feature. Taking region 1 as an example, the 4 parts of the third sub-feature are respectively at position 1, position 2, position 3, and position 4.
[0179] S405, perform cross-attention fusion on the BEV updated query feature, key feature, and value feature to obtain the query fusion feature of BEV.
[0180] After completing the global window division of each feature, the 4 parts of sub-features at the same position in different regions of each feature can perform cross-attention operations. In this way, the size of self-attention fusion can also be reduced, thereby improving the efficiency of self-attention fusion. For example, when N = 4, the size of the BEV updated query feature will change from (1, D, BEV_H, BEV_W) to (4, D, BEV_H / 2, BEV_W / 2), and its time complexity will be reduced to 1 / 4 of the original global attention operation.
[0181] For example, the 4 parts of the fourth sub-feature at the same position in different regions of the BEV updated query feature can be a part of the fourth sub-feature at position 1 in region 1, a part of the fourth sub-feature at position 1 in region 2, a part of the fourth sub-feature at position 1 in region 3, and a part of the fourth sub-feature at position 1 in region 4, with a total of 4 parts of the fourth sub-feature.
[0182] For the 4 parts of the second sub-feature at the same position in different regions of the key feature, they can be a part of the second sub-feature at position 1 in region 1, a part of the second sub-feature at position 1 in region 2, a part of the second sub-feature at position 1 in region 3, and a part of the second sub-feature at position 1 in region 4, with a total of 4 parts of the second sub-feature.
[0183] For the 4 parts of the third sub-feature at the same position in different regions of the value feature, they can be a part of the third sub-feature at position 1 in region 1, a part of the third sub-feature at position 1 in region 2, a part of the third sub-feature at position 1 in region 3, and a part of the third sub-feature at position 1 in region 4, with a total of 4 parts of the third sub-feature.
[0184] In some embodiments of the present application, the process of self-attention fusion in the transformer hierarchical self-attention module is the process of fusing the four parts of the fourth sub-feature at position 1 of the above-mentioned 4 regions, the four parts of the second sub-feature at position 1 of the 4 regions, and the four parts of the third sub-feature at position 1 of the 4 regions. Similarly, self-attention fusion is also required for the four parts of the fourth sub-feature, the four parts of the second sub-feature, and the four parts of the third sub-feature of different regions at other same positions.
[0185] It can be understood that during the self-attention fusion process, each feature is divided into 4 regions and each region is divided into 4 parts of sub-features. In other embodiments, each feature can also be divided into more or fewer regions, and the number of regions is the same as the number of parts of the sub-features in each region. The embodiments of the present application do not limit the number of regions into which each feature is divided and the number of parts of the sub-features within each region.
[0186] S406, perform global window restoration on the query fusion feature of BEV.
[0187] Exemplarily, after executing step S405, the transformer hierarchical self-attention module restores the output of self-attention optimization, and the output of self-attention optimization can be used as the query fusion feature of BEV. It can be understood that the query fusion feature of BEV output by self-attention optimization is also globally windowed, and the query fusion feature of BEV can be restored from (4, D, BEV_H / 2, BEV_W / 2) to (1, D, BEV_H, BEV_W). It can be understood that the query fusion feature of BEV after restoration is the third feature data.
[0188] It can be understood that in some embodiments of the present application, the vehicle-mounted device divides the original transformer attention into two hierarchical attention mechanisms through the transformer hierarchical self-attention module, greatly reducing the computational time complexity of the attention mechanism. The reduction of the computational time complexity is related to the number of regions N divided. When the value of N is larger, the computational complexity is reduced to 2 / N of the original. At the vehicle end, when N = 16, a balance between computational efficiency and the detection performance of the BEV multi-task model can be achieved.
[0189] The vehicle-mounted device involved in each of the above embodiments is introduced below.
[0190] For example, Figure 6 According to some embodiments of the present application, a schematic structural diagram of a vehicle-mounted device 100 is shown.
[0191] The vehicle-mounted device 100 can be used to implement the image processing method provided in the foregoing embodiments.
[0192] As shown Figure 6 in the figure, the in-vehicle device 100 includes one or more processors 101, a system memory 102, a non-volatile memory (NVM) 103, a communication interface 104, an input / output device 105, and a system control logic unit 106 for coupling the processor 101, the system memory 102, the non-volatile memory 103, the communication interface 104, and the input / output device 105. Among them:
[0193] The processor 101 may include one or more processing units. For example, it may include a central processing unit (CPU), a graphics processing unit (GPU), a digital signal processor (DSP), a micro-programmed control unit (MCU), an artificial intelligence (AI) processor, or a field programmable gate array (FPGA), a neural-network processing unit (NPU), etc. The processing module or processing circuit may include one or more single-core or multi-core processors. In some embodiments, the CPU may be used to optimize the neural network model to be run, and the NPU may be used to run the neural network model to be run.
[0194] The system memory 102 is a volatile memory, such as a random-access memory (RAM), a double data rate synchronous dynamic random access memory (DDR SDRAM), etc. The system memory is used to temporarily store data and / or instructions. For example, in some embodiments, the system memory 102 may be used to store the data provided by the foregoing different services, such as sensor data, image data, or video data, etc., and may also be used to store the instructions of the image generation method provided by the foregoing embodiments.
[0195] The non-volatile memory 103 may include one or more tangible, non-transitory computer-readable media for storing data and / or instructions. In some embodiments, the non-volatile memory 103 may include any suitable non-volatile memory such as flash memory and / or any suitable non-volatile storage device, such as a hard disk drive (HDD), a compact disc (CD), a digital versatile disc (DVD), a solid-state drive (SSD), etc. In some embodiments, the non-volatile memory 103 may also be a removable storage medium, such as a secure digital (SD) memory card, etc. In other embodiments, the non-volatile memory 103 may be used to store instructions of the image generation method provided in the foregoing embodiments, etc.
[0196] In particular, the system memory 102 and the non-volatile memory 103 may respectively include: a temporary copy and a permanent copy of the instructions 107. The instructions 107 may include: when executed by at least one of the processors 101, enabling the vehicle-mounted device 100 to implement the image generation method provided in the embodiments of the present application.
[0197] The communication interface 104 may include a transceiver for providing a wired or wireless communication interface for the vehicle-mounted device 100, and then communicating with any other suitable device through one or more networks. In some embodiments, the communication interface 104 may be integrated into other components of the vehicle-mounted device 100. For example, the communication interface 104 may be integrated into the processor 101. In some embodiments, the vehicle-mounted device 100 may communicate with other devices through the communication interface 104. For example, the vehicle-mounted device 100 may obtain corresponding data from other devices through the communication interface 104.
[0198] The input / output device 105 may include an input device such as a keyboard, a mouse, etc., and an output device such as a display, etc. The user may interact with the vehicle-mounted device 100 through the input / output device 105.
[0199] The system control logic unit 106 may include any suitable interface controller to provide any suitable interface for other modules of the vehicle-mounted device 100. For example, in some embodiments, the system control logic unit 106 may include one or more memory controllers to provide an interface connected to the system memory 102 and the non-volatile memory 103.
[0200] In some embodiments, at least one of the processors 101 may be logically packaged with one or more controllers for the system control logic unit 106 to form a system in package (SiP). In other embodiments, at least one of the processors 101 may also be integrated with the logic of one or more controllers for the system control logic unit 106 on the same chip to form a system-on-chip (SoC).
[0201] It can be understood that Figure 6 The structure of the in-vehicle device 100 shown is only an example. In other embodiments, the in-vehicle device 100 may include more or fewer components than shown, or combine certain components, or split certain components, or have different component arrangements. The illustrated components can be implemented in hardware, software, or a combination of software and hardware.
[0202] It can be understood that the in-vehicle device 100 can be any device configured on a vehicle, including but not limited to mobile phones, in-vehicle computers, terminals in self-driving, wireless terminals in transportation safety, terminals in a smart city, and so on.
[0203] The embodiments of the present application also provide a program product. When executed on the in-vehicle device, the program product can enable the in-vehicle device to implement the image processing methods provided in the foregoing embodiments.
[0204] The embodiments of the present application also provide a readable storage medium. One or more programs are stored in the readable storage medium. When the one or more programs are executed by the in-vehicle device, the in-vehicle device is enabled to implement the image processing methods provided in the foregoing embodiments.
[0205] The embodiments of the mechanism disclosed in the present application can be implemented in hardware, software, firmware, or a combination of these implementation methods. The embodiments of the present application can be implemented as a computer program or program code executed on a programmable system. The programmable system includes at least one processor, a storage system (including volatile and non-volatile memories and / or storage elements), at least one input device, and at least one output device.
[0206] The program code can be applied to the input instructions to execute the various functions described in the present application and generate output information. The output information can be applied to one or more output devices in a known manner. For the purposes of the present application, the processing system includes any system having a processor such as, for example, a digital signal processor, a microcontroller, an application specific integrated circuit, or a microprocessor.
[0207] The program code can be implemented in a high-level procedural language or an object-oriented programming language to communicate with the processing system. When needed, the program code can also be implemented in assembly language or machine language. In fact, the mechanisms described in this application are not limited to the scope of any specific programming language. In any case, the language can be a compiled language or an interpreted language.
[0208] In some cases, the disclosed embodiments can be implemented in hardware, firmware, software, or any combination thereof. The disclosed embodiments can also be implemented as instructions carried or stored on one or more transient or non-transitory machine-readable (e.g., computer-readable) storage media, which can be read and executed by one or more processors. For example, the instructions can be distributed via a network or via other computer-readable media. Thus, machine-readable media can include any mechanism for storing or transmitting information in a form readable by a machine (e.g., a computer), including but not limited to, floppy disks, optical disks, optical discs, compact disc-read only memory (CD-ROMs), magneto-optical disks, read only memory (ROM), random-access memory (RAM), erasable programmable read only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), magnetic or optical cards, flash memory, or tangible machine-readable memories for transmitting information (e.g., carrier waves, infrared signals, digital signals, etc.) in electrical, optical, acoustic, or other forms using the Internet. Thus, machine-readable media include any type of machine-readable media suitable for storing or transmitting electronic instructions or information in a form readable by a machine (e.g., a computer).
[0209] In the drawings, some structural or method features may be shown in a particular arrangement and / or order. However, it should be understood that such a particular arrangement and / or ordering may not be required. Rather, in some embodiments, these features may be arranged in a manner and / or order different from that shown in the illustrative drawings. Additionally, the inclusion of a structural or method feature in a particular figure does not imply that such a feature is required in all embodiments, and in some embodiments, these features may not be included or may be combined with other features.
[0210] It should be noted that each unit / module mentioned in the device embodiments of the present application is a logical unit / module. Physically, a logical unit / module can be a physical unit / module, a part of a physical unit / module, or can be implemented as a combination of multiple physical units / module. The physical implementation manner of these logical units / modules themselves is not the most important. The combination of the functions implemented by these logical units / modules is the key to solving the technical problems proposed by the present application. In addition, in order to highlight the innovative part of the present application, the above-mentioned device embodiments of the present application do not introduce units / modules that are not closely related to solving the technical problems proposed by the present application, which does not mean that there are no other units / modules in the above-mentioned device embodiments.
[0211] It should be noted that in the examples and the description of the present patent, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variation thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including one" does not exclude the presence of additional identical elements in the process, method, article or device including the said element.
[0212] Although the present application has been illustrated and described by referring to some preferred embodiments of the present application, those of ordinary skill in the art should understand that various changes can be made in form and detail without departing from the scope of the present application.
Claims
1. An image processing method, applied to an on-board device on a vehicle, characterized in that: The method comprises: In response to a request to detect the surrounding environment of the vehicle, a plurality of images are acquired, where the plurality of images are images from different perspectives acquired by a plurality of sensors of the vehicle; Sampling the multiple images based on multiple sampling rates to obtain multiple feature maps of each image, where the multiple feature maps have different sizes; Projecting a first feature map among the plurality of feature maps into a set plane to obtain a first plane image; Performing feature extraction on the first plane image to obtain first feature data; Performing feature extraction on feature maps other than the first feature map among the plurality of feature maps to obtain second feature data; The first feature data includes N regions, each region includes N first sub-features, and the positions of the N first sub-features in the same region are different; The second feature data includes a first key feature, wherein the first key feature includes N regions, each region includes N second sub-features, and the positions of the N second sub-features in the same region are different; The second feature data includes a first value feature, wherein the first value feature includes N regions, each region includes N third sub-features, and the positions of the N third sub-features in the same region are different; The N parts of the first sub-features, the N parts of the second sub-features, and the N parts of the second sub-features in the same area are fused based on a self-attention mechanism to obtain a first query feature, and third feature data is determined based on the first query feature, where N is greater than or equal to 2; Based on the third characteristic data, the surrounding environment of the vehicle is detected.
2. The method according to claim 1, characterized in that The step of projecting a first feature map from the plurality of feature maps onto a set plane to obtain a first plane image includes: Acquiring acquisition parameters of the multiple sensors; Based on the acquisition parameters, an inverse perspective transformation is performed on the first feature map to generate a first planar image.
3. The method according to claim 1, characterized in that The extracting features from the feature maps other than the first feature map among the plurality of feature maps to obtain second feature data includes: Aligning the feature maps with different sampling rates among the multiple feature maps except the first feature map to obtain a second feature map; Acquiring acquisition parameters of the multiple sensors, and encoding the acquisition parameters based on the size of the second feature graph to obtain sensor features; Performing feature extraction on the second feature map to obtain fourth feature data; The fourth feature data and the sensor feature are fused to determine the second feature data.
4. The method according to claim 2 or 3, characterized in that: The multiple sensors include a camera, and the acquisition parameters include internal parameters and external parameters of the camera.
5. The method according to claim 1, characterized in that: The first query feature includes N regions, each region includes N fourth sub-features, and the positions of the N fourth sub-features in the same region are different; The determining the third feature data based on the first query feature includes: The N parts of the fourth sub-features, the N parts of the second sub-features, and the N parts of the third sub-features at the same position of N different regions are fused based on a self-attention mechanism to obtain the third feature data.
6. A vehicle-mounted device, characterized in that: include: A memory for storing instructions; At least one processor is used to execute the instructions so that the in-vehicle device implements the method according to any one of claims 1 to 5.
7. A computer-readable storage medium, characterized in that: The readable storage medium stores instructions, and when the instructions are executed on a computer, the computer is caused to execute the method according to any one of claims 1 to 5.
8. A computer program product, characterized in that When the computer program product is executed on a device, the device is caused to execute the method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Roadside aerial view target detection method and device, computing equipment and storage medium
CN117994748A
Image processing method, apparatus and device, and computer-readable storage medium
US20230316742A1