Image processing method, readable storage medium, program product and vehicle-mounted equipment
By sampling and feature extraction and fusion of multiple sampling rates for multi-view images, more accurate third feature data is generated, which solves the problem of insufficient BEV space in the prior art, and improves the accuracy and safety of vehicle environment detection.
Patent Information
- Application Number
- CN202510501624.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-21
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-04-21
AI Technical Summary
The prior art fails to fully utilize the features of the multi-view image when projecting a multi-view image onto a BEV space, resulting in the generated BEV space being inaccurate enough.
By sampling multiple images at multiple sampling rates, multiple feature maps of different sizes are generated, feature extraction and fusion are performed, and more accurate third feature data is generated to improve the accuracy of detection of the vehicle's surrounding environment.
By fusing the feature information of multi-view images at different sampling rates, the generated third feature data can retain more information, improve the accuracy of vehicle environment detection, and enhance the safety of vehicle driving.
Smart Images

Figure CN120014605A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image processing technology, and in particular to an image processing method, a readable storage medium, a program product, and a vehicle-mounted device. Background Art
[0002] Multi-view image three-dimensional object detection is currently widely used in many fields and application scenarios, such as autonomous driving scenarios, intelligent transportation scenarios, industrial automation, as well as virtual reality and augmented reality scenarios.
[0003] For example, in assisted driving technology, bird's eye view (BEV) technology can provide multi-perspective image features of the vehicle's surroundings projected into the BEV space, and then estimate the global perspective of the position, size and attributes of dynamic obstacles, thereby facilitating path planning and decision-making for users.
[0004] However, currently, in the process of projecting multi-view images into the BEV space, the features of the multi-view images are not fully utilized, resulting in the BEV space obtained based on the multi-view images being not accurate enough. Summary of the invention
[0005] Embodiments of the present application provide an image processing method, a readable storage medium, a program product, and a vehicle-mounted device.
[0006] In a first aspect, an embodiment of the present application provides an image processing method, which is applied to an on-board device on a vehicle. The method includes: the on-board device obtains multiple images in response to a request to detect the surrounding environment of the vehicle, wherein the multiple images are images of different perspectives collected by multiple sensors of the vehicle. The on-board device samples the multiple images based on multiple sampling rates to obtain multiple feature maps of each image, and the sizes of the multiple feature maps are different. The on-board device projects the first feature map of the multiple feature maps into a set plane to obtain a first plane image. The on-board device extracts features from the first plane image to obtain first feature data. The on-board device extracts features from feature maps other than the first feature map in the multiple feature maps to obtain second feature data. The on-board device fuses the first feature data and the second feature data to obtain third feature data. Then, the on-board device detects the surrounding environment of the vehicle based on the third feature data.
[0007] In some embodiments of the present application, after acquiring multiple images of different perspectives of the vehicle, the vehicle-mounted device can sample each image based on multiple sampling rates, so that each image has feature maps of multiple sizes. For example, the original image can be downsampled 4 times, 32 times, and 64 times. In the process of generating the first environmental data, a first feature map of a size can be used to generate a first plane image. The first feature map is, for example, a feature map corresponding to downsampling 4 times, and the first plane image is, for example, an initial BEV space image. Then the vehicle-mounted device can extract the feature data of the first plane image, and the second feature data of the feature map other than the first feature map in the feature maps of multiple sizes, and fuse the first feature data and the second feature data to obtain the third feature data. It can be understood that the third feature data includes information obtained by sampling multiple images at different sampling rates. Therefore, the third feature data can retain more information in multiple images, so that the vehicle-mounted device can detect the surrounding environment of the vehicle more accurately based on the third feature data. For example, the vehicle-mounted device can perform more accurate detection of pedestrians, other vehicles, lane lines, etc. around the vehicle based on the third feature data, thereby improving the safety of vehicle driving.
[0008] In a possible implementation of the first aspect above, the first feature map among the multiple feature maps is projected into a set plane to obtain a first plane image, including: the vehicle-mounted equipment obtains acquisition parameters of multiple sensors, and based on the acquisition parameters, performs an inverse perspective transformation on the first feature map to generate a first plane image.
[0009] In some embodiments of the present application, the process of projecting the first feature map is, for example, projecting the first feature map to the BEV space through an inverse perspective transformation to obtain an initial feature map of the BEV space.
[0010] In a possible implementation of the first aspect, the feature extraction of feature maps other than the first feature map among the multiple feature maps to obtain the second feature data includes: the vehicle-mounted device aligns feature maps with different sampling rates among the multiple feature maps other than the first feature map to obtain the second feature map. The vehicle-mounted device obtains acquisition parameters of multiple sensors, and encodes the acquisition parameters based on the size of the second feature map to obtain sensor features. The vehicle-mounted device extracts features from the second feature map to obtain fourth feature data. The vehicle-mounted device fuses the fourth feature data with the sensor features to determine the second feature data.
[0011] In some embodiments of the present application, the vehicle-mounted device can align feature maps of different sampling rates in multiple feature maps except the first feature map through a pyramid feature network to generate a second feature map. It can be understood that the sizes of the second feature maps are the same. In the process of generating the second feature data, the vehicle-mounted device can also fuse the encoding corresponding to the acquisition parameters in the sensor on the vehicle with the fourth feature data extracted from the second feature map to obtain the second feature data. It can be understood that the second feature data also includes the acquisition data of the sensor so as to obtain more information based on the image acquired by the sensor.
[0012] In a possible implementation of the first aspect, the multiple sensors include a camera, and the acquisition parameters include internal parameters and external parameters of the camera.
[0013] It can be understood that in some embodiments of the present application, the sensor may be a camera, and the acquisition parameters of the sensor include internal parameters and external parameters of the camera, and the internal parameters may include, for example, focal length, distortion coefficient, etc. The external parameters may include the position and direction of the camera, and the internal parameters and external parameters of the camera are used to describe the internal characteristics of the camera and the spatial information relative to the external coordinates of the vehicle origin, respectively.
[0014] In a possible implementation of the first aspect, the fusing of the first feature data and the second feature data to obtain the third feature data includes: the vehicle-mounted device fuses the first feature data and the second feature data based on a self-attention mechanism to obtain the third feature data.
[0015] In some embodiments of the present application, the first feature data and the second feature data can be fused through a self-attention mechanism. The self-attention mechanism has a global vision and can capture the connection between different positions in the input features, thereby considering global information to improve the accuracy of the third feature data.
[0016] In a possible implementation of the first aspect, the first feature data and the second feature data are fused based on a self-attention mechanism, including: the first feature data includes N parts of regions, each region includes N parts of first sub-features, and the positions of the N parts of first sub-features in the same region are different. The second feature data includes a first key feature, wherein the first key feature includes N parts of regions, each region includes N parts of second sub-features, and the positions of the N parts of second sub-features in the same region are different. The second feature data includes a first value feature, wherein the first value feature includes N parts of regions, each region includes N parts of third sub-features, and the positions of the N parts of third sub-features in the same region are different. The vehicle-mounted device fuses the N parts of the first sub-features, the N parts of the second sub-features, and the N parts of the second sub-features in the same region based on a self-attention mechanism to obtain a first query feature, and determines the third feature data based on the first query feature, wherein N is greater than or equal to 2.
[0017] In some embodiments of the present application, in order to increase the speed of acquiring the third feature data, in the process of fusing the first feature data and the second feature data, the first feature data and the second feature data can be fused in different regions, thereby increasing the speed of fusion of the self-attention mechanism.
[0018] In a possible implementation of the first aspect, the first query feature includes N regions, each region includes N fourth sub-features, and the positions of the N fourth sub-features in the same region are different. Determining the third feature data based on the second query feature includes: the vehicle-mounted device fuses the N fourth sub-features, the N second sub-features, and the N third sub-features at the same positions of the N different regions based on a self-attention mechanism to obtain the third feature data.
[0019] In some embodiments of the present application, in order to ensure the effect of the self-attention mechanism fusing the first feature data and the second feature data, the result of the first feature data and the second feature data being fused for the first time based on the self-attention mechanism can be used as new data, i.e., the first query data, and cross-fused with the second feature data again based on the self-attention mechanism, and the corresponding partitions are different during the cross-fusion process, which can further improve the accuracy of obtaining the third feature data through the self-attention mechanism, so that the third feature data retains more information.
[0020] In a second aspect, the present application provides a vehicle-mounted device, comprising: a memory for storing instructions; and at least one processor for executing instructions to enable the device to implement the first aspect and any possible implementation method of the first aspect. The beneficial effects that can be achieved in the second aspect can refer to the beneficial effects of the method provided in any implementation of the first aspect, and will not be repeated here.
[0021] In a third aspect, the present application provides a computer-readable storage medium, in which instructions are stored, and when the instructions are executed by a device, the computer implements the method provided in the first aspect and any possible implementation of the first aspect. The beneficial effects that can be achieved in the third aspect can refer to the beneficial effects of the method provided in any implementation of the first aspect, and will not be repeated here.
[0022] In a fourth aspect, the present application provides a computer program product, which, when running on a device, enables the device to implement the method provided in the first aspect and any possible implementation of the first aspect. The beneficial effects that can be achieved in the fourth aspect can refer to the beneficial effects of the method provided in any implementation of the first aspect, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] Figure 1A A flow chart for generating a BEV space is shown; Figure 1B A schematic diagram of a vehicle in the driving process is shown; Figure 1C A flow chart of generating a BEV space based on IPM is shown; Figure 2 According to an embodiment of the present application, a flowchart of an implementation method of an image processing method is shown; Figure 3 According to some embodiments of the present application, a schematic diagram of generating third characteristic data is shown; Figure 4 A schematic diagram of a cross-attention mechanism is shown according to some embodiments of the present application; Figure 5A According to some embodiments of the present application, a schematic diagram of local window division is shown; Figure 5B According to an embodiment of the present application, a schematic diagram of global window division is shown; Figure 6 According to some embodiments of the present application, a schematic structural diagram of a vehicle-mounted device is shown. DETAILED DESCRIPTION
[0024] The illustrative embodiments of the present application include, but are not limited to, an image processing method, a readable storage medium, a program product, and a vehicle-mounted device.
[0025] In order to make the purpose, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be described in detail below in conjunction with the accompanying drawings and specific implementation methods.
[0026] To facilitate understanding, some of the terms and related technologies involved in this application are explained below.
[0027] 1.BEV technology: BEV technology is a technology that uses neural networks to convert image information from image space to BEV space. BEV technology can simplify complex three-dimensional environments into two-dimensional images. For example, in the field of autonomous driving, BEV technology can generate a panoramic view from above the vehicle based on the spatial information around the vehicle, thereby displaying the environment around the vehicle in all directions, including the front, back, left, right and top. This enables the autonomous driving system to better understand the surrounding environment and improve the accuracy of perception and decision-making.
[0028] Next, the process of vehicle generating BEV space is introduced.
[0029] For example, Figure 1A A flow chart for generating a BEV space is shown.
[0030] It can be understood that the following processes can be executed by a vehicle-mounted device, and the vehicle-mounted device in the embodiment of the present application can also be referred to as a vehicle-mounted terminal. The vehicle-mounted terminal can be a mobile phone, a vehicle computer, a terminal in self-driving, a wireless terminal in transportation safety, a terminal in a smart city, etc. The following will be explained using a vehicle computer as an example. However, it can be understood that the technical solution described in the present application is applicable to the various vehicle-mounted devices for three-dimensional target detection mentioned above, not limited to vehicle computers.
[0031] like Figure 1A As shown, the process includes: S101, acquiring images of a vehicle from multiple perspectives.
[0032] For example, Figure 1B A schematic diagram of a vehicle in the driving process is shown.
[0033] It is understood that the vehicle 10 is usually equipped with multiple cameras (for example, including front view, side view, rear view cameras, etc.). Figure 1A As shown, during the driving process of the vehicle 10 , multiple cameras can be used to capture images of multiple perspectives of the vehicle 10 at the same time through a timestamp synchronization mechanism or a hardware synchronization mechanism.
[0034] For example, the vehicle 10 may capture images of the vicinity of the vehicle 10 using a plurality of cameras. Figure 1B The image captured by the camera may include, for example, a first car 01, a pedestrian 02, a first building 03, a second car 04 and a second building 05.
[0035] The vehicle computer deployed on the vehicle 10 can obtain pictures taken by multiple cameras on the vehicle 10 at the same time, and use the pictures as images of multiple perspectives of the vehicle 10.
[0036] S102, extracting feature maps from multi-view images based on a feature extraction module.
[0037] For example, an image extraction module may be deployed on the vehicle computer, and features of multi-view images (i.e., original images) may be extracted through the image extraction module, thereby generating feature maps of images of each view. For example, the feature map obtained by the feature extraction module may be a feature map that can be downsampled n times (e.g., 4 times) compared to the original image.
[0038] Among them, downsampling (such as pooling operation) is a method used to reduce data dimensions in convolutional neural networks. When an image is downsampled 4 times, it means that the resolution of the data is reduced by 4 times in a certain dimension of the feature map (usually width and height). This operation is usually used to reduce the amount of calculation, extract more abstract features, and increase the robustness of the model. The feature map after downsampling 4 times retains the important information in the original image, but the details are ignored or simplified. In an embodiment of the present application, the image downsampled 4 times can be referred to as a stride4 feature map. It can be understood that the image features of images downsampled 8 times, 16 times, 32 times, 64 times, etc. from multiple perspectives can also be extracted through the feature extraction module. This application does not limit the process of image feature extraction.
[0039] S103, performing BEV space feature fusion on feature maps of multi-view images based on a BEV space feature fusion module to obtain a BEV space.
[0040] For example, after extracting feature maps of images from multiple perspectives, the vehicle computer can perform inverse perspective mapping (IPM) on the image features according to the internal and external parameters of each camera, thereby fusing the BEV spatial features to obtain the BEV spatial features.
[0041] Among them, IPM is based on the geometric principle of perspective projection, and uses mathematical algorithms to convert a two-dimensional image containing three-dimensional information into a top view or bird's-eye view. This process relies on a transformation matrix, which is usually determined by the camera's internal parameters (such as focal length, principal point, etc.) and external parameters (such as the position and direction of the camera).
[0042] For example, Figure 1C A flow chart of generating BEV space based on IPM is shown.
[0043] like Figure 1CAs shown, exemplarily, after the vehicle computer extracts the feature map of the image of multiple perspectives of the vehicle through the feature extraction module, the feature map can be IPM transformed. For example, in some embodiments of the present application, the image feature extracted by the vehicle computer is the feature map of stride4. The vehicle computer can input the feature map of stride4 into the spatial feature fusion module for IPM transformation. Then, the vehicle computer can convert the two-dimensional image containing three-dimensional information into a top view or a bird's-eye view, thereby generating a BEV space.
[0044] S104, fusion of BEV space based on the time series fusion module.
[0045] For example, the temporal fusion module integrates temporal and spatial information, which can make up for the limitations of single-frame BEV spatial perception, thereby increasing the receptive field, improving the inter-frame jump and target occlusion problems of target detection, and more accurately judging the target movement speed. It also plays an important role in target prediction and tracking. By introducing temporal information, the temporal fusion module makes the perception results more stable, improves the accuracy of the vehicle's judgment of road conditions, and thus enhances the safety of autonomous driving.
[0046] In some embodiments, by integrating the vehicle's own posture and position information into the time series fusion module, objects and obstacles in the surrounding environment can be more accurately perceived. This information helps the model better understand the relative relationship between the vehicle and the surrounding environment, thereby improving the accuracy of perception.
[0047] S105, optimizing the BEV space by acquiring deeper feature information of the image based on the BEV deep feature extraction module.
[0048] For example, in some embodiments, in the BEV space, the deep feature extraction module can use deep learning algorithms (such as convolutional neural networks (CNNs) etc.) to extract deep feature information from the converted images. These feature information include the contour, shape, texture, etc. of the object, which is the basis for subsequent recognition and tracking tasks. For example, semantic information such as lane lines, obstacles, pedestrians, and vehicles can be annotated in the BEV image. And dynamic update, dynamically adjust the BEV space according to the real-time update of sensor data. For example, the BEV space generated by the vehicle computer can also refer to Figure 1B of the view.
[0049] S106, performing multi-task detection based on the BEV space.
[0050] For example, after obtaining the BEV space, the vehicle computer can perform multi-task detection based on the BEV space, such as vehicle 3D detection, pedestrian / cyclist 3D detection, lane line detection, ground discrete element detection, ground passable area detection or other detection.
[0051] It can be understood that in the above process of generating the BEV space, the BEV space feature fusion is mainly performed based on the BEV space feature fusion module in S103 to obtain the BEV space. However, when performing the IPM transformation, only one downsampled feature map of the multi-view image (such as the stride4 feature map) is used, which results in the loss of some information in the image during the sampling process, resulting in the BEV space obtained by the IPM transformation being inaccurate.
[0052] As mentioned above, in the process of processing multi-view images (such as performing IPM transformation to generate BEV space), the information of the multi-view images is not fully utilized, resulting in inaccurate spatial information generated based on the multi-view images.
[0053] In order to solve the problem that the spatial information generated based on an image feature is not accurate enough, the present application proposes an image processing method, which obtains multiple images in response to a request to detect the surrounding environment of a vehicle, and the multiple images are images of different perspectives collected by multiple sensors of the vehicle. The multiple images are sampled based on multiple sampling rates to obtain multiple feature maps of each image, and the sizes of the multiple feature maps are different. The first feature map (for example, the feature map of stride4) of the multiple feature maps is projected into a set plane to obtain a first plane image (for example, to generate a bird's-eye view of the BEV).
[0054] Perform feature extraction on the first plane image to obtain first feature data (features corresponding to the BEV bird's-eye view, such as the query feature of BEV mentioned below), perform feature extraction on feature maps other than the first feature map (such as the feature map of stride16 and the feature map of stride32) among the multiple feature maps to obtain second feature data (such as the key feature and value feature of multiple perspectives mentioned below), fuse the first feature data and the second feature data to obtain third feature data (such as the spatial fusion feature of BEV mentioned below); and detect the surrounding environment of the vehicle based on the third feature data.
[0055] Through the above scheme, the vehicle-mounted equipment can extract feature maps of multiple scales from images of multiple perspectives, and then fuse the features corresponding to the feature maps of different sizes to obtain third feature data, so that the third feature data retains more information of the multi-perspective images, so as to improve the accuracy of the vehicle-mounted equipment in monitoring the vehicle's surrounding environment based on the third feature data.
[0056] In some embodiments of the present application, the first feature data and the second feature data may be fused through a self-attention mechanism. The self-attention mechanism can capture long-range dependencies in the image, thereby improving the accuracy of generating the first environment image.
[0057] Next, the image processing method in the embodiment of the present application is introduced.
[0058] For example, Figure 2 According to an embodiment of the present application, a flowchart of an implementation of an image processing method is shown.
[0059] Exemplarily, the execution entities of the following processes may all be vehicle-mounted devices.
[0060] like Figure 2 As shown, the method includes: S201, in response to a request to detect the surrounding environment of a vehicle, a plurality of images are acquired, where the plurality of images are images of different perspectives acquired by a plurality of sensors of the vehicle.
[0061] For example, in some embodiments of the present application, a vehicle can collect images of different perspectives of the vehicle's surrounding environment through multiple sensors during driving. For example, the vehicle can be equipped with front-view, side-view, and rear-view cameras, and the cameras on the vehicle can collect images of different perspectives at the same time.
[0062] After starting the assisted driving or automatic driving function, the on-board equipment on the vehicle can detect a request to detect the vehicle's surrounding environment. After the on-board equipment responds to the request to detect the vehicle's surrounding environment, it can obtain multiple images from different perspectives captured by various cameras on the vehicle at the same time.
[0063] S202, sampling the multiple images based on multiple sampling rates to obtain multiple feature maps of each image, where the sizes of the multiple feature maps are different.
[0064] Exemplarily, in some embodiments of the present application, after the vehicle-mounted device acquires multiple images (hereinafter referred to as multi-view images), the multi-view images can be downsampled using a variety of different sampling rates to obtain feature maps of different sizes.
[0065] For example, the vehicle-mounted device may perform downsampling through a neural network (backbone network), and the downsampling sampling rate may include, for example, downsampling by 4 times, downsampling by 32 times, downsampling by 64 times, etc. The embodiment of the present application does not limit the downsampling sampling rate.
[0066] S203, projecting a first feature map among the plurality of feature maps onto a set plane to obtain a first plane image.
[0067] In some embodiments of the present application, the vehicle-mounted device may project a first feature map from a plurality of feature maps into a set plane to obtain a first plane image. For example, the vehicle-mounted device may obtain acquisition parameters of a plurality of sensors; based on the acquisition parameters, the first feature map may be subjected to an inverse perspective transformation to generate a first plane image.
[0068] Exemplarily, the on-board device may select a first feature map of a size from feature maps of different sizes, and perform an inverse perspective transformation on the first feature map to generate a first plane map centered on the vehicle, which may be an initial spatial feature map of the BEV.
[0069] For example, the first feature map may be a feature map that is downsampled 4 times, that is, a feature map of stride 4. The vehicle-mounted device may obtain internal and external parameters of each camera on the vehicle, and perform an inverse perspective transformation on the feature map of stride 4 based on the internal and external parameters of each camera, thereby generating a BEV initial spatial feature map.
[0070] In other embodiments, the first feature map may also be a feature map downsampled 16 times, a feature map downsampled 32 times, or a feature map downsampled 64 times, etc. The embodiments of the present application do not limit the specific sampling rate of the first feature map.
[0071] S204: extract features from the first plane image to obtain first feature data.
[0072] In some embodiments of the present application, the vehicle-mounted device may perform feature extraction on the first plane image to obtain first feature data.
[0073] For example, the vehicle-mounted device downsamples the BEV initial spatial feature map, and the downsampling sampling rate may be, for example, 32 times. It can be understood that the downsampling sampling rate of the BEV initial spatial feature map is a target sampling rate preset by the vehicle-mounted device. In other embodiments, the sampling rate may also be 16 times, 8 times, 64 times, etc. After the BEV initial spatial feature map is downsampled, feature extraction is performed on the BEV initial spatial feature map to generate first feature data.
[0074] In other embodiments, the first feature data also includes the BEV spatial position code of the BEV initial spatial feature map. For example, in some embodiments, the downsampled BEV initial spatial feature map can be encoded by the BEV spatial position coding module to obtain the BEV spatial position code, and the BEV spatial position code is added and fused with the features corresponding to the downsampled BEV initial spatial feature map to obtain the first feature data of the BEV initial spatial feature map. In some embodiments, the first spatial data can be used as the query feature of the BEV to perform self-attention fusion with other features, and the self-attention fusion process is described in detail below.
[0075] S205, performing feature extraction on feature maps other than the first feature map among the plurality of feature maps to obtain second feature data.
[0076] In some embodiments of the present application, since the first feature map is used to convert into a BEV initial spatial feature map, the on-board device can perform feature extraction on the remaining feature maps of different sizes in the multiple feature maps to obtain second feature data.
[0077] In some embodiments, feature extraction may be performed on feature maps other than the first feature map among the plurality of feature maps in the following manner to obtain second feature data: The feature maps with different sampling rates among the multiple feature maps except the first feature map are aligned to obtain a second feature map.
[0078] For example, in some embodiments of the present application, since the multiple feature maps include a feature map of stride4, a feature map of stride32, and a feature map of stride64. The feature map of stride4 is used to convert into a BEV initial spatial feature map, the remaining feature map of stride64 and the feature map of stride32 can be aligned, and then the aligned feature map is determined. For example, the feature map of stride64 can be upsampled by a pyramid feature module, and then fused with the feature map of stride32 to obtain a multi-view channel alignment feature map of stride32, and the multi-view channel alignment feature map is the second feature map.
[0079] In some embodiments of the present application, the vehicle-mounted device may also fuse the features of the acquisition parameters of multiple sensors with the features corresponding to the second feature graph to generate second feature data.
[0080] For example, the vehicle-mounted device may obtain acquisition parameters of multiple sensors, and encode the acquisition parameters based on the size of the second feature map to obtain sensor features.
[0081] In some embodiments, the sensor of the vehicle-mounted device may include a camera, for example, and the acquisition parameters of the camera may be, for example, internal parameters and external parameters of the camera. The vehicle-mounted device may encode the internal parameters and external parameters of the camera to obtain sensor characteristics.
[0082] The vehicle-mounted device may extract features from the second feature map to obtain fourth feature data. The vehicle-mounted device fuses the fourth feature data with the sensor feature to determine the second feature data.
[0083] It can be understood that the second feature map is the multi-view channel alignment feature map mentioned above, and the vehicle-mounted device can extract features from the multi-view channel alignment feature map to obtain the fourth feature data. It can be understood that the size of the multi-channel alignment feature map is the same as the size of the feature map corresponding to the target sampling rate preset by the vehicle-mounted device, that is, the feature map of stride32. The vehicle-mounted device can also fuse the sensor features and the fourth feature data to determine the second feature data. It can be understood that the second feature data includes information in the feature map of stride32 and the feature map of stride64.
[0084] It can be understood that the second feature data can also be generated by feature graphs with more sampling rates or feature graphs with fewer sampling rates. The embodiments of the present application do not limit the types and numbers of sampling rates of feature graphs for generating the second feature data.
[0085] In some embodiments of the present application, the second feature data may include a key feature (as an example of the first key feature) and a value feature (as an example of the first value feature), wherein the key feature is obtained by fusing the feature obtained by the second feature graph based on a sampling method and the sensor feature. The value feature is a feature obtained by the second feature graph based on another sampling method. The acquisition process of the second feature data is described in detail below.
[0086] S206: Merge the first feature data and the second feature data to obtain third feature data.
[0087] In some embodiments of the present application, the vehicle-mounted device may fuse the first feature data and the second feature data through a self-attention mechanism to obtain third feature data.
[0088] For example, the first feature data is linearly transformed to obtain query features, and the second feature data is linearly transformed differently to obtain corresponding key features and value features.
[0089] Then, the vehicle-mounted device performs calculations based on the query feature, key feature, and value feature using formula (1): (one); Among them, Q is the query feature, K is the key feature, and V is the value feature. Indicates self-attention operation on Q, K and V. It means that after multiplying the Q and K matrices, the normalized attention weights are obtained through the softmax operator. It can be understood that The output feature can be the feature of the query feature fusion of BEV. The feature of the query feature fusion of BEV is upsampled 32 times to obtain the third feature data. It can be understood that since the first feature map in S201 is obtained by downsampling, and in S202, the first feature data of the query is obtained by the BEV spatial encoding module extracting the feature of the downsampled BEV initial spatial feature map, it is necessary to restore the feature of the query feature fusion of BEV by upsampling to obtain the third feature data. That is to say, the third feature data includes the information in the feature map of stride4, the feature map of stride32, and the feature map of stride64 of the multi-view image, so that the information in the multi-view image is retained more completely.
[0090] S207: Detect the surrounding environment of the vehicle based on the third characteristic data.
[0091] For example, in some embodiments of the present application, after the vehicle-mounted device acquires the third characteristic data, the surrounding environment of the vehicle can be detected according to the third characteristic data.
[0092] It can be understood that the third feature data includes information retained by sampling the multi-view images at different sampling rates, so the third feature data can retain more data in the multi-view images. Therefore, the accuracy of detecting the surrounding environment of the vehicle through the third feature data is higher.
[0093] For example, the vehicle-mounted equipment can more accurately detect other vehicles, pedestrians, road lines, traffic signs, etc. around the vehicle based on the third feature, and then provide various services based on the detection results, such as automatic driving, assisted driving, road planning, etc., thereby improving the safety of vehicle driving.
[0094] Next, the process of obtaining the third characteristic data in the embodiment of the present application is introduced.
[0095] For example, Figure 3 According to some embodiments of the present application, a schematic diagram of generating third characteristic data is shown.
[0096] like Figure 3 As shown, the vehicle-mounted device can be configured with an IPM inverse perspective transformation module 305, an internal and external parameter encoding module 306, a pyramid feature module, a BEV feature downsampling module 309, a BEV spatial position encoding module 310, a multi-view key feature and value feature generation module 314, a transformer hierarchical self-attention module 316, and a BEV feature upsampling module 318. The following introduces the process of generating the third feature data.
[0097] Multi-view stride4 image feature map 301.
[0098] In some embodiments of the present application, the multi-view stride4 image feature map (as an example of the first feature map) 301 is obtained by downsampling each view image by 4 times through the image feature extraction module. Exemplarily, the size of each multi-view image is (3, img_h, img_w), where 3 represents the number of channels, which corresponds to the three color channels of RGB (red, green, and blue). img_h represents the height of the image, that is, the number of pixels in the vertical direction of the image, and img_w represents the width of the image, that is, the number of pixels in the horizontal direction of the image.
[0099] The feature map size after downsampling the multi-view image by 4 times is (c4, img_h / 4, img_w / 4), where c4 represents the number of channels of the multi-view stride4 image feature map 301. It can be understood that if a convolution layer is applied during the downsampling process, the value of c4 will depend on the number of output channels of the convolution layer. For example, if a convolution layer with 16 output channels is applied, the value of c4 will be 16.
[0100] Camera internal and external reference 302.
[0101] The camera internal and external parameters 302 include the internal parameters and external parameters of the camera. Among them, the internal parameters are related to the hardware structure and imaging principle of the camera, and are used to represent the mapping relationship of light from three-dimensional space to a two-dimensional plane through the camera. The internal parameters are represented by Intri_mat, a 3×3 matrix. In some embodiments, the internal parameters also include camera distortion coefficients, and the camera external parameters are represented by Extri_mat, a 4×4 matrix. One viewing angle image corresponds to one camera internal and external parameters 302, which are used to describe the internal characteristics of the camera and the external spatial information relative to the vehicle origin coordinates.
[0102] Multi-view stride32 image feature map 303.
[0103] Similar to the multi-view stride4 image feature map 301, the size of the feature map obtained after the multi-view image is downsampled 32 times by the image feature extraction module is (c32, img_h / 32, img_w / 32), that is, the multi-view stride32 image feature map 303 is generated. It can be understood that the multi-view stride32 image feature map 303 is one of the feature maps except the first feature map.
[0104] Multi-view stride64 image feature map 304 .
[0105] Similar to the multi-view stride4 image feature map 301, the multi-view stride64 image feature map 304 is obtained after the multi-view image is downsampled 64 times by the image feature extraction module, and its feature map size is (c64, img_h / 64, img_w / 64). It can be understood that the multi-view stride64 image feature map 304 is also one of the feature maps except the first feature map.
[0106] IPM inverse perspective transformation module 305 .
[0107] The IPM inverse perspective transformation module 305 can obtain the camera internal and external parameters 302, so as to perform an IPM inverse perspective transformation on the multi-view stride4 image feature map 301.
[0108] Referring to the process of S203, the vehicle-mounted equipment can perform IPM transformation through the multi-view stride4 image feature map 301 and the internal and external parameters 302 of the camera to generate a BEV initial spatial feature map (as an example of the first plane image) 308. Based on this initialization method, it is guaranteed that the attention mechanism of the subsequent transformation model (transformer) has better convergence and the model training is more stable. After IPM, the BEV initial spatial feature map 308 is obtained, and its size is (bev_c, bev_h, bev_w), where bev_c, bev_h, bev_w are the number of channels, height and width of the BEV initial spatial feature map 308 respectively. It can be understood that the first feature data can be obtained by extracting the features of the BEV initial spatial feature map 308.
[0109] Internal and external parameter encoding module 306.
[0110] In some embodiments of the present application, the internal and external parameter encoding module 306 inputs the corresponding internal parameter matrix intri_mat and the external parameter matrix extri_mat for each view, and generates a multi-view camera encoding feature 312 for each stride32 image feature. It can be understood that since the target sampling rate of the vehicle-mounted device is stride32, that is, downsampling by 32 times, the internal and external parameter encoding module 306 can align with the stride32 image feature when encoding the internal and external parameters of the camera.
[0111] For example, the internal and external parameter encoding module 306 can generate a coordinate point matrix img_plane with a size of (3, img_h / 32, img_w / 32) based on the pixel positions corresponding to the multi-view stride32 image feature map 303, where 3 represents three channels, wherein channels 1 and 2 respectively represent the x (e.g., horizontal coordinate) and y (e.g., vertical coordinate) points of each pixel point in the image coordinate system of the multi-view stride32 image feature map 303, and channel 3 is a fixed fill value of 1.
[0112] The internal and external parameter encoding module 306 obtains intri_mat_inv and extri_mat_inv by performing inverse transformation on the internal parameter matrix intri_mat and the external parameter matrix extri_mat. Then, img_pos is generated by matrix multiplication and fixed padding value operation, where img_pos = extri_mat_inv × pad(intri_mat_inv × img_plane). The × represents matrix multiplication, pad represents the alignment of dimensions by fixed value padding, and img_pos is an image space position matrix with a size of (4, img_h / 32, img_w / 32), which represents the position correspondence of each image feature pixel in the vehicle coordinate system (VCS) of the vehicle.
[0113] Finally, img_pos is transformed through the fully connected network and the feature dimension is transformed to obtain the multi-view camera encoding feature 312, whose dimension is (d, img_h / 32, img_w / 32). The same process operation is performed on the camera of each view to obtain the multi-view camera encoding feature 312.
[0114] Feature pyramid network (FPN) 307.
[0115] FPN 307 can fuse feature maps of different scales to achieve effective detection of objects of different sizes. For example, in some embodiments of the present application, two downsampling scales, including image features of stride32 and stride64, can be fused through FPN 307 to obtain channel-aligned image features after scale fusion, whose size is (d, img_h / 32, img_w / 32), and whose length and width scales are consistent with the stride32 image feature map 301, and the channel dimension is consistent with the camera encoding feature. The feature maps of different scales of each perspective are fused separately, and a multi-perspective channel alignment feature map (as an example of a second feature map) 313 can be obtained. It can be understood that feature extraction of the multi-perspective channel alignment feature map 313 can obtain the second feature data.
[0116] BEV feature downsampling module 309 .
[0117] In some embodiments of the present application, the BEV initial spatial feature map 308 is a dense BEV space representation (a BEV space that contains information about each pixel of the feature map), and the computational time complexity of the transformer attention operation in this dense space is large. The BEV feature downsampling module 309 can transform the BEV initial spatial feature map 308 of size (bev_c, bev_h, bev_w) into a downsampled BEV feature of size (d, bev_h / 4, bev_w / 4) by stacking a convolutional network and a PixelUnshuffle operation, thereby reducing the pixel resolution in the BEV space, while keeping the channel dimension consistent with the multi-view camera encoding feature 312.
[0118] BEV spatial position encoding module 310.
[0119] The encoding process of the BEV spatial position encoding module 310 is similar to the img_pos encoding in the internal and external parameter encoding module 306. Each pixel point of the downsampled BEV initial spatial feature map corresponds to the specific position point coordinates in the VCS coordinate system of the vehicle, represented by the bev_pos matrix, whose size is (3, bev_h / 4, bev_w / 4), where channels 1 and 2 represent the x and y points of each pixel point in the VCS left system respectively, and channel 3 is the sqrt(x×x + y×y) value, which represents the actual distance of the pixel point from the origin coordinate in the VCS coordinate system. The bev_pos matrix is passed through a fully connected network to align the channel dimension and obtain a BEV spatial position encoding with a size of (d, bev_h / 4, bev_w / 4). Finally, the BEV spatial position encoding is added and fused with the features corresponding to the downsampled BEV initial spatial feature map to obtain the BEV query feature (as an example of the first feature data) 311. The BEV query feature 311 can be used as one of the key inputs of the transformer hierarchical self-attention module 316.
[0120] Multi-view key feature and value feature generation module 314.
[0121] In some embodiments of the present application, the multi-view channel alignment feature map 313 can be processed in two ways. For example, in one processing method, the multi-view channel alignment feature map 313 first passes through a fully connected network to perform feature transformation, and is added and fused with the multi-view camera encoding feature 312 to generate a multi-view key feature, whose size remains unchanged, still (d, img_h / 32, img_w / 32). In another processing, the multi-view channel alignment feature map 313 passes through another fully connected network to perform feature transformation to generate a multi-view value feature. It can be understood that the key feature (as an embodiment of the first key feature) and the value feature (as an embodiment of the first value feature) can be used as an embodiment of the second feature data in the embodiments of the present application, that is, the second feature data includes the multi-view key feature and value feature 315.
[0122] Transformer level self-attention module 316.
[0123] In some embodiments of the present application, the query feature 311 of BEV and the key feature and value feature 315 of multi-view are used as inputs, and an attention operation is performed to obtain the query fusion feature 317 of BEV. It can be understood that a better BEV feature expression is obtained by fusing multi-view image features through the self-attention mechanism. Exemplarily, the process of self-attention mechanism fusion is described in detail below.
[0124] BEV feature upsampling module 318 .
[0125] Since the BEV initial spatial feature map 308 has been downsampled before the transformer level self-attention module 316, it is necessary to restore the resolution size of the BEV query fusion feature 317 in the VCS space through the convolutional network and Upsample operation. After the upsampling module, the BEV spatial fusion feature 319 with the same size as the BEV initial spatial feature map 308 is obtained. The BEV spatial fusion feature 319 can be used as the final output of the transformer level self-attention module 316. It can be understood that the BEV spatial fusion feature 319 can be used as an embodiment of the third feature data.
[0126] Exemplarily, the self-attention operation can usually be processed by formula (1). In some embodiments of the present application, the number of query features of BEV is BEV_H×BEV_W (where BEV_H is bev_h / 4, BEV_W is bev_w / 4), the number of cameras is M, the size of a single image feature map is D×H×W, the time complexity is O(BEV_H×BEV_W×M×H×W×D), and the space complexity is O(BEV_H×BEV_W×N×H×W + (BEV_H×BEV_W+M×H×W)×D). The computational complexity of the transformer-level self-attention module 316 is strongly related to the size of the BEV query feature and the size of the key feature and value feature of the multi-view image. Therefore, the way to reduce its computational complexity is to reduce the size of the BEV query feature and the size of the key feature and value feature of the multi-view image for the attention calculation.
[0127] For example, in some embodiments of the present application, the query feature 311 of the BEV and the key features and value features of the multi-view image are simultaneously divided into different areas through the transformer hierarchical self-attention module 316, and cross-attention operations are performed between the areas, which greatly reduces the computational complexity of the attention operation, so that the transformer hierarchical self-attention module 316 is deployed on the chip of the vehicle-mounted equipment (such as the J6 chip) with a higher frame rate, meeting the real-time detection efficiency.
[0128] Therefore, in some embodiments of the present application, the vehicle-mounted device may fuse the first feature data and the second feature data based on the self-attention mechanism. For example, the vehicle-mounted device may divide the first feature data into N regions, each region including N first sub-features, and the positions of the N first sub-features in the same region are different.
[0129] It can be understood that the first feature data is the feature data obtained by extracting the BEV initial spatial feature map 308, that is, the BEV query feature 311. In some embodiments of the present application, in order to implement the cross-attention operation, the BEV query feature 311 can be divided into N parts, each area includes N parts of the first sub-features, and the positions of the N parts of the first sub-features in the same area are different.
[0130] In some embodiments of the present application, the key feature in the second feature data also includes N parts of regions, each region includes N parts of second sub-features, and the positions of the N parts of second sub-features in the same region are different.
[0131] The value feature in the second feature data also includes N regions, each region includes N third sub-features, and the positions of the N third sub-features in the same region are different. It can be understood that the process of extracting key features and extracting value features can refer to the part of the multi-view key feature and value feature generation module 314.
[0132] Then, N parts of the first sub-features, N parts of the second sub-features, and N parts of the second sub-features of the same area are fused based on the self-attention mechanism to obtain the second query feature, and the third feature data is determined based on the second query feature, where N is greater than or equal to 2.
[0133] It can be understood that by dividing the query feature, key feature, and value feature of BEV into multiple regions, and performing self-attention optimization on multiple sub-features (including the first sub-feature, the second sub-feature, and the third sub-feature) in each region, the size of the self-attention optimization can be reduced, thereby improving the efficiency of the self-attention optimization.
[0134] For example, Figure 4 According to some embodiments of the present application, a flowchart for implementing a cross-attention mechanism is shown.
[0135] Exemplarily, the execution entities of the following processes are all transformer-level self-attention modules.
[0136] like Figure 4 As shown, the process includes: S401A, perform local window division on the query feature of BEV.
[0137] In some embodiments of the present application, the transformer hierarchical self-attention module divides the query feature of BEV (as an embodiment of the first feature data) equally into N windows (i.e., into N partial regions) in local areas, and promotes the window to the Batch dimension, i.e., each window includes N parts of the first sub-features, and each part of the first sub-features of the same window is at a different position of the window.
[0138] For example, Figure 5A According to some embodiments of the present application, a schematic diagram of local window division is shown.
[0139] like Figure 5AAs shown, in some embodiments of the present application, the query feature of the original feature BEV is taken as an example, where N is 4, that is, the query feature of BEV can be divided into 4 regions, for example, region 1, region 2, region 3 and region 4. Each region includes 4 sub-features. When the original feature is the query feature of BEV, the sub-feature corresponding to the original feature is the first sub-feature. Taking region 1 as an example, the 4 first sub-features are respectively at position 1, position 2, position 3 and position 4.
[0140] S401B, perform local window division on the key feature and the value feature.
[0141] Exemplarily, the local window division of the key feature and the value feature is similar to the local window division of the query feature of BEV.
[0142] For example, refer to Figure 5A , take the original feature as the key feature as an example, where the key feature is divided into 4 regions, for example, region 1, region 2, region 3 and region 4. Each region includes 4 sub-features. When the original feature is the key feature, the sub-feature corresponding to the original feature is the second sub-feature. Taking region 1 as an example, the 4 second sub-features are at position 1, position 2, position 3 and position 4 respectively.
[0143] Reference Figure 5A , take the original feature as value feature as an example, where the value feature is divided into 4 regions, for example, region 1, region 2, region 3 and region 4. Each region includes 4 sub-features. When the original feature is value feature, the sub-feature corresponding to the original feature is the third sub-feature. Taking region 1 as an example, the 4 third sub-features are respectively at position 1, position 2, position 3 and position 4.
[0144] S402, cross-attention fusion is performed on the query feature, key feature and value feature of BEV to obtain the BEV updated query feature.
[0145] After completing the local window division of each feature, the 4 sub-features of each feature in the same area can be cross-attention operated, and the sub-features in different areas will not be cross-attention operated. In this way, the size of self-attention fusion can be reduced, thereby improving the efficiency of self-attention fusion. For example, when N=4, the size of BEV's query feature will change from (1, D, BEV_H, BEV_W) to (4, D, BEV_H / 2, BEV_W / 2), and its time complexity will be reduced to 1 / 4 of the original global attention operation, thereby improving the efficiency of self-attention fusion.
[0146] S403A, perform local window restoration on the BEV updated query feature.
[0147] Exemplarily, after executing the process of S402, the transformer hierarchical self-attention module restores the output of the self-attention optimization, and the output of the self-attention optimization can be used as the BEV update query feature. It can be understood that the BEV update query feature output by the self-attention optimization is also windowed, and the BEV update query feature can be restored from (4, D, BEV_H / 2, BEV_W / 2) to (1, D, BEV_H, BEV_W). It can be understood that the query feature of BEV is restored to the first query feature.
[0148] S403B, perform local window restoration on the key feature and the value feature.
[0149] For example, the transformer-level self-attention module can restore key features and value features by restoring BEV to update query features.
[0150] In some embodiments of the present application, after the vehicle-mounted device obtains the first query feature through the transformer hierarchical self-attention module, it can continue to perform cross-attention optimization.
[0151] For example, the first query feature includes N regions, each region includes N fourth sub-features, and the positions of the N fourth sub-features in the same region are different.
[0152] In some embodiments of the present application, the first query feature is a BEV update query feature, and the BEV update query feature may also include N partial regions, each region including N partial fourth sub-features of the BEV update query feature. The N partial fourth sub-features of the same region have different positions.
[0153] In some embodiments, the third feature data may be determined based on the second query feature in the following manner: The vehicle-mounted device can fuse N parts of the fourth sub-features, N parts of the second sub-features, and N parts of the third sub-features at the same position of N different areas based on the self-attention mechanism to obtain the third feature data.
[0154] For example, the following introduces the fusion process of the fourth sub-feature of different regions with the corresponding second sub-feature and the third sub-feature based on the self-attention mechanism.
[0155] S404A, perform global window division on the BEV update query feature.
[0156] Exemplarily, the transformer hierarchical self-attention module divides the BEV updated query feature into N windows. Similar to 401A, the BEV updated query feature is equally divided into N windows (i.e., divided into N parts), and the window is promoted to the Batch dimension, i.e., each window includes N parts of the fourth sub-features, and each part of the fourth sub-features is at a different position in the window.
[0157] For example, Figure 5B According to an embodiment of the present application, a schematic diagram of global window division is shown.
[0158] like Figure 5B As shown, in some embodiments of the present application, the original feature is a BEV update query feature as an example, where N is 4, that is, the BEV update query feature is divided into 4 regions, for example, region 1, region 2, region 3, and region 4. Each region includes 4 sub-features. When the original feature is a BEV update query feature, the sub-feature corresponding to the original feature is the fourth sub-feature. Taking region 1 as an example, the 4 fourth sub-features are respectively at position 1, position 2, position 3, and position 4.
[0159] S404B, perform global window division on the key feature and the value feature.
[0160] Exemplarily, the local window division of the key feature and the value feature is similar to the local window division of the query feature of BEV.
[0161] For example, refer to Figure 5B , the original feature is the key feature, and the key feature is divided into 4 regions, for example, region 1, region 2, region 3, and region 4. Each region includes 4 sub-features. When the original feature is the key feature, the sub-feature corresponding to the original feature is the second sub-feature. Taking region 1 as an example, the 4 second sub-features are at position 1, position 2, position 3, and position 4 respectively.
[0162] Reference Figure 5B , the original feature is the value feature, and the value feature is divided into 4 regions, for example, region 1, region 2, region 3, and region 4. Each region includes 4 sub-features. When the original feature is the value feature, the sub-feature corresponding to the original feature is the third sub-feature. Taking region 1 as an example, the 4 third sub-features are at position 1, position 2, position 3, and position 4 respectively.
[0163] S405, performing cross-attention fusion on the BEV updated query feature, key feature and value feature to obtain the BEV query fusion feature.
[0164] After completing the global window division of each feature, the four sub-features at the same position in different regions of each feature can be cross-attention operated, which can also reduce the size of self-attention fusion, thereby improving the efficiency of self-attention fusion. For example, when N=4, the size of BEV update query feature will change from (1, D, BEV_H, BEV_W) to (4, D, BEV_H / 2, BEV_W / 2), and its time complexity will be reduced to 1 / 4 of the original global attention operation.
[0165] For example, the four fourth sub-features at the same positions of different regions of the BEV updated query feature may be a fourth sub-feature of a portion of position 1 in region 1, a fourth sub-feature of a portion of position 1 in region 2, a fourth sub-feature of a portion of position 1 in region 3, and a fourth sub-feature of a portion of position 1 in region 4, for a total of four fourth sub-features.
[0166] For the 4 partial second sub-features at the same position in different regions of the key feature, they can be a partial second sub-feature at position 1 in region 1, a partial second sub-feature at position 1 in region 2, a partial second sub-feature at position 1 in region 3, and a partial second sub-feature at position 1 in region 4, for a total of 4 partial second sub-features.
[0167] For the 4 third sub-features at the same position in different regions of the value feature, they may be a part of the third sub-feature at position 1 in region 1, a part of the third sub-feature at position 1 in region 2, a part of the third sub-feature at position 1 in region 3, and a part of the third sub-feature at position 1 in region 4, for a total of 4 third sub-features.
[0168] In some embodiments of the present application, the self-attention fusion process of the transformer hierarchical self-attention module is to fuse the 4 parts of the fourth sub-features at position 1 of the above 4 regions, the 4 parts of the second sub-features at position 1 of the 4 regions, and the 4 parts of the third sub-features at position 1 of the 4 regions. Similarly, the 4 parts of the fourth sub-features, the 4 parts of the second sub-features, and the 4 parts of the third sub-features at position 1 of other regions at the same position also need to be self-attention fused.
[0169] It can be understood that in the self-attention fusion process, each feature is divided into 4 regions and each region is divided into 4 parts of sub-features. In other embodiments, each feature can also be divided into more or fewer regions, and the number of regions is the same as the number of parts of sub-features in each region. The embodiments of the present application do not limit the number of regions into which each feature is divided and the number of parts of sub-features in each region.
[0170] S406, performing global window restoration on the query fusion features of BEV.
[0171] Exemplarily, after executing step S405, the transformer hierarchical self-attention module restores the output of the self-attention optimization, and the output of the self-attention optimization can be used as the query fusion feature of BEV. It can be understood that the query fusion feature of BEV output by the self-attention optimization is also globally windowed, and the query fusion feature of BEV can be restored from (4, D, BEV_H / 2, BEV_W / 2) to (1, D, BEV_H, BEV_W). It can be understood that the query fusion feature of BEV is restored to the third feature data.
[0172] It can be understood that in some embodiments of the present application, the vehicle-mounted device divides the original transformer attention into two hierarchical attention mechanisms through the transformer hierarchical self-attention module, which greatly reduces the computational time complexity of the attention mechanism. The reduction in computational time complexity is related to the number of divided areas N. When the value of N is larger, the computational complexity is reduced to the original 2 / N. On the vehicle side, taking N=16 can achieve a balance between computational efficiency and BEV multi-task model detection performance.
[0173] The vehicle-mounted devices involved in the above embodiments are introduced below.
[0174] For example, Figure 6 According to some embodiments of the present application, a schematic structural diagram of a vehicle-mounted device 100 is shown.
[0175] The vehicle-mounted device 100 can be used to implement the image processing methods provided in the aforementioned embodiments.
[0176] like Figure 6 As shown, the vehicle-mounted device 100 includes one or more processors 101, a system memory 102, a non-volatile memory (NVM) 103, a communication interface 104, an input / output device 105, and a system control logic unit 106 for coupling the processor 101, the system memory 102, the non-volatile memory 103, the communication interface 104 and the input / output device 105. Among them: The processor 101 may include one or more processing units, for example, a central processing unit (CPU), a graphics processing unit (GPU), a digital signal processor (DSP), a micro-programmed control unit (MCU), an artificial intelligence (AI) processor or a programmable logic device (field programmable gate array, FPGA), a neural network processor (neural-network processing unit, NPU), etc. The processing module or processing circuit may include one or more single-core or multi-core processors. In some embodiments, the CPU can be used to optimize the neural network model to be run, and the NPU can be used to run the neural network model to be run.
[0177] The system memory 102 is a volatile memory, such as a random-access memory (RAM), a double data rate synchronous dynamic random access memory (DDR SDRAM), etc. The system memory is used to temporarily store data and / or instructions. For example, in some embodiments, the system memory 102 can be used to store the data provided by the aforementioned different services, such as sensor data, image data or video data, etc., and can also be used to store instructions of the image generation method provided by the aforementioned embodiments.
[0178] The non-volatile memory 103 may include one or more tangible, non-transitory computer-readable media for storing data and / or instructions. In some embodiments, the non-volatile memory 103 may include any suitable non-volatile memory such as flash memory and / or any suitable non-volatile storage device, such as a hard disk drive (HDD), a compact disc (CD), a digital versatile disc (DVD), a solid-state drive (SSD), etc. In some embodiments, the non-volatile memory 103 may also be a removable storage medium, such as a secure digital (SD) memory card, etc. In other embodiments, the non-volatile memory 103 may be used to store instructions of the image generation method provided in the aforementioned embodiments, etc.
[0179] In particular, the system memory 102 and the non-volatile memory 103 may respectively include: a temporary copy and a permanent copy of the instruction 107. The instruction 107 may include: when executed by at least one of the processors 101, the vehicle-mounted device 100 implements the image generation method provided by each embodiment of the present application.
[0180] The communication interface 104 may include a transceiver for providing a wired or wireless communication interface for the vehicle-mounted device 100, and then communicating with any other suitable device through one or more networks. In some embodiments, the communication interface 104 may be integrated into other components of the vehicle-mounted device 100, for example, the communication interface 104 may be integrated into the processor 101. In some embodiments, the vehicle-mounted device 100 may communicate with other devices through the communication interface 104, for example, the vehicle-mounted device 100 may obtain corresponding data from other devices through the communication interface 104.
[0181] The input / output device 105 may be an input device such as a keyboard, a mouse, etc., and an output device such as a display, etc. The user may interact with the vehicle-mounted device 100 through the input / output device 105 .
[0182] The system control logic unit 106 may include any suitable interface controller to provide any suitable interface with other modules of the vehicle-mounted device 100. For example, in some embodiments, the system control logic unit 106 may include one or more memory controllers to provide interfaces connected to the system memory 102 and the non-volatile memory 103.
[0183] In some embodiments, at least one of the processors 101 may be packaged together with the logic of one or more controllers for the system control logic unit 106 to form a system in package (SiP). In other embodiments, at least one of the processors 101 may also be integrated with the logic of one or more controllers for the system control logic unit 106 on the same chip to form a system-on-chip (SoC).
[0184] Understandably, Figure 6 The structure of the vehicle-mounted device 100 shown is only an example. In other embodiments, the vehicle-mounted device 100 may include more or fewer components than shown, or combine some components, or split some components, or arrange the components differently. The components shown in the figure may be implemented in hardware, software, or a combination of software and hardware.
[0185] It can be understood that the vehicle-mounted device 100 can be any device configured on the vehicle, including but not limited to mobile phones, car computers, terminals in self-driving, wireless terminals in transportation safety, terminals in smart cities, etc.
[0186] An embodiment of the present application also provides a program product, which, when executed on a vehicle-mounted device, can enable the vehicle-mounted device to implement the image processing methods provided by the aforementioned embodiments.
[0187] An embodiment of the present application also provides a readable storage medium, in which one or more programs are stored. When the one or more programs are executed by a vehicle-mounted device, the vehicle-mounted device implements the image processing method provided by the aforementioned embodiments.
[0188] The various embodiments of the mechanism disclosed in the present application can be implemented in hardware, software, firmware or a combination of these implementation methods. The embodiments of the present application can be implemented as a computer program or program code executed on a programmable system, which includes at least one processor, a storage system (including volatile and non-volatile memory and / or storage elements), at least one input device and at least one output device.
[0189] Program code can be applied to input instructions to perform the functions described in this application and generate output information. The output information can be applied to one or more output devices in a known manner. For the purposes of this application, a processing system includes any system having a processor such as, for example, a digital signal processor, a microcontroller, an application specific integrated circuit, or a microprocessor.
[0190] Program code can be implemented with high-level programming language or object-oriented programming language to communicate with the processing system. When necessary, program code can also be implemented with assembly language or machine language. In fact, the mechanism described in this application is not limited to the scope of any specific programming language. In either case, the language can be a compiled language or an interpreted language.
[0191] In some cases, the disclosed embodiments may be implemented in hardware, firmware, software, or any combination thereof. The disclosed embodiments may also be implemented as instructions carried by or stored on one or more temporary or non-temporary machine-readable (e.g., computer-readable) storage media, which may be read and executed by one or more processors. For example, the instructions may be distributed over a network or through other computer-readable media. Therefore, a machine-readable medium may include any mechanism for storing or transmitting information in a form readable by a machine (e.g., a computer), including, but not limited to, floppy disks, optical disks, optical discs, compact disc-read only memory (CD-ROMs), magneto-optical disks, read-only memory (ROM), random-access memory (RAM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), magnetic or optical cards, flash memory, or a tangible machine-readable memory for transmitting information using the Internet in an electrical, optical, acoustic or other form of propagation signal (e.g., carrier wave, infrared signal, digital signal, etc.). Therefore, a machine-readable medium includes any type of machine-readable medium suitable for storing or transmitting electronic instructions or information in a form readable by a machine (e.g., a computer).
[0192] In the accompanying drawings, some structural or method features may be shown in a specific arrangement and / or order. However, it should be understood that such a specific arrangement and / or order may not be required. Instead, in some embodiments, these features may be arranged in a manner and / or order different from that shown in the illustrative drawings. In addition, the inclusion of structural or method features in a particular figure does not mean that such features are required in all embodiments, and in some embodiments, these features may not be included or may be combined with other features.
[0193] It should be noted that the units / modules mentioned in the various device embodiments of the present application are all logical units / modules. Physically, a logical unit / module can be a physical unit / module, or a part of a physical unit / module, or can be implemented as a combination of multiple physical units / modules. The physical implementation method of these logical units / modules themselves is not the most important. The combination of functions implemented by these logical units / modules is the key to solving the technical problems proposed by the present application. In addition, in order to highlight the innovative part of the present application, the above-mentioned device embodiments of the present application do not introduce units / modules that are not closely related to solving the technical problems proposed by the present application, which does not mean that there are no other units / modules in the above-mentioned device embodiments.
[0194] It should be noted that in the examples and description of this patent, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "including one" do not exclude the existence of other identical elements in the process, method, article or device including the elements.
[0195] Although the present application has been illustrated and described with reference to certain preferred embodiments thereof, it will be understood by those skilled in the art that various changes in form and details may be made therein without departing from the scope of the present application.
Claims
1. An image processing method, applied to an on-board device on a vehicle, characterized in that: The method comprises: In response to a request to detect the surrounding environment of the vehicle, a plurality of images are acquired, where the plurality of images are images from different perspectives acquired by a plurality of sensors of the vehicle; Sampling the multiple images based on multiple sampling rates to obtain multiple feature maps of each image, where the multiple feature maps have different sizes; Projecting a first feature map among the plurality of feature maps of each of the images into a set plane to obtain a first plane image, wherein the first feature map of each of the images has the same size; Performing feature extraction on the first plane image to obtain first feature data; Performing feature extraction on feature maps other than the first feature map among the plurality of feature maps to obtain second feature data; Merging the first feature data and the second feature data to obtain third feature data; Based on the third characteristic data, the surrounding environment of the vehicle is detected.
2. The image processing method according to claim 1, characterized in that: Projecting a first feature map among the plurality of feature maps onto a set plane to obtain a first plane image includes: Acquiring acquisition parameters of the multiple sensors; Based on the acquisition parameters, an inverse perspective transformation is performed on the first feature map to generate a first planar image.
3. The image processing method according to claim 1, characterized in that: The extracting features from the feature maps other than the first feature map among the plurality of feature maps to obtain second feature data includes: Aligning the feature maps with different sampling rates among the multiple feature maps except the first feature map to obtain a second feature map; Acquiring acquisition parameters of the multiple sensors, and encoding the acquisition parameters based on the size of the second feature graph to obtain sensor features; Performing feature extraction on the second feature map to obtain fourth feature data; The fourth feature data and the sensor feature are fused to determine the second feature data.
4. The image processing method according to claim 2 or 3, characterized in that: The multiple sensors include a camera, and the acquisition parameters include internal parameters and external parameters of the camera.
5. The image processing method according to claim 1, characterized in that: The fusing the first feature data and the second feature data to obtain third feature data includes: The first feature data and the second feature data are fused based on a self-attention mechanism to obtain the third feature data.
6. The image processing method according to claim 5, characterized in that: The fusing the first feature data and the second feature data based on a self-attention mechanism includes: The first feature data includes N regions, each region includes N first sub-features, and the positions of the N first sub-features in the same region are different; The second feature data includes a first key feature, wherein the first key feature includes N regions, each region includes N second sub-features, and the positions of the N second sub-features in the same region are different; The second feature data includes a first value feature, wherein the first value feature includes N regions, each region includes N third sub-features, and the positions of the N third sub-features in the same region are different; The N parts of the first sub-features, the N parts of the second sub-features, and the N parts of the second sub-features of the same area are fused based on the self-attention mechanism to obtain a first query feature, and the third feature data is determined based on the first query feature, where N is greater than or equal to 2.
7. The image processing method according to claim 6, characterized in that: The first query feature includes N regions, each region includes N fourth sub-features, and the positions of the N fourth sub-features in the same region are different; The determining the third feature data based on the first query feature includes: The N parts of the fourth sub-features, the N parts of the second sub-features, and the N parts of the third sub-features at the same position of N different regions are fused based on a self-attention mechanism to obtain the third feature data.
8. A vehicle-mounted device, characterized in that: include: A memory for storing instructions; At least one processor is used to execute the instructions so that the vehicle-mounted device implements the image processing method according to any one of claims 1 to 7.
9. A computer-readable storage medium, characterized in that: The readable storage medium stores instructions, and when the instructions are executed on a computer, the computer is caused to execute the image processing method according to any one of claims 1 to 7.
10. A computer program product, characterized in that When the computer program product is executed on a device, the device is enabled to execute the image processing method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Roadside aerial view target detection method and device, computing equipment and storage medium
CN117994748A
BEV image construction method, device and system and storage medium
CN118780979A
Intelligent driving planning method for pedestrian interaction scene
CN119296077A
Image processing method, apparatus and device, and computer-readable storage medium
US20230316742A1
Image processing method and apparatus, and electronic device and storage medium
WO2024098941A1