Bird-eye view feature fusion method and device, vehicle and storage medium

By acquiring and fusing the bird's-eye view features of multiple historical visual frames, the information loss problem caused by single-frame fusion is solved, and the accuracy of autonomous driving environment perception is improved.

CN120726437APending Publication Date: 2025-09-30GREAT WALL MOTOR CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510889703.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-30
Publication Date
2025-09-30

AI Technical Summary

Technical Problem

In the existing technology, the bird's-eye view feature fusion method of a single historical visual frame causes information loss and cannot effectively utilize the information of multiple historical visual frames, resulting in a decrease in the accuracy of autonomous driving environment perception.

Method used

By obtaining the first bird's-eye view feature of the vehicle's current visual frame and multiple second bird's-eye view features of multiple historical visual frames adjacent to the current visual frame, and converting them to the current visual frame coordinate system for feature alignment and fusion, the occluded information in the first bird's-eye view feature is supplemented.

Benefits of technology

The accuracy of bird's-eye view features has been improved, enabling more accurate identification of the vehicle's environment and enhancing the perception accuracy of autonomous driving.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120726437A_ABST
    Figure CN120726437A_ABST
Patent Text Reader

Abstract

The invention provides an aerial view feature fusion method and device, a vehicle and a storage medium, the method is applied to the field of point cloud data processing, and the method comprises the steps: obtaining a first aerial view feature of a current visual frame of the vehicle and a plurality of second aerial view features of a plurality of historical visual frames adjacent to the current visual frame; converting the plurality of second aerial view features into a current visual frame coordinate system where the first aerial view features are located, and shifting the plurality of second aerial view features to the positions where the first aerial view features are located, so that the aerial view features are aligned; and carrying out feature fusion on the plurality of second aerial view features and the first aerial view feature to obtain a target aerial view feature, so that the environment where the vehicle is located is identified through the target aerial view feature. According to the method, the plurality of second aerial view features can be fused with the first aerial view feature, so that lost information in the first aerial view feature is supplemented through the plurality of second aerial view features, and the aerial view features are more accurate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of point cloud data processing, and more specifically, to a method, device, vehicle, and storage medium for fusion of bird's-eye view features in the field of point cloud data processing. Background Art

[0002] Autonomous driving technology has been fully integrated into various vehicle applications. In practical applications, bird's-eye view (BEV) perception algorithms and time series fusion algorithms are required to improve the accuracy of dynamic environment perception.

[0003] Related techniques involve correcting the bird's-eye view features of a single historical visual frame to the bird's-eye view space of the current visual frame corresponding to that historical frame, and then fusing them with the bird's-eye view features of the current visual frame. However, in the temporal fusion process, since only a single historical visual frame is used for fusion, the entire information in the historical visual frame cannot be fully utilized, and useful information is easily lost. Summary of the Invention

[0004] The present application provides a method, device, vehicle and storage medium for fusing features of a bird's-eye view. The method can fuse multiple second bird's-eye view features with a first bird's-eye view feature, thereby supplementing the lost information in the first bird's-eye view feature through the multiple second bird's-eye view features, making the bird's-eye view feature more accurate.

[0005] In a first aspect, a method for fusing features of a bird's-eye view is provided, the method comprising:

[0006] Acquire a first bird's-eye view feature of a current visual frame of the vehicle and a plurality of second bird's-eye view features of a plurality of historical visual frames adjacent to the current visual frame;

[0007] Converting the plurality of second bird's-eye view features to the current visual frame coordinate system where the first bird's-eye view feature is located, and shifting the plurality of second bird's-eye view features to the position where the first bird's-eye view feature is located, so that the plurality of second bird's-eye view features are aligned with the first bird's-eye view feature;

[0008] The plurality of second bird's-eye view features are fused with the first bird's-eye view feature to obtain a target bird's-eye view feature, so that the environment in which the vehicle is located can be identified through the target bird's-eye view feature.

[0009] Through the above scheme, the method can obtain a first bird's-eye view feature of the vehicle's current visual frame and multiple second bird's-eye view features from multiple historical visual frames adjacent to the current visual frame. In the real world, a vehicle is in motion while driving, as are other vehicles and pedestrians around it. However, the second and first bird's-eye view features are generated based on the visual frame coordinate systems at different times. If the first bird's-eye view feature contains partially obscured information, the second bird's-eye view features can be used to supplement the information in the first bird's-eye view feature. To effectively integrate past information to determine the vehicle's current scene, the bird's-eye view features need to be fused. The multiple second bird's-eye view features are converted to the current visual frame coordinate system where the first bird's-eye view feature resides and then shifted to the position of the first bird's-eye view feature to align them with the first bird's-eye view feature. The multiple second bird's-eye view features are then fused with the first bird's-eye view feature to obtain a target bird's-eye view feature, which can then be used to identify the vehicle's environment. That is, multiple second bird's-eye view features are fused with the first bird's-eye view feature, and the visible information in the multiple second bird's-eye view features is used to supplement the partially obscured information in the first bird's-eye view feature, thereby increasing the accuracy of the bird's-eye view feature.

[0010] In conjunction with the first aspect, in some possible implementations, obtaining a first bird's-eye view feature of a current visual frame of the vehicle and a plurality of second bird's-eye view features of a plurality of historical visual frames adjacent to the current visual frame includes:

[0011] Acquiring first visual data of a current visual frame of the vehicle; performing feature extraction on the first visual data to obtain a first bird's-eye view feature corresponding to the first visual data;

[0012] According to the time sequence of the multiple historical visual frames, multiple second bird's-eye view features of the multiple historical visual frames are obtained from the queue.

[0013] Through the above method, the first bird's-eye view feature of the vehicle's current visual frame and multiple second bird's-eye view features of multiple historical visual frames adjacent to the current visual frame are obtained, so as to predict the vehicle's surrounding environment based on the multiple second bird's-eye view features and the first bird's-eye view features.

[0014] In conjunction with the first aspect, in some possible implementations, before obtaining the plurality of second bird's-eye view features of the plurality of historical visual frames in a queue form according to the chronological order of the plurality of historical visual frames, the method includes:

[0015] Based on the temporal sequence of the current visual frame and the multiple historical visual frames, the multiple second bird's-eye view features of the multiple historical visual frames are cached in the queue form.

[0016] Through the above method, based on the current visual frame and the time sequence of the multiple historical visual frames, the multiple second bird's-eye view features of the multiple historical visual frames are cached in the form of the queue, so that the vehicle's surrounding environment during the vehicle's driving process can be supplemented based on the second bird's-eye view features of the multiple historical visual frames.

[0017] In conjunction with the first aspect, in some possible implementations, converting the plurality of second bird's-eye view features to the current visual frame coordinate system where the first bird's-eye view feature is located includes:

[0018] Converting the plurality of second bird's-eye view features from respective historical visual frame coordinate systems corresponding to respective second bird's-eye view features to a world coordinate system;

[0019] The plurality of second bird's-eye view features are transformed from the world coordinate system to the current visual frame coordinate system.

[0020] Through the above method, since the historical visual frame coordinate systems corresponding to each of the second bird's-eye view features are different from the current visual frame coordinate system where the first bird's-eye view feature is located, in this case, the first bird's-eye view feature cannot be fused with the second bird's-eye view feature. Therefore, it is necessary to convert multiple second bird's-eye view features into the current visual frame coordinate system to facilitate feature fusion.

[0021] In conjunction with the first aspect, in some possible implementations, converting the plurality of second bird's-eye view features from the historical visual frame coordinate system to the world coordinate system includes:

[0022] Obtain a historical visual frame coordinate transformation matrix of the plurality of second bird's-eye view features in the historical visual frame coordinate system;

[0023] Converting the historical visual frame coordinate transformation matrix to the world coordinate system based on homogeneous coordinate transformation;

[0024] The step of converting the plurality of second bird's-eye view features from the world coordinate system to the current visual frame coordinate system includes:

[0025] Obtaining a current visual frame coordinate transformation matrix of the first bird's-eye view feature in the current visual frame coordinate system;

[0026] Calculate the inverse transformation matrix of the current visual frame coordinate transformation matrix;

[0027] The historical visual frame coordinate transformation matrix in the world coordinate system is converted to the current visual frame coordinate system based on the homogeneous coordinate transformation.

[0028] Through the above method, the current visual frame coordinate transformation matrix in the current visual frame coordinate system is obtained, that is, the matrix for transforming the coordinate point from the world coordinate system to the current visual frame coordinate system is obtained. The inverse transformation matrix of the current visual frame coordinate transformation matrix is ​​calculated, that is, the matrix for transforming the coordinate point from the current visual frame coordinate system to the world coordinate system is obtained. Based on the homogeneous coordinate transformation, the historical frame coordinate transformation matrix can be directly transformed into the current visual frame coordinate system, so that the second bird's-eye view feature of the time history visual frame can be aligned with the first bird's-eye view feature of the current visual frame, solving the problem of inconsistent coordinate systems of perceived data at different times caused by the movement of the sound vehicle.

[0029] In conjunction with the first aspect, in some possible implementations, displacing the plurality of second bird's-eye view features to the location of the first bird's-eye view feature includes:

[0030] In the current visual frame coordinate system, obtaining the time interval between the current visual frame and each of the historical visual frames;

[0031] Calculate the weight corresponding to each of the second bird's-eye view features based on the time interval;

[0032] The plurality of second bird's-eye view features are rotated and translated to the first coordinate of the first bird's-eye view feature in the current visual frame coordinate system based on the respective weights.

[0033] In practical applications, the above method may cause missing features in the current visual frame due to occlusion or illumination changes. In the coordinate system of the current visual frame, the time interval between the current visual frame and each of the historical visual frames is obtained; the weight corresponding to each of the second bird's-eye view features is calculated based on the time interval; and based on each of the weights, the multiple second bird's-eye view features are rotated and translated to the first coordinate of the first bird's-eye view feature in the coordinate system of the current visual frame. In other words, the first bird's-eye view feature is processed using the second bird's-eye view feature of the historical visual frame, and fused through time weights to make the changes of adjacent second bird's-eye view features smoother, thereby optimizing the continuity of the vehicle's motion trajectory.

[0034] In conjunction with the first aspect, in some possible implementations, the step of fusing the plurality of second bird's-eye view features with the first bird's-eye view feature to obtain a target bird's-eye view feature includes:

[0035] The multiple second bird's-eye view features are fused with the first bird's-eye view feature through a convolutional network to obtain a target bird's-eye view feature.

[0036] Through the above method, since the first bird's-eye view feature of the current visual frame and the second bird's-eye view features of multiple historical visual frames are fused, the vehicle is always moving within these visual frames, and the vehicle can scan the surrounding environment at different positions. If a target point is occluded in the current visual frame but not in the previous historical visual frames, the target point will appear in the fused target bird's-eye view feature, thereby overcoming the problem of local information loss due to vehicle occlusion.

[0037] In a second aspect, a feature fusion device for a bird's-eye view is provided, characterized in that the device comprises:

[0038] an acquisition module, configured to acquire a first bird's-eye view feature of a current visual frame of the vehicle and a plurality of second bird's-eye view features of a plurality of historical visual frames adjacent to the current visual frame;

[0039] a conversion and displacement module, configured to convert the plurality of second bird's-eye view features to the current visual frame coordinate system where the first bird's-eye view feature is located, and to displace the plurality of second bird's-eye view features to the position where the first bird's-eye view feature is located, so that the plurality of second bird's-eye view features are aligned with the first bird's-eye view feature;

[0040] The fusion module is used to perform feature fusion on the multiple second bird's-eye view features and the first bird's-eye view features to obtain a target bird's-eye view feature, so that the environment in which the vehicle is located can be identified through the target bird's-eye view feature.

[0041] In a third aspect, a vehicle is provided, comprising a memory and a processor. The memory is configured to store executable program code, and the processor is configured to call and execute the executable program code from the memory, so that the vehicle executes the method performed by the above-mentioned bird's-eye view feature fusion method.

[0042] In a fourth aspect, a computer program product is provided, which includes: a computer program code, which, when executed on a computer, enables the computer to execute the method executed by the above-mentioned bird's-eye view feature fusion method.

[0043] In a fifth aspect, a computer-readable storage medium is provided, which stores a computer program code. When the computer program code runs on a computer, the computer executes the method executed by the above-mentioned bird's-eye view feature fusion method. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] Figure 1 Schematic diagram of an implementation environment of a feature fusion method for a bird's-eye view provided in an embodiment of the present application;

[0045] Figure 2This is a schematic flow chart of a feature fusion method for a bird's-eye view provided in an embodiment of the present application;

[0046] Figure 3 is a schematic flow chart of another method for fusion of features of a bird's-eye view provided in an embodiment of the present application;

[0047] Figure 4 This is a schematic structural diagram of a feature fusion device for a bird's-eye view provided in an embodiment of the present application;

[0048] Figure 5 It is a structural schematic diagram of a vehicle provided in an embodiment of the present application. DETAILED DESCRIPTION

[0049] The following will clearly and thoroughly describe the technical solutions in this application in conjunction with the accompanying drawings. In the description of the embodiments of this application, unless otherwise specified, " / " means or, for example, A / B can mean A or B: "and / or" in the text is only a description of the association relationship of associated objects, indicating that there can be three relationships, for example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. In addition, in the description of the embodiments of this application, "multiple" means two or more than two.

[0050] In the following, the terms "first" and "second" are used for descriptive purposes only and should not be understood to imply or suggest relative importance or implicitly indicate the number of technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the features.

[0051] In the related art, during the process of fusing bird's-eye view features, only the historical bird's-eye view features of one historical visual frame are saved, the historical bird's-eye view features are mapped to the current visual frame coordinate system, and the historical bird's-eye view features are fused with the current bird's-eye view features of the current visual frame. However, when a vehicle is driving, one historical visual frame cannot depict the complete environmental changes around the vehicle during the driving process. For example, when a vehicle is moving, other pedestrians and vehicles around the vehicle are also moving. Therefore, if only one historical visual frame is used for fusion, it is easy to cause the loss of useful information from multiple historical visual frames.

[0052] To solve the above problems, an embodiment of the present application provides a bird's-eye view feature fusion method, which can obtain a first bird's-eye view feature of a vehicle's current visual frame and multiple second bird's-eye view features of multiple historical visual frames adjacent to the current visual frame. In the real world, when a vehicle is driving, the vehicle itself is moving, and other vehicles and pedestrians around the vehicle are also moving. However, the second bird's-eye view feature and the first bird's-eye view feature are generated based on the visual frame coordinate system at different times. If there is partially obscured information in the first bird's-eye view feature, the information in the first bird's-eye view feature can be supplemented by the second bird's-eye view feature. In order to effectively fuse past information to judge the scene of the vehicle at the current moment, it is necessary to fuse the bird's-eye view features, convert the multiple second bird's-eye view features to the current visual frame coordinate system where the first bird's-eye view feature is located, and shift the multiple second bird's-eye view features to the position where the first bird's-eye view feature is located, so that the multiple second bird's-eye view features are feature-aligned with the first bird's-eye view feature. The plurality of second bird's-eye view features are fused with the first bird's-eye view feature to obtain a target bird's-eye view feature, so that the target bird's-eye view feature can be used to identify the environment in which the vehicle is located. In other words, the plurality of second bird's-eye view features are fused with the first bird's-eye view feature to supplement the obscured information in the first bird's-eye view feature with the visible information in the plurality of second bird's-eye view features, thereby increasing the accuracy of the bird's-eye view feature.

[0053] Figure 1 This is a schematic diagram of the implementation environment of a feature fusion method for a bird's-eye view provided in an embodiment of the present application.

[0054] For example, Figure 1 As shown, the implementation environment includes a vehicle controller 110 and a multimodal sensor 120 .

[0055] The vehicle controller 110 is a key control unit of the vehicle. It can obtain relevant vehicle data and control the vehicle to perform corresponding operations based on this relevant data. For example, the vehicle controller 110 can obtain visual data of the vehicle's surroundings and determine the vehicle's current environment based on this visual data.

[0056] The multimodal sensor 120 is used to acquire visual data of the vehicle and transmit the visual data to the vehicle controller 110. For example, if the multimodal sensor 120 is a lidar, the multimodal sensor 120 can acquire point cloud data and transmit the point cloud data to the vehicle controller 110. Alternatively, if the multimodal sensor 120 is an image sensor, the multimodal sensor can acquire image data of the vehicle and transmit the image data to the vehicle controller 110.

[0057] Figure 2This is a schematic flowchart of a feature fusion method for a bird's-eye view provided in an embodiment of the present application.

[0058] For example, Figure 2 As shown, taking the execution subject as the vehicle controller as an example, a feature fusion method of a bird's-eye view provided by the present application is described. The method 200 includes the following steps 201 to 203.

[0059] Step 201 : Acquire a first bird's-eye view feature of a current visual frame of a vehicle and a plurality of second bird's-eye view features of a plurality of historical visual frames adjacent to the current visual frame.

[0060] A bird's-eye view in a vehicle is a bird's-eye view projection that is synthesized after processing, using images of the vehicle's surroundings collected by multiple cameras or point cloud data collected by lidar. Bird's-eye views are widely used in the field of autonomous driving. For example, when driving or parking, a bird's-eye view can present the environment around the vehicle in real time. Since a bird's-eye view is a single-frame image, a single-frame image can only provide the instantaneous position of an object and cannot infer the vehicle's direction of movement or speed. In other words, a single-frame bird's-eye view cannot capture the overall movement trend of the vehicle. However, since the objects around the vehicle are constantly moving during the vehicle's movement, it is necessary to obtain not only the vehicle's complete movement trajectory, but also the operating environment around the vehicle during the vehicle's movement. However, a single-frame bird's-eye view cannot describe the vehicle's operating environment during continuous driving. Therefore, it is necessary to obtain the first bird's-eye view feature (BEVFeature) of the vehicle's current visual frame and multiple second bird's-eye view features of multiple historical visual frames adjacent to the current visual frame.

[0061] The first bird's-eye view feature represents the projection of the vehicle's three-dimensional world onto a two-dimensional top-down plane. This feature can dynamically identify the position and movement of vehicles, pedestrians, and non-motorized vehicles, and is widely used in the field of autonomous driving. The second bird's-eye view feature is generated from historical visual frames adjacent to the current visual frame.

[0062] A visual frame is a discrete sampling of continuous visual signals by a vehicle's multimodal sensor. A visual frame is a projection of the three-dimensional real world onto a two-dimensional plane. The current visual frame is the visual frame captured by the multimodal sensor at the current time, and the multiple historical visual frames adjacent to the current visual frame are the visual frames captured by the multimodal sensor before the current time. That is, the current visual frame is the latest frame of the bird's-eye view, and the multiple historical visual frames adjacent to the current visual frame refer to several frames immediately before the current frame on the timeline. For example, in a video captured by a multimodal sensor at 30 frames per second, the latest frame 30 is the current visual frame, and the visual frames from frame 1 to frame 29 are multiple historical visual frames. The multiple historical visual frames adjacent to the current visual frame indicate that there is a continuous time sequence between the multiple historical visual frames and the current visual frame. For example, if the timestamp of the current visual frame is 21:22:21, then the multiple historical visual frames adjacent to the current visual frame are the first bird's-eye view frame, the second bird's-eye view frame, and the third bird's-eye view frame. The timestamp corresponding to the first bird's-eye view frame is 21:22:18, the timestamp corresponding to the second bird's-eye view frame is 21:22:19, and the timestamp corresponding to the third bird's-eye view frame is 21:22:20. In other words, the first bird's-eye view frame, the second bird's-eye view frame, the third bird's-eye view frame, and the current visual frame are temporally continuous.

[0063] Step 202: convert the plurality of second bird's-eye view features to the current visual frame coordinate system where the first bird's-eye view feature is located, and shift the plurality of second bird's-eye view features to the position where the first bird's-eye view feature is located, so that the plurality of second bird's-eye view features are aligned with the first bird's-eye view feature.

[0064] It should be understood that in the real world, when a vehicle is driving, other vehicles and pedestrians around it are also moving. However, the second bird's-eye view features of different historical frames collected at different times are generated based on the visual frame coordinate system at different times, and the vehicle's own coordinate system is constantly changing with the movement of the vehicle. In order to effectively integrate past information to judge the scene of the vehicle at the current moment, the bird's-eye view features need to be unified in the same spatial reference system for comparison and aggregation, that is, the multiple second bird's-eye view features of the historical visual frame need to be converted to the current visual frame coordinate system where the first bird's-eye view features are located. The above conversion process is called align.

[0065] It should also be understood that during the vehicle's own movement, the objects around the vehicle are also moving. The second bird's-eye view feature of the historical frame is generated in the historical frame coordinate system at the historical moment, and the first bird's-eye view feature of the current frame is generated in the current frame coordinate system at the current moment. That is, the multiple second bird's-eye view features of multiple historical frames and the first bird's-eye view feature of the current frame are in different visual frame coordinate systems. If the second bird's-eye view features are not aligned with the first bird's-eye view features, it is easy for the objects detected in the historical frame to appear in the wrong position on the bird's-eye view. For example, because the vehicle is moving forward, the front vehicle is located directly in front of the vehicle in the historical frame, but it will "erroneously" appear behind the vehicle in the current frame coordinate system, making it impossible to distinguish between the actual movement of the vehicle and the position movement caused by the vehicle movement. In this case, it is necessary to move the multiple second bird's-eye view features to the position of the first bird's-eye view feature so that the multiple second bird's-eye view features are aligned with the first bird's-eye view feature.

[0066] Step 203 : Fusing the plurality of second bird's-eye view features with the first bird's-eye view feature to obtain a target bird's-eye view feature, so as to identify the environment in which the vehicle is located through the target bird's-eye view feature.

[0067] Among them, the target bird's-eye view features are used to identify the environment in which the vehicle is located in the subsequent perception processing of the vehicle.

[0068] It should be understood that since multiple second bird's-eye view features contain motion trajectory information of the vehicle and objects around the vehicle, the first bird's-eye view feature of the current frame may cause errors due to environmental interference. Therefore, fusing the first bird's-eye view feature with the second bird's-eye view feature can restore some of the obscured information in the first bird's-eye view feature.

[0069] An embodiment of the present application provides a bird's-eye view feature fusion method, which can obtain a first bird's-eye view feature of a vehicle's current visual frame and multiple second bird's-eye view features of multiple historical visual frames adjacent to the current visual frame. In the real world, when a vehicle is driving, the vehicle itself is moving, and other vehicles and pedestrians around the vehicle are also moving. However, the second bird's-eye view feature and the first bird's-eye view feature are generated based on the visual frame coordinate system at different times. If there is partially obscured information in the first bird's-eye view feature, the information in the first bird's-eye view feature can be supplemented by the second bird's-eye view feature. In order to effectively fuse past information to judge the scene of the vehicle at the current moment, it is necessary to fuse the bird's-eye view features, convert the multiple second bird's-eye view features to the current visual frame coordinate system where the first bird's-eye view feature is located, and shift the multiple second bird's-eye view features to the position where the first bird's-eye view feature is located, so that the multiple second bird's-eye view features are feature-aligned with the first bird's-eye view feature. The plurality of second bird's-eye view features are fused with the first bird's-eye view feature to obtain a target bird's-eye view feature, so that the target bird's-eye view feature can be used to identify the environment in which the vehicle is located. In other words, the plurality of second bird's-eye view features are fused with the first bird's-eye view feature to supplement the obscured information in the first bird's-eye view feature with the visible information in the plurality of second bird's-eye view features, thereby increasing the accuracy of the bird's-eye view feature.

[0070] It should be noted that the above steps 201-203 are a simple description of a feature fusion method for a bird's-eye view provided in an embodiment of the present application. The following will provide a more detailed description of the feature fusion method for a bird's-eye view provided in an embodiment of the present application with reference to some examples. Figure 3 Taking the execution subject as the vehicle controller as an example, the method includes the following steps 301 to 304.

[0071] Step 301 : Based on the temporal sequence of the current visual frame and the multiple historical visual frames, cache the multiple second bird's-eye view features of the multiple historical visual frames in the queue form.

[0072] It should be understood that a bird's-eye view in a vehicle is a bird's-eye view projection created by processing and synthesizing images of the vehicle's surroundings using multiple cameras or point cloud data collected by lidar. Bird's-eye views are widely used in the field of autonomous driving. For example, a bird's-eye view can provide a real-time view of the vehicle's surroundings while driving or parking.

[0073] The second bird's-eye view feature is generated from a historical visual frame adjacent to the current visual frame. The first bird's-eye view feature represents the vehicle's three-dimensional world projected onto a two-dimensional top-down plane. The bird's-eye view feature can dynamically identify the position and movement trajectory of vehicles, pedestrians, and non-motorized vehicles.

[0074] In one possible implementation, first visual data of a current visual frame and second visual data of multiple historical frames are acquired through a multimodal sensor; feature recognition is performed on the second point cloud data or second image data of the multiple historical visual frames to obtain multiple second bird's-eye view features; and based on the chronological sequence of the current visual frame and the multiple historical visual frames, the multiple second bird's-eye view features of the multiple historical visual frames are cached in the form of the queue.

[0075] The multimodal sensor includes a laser radar and an image sensor. The first visual data includes first point cloud data or first image data. The second visual data includes second point cloud data or second image data. The queue indicates storage in a first-in, first-out manner.

[0076] It should be understood that the movement of a vehicle follows the law of temporal causality, and the sequence of multiple historical frames can directly reflect the vehicle's movement trajectory and state changes. If the multiple historical visual frames are stored in a disordered order or the timing information is lost, it will lead to errors in the vehicle's motion estimation.

[0077] Under this implementation, feature recognition is performed on the second point cloud data or second image data of the multiple historical visual frames to obtain second bird's-eye view features; based on the current visual frame and the chronological order of the multiple historical visual frames, the multiple second bird's-eye view features of the multiple historical visual frames are cached in the form of the queue, so that the historical driving trajectory of the vehicle and the surrounding environment of the vehicle during driving can be supplemented based on the second bird's-eye view features of the multiple historical visual frames, and the surrounding environment of the vehicle can be more accurately predicted based on the second bird's-eye view.

[0078] In order to explain the above embodiment in more detail, the following describes the above embodiment in several parts.

[0079] The first part describes the contents of first visual data of a current visual frame and second visual data of multiple historical frames obtained by a multimodal sensor.

[0080] In some embodiments, first point cloud data of a current visual frame and second point cloud data of multiple historical visual frames are acquired through a laser radar.

[0081] In some embodiments, first image data of a current visual frame and second image data of a plurality of historical visual frames are acquired through an image sensor.

[0082] In the second part, feature recognition is performed on the second point cloud data or the second image data of multiple historical visual frames to obtain the content of multiple second bird's-eye view features.

[0083] In a possible implementation, feature recognition is performed on the second point cloud data of multiple historical visual frames to obtain multiple second bird's-eye view features.

[0084] In some embodiments, the point cloud is discretized into a regular three-dimensional voxel grid. For example, the space is divided into small cubes of equal size (X, Y, Z). For each point within a non-empty voxel, a PointNet or similar structure is used to extract local features. An efficient 3D sparse convolutional network is directly applied to the entire sparse voxel grid to extract multi-scale features. The multi-scale features are compressed and aggregated to obtain a two-dimensional second bird's-eye view feature.

[0085] In this implementation, the two-dimensional second bird's-eye view features can be directly projected from the three-dimensional world through point cloud data, and the calculation is relatively simple.

[0086] In a possible implementation, feature recognition is performed on the second image data of multiple historical visual frames to obtain multiple second bird's-eye view features.

[0087] In some embodiments, the second image data is extracted through a two-dimensional convolutional neural network to obtain image features (ImgFeature) of the second image of the second image data, and the depth distribution of each second image pixel or feature point is predicted. For each image feature point and the depth information corresponding to each image feature point, the coordinates of each image feature point in the three-dimensional space are calculated according to the intrinsic and extrinsic parameters of the image sensor, the image features are mapped to the three-dimensional space according to the identity information, and the features in the three-dimensional space are vertically projected onto the bird's-eye view network to obtain the second bird's-eye view features.

[0088] In this embodiment, feature recognition is performed on the second image data through a convolutional neural network to obtain multiple second bird's-eye view features, which can extract more powerful features at a lower computational cost.

[0089] The third part describes the content of caching the multiple second bird's-eye view features of the multiple historical visual frames in the queue form based on the temporal sequence of the current visual frame and the multiple historical visual frames.

[0090] In some embodiments, caching the plurality of second bird's-eye view features of the plurality of historical visual frames in the queue format means storing the plurality of second bird's-eye view features in visual frames of a preset length. For example, the preset length is 5 frames. That is, if 5 historical visual frames are cached, when a new visual frame arrives, the first frame can be directly removed and the latest visual frame can be added.

[0091] Step 302 : Acquire a first bird's-eye view feature of a current visual frame of the vehicle and a plurality of second bird's-eye view features of a plurality of historical visual frames adjacent to the current visual frame.

[0092] In one possible implementation, first visual data of the current visual frame of the vehicle is obtained; feature extraction is performed on the first visual data to obtain a first bird's-eye view feature corresponding to the first visual data; and multiple second bird's-eye view features of the multiple historical visual frames are obtained from a queue in a chronological order of the multiple historical visual frames.

[0093] There are a plurality of second bird's-eye view features cached in the queue, and the second bird's-eye view features are stored in a temporal order of a plurality of historical visual frames.

[0094] In this embodiment, a first bird's-eye view feature of a current visual frame of the vehicle and multiple second bird's-eye view features of multiple historical visual frames adjacent to the current visual frame are obtained to facilitate prediction of the vehicle's surrounding environment based on the multiple second bird's-eye view features and the first bird's-eye view features.

[0095] In order to explain the above embodiment in more detail, the following describes the above embodiment in several parts.

[0096] The first part describes the content of obtaining the first visual data of the current visual frame of the vehicle.

[0097] In some embodiments, the first visual data of the current visual frame is acquired through a laser radar or an image sensor.

[0098] In the second part, feature extraction is performed on the first visual data to obtain the content of the first bird's-eye view feature corresponding to the first visual data.

[0099] In a possible implementation, feature recognition is performed on the first visual data to obtain a first bird's-eye view feature.

[0100] In some embodiments, the point cloud is discretized into a regular three-dimensional voxel grid. For example, the space is divided into small cubes of equal size (X, Y, Z). For each point within a non-empty voxel, a PointNet or similar structure is used to extract local features. An efficient 3D sparse convolutional network is directly applied to the entire sparse voxel grid to extract multi-scale features. The multi-scale features are compressed and aggregated to obtain a two-dimensional first bird's-eye view feature.

[0101] In this implementation, the two-dimensional first bird's-eye view features can be directly projected from the three-dimensional world through point cloud data, and the calculation is relatively simple.

[0102] In some embodiments, first visual data is extracted through a two-dimensional convolutional neural network to obtain first image pixels or feature points of the first visual data, and the depth distribution of each first image pixel or feature point is predicted. For each image feature point and the depth information corresponding to each image feature point, the coordinates of each image feature point in three-dimensional space are calculated based on the intrinsic and extrinsic parameters of the image sensor, the image features are mapped to the three-dimensional space based on the identity information, and the features in the three-dimensional space are vertically projected onto the bird's-eye view network to obtain the first bird's-eye view features.

[0103] In this embodiment, feature recognition is performed on the first visual data through a convolutional neural network to obtain first bird's-eye view features, which can extract more powerful features at a lower computational cost.

[0104] The second part describes the content of obtaining the multiple second bird's-eye view features of the multiple historical visual frames from the queue according to the time sequence of the multiple historical visual frames.

[0105] In some embodiments, the plurality of second bird's-eye view features are obtained from a queue based on a controller area network (CAN).

[0106] Step 303: convert the multiple second bird's-eye view features to the current visual frame coordinate system where the first bird's-eye view feature is located, and shift the multiple second bird's-eye view features to the position where the first bird's-eye view feature is located, so that the multiple second bird's-eye view features are feature-aligned with the first bird's-eye view feature.

[0107] It should be understood that in the real world, when a vehicle is driving, the vehicle itself is moving, and other vehicles and pedestrians around the vehicle are also moving. However, the second bird's-eye view features of different historical frames collected at different times are generated based on the visual frame coordinate system at different times, and the vehicle's own coordinate system is constantly changing with the movement of the vehicle. In order to effectively integrate past information to judge the scene of the vehicle at the current moment, the bird's-eye view features need to be unified in the same spatial reference system for comparison and aggregation, that is, multiple second bird's-eye view features of the historical visual frame need to be converted to the current visual frame coordinate system where the first bird's-eye view features are located.

[0108] The following describes the content of converting the multiple second bird's-eye view features into the current visual frame coordinate system where the first bird's-eye view feature is located.

[0109] In a possible implementation, the multiple second bird's-eye view features are transformed from the respective historical visual frame coordinate systems corresponding to the respective second bird's-eye view features to the world coordinate system; and the multiple second bird's-eye view features are transformed from the world coordinate system to the current visual frame coordinate system.

[0110] The world coordinate system is a reference coordinate system used to describe the absolute position of a vehicle in three-dimensional space. It is also called the global coordinate system or the cosmic coordinate system. The world coordinate system provides a unified reference framework for all other local coordinate systems (such as the camera coordinate system and the object coordinate system).

[0111] In this embodiment, since the historical visual frame coordinate systems corresponding to each of the second bird's-eye view features are different from the current visual frame coordinate system where the first bird's-eye view feature is located, in this case, the first bird's-eye view feature cannot be fused with the second bird's-eye view feature. Therefore, multiple second bird's-eye view features need to be converted to the current visual frame coordinate system to facilitate feature fusion.

[0112] In order to explain the above embodiment in more detail, the following describes the above embodiment in several parts.

[0113] The first part describes the content of converting the plurality of second bird's-eye view features from the respective historical visual frame coordinate systems corresponding to the respective second bird's-eye view features to the world coordinate system.

[0114] In some embodiments, a historical visual frame coordinate transformation matrix of the plurality of second bird's-eye view features in the historical visual frame coordinate system is obtained; and the historical visual frame coordinate transformation matrix is ​​converted to the world coordinate system based on a homogeneous coordinate transformation.

[0115] The historical visual frame coordinate transformation matrix represents the transformation matrix from the historical visual frame coordinate system at that historical moment to the world coordinate system. Homogeneous coordinate transformations are used to represent the expansion of points and directions in a three-dimensional void. Homogeneous coordinate transformations enable efficient rotation and translation of points or vectors.

[0116] It should be understood that the historical visual frame coordinate transformation matrix includes a rotation matrix and a translation vector. The historical visual frame coordinate transformation matrix represents the orientation transformation of the historical visual frame coordinate system relative to the world coordinate system. The translation vector represents the position of the origin of the historical visual frame coordinate system in the world coordinate system. In this case, after converting the bird's-eye view features perceived by the historical frame to the world coordinate system, it is possible to accumulate and construct an environmental map that conforms to the vehicle's motion pattern. For example, after converting the historical trajectory features of a dynamic vehicle to the world coordinate system, the long-term motion pattern of the dynamic vehicle can be analyzed in a unified world coordinate system.

[0117] The second part describes the content of converting the plurality of second bird's-eye view features from the world coordinate system to the current visual frame coordinate system.

[0118] It should be understood that during the vehicle's own movement, the objects around the vehicle are also moving. The second bird's-eye view feature of the historical frame is generated in the historical frame coordinate system at the historical moment, and the first bird's-eye view feature of the current frame is generated in the current frame coordinate system at the current moment. That is, the multiple second bird's-eye view features of multiple historical frames and the first bird's-eye view feature of the current frame are in different visual frame coordinate systems. If the second bird's-eye view features are not aligned with the first bird's-eye view features, it is easy for the objects detected in the historical frame to appear in the wrong position on the bird's-eye view. For example, because the vehicle is moving forward, the front vehicle is located directly in front of the vehicle in the historical frame, but it will "erroneously" appear behind the vehicle in the current frame coordinate system, making it impossible to distinguish between the real movement of the vehicle and the position movement caused by the vehicle movement. In this case, it is necessary to move the multiple second bird's-eye view features to the position of the first bird's-eye view feature so that the multiple second bird's-eye view features are aligned with the first bird's-eye view feature.

[0119] In some embodiments, the current visual frame coordinate transformation matrix of the first bird's-eye view feature in the current visual frame coordinate system is obtained; the inverse transformation matrix of the current visual frame coordinate transformation matrix is ​​calculated; and the historical visual frame coordinate transformation matrix in the world coordinate system is converted to the current visual frame coordinate system based on the homogeneous coordinate transformation.

[0120] The current visual frame coordinate transformation matrix represents the transformation matrix from the current visual frame coordinate system to the world coordinate system at the current moment. The inverse transformation matrix represents the conversion from the world coordinate system to the current visual frame coordinate system. The inverse transformation matrix represents the transformation in the opposite direction.

[0121] It should be understood that obtaining the current visual frame coordinate transformation matrix in the current visual frame coordinate system is equivalent to obtaining the matrix that transforms the coordinate points from the world coordinate system to the current visual frame coordinate system. Calculating the inverse transformation matrix of the current visual frame coordinate transformation matrix is ​​equivalent to obtaining the matrix that transforms the coordinate points from the current visual frame coordinate system to the world coordinate system. Based on the homogeneous coordinate transformation, the historical frame coordinate transformation matrix can be directly transformed into the current visual frame coordinate system.

[0122] The following describes how the plurality of second bird's-eye view features are displaced to the location of the first bird's-eye view feature so as to align the plurality of second bird's-eye view features with the first bird's-eye view feature.

[0123] The position of the first bird's-eye view feature is the coordinate of the first bird's-eye view feature in the current visual frame coordinate system.

[0124] It should be understood that when converting multiple second bird's-eye view features to the current visual frame coordinate system where the first bird's-eye view feature is located, if the second bird's-eye view features are directly superimposed on the first bird's-eye view features, it will cause ghosting between the second bird's-eye view features and the first bird's-eye view features. Therefore, it is necessary to align the second bird's-eye view features with the first bird's-eye view features.

[0125] In one possible implementation, in the current visual frame coordinate system, the time interval between the current visual frame and each of the historical visual frames is obtained; the weight corresponding to each of the second bird's-eye view features is calculated based on the time interval; and based on each of the weights, the multiple second bird's-eye view features are rotated and translated to the first coordinate of the first bird's-eye view feature in the current visual frame coordinate system.

[0126] In some embodiments, a timestamp generated for each visual frame is collected; and a difference between a first timestamp of a current visual frame and each timestamp of each historical visual frame is determined as a time interval between the current visual frame and each of the historical visual frames.

[0127] For example, the first timestamp of the current visual frame is T, and the first timestamp of the i-th historical visual frame is t i , then the time interval is Tt i .

[0128] It should be understood that if the time interval is smaller, it means that the environmental correlation between the historical visual frame and the current visual frame is higher, and the weight should be larger.

[0129] In some embodiments, each second bird's-eye view feature is multiplied by each weight, thereby performing feature alignment between the second bird's-eye view feature and the first bird's-eye view feature.

[0130] In this embodiment, in actual applications, the current visual frame may be missing features due to occlusion or lighting changes. In the coordinate system of the current visual frame, the time interval between the current visual frame and each of the historical visual frames is obtained; the weight corresponding to each of the second bird's-eye view features is calculated based on the time interval; and based on each of the weights, the multiple second bird's-eye view features are rotated and translated to the first coordinate of the first bird's-eye view feature in the coordinate system of the current visual frame. In other words, the first bird's-eye view feature is processed using the second bird's-eye view feature of the historical visual frame, and through time weight fusion, the changes of adjacent second bird's-eye view features are made smoother, optimizing the continuity of the vehicle's motion trajectory.

[0131] It should be understood that in actual applications, after the second bird's-eye view feature undergoes rotation and translation, the size or resolution of the second bird's-eye view feature may be inconsistent with that of the first bird's-eye view feature. Therefore, it is necessary to adjust the second bird's-eye view feature through interpolation or cropping so that the second bird's-eye view feature and the first bird's-eye view feature are in the same spatial dimension. For example, the second bird's-eye view feature may be adjusted through bilinear interpolation.

[0132] Step 304: The plurality of second bird's-eye view features are fused with the first bird's-eye view feature through a convolutional network to obtain a target bird's-eye view feature.

[0133] It should be understood that if the number of historical visual frames is large, 3D convolution can be used to process both spatial and temporal dimensions simultaneously.

[0134] Among them, the target bird's-eye view feature is used to convert multi-view images and other information into a unified 2D bird's-eye view network, so that the position, size and orientation of objects are more consistent with the geometric relationship of the real world, eliminating the occlusion problem caused by perspective projection.

[0135] In some embodiments, the second bird's-eye view feature and the first bird's-eye view feature are concatenated in the channel dimension and input into the convolution layer, the second bird's-eye view feature is convolved to generate a residual term, which is added to the first bird's-eye view feature to obtain the target bird's-eye view feature.

[0136] It should be understood that the second bird's-eye view feature and the first bird's-eye view feature can be fused through the above-mentioned splicing fusion and residual fusion, so that the lost information of the first bird's-eye view feature can be supplemented by the second bird's-eye view feature.

[0137] In some embodiments, lane information and road boundaries of the vehicle and its surroundings are detected based on the target bird's-eye view features, and a local map corresponding to the vehicle's driving trajectory is constructed based on the lane information and road boundaries.

[0138] It should be understood that since the first bird's-eye view feature of the current visual frame and the second bird's-eye view feature of multiple historical visual frames are fused, the vehicle is always moving within these visual frames, and the vehicle can scan the surrounding environment at different positions. If a target point is blocked in the current visual frame but not in the previous historical visual frames, the target point will appear in the fused target bird's-eye view feature, thereby overcoming the problem of local information loss due to vehicle occlusion.

[0139] An embodiment of the present application provides a bird's-eye view feature fusion method, which can obtain a first bird's-eye view feature of a vehicle's current visual frame and multiple second bird's-eye view features of multiple historical visual frames adjacent to the current visual frame. In the real world, when a vehicle is driving, the vehicle itself is moving, and other vehicles and pedestrians around the vehicle are also moving. However, the second bird's-eye view feature and the first bird's-eye view feature are generated based on the visual frame coordinate system at different times. If there is partially obscured information in the first bird's-eye view feature, the information in the first bird's-eye view feature can be supplemented by the second bird's-eye view feature. In order to effectively fuse past information to judge the scene of the vehicle at the current moment, it is necessary to fuse the bird's-eye view features, convert the multiple second bird's-eye view features to the current visual frame coordinate system where the first bird's-eye view feature is located, and shift the multiple second bird's-eye view features to the position where the first bird's-eye view feature is located, so that the multiple second bird's-eye view features are feature-aligned with the first bird's-eye view feature. The plurality of second bird's-eye view features are fused with the first bird's-eye view feature to obtain a target bird's-eye view feature, so that the target bird's-eye view feature can be used to identify the environment in which the vehicle is located. In other words, the plurality of second bird's-eye view features are fused with the first bird's-eye view feature to restore partially obscured information in the first bird's-eye view feature using visible information in the plurality of second bird's-eye view features, thereby increasing the accuracy of the bird's-eye view feature.

[0140] Figure 4 It is a structural schematic diagram of a feature fusion device for a bird's-eye view provided in an embodiment of the present application.

[0141] For example, Figure 4 As shown, the apparatus 400 includes:

[0142] An acquisition module 401 is configured to acquire a first bird's-eye view feature of a current visual frame of a vehicle and a plurality of second bird's-eye view features of a plurality of historical visual frames adjacent to the current visual frame;

[0143] a conversion and displacement module 402 for converting the plurality of second bird's-eye view features to the current visual frame coordinate system where the first bird's-eye view feature is located, and displacing the plurality of second bird's-eye view features to the position where the first bird's-eye view feature is located, so as to align the plurality of second bird's-eye view features with the first bird's-eye view feature;

[0144] The fusion module 403 is configured to fuse the plurality of second bird's-eye view features with the first bird's-eye view feature to obtain a target bird's-eye view feature, so as to identify the environment in which the vehicle is located through the target bird's-eye view feature.

[0145] In one possible implementation, the device 400 includes:

[0146] An acquisition module 401 is configured to acquire first visual data of a current visual frame of the vehicle;

[0147] a feature extraction module, configured to extract features from the first visual data to obtain first bird's-eye view features corresponding to the first visual data;

[0148] The acquisition module 401 is configured to acquire, from a queue, a plurality of second bird's-eye view features of the plurality of historical visual frames according to a time sequence of the plurality of historical visual frames.

[0149] In one possible implementation, the device 400 includes:

[0150] The cache module is configured to cache the second bird's-eye view features of the multiple historical visual frames in the queue form based on a temporal sequence of the current visual frame and the multiple historical visual frames.

[0151] In one possible implementation, the device 400 includes:

[0152] A conversion and displacement module 402 is configured to convert the plurality of second bird's-eye view features from respective historical visual frame coordinate systems corresponding to the respective second bird's-eye view features to a world coordinate system;

[0153] The conversion and displacement module 402 is configured to convert the plurality of second bird's-eye view features from the world coordinate system to the current visual frame coordinate system.

[0154] In one possible implementation, the device 400 includes:

[0155] The conversion and displacement module 402 is configured to obtain a historical visual frame coordinate transformation matrix of the plurality of second bird's-eye view features in the historical visual frame coordinate system;

[0156] A conversion and displacement module 402 is used to convert the historical visual frame coordinate transformation matrix into the world coordinate system based on homogeneous coordinate transformation;

[0157] In one possible implementation, the device 400 includes:

[0158] The conversion and displacement module 402 is used to obtain a current visual frame coordinate transformation matrix of the first bird's-eye view feature in the current visual frame coordinate system;

[0159] The conversion and displacement module 402 is used to calculate the inverse transformation matrix of the current visual frame coordinate transformation matrix;

[0160] The conversion and displacement module 402 is configured to convert the historical visual frame coordinate transformation matrix in the world coordinate system to the current visual frame coordinate system based on homogeneous coordinate transformation.

[0161] In one possible implementation, the device 400 includes:

[0162] An acquisition module 401 is configured to acquire, in the current visual frame coordinate system, a time interval between the current visual frame and each of the historical visual frames;

[0163] A weight calculation module, configured to calculate a weight corresponding to each of the second bird's-eye view features based on the time interval;

[0164] The conversion and displacement module 402 is configured to rotate and translate the plurality of second bird's-eye view features to the first coordinate of the first bird's-eye view feature in the current visual frame coordinate system based on the weights.

[0165] In one possible implementation, the device 400 includes:

[0166] The fusion module 403 is configured to fuse the plurality of second bird's-eye view features with the first bird's-eye view feature through a convolutional network to obtain a target bird's-eye view feature.

[0167] Figure 5 It is a structural schematic diagram of a vehicle provided in an embodiment of the present application.

[0168] For example, Figure 5 As shown, the vehicle 500 includes: a memory 501 and a processor 502, wherein the memory 501 stores an executable program code 503, and the processor 502 is used to call and execute the executable program code 503 to perform a feature fusion method of a bird's-eye view.

[0169] In addition, an embodiment of the present application also protects a device, which may include a memory and a processor, wherein the memory stores executable program code, and the processor is used to call and execute the executable program code to perform a feature fusion method of a bird's-eye view provided in an embodiment of the present application.

[0170] In this embodiment, the device can be divided into functional modules based on the above-described method examples. For example, each functional module can be mapped to a specific functional module, or two or more functions can be integrated into a single processing module. The integrated module can be implemented in hardware. It should be noted that the module division in this embodiment is illustrative and represents only a logical functional division. In actual implementation, other division methods may be used.

[0171] In the case of dividing each functional module into corresponding functional modules, the device may further include a weight calculation module, a feature extraction module, a cache module, etc. It should be noted that all relevant contents of each step involved in the above method embodiment can be referred to the functional description of the corresponding functional module and will not be repeated here.

[0172] It should be understood that the device provided in this embodiment is used to execute the above-mentioned feature fusion method of the bird's-eye view, and thus can achieve the same effect as the above-mentioned implementation method.

[0173] In the case of an integrated unit, the device may include a processing module and a storage module. When the device is used in a vehicle, the processing module may be used to control and manage the vehicle's movements, while the storage module may be used to support the vehicle's execution of relevant program codes.

[0174] The processing module may be a processor or controller that implements or executes the various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure of this application. The processor may also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a digital signal processor (DSP) and a microprocessor, and the storage module may be a memory.

[0175] In addition, the device provided in the embodiments of the present application can specifically be a chip, component or module, and the chip may include a connected processor and memory; wherein the memory is used to store instructions, and when the processor calls and executes the instructions, the chip can execute a feature fusion method of a bird's-eye view provided in the above embodiment.

[0176] This embodiment also provides a computer-readable storage medium, which stores computer program code. When the computer program code runs on a computer, the computer executes the above-mentioned related method steps to implement a feature fusion method of a bird's-eye view provided in the above embodiment.

[0177] This embodiment also provides a computer program product. When the computer program product is run on a computer, it enables the computer to execute the above-mentioned related steps to implement the feature fusion method of the bird's-eye view provided by the above embodiment.

[0178] Among them, the device, computer-readable storage medium, computer program product or chip provided in this embodiment are all used to execute the corresponding methods provided above. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects in the corresponding methods provided above, and will not be repeated here.

[0179] Through the description of the above implementation methods, technical personnel in the relevant field can understand that for the convenience and simplicity of description, only the division of the above-mentioned functional modules is used as an example. In actual applications, the above-mentioned functions can be distributed and completed by different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.

[0180] In the embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0181] The above content is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.

Claims

1. A feature fusion method for a bird's-eye view image, characterized in that: The method comprises: Acquire a first bird's-eye view feature of a current visual frame of the vehicle and a plurality of second bird's-eye view features of a plurality of historical visual frames adjacent to the current visual frame; Converting the plurality of second bird's-eye view features to the current visual frame coordinate system where the first bird's-eye view feature is located, and shifting the plurality of second bird's-eye view features to the position where the first bird's-eye view feature is located, so that the plurality of second bird's-eye view features are aligned with the first bird's-eye view feature; The plurality of second bird's-eye view features are fused with the first bird's-eye view features to obtain a target bird's-eye view feature, so that the environment in which the vehicle is located can be identified through the target bird's-eye view feature.

2. The method according to claim 1, characterized in that The acquiring of a first bird's-eye view feature of a current visual frame of the vehicle and a plurality of second bird's-eye view features of a plurality of historical visual frames adjacent to the current visual frame comprises: Acquiring first visual data of a current visual frame of the vehicle; performing feature extraction on the first visual data to obtain a first bird's-eye view feature corresponding to the first visual data; According to the time sequence of the multiple historical visual frames, multiple second bird's-eye view features of the multiple historical visual frames are obtained from the queue.

3. The method according to claim 2, characterized in that Before obtaining the plurality of second bird's-eye view features of the plurality of historical visual frames in a queue form according to the time sequence of the plurality of historical visual frames, the method includes: Based on the temporal sequence of the current visual frame and the multiple historical visual frames, the multiple second bird's-eye view features of the multiple historical visual frames are cached in the queue form.

4. The method according to claim 1, wherein The converting the plurality of second bird's-eye view features into the current visual frame coordinate system where the first bird's-eye view features are located comprises: Converting the plurality of second bird's-eye view features from respective historical visual frame coordinate systems corresponding to respective second bird's-eye view features to a world coordinate system; The plurality of second bird's-eye view features are transformed from the world coordinate system to the current visual frame coordinate system.

5. The method according to claim 4, characterized in that The converting the plurality of second bird's-eye view features from the historical visual frame coordinate system to the world coordinate system comprises: Obtaining a historical visual frame coordinate transformation matrix of the plurality of second bird's-eye view features in the historical visual frame coordinate system; Converting the historical visual frame coordinate transformation matrix to the world coordinate system based on homogeneous coordinate transformation; The converting the plurality of second bird's-eye view features from the world coordinate system to the current visual frame coordinate system comprises: Obtaining a current visual frame coordinate transformation matrix of the first bird's-eye view feature in the current visual frame coordinate system; Calculating an inverse transformation matrix of the current visual frame coordinate transformation matrix; The historical visual frame coordinate transformation matrix in the world coordinate system is converted to the current visual frame coordinate system based on homogeneous coordinate transformation.

6. The method according to claim 5, characterized in that The step of shifting the plurality of second bird's-eye view features to positions where the first bird's-eye view features are located comprises: In the current visual frame coordinate system, obtaining the time interval between the current visual frame and each of the historical visual frames; Calculating the weight corresponding to each of the second bird's-eye view features based on the time interval; The plurality of second bird's-eye view features are rotated and translated to the first coordinates of the first bird's-eye view feature in the current visual frame coordinate system based on the respective weights.

7. The method according to claim 1, characterized in that The step of fusing the plurality of second bird's-eye view features with the first bird's-eye view features to obtain a target bird's-eye view feature includes: The multiple second bird's-eye view features are fused with the first bird's-eye view features through a convolutional network to obtain a target bird's-eye view feature.

8. A feature fusion device for a bird's-eye view, characterized in that: The device comprises: an acquisition module, configured to acquire a first bird's-eye view feature of a current visual frame of the vehicle and a plurality of second bird's-eye view features of a plurality of historical visual frames adjacent to the current visual frame; a conversion and displacement module, configured to convert the plurality of second bird's-eye view features to a current visual frame coordinate system where the first bird's-eye view feature is located, and to displace the plurality of second bird's-eye view features to a position where the first bird's-eye view feature is located, so that the plurality of second bird's-eye view features are aligned with the first bird's-eye view feature; A fusion module is used to perform feature fusion on the multiple second bird's-eye view features and the first bird's-eye view features to obtain a target bird's-eye view feature, so that the environment in which the vehicle is located can be identified through the target bird's-eye view feature.

9. A vehicle, characterized in that: The vehicle comprises: a memory for storing executable program code; A processor is configured to call and run the executable program code from the memory, so that the vehicle executes the method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed, the method according to any one of claims 1 to 7 is implemented.