Method and apparatus for generating aerial view of vehicle, storage medium, electronic device, and vehicle

Through the combination of the convolutional network and the multi-head cross-attention module, a bird's-eye view of lightweight vehicles is generated, which solves the problems of large calculation volume and low generalization, and realizes efficient environmental perception in complex traffic scenarios.

WO2025145988A1PCT designated stage expired Publication Date: 2025-07-10CHONGQING CHANGAN AUTOMOBILE CO LTD +1

Patent Information

Application Number
PCT/CN2024/143374
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-02
Filing Date
2024-12-27
Publication Date
2025-07-10

AI Technical Summary

Technical Problem

In the prior art, the aerial view generation method of vehicles based on pure vision has a large amount of calculation and low generalization, and cannot adapt to complex and changeable real traffic scenarios.

Method used

A convolutional network is used to extract multi-scale image features, combine the camera parameters and three-dimensional coordinate grid of the vehicle camera to generate a bird's-eye view of the vehicle environment, and feature fusion is performed using a lightweight model and a multi-head cross attention module.

Benefits of technology

It reduces the computational volume, improves the generalization and robustness of the model, can better represent the surrounding environment of the car, and adapt to complex and changeable real traffic scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024143374_10072025_PF_FP_ABST
    Figure CN2024143374_10072025_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure provides a method and apparatus for generating an aerial view of a vehicle, a storage medium, an electronic device, and a vehicle. The method comprises: obtaining a road scenario image set of a plurality of directions acquired by a vehicle-mounted camera of a target vehicle; preprocessing the road scenario image set into a target image set of a target size, and obtaining camera parameters of the vehicle-mounted camera; using a convolutional network to extract multi-scale image features of the target image set, and using the central point of the target vehicle as an original point to create a three-dimensional coordinate grid; and using the multi-scale image features, the camera parameters, and the three-dimensional coordinate grid to generate an environment aerial view of the target vehicle.
Need to check novelty before this filing date? Find Prior Art

Description

Method, device, storage medium, electronic device, and vehicle for generating a bird's-eye view of a vehicle

[0001] This application claims priority to Chinese application No. 202410004732.1, filed on January 2, 2024, entitled "Method and device for generating a bird's-eye view of a vehicle, storage medium, and electronic device," the entire contents of which are incorporated herein by reference. Technical Field

[0002] The embodiments of the present disclosure relate to, but are not limited to, the field of intelligent driving. Specifically, they relate to a method for generating a bird's-eye view of a vehicle, a generating device, a storage medium, an electronic device, and a vehicle. Background Art

[0003] Among related technologies, autonomous (and assisted) driving has become a mainstream development in the automotive industry. Compared to lidar, camera-based approaches offer lower costs, simpler deployment, and less computational effort. Furthermore, cameras provide rich semantic information and can identify target types. Furthermore, mature camera-based perception methods can provide rich and accurate semantic information when fused with lidar data.

[0004] In related technologies, purely visual perception methods mostly use a bird's-eye view as a method for environmental representation, as it clearly and intuitively depicts the road environment surrounding the vehicle and effectively addresses object occlusion. Methods can be categorized into geometry-based and network-based approaches based on how the perspective image acquired by the camera is transformed into a bird's-eye view. Geometry-based methods use inverse perspective mapping to assign depth information to camera images for view transformation. However, these methods suffer from high computational complexity and low generalizability, making them incapable of adapting to complex and changing real-world traffic scenarios.

[0005] For the above-mentioned problems existing in related technologies, no efficient and accurate solutions have been found yet. Technical Solutions

[0006] The following is a summary of the subject matter described in detail herein. This summary is not intended to limit the scope of the claims.

[0007] The embodiments of the present disclosure provide a method for generating a bird's-eye view of a vehicle, a generating device, a storage medium, an electronic device, and a vehicle to solve technical problems in related technologies.

[0008] According to one embodiment of the present disclosure, a method for generating a bird's-eye view of a vehicle is provided, comprising: obtaining a set of road scene images from multiple orientations captured by an on-board camera of a target vehicle; preprocessing the road scene image set into a target image set of a target size, and obtaining camera parameters of the on-board camera; extracting multi-scale image features of the target image set using a convolutional network, and creating a three-dimensional coordinate grid with the center point of the target vehicle as the origin; and generating a bird's-eye view of the environment of the target vehicle using the multi-scale image features, the camera parameters, and the three-dimensional coordinate grid.

[0009] Optionally, obtaining the camera parameters of the vehicle-mounted camera includes: reading the scaling matrix and translation matrix from the pixel coordinate system to the camera coordinate system from the preset calibration parameters of the vehicle-mounted camera, and determining the scaling matrix and the translation matrix as camera intrinsic parameters; obtaining the first rotation matrix and the first translation matrix of the camera relative to the target vehicle according to the camera installation position of the vehicle-mounted camera, calculating the third rotation matrix and the third translation matrix of the vehicle yaw angle according to the second rotation matrix and the second translation matrix of the vehicle position at the current moment relative to the vehicle position in the initial state, and calculating the camera extrinsic parameters through the first rotation matrix, the first translation matrix, the second rotation matrix, the second translation matrix, the third rotation matrix and the third translation matrix; wherein, the camera parameters include the camera intrinsic parameters and the camera extrinsic parameters.

[0010] Optionally, using a convolutional network to extract multi-scale image features of the target image set includes: inputting the target image set into a pre-trained convolutional network, wherein the convolutional network includes N layers of network structures connected in series, and N is greater than 5; outputting first-scale image features and second-size image features from the Mth and Pth layer network structures of the convolutional network, respectively, and determining the first-scale image features and the second-size image features as multi-scale image features of the target image set, wherein M and P are both less than N.

[0011] Optionally, using the multi-scale image features, the camera parameters, and the three-dimensional coordinate grid to generate a bird's-eye view of the environment of the target vehicle includes: using the three-dimensional coordinate grid to generate a first position code of the drivable road and a second position code of the target vehicle; projecting each element in the multi-scale image features from a pixel coordinate system to a feature vector of a world coordinate system based on the camera parameters; using the feature vector to generate a third position code of the road scene image set; and using the first position code, the second position code, and the third position code to generate a bird's-eye view of the environment of the target vehicle.

[0012] Optionally, using the first position code, the second position code, and the third position code to generate a bird's-eye view of the environment of the target vehicle includes: reading the first model parameters and the second model parameters of the convolutional network; embedding the first model parameters and the second model parameters into the first position code and the second position code respectively to obtain a first query feature and a second query feature, wherein the first query feature is used to characterize the shape features and position features of the static road, and the second query feature is used to characterize the shape features and position features of the vehicle; using a projection network to map the multi-scale image features into a first intermediate feature, embedding the intermediate feature into the third position code to obtain a second intermediate feature, and configuring the first intermediate feature as a value and the second intermediate feature as a key; using the first intermediate feature and the second intermediate feature to obtain a feature map of the bird's-eye view to be generated; and using the feature map to generate a bird's-eye view of the environment of the target vehicle.

[0013] Optionally, using the first intermediate feature and the second intermediate feature to obtain a feature map of the bird's-eye view to be generated includes: inputting the first query feature and the second query feature with the key and value into a multi-head cross-attention module respectively to obtain a third query feature and a fourth query feature; upsampling and decoding the third query feature and the fourth query feature respectively to obtain a third intermediate feature and a fourth intermediate feature corresponding to the static road and the target vehicle respectively; and fusing the third intermediate feature and the fourth intermediate feature into a feature map of the bird's-eye view to be generated.

[0014] Optionally, using the feature map to generate a bird's-eye view of the environment of the target vehicle includes: creating a blank bird's-eye view; reducing the dimension of the feature map to three dimensions to obtain target elements of three dimensions, wherein the three dimensions include a static road dimension, other vehicle dimension, and own vehicle dimension; traversing each target element in each dimension respectively to filter out designated targets whose probability of occurrence is greater than a set probability threshold; searching for a target color that matches the dimension of the designated target; and filling the target color in the corresponding element position of the blank bird's-eye view to obtain a bird's-eye view of the environment of the target vehicle.

[0015] According to another embodiment of the present disclosure, a device for generating a vehicle bird's-eye view is provided, including: an acquisition module for acquiring a set of road scene images in multiple orientations captured by an on-board camera of a target vehicle; a processing module for preprocessing the road scene image set into a target image set of a target size and acquiring camera parameters of the on-board camera; an extraction module for extracting multi-scale image features of the target image set using a convolutional network and creating a three-dimensional coordinate grid with the center point of the target vehicle as the origin; and a generation module for generating a bird's-eye view of the environment of the target vehicle using the multi-scale image features, the camera parameters, and the three-dimensional coordinate grid.

[0016] Optionally, the processing module includes: a reading unit, used to read the scaling matrix and translation matrix from the pixel coordinate system to the camera coordinate system from the preset calibration parameters of the vehicle-mounted camera, and determine the scaling matrix and the translation matrix as camera intrinsic parameters; a calculation unit, used to obtain the first rotation matrix and the first translation matrix of the camera relative to the target vehicle according to the camera installation position of the vehicle-mounted camera, and calculate the third rotation matrix and the third translation matrix of the vehicle yaw angle according to the second rotation matrix and the second translation matrix of the vehicle position at the current moment relative to the vehicle position in the initial state, and calculate the camera extrinsic parameters through the first rotation matrix, the first translation matrix, the second rotation matrix, the second translation matrix, the third rotation matrix and the third translation matrix; wherein, the camera parameters include the camera intrinsic parameters and the camera extrinsic parameters.

[0017] Optionally, the extraction module includes: an input unit for inputting the target image set into a pre-trained convolutional network, wherein the convolutional network includes N layers of network structures connected in series, and N is greater than 5; an output unit for outputting first-scale image features and second-scale image features from the M-th layer and P-th layer network structures of the convolutional network, respectively, and determining the first-scale image features and the second-scale image features as multi-scale image features of the target image set, wherein M and P are both less than N.

[0018] Optionally, the generation module includes: a first generation unit, used to generate a first position code of the drivable road and a second position code of the target vehicle using the three-dimensional coordinate grid; a projection unit, used to project each element in the multi-scale image feature from a pixel coordinate system to a feature vector of a world coordinate system based on the camera parameters; a second generation unit, used to generate a third position code of the road scene image set using the feature vector; and a third generation unit, used to generate a bird's-eye view of the environment of the target vehicle using the first position code, the second position code, and the third position code.

[0019] Optionally, the third generation unit includes: a reading subunit, used to read the first model parameters and the second model parameters of the convolutional network; an embedding subunit, used to embed the first model parameters and the second model parameters into the first position code and the second position code respectively to obtain a first query feature and a second query feature, wherein the first query feature is used to characterize the shape feature and position feature of the static road, and the second query feature is used to characterize the shape feature and position feature of the vehicle; a processing subunit, used to use a projection network to map the multi-scale image feature into a first intermediate feature, embed the intermediate feature into the third position code to obtain a second intermediate feature, and configure the first intermediate feature as a value and the second intermediate feature as a key; a first generation subunit, used to use the first intermediate feature and the second intermediate feature to obtain a feature map of the bird's-eye view to be generated; a second generation subunit, used to use the feature map to generate an environmental bird's-eye view of the target vehicle.

[0020] Optionally, the first generation sub-unit is also used to: input the first query feature and the second query feature into the multi-head cross attention module with the key and value respectively to obtain a third query feature and a fourth query feature; upsample and decode the third query feature and the fourth query feature respectively to obtain a third intermediate feature and a fourth intermediate feature corresponding to the static road and the target vehicle respectively; and fuse the third intermediate feature and the fourth intermediate feature into a feature map of the bird's-eye view to be generated.

[0021] Optionally, the second generating sub-unit is also used to: create a blank bird's-eye view; reduce the dimension of the feature map to three dimensions to obtain target elements of three dimensions, wherein the three dimensions include a static road dimension, other vehicle dimension, and own vehicle dimension; traverse each target element in each dimension respectively to filter out designated targets whose probability of occurrence is greater than a set probability threshold; find a target color that matches the dimension of the designated target; fill the target color in the corresponding element position of the blank bird's-eye view to obtain an environmental bird's-eye view of the target vehicle.

[0022] According to another aspect of an embodiment of the present disclosure, a storage medium is further provided. The storage medium includes a stored program, and the above steps are executed when the program is run.

[0023] According to another aspect of an embodiment of the present disclosure, an electronic device is also provided, including a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other through the communication bus; wherein: the memory is used to store computer programs; and the processor is used to execute the steps in the above method by running the program stored in the memory.

[0024] According to another aspect of an embodiment of the present application, a vehicle is provided, comprising the above-mentioned electronic device.

[0025] The embodiment of the present disclosure further provides a computer program product comprising instructions, which, when executed on a computer, enables the computer to execute the steps in the above-mentioned method for generating a bird's-eye view of a vehicle. Beneficial effects

[0026] The beneficial effects of the present disclosure are:

[0027] 1. A bird's-eye view perception method based on a lightweight pure vision model can reduce model size while maintaining model accuracy, making it easier to deploy the network in vehicle-mounted equipment and controlling network size;

[0028] 2. Using two different bird's-eye view queries for static road and vehicle identification can solve the problem that small model network features cannot simultaneously carry multiple target information;

[0029] 3. By fusing features, the model can utilize the similarity between static road features and vehicle features to improve model accuracy, effectively represent the environment around the car, and have good generalization and robustness.

[0030] The above description is only an overview of the technical solution of the present disclosure. In order to more clearly understand the technical means of the present disclosure, it can be implemented in accordance with the contents of the specification. In order to make the above and other purposes, features and advantages of the present disclosure more obvious and easy to understand, the specific implementation methods of the present disclosure are listed below. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure or related technologies, the following briefly introduces the drawings required for describing the embodiments.

[0032] FIG1 is a block diagram of the hardware structure of a vehicle according to an embodiment of the present disclosure;

[0033] FIG2 is a flow chart of a method for generating a bird's-eye view of a vehicle according to an embodiment of the present disclosure;

[0034] FIG3 is a schematic diagram of the principle of sensing vehicle environment in an embodiment of the present disclosure;

[0035] FIG4 is a schematic diagram of the principle of upsampling in an embodiment of the present disclosure;

[0036] FIG5 is a flowchart of a bird's-eye view perception based on a lightweight model according to an embodiment of the present disclosure;

[0037] FIG6 is a structural block diagram of a device for generating a bird's-eye view of a vehicle according to an embodiment of the present disclosure.

[0038] Implementation of the present disclosure

[0039] In order to enable those skilled in the art to better understand the solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the drawings in the embodiments of the present disclosure. Obviously, the described embodiments are only embodiments of a part of the present disclosure, not all embodiments. Based on the embodiments in the present disclosure, all other embodiments obtained by ordinary technicians in this field without making creative work should fall within the scope of protection of the present disclosure. It should be noted that the embodiments in the present disclosure and the features in the embodiments can be combined with each other in the absence of conflict.

[0040] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the numbers used in this way can be interchanged where appropriate, so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0041] The method embodiments provided by the embodiments of the present disclosure can be executed in an on-board chip, an autonomous driving module, a vehicle safety module, an AEB component, a vehicle or a similar processing device. Taking operation on a vehicle as an example, FIG1 is a hardware structure block diagram of a vehicle according to an embodiment of the present disclosure. As shown in FIG1 , the vehicle may include one or more (only one is shown in FIG1 ) processors 102 (the processor 102 may include but is not limited to a processing device such as a microprocessor MCU or a programmable logic device FPGA) and a memory 104 for storing data. Optionally, the above-mentioned vehicle may also include a transmission device 106 and an input and output device 108 for communication functions. It will be understood by those skilled in the art that the structure shown in FIG1 is only for illustration and does not limit the structure of the above-mentioned vehicle. For example, the vehicle may also include more or fewer components than those shown in FIG1 , or have a configuration different from that shown in FIG1 .

[0042] The memory 104 can be used to store vehicle programs, for example, software programs and modules of application software, such as the vehicle program corresponding to the method for generating a bird's-eye view of a vehicle in an embodiment of the present disclosure. The processor 102 executes various functional applications and data processing by running the vehicle program stored in the memory 104, that is, implementing the above-mentioned method. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include a memory remotely located relative to the processor 102, and these remote memories may be connected to the vehicle via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0043] Transmission device 106 is used to receive or transmit data via a network. A specific example of such a network may include a wireless network provided by the vehicle's communications provider. In one embodiment, transmission device 106 includes a network interface controller (NIC), which can connect to other network devices via a base station to enable communication with the Internet. In another embodiment, transmission device 106 may be a radio frequency (RF) module for wireless communication with the Internet.

[0044] In this embodiment, a method for generating a vehicle bird's-eye view is provided. FIG2 is a flow chart of a method for generating a vehicle bird's-eye view according to an embodiment of the present disclosure. As shown in FIG2 , the process includes the following steps:

[0045] Step S202, obtaining a set of road scene images in multiple directions captured by a vehicle-mounted camera of a target vehicle;

[0046] Optionally, the target vehicle is equipped with 6 cameras, which respectively capture images of the front, left front, right front, rear, left rear and right rear of the vehicle, covering a 360° range around the vehicle. The vehicle collects road scene images in real time through the 6 cameras on board.

[0047] When pre-training the neural network model used in the solution of this embodiment, it is necessary to mark the vehicles and road targets in the image, and record the coordinate positions of the targets relative to the current car to generate a label file.

[0048] Step S204: preprocessing the road scene image set into a target image set of a target size, and obtaining camera parameters of the vehicle-mounted camera;

[0049] The collected image is scaled and cropped to obtain Where H and W are the image sizes, corresponding to length and width respectively.

[0050] Step S206: extracting multi-scale image features of the target image set using a convolutional network, and creating a three-dimensional coordinate grid with the center point of the target vehicle as the origin;

[0051] Optionally, a three-dimensional coordinate grid g∈[0,1] is created with the target vehicle itself as the origin. h×w×3 Used to represent the 100m×100m area around the vehicle, where each element g i Represents the three-dimensional coordinates (x i ,y i ,z i ).

[0052] Step S208 : Generate a bird's-eye view of the target vehicle's environment using multi-scale image features, camera parameters, and a three-dimensional coordinate grid.

[0053] Through the above steps, a set of road scene images from multiple directions captured by the on-board camera of the target vehicle is obtained; the road scene image set is preprocessed into a target image set of target size, and the camera parameters of the on-board camera are obtained; a convolutional network is used to extract multi-scale image features of the target image set, and a three-dimensional coordinate grid is created with the center point of the target vehicle as the origin; a bird's-eye view of the environment of the target vehicle is generated using the multi-scale image features, camera parameters, and the three-dimensional coordinate grid. By collecting visual images from multiple directions, multi-scale features are extracted from them, and a bird's-eye view of the vehicle's driving environment is generated with the three-dimensional coordinate grid, thereby realizing a purely visual environmental perception solution, solving the technical problem of high computational complexity in related technologies of perceiving the vehicle environment based on purely visual information, reducing the computational complexity of the vehicle's perception of the environment, improving generalization, and being able to adapt to complex and changeable real traffic scenarios.

[0054] In one implementation of this embodiment, obtaining the camera parameters of the vehicle-mounted camera includes: reading the scaling matrix and translation matrix from the pixel coordinate system to the camera coordinate system from the preset calibration parameters of the vehicle-mounted camera, and determining the scaling matrix and translation matrix as camera intrinsic parameters; obtaining the first rotation matrix and first translation matrix of the camera relative to the target vehicle according to the camera installation position of the vehicle-mounted camera, calculating the third rotation matrix and third translation matrix of the vehicle yaw angle according to the second rotation matrix and second translation matrix of the vehicle position at the current moment relative to the vehicle position in the initial state, and calculating the camera extrinsic parameters through the first rotation matrix, the first translation matrix, the second rotation matrix, the second translation matrix, the third rotation matrix and the third translation matrix; wherein, the camera parameters include camera intrinsic parameters and camera extrinsic parameters.

[0055] In this embodiment, the camera intrinsic parameters are obtained by camera calibration. The camera external parameters are calculated by the camera installation position and vehicle posture

[0056] The camera internal parameters are from the pixel coordinate system (x c ,y c ,z c ) to the camera coordinate system (u, v), including the matrices for scaling and translation transformations, as shown in formula (1).

[0057] Where, f is the focal length of the camera, dX, dY is the length of one pixel (mm), c x ,c y is the translation distance from the pixel origin to the camera origin (mm).

[0058] In this embodiment, the camera extrinsic parameter (R|t) used is based on the camera's current time t n The rotation amount compared to the initial state t0 and translation Get, used to represent the camera position in the world coordinate system. First, obtain the rotation matrix R of the camera relative to the vehicle based on the camera installation position c and the translation matrix t c , then according to the current time t n The rotation matrix R of the vehicle position relative to the initial state t0 w and the translation matrix t w , and finally calculate the rotation matrix R of the vehicle yaw angle v and the translation matrix t v The camera extrinsic parameters are calculated using the above three rotation matrices and translation matrices, as shown in formula (2):

[0059] In an example of this embodiment, extracting multi-scale image features of a target image set using a convolutional network includes: inputting the target image set into a pre-trained convolutional network, wherein the convolutional network includes N layers of network structures connected in series, where N is greater than 5; outputting first-scale image features and second-size image features from the Mth and Pth layer network structures of the convolutional network, respectively, and determining the first-scale image features and the second-size image features as multi-scale image features of the target image set, wherein M and P are both less than N.

[0060] The preprocessed image (target image set) is input into a lightweight EfficientNet-B4 (e.g., using the first 5 layers of the original EfficientNet-B4 network structure) for feature extraction to obtain multi-scale image features F i ={F1,F2}, where hi ,w i is the resolution of the feature.

[0061] In one implementation of this embodiment, generating a bird's-eye view of the target vehicle's environment using multi-scale image features, camera parameters, and a three-dimensional coordinate grid includes:

[0062] S11, generating a first position code of a drivable road and a second position code of a target vehicle using a three-dimensional coordinate grid;

[0063] Pass g (3D coordinate grid) through two multi-layer perceptrons (MLP) to generate bird’s-eye view position codes for drivable roads and vehicles respectively.

[0064] S12, projecting each element in the multi-scale image feature from the pixel coordinate system to a feature vector in the world coordinate system based on the camera parameters;

[0065] S13, generating a third position code of the road scene image set using the feature vector;

[0066] The multi-scale image features F are transformed using the camera’s internal and external parameters i Each element in is represented by pixel coordinates The eigenvector d projected into a vector representation of world coordinates i,j , as shown in formula (3):

[0067] Finally, d i,j The position encoding δ of the image is obtained by MLP i,j .

[0068] S14, using the first position code, the second position code, and the third position code to generate a bird's-eye view of the environment of the target vehicle.

[0069] In one example, using a first position code, a second position code, and a third position code to generate a bird's-eye view of the environment of a target vehicle includes: reading a first model parameter and a second model parameter of a convolutional network; embedding the first model parameter and the second model parameter into the first position code and the second position code, respectively, to obtain a first query feature and a second query feature, wherein the first query feature is used to characterize the shape feature and position feature of a static road, and the second query feature is used to characterize the shape feature and position feature of a vehicle; using a projection network to map multi-scale image features into a first intermediate feature, embedding the intermediate feature into a third position code to obtain a second intermediate feature, and configuring the first intermediate feature as a value and the second intermediate feature as a key; using the first intermediate feature and the second intermediate feature to obtain a feature map of a bird's-eye view to be generated; and using the feature map to generate a bird's-eye view of the environment of the target vehicle.

[0070] In this example, the learnable parameters p of the same size are generated according to the feature size of the bird's-eye view finally generated by the view conversion module (convolutional network). r ,p v (First model parameter and second model parameter). The learnable parameters generated by the following model are embedded into the bird’s-eye view position encoding c r ,c v Then get the bird's eye view query (First query feature and second query feature), which are used to record the shape features and position features of static roads and vehicles in a specific area respectively. In addition, the image feature F1 is used as After F1 passes through the projection network, it is embedded into the image position code δ as As shown in formula (4):

[0071] Where: B represents batch normalization, the query obtained r ,query v Input the key and value into the multi-head cross attention module to calculate query′ r ,query′ v , as shown in formula (5):

[0072] Where: h represents the number of attention heads, W q ,W k ,W v Represents the projection weight matrix of query, key, and value, W h Represents the linear projection matrix of multi-head attention, d k Indicates the dimensions of query and key.

[0073] At the same time, the image feature F2 is obtained through the above similar steps and Then with query′ r ,query′ v Perform multi-head cross attention separately, as shown in formula (5) (the third query feature and the fourth query feature); Figure 3 is a schematic diagram of the principle of perceiving the vehicle environment in an embodiment of the present disclosure, which overall includes three major parts: multi-camera image acquisition, preprocessing, environmental perception, and visualization processing. The environmental perception part includes: a feature extraction network to realize image feature extraction, a view conversion network to realize conversion from perspective view to bird's-eye view, an upsampling network to realize feature size enhancement, and a feature fusion network to realize the fusion of multiple features.

[0074] Optionally, using the first intermediate feature and the second intermediate feature to obtain a feature map of the bird's-eye view to be generated includes: inputting the first query feature and the second query feature with the key and the value into a multi-head cross-attention module to obtain a third query feature and a fourth query feature; upsampling and decoding the third query feature and the fourth query feature to obtain a third intermediate feature and a fourth intermediate feature corresponding to the static road and the target vehicle, respectively; and fusing the third intermediate feature and the fourth intermediate feature into a feature map of the bird's-eye view to be generated.

[0075] query″ r ,query″ v Perform upsampling decoding respectively to improve the feature size. Upsampling decoding performs a total of log2((H B ×W B ) / (h×w)) times, where H B ,W B is the size of the bird’s-eye view. Finally, we get the corresponding static road and vehicle The structure of the upsampling module is shown in FIG4 , which is a schematic diagram of the principle of upsampling in an embodiment of the present disclosure.

[0076] B r ,B v The feature fusion module is used to perform feature fusion. First, B r ,B v Cascade and use 1×1 convolution for dimensionality reduction. The reduced dimensionality features are input into the channel attention module SE to obtain Finally, B r ,B v Multiply element-by-element with B′,1-B′ and add them together to get the fused bird's-eye view feature map As shown in formula (6):

[0077] Optionally, using a feature map to generate a bird's-eye view of the target vehicle's environment includes: creating a blank bird's-eye view; reducing the feature map to three dimensions to obtain target elements of three dimensions, where the three dimensions include a static road dimension, other vehicle dimension, and own vehicle dimension; traversing each target element in each dimension separately to filter out designated targets whose probability of occurrence is greater than a set probability threshold; finding a target color that matches the dimension of the designated target; filling the target color in the corresponding element position of the blank bird's-eye view to obtain a bird's-eye view of the target vehicle's environment.

[0078] Create a blank bird's-eye view image with a size of 200×200. f A 1×1 convolution is used to reduce the dimensionality to 3 dimensions. After normalizing the three dimensions, a bird's-eye view representation is obtained. The three dimensions represent the static road representation, the vehicle representation, and the vehicle itself. This representation is 200×200 pixels, representing a 100m×100m area around the vehicle. Each element represents the probability of an object at that location. If an element in the representation exceeds a set probability threshold, it indicates that an object (road or vehicle) is present at that location. The corresponding blank pixel in the bird's-eye view is then filled with the color corresponding to the object. After iterating through all the representation elements, the final bird's-eye view is obtained.

[0079] The solution of this embodiment provides a bird's-eye view perception method based on a lightweight pure visual model. The embodiment of the present disclosure provides a bird's-eye view perception method based on a lightweight model. FIG5 is a flowchart of the bird's-eye view perception method based on a lightweight model in the embodiment of the present disclosure, which includes the following steps:

[0080] The six onboard cameras capture road scene images {f1,f2,f3,f4,f5,f6}. When training the model, it is necessary to annotate the vehicles and roads in the images and generate a label file.

[0081] Each image is scaled and cropped to obtain {f′1,f′2,f′3,f′4,f′5,f′6}. The camera intrinsic parameter K is obtained through camera calibration, and the camera extrinsic parameter (R|t) is calculated by the camera installation position and vehicle posture.

[0082] The preprocessed image is input into the lightweight EfficientNet-B4 (only the first 5 layers of the original EfficientNet-B4 network structure are used) for feature extraction. The features of the 3rd and 5th layers of the network are respectively taken to obtain the multi-scale image features F i ={F1,F2};

[0083] Establish a bird's-eye view three-dimensional coordinate grid g∈[0,1] with the vehicle itself as the origin h×w×3Used to represent the 100m×100m area around the vehicle. G is passed through two MLPs to generate bird’s-eye view position codes for drivable roads and vehicles respectively {c r ,c v Then, the image feature F is transformed into i Each element in is represented by pixel coordinates The vector projected into world coordinates is represented by d i,j , and finally d i,j The position encoding δ of the image is obtained by MLP i,j ;

[0084] The learnable parameters p generated by the model will be r ,p v Embedding bird's-eye view position encoding c r ,c v Then get the bird's-eye view query r ,query v In addition, the image feature F1 is used as the value after passing through the projection network, and F1 is embedded in the image position code δ after passing through the projection network as the key. r ,query v Input the key and value into the multi-head cross attention module to calculate query′ r ,query′ v ;

[0085] Use the same method as S5 to get key' and value' through image feature F2, and compare them with query' r ,query′ v Perform multi-head cross attention to obtain query″ r ,query″ v ;

[0086] query″ r ,query″ v Perform upsampling decoding respectively to increase the feature size and obtain B corresponding to the static road and vehicle respectively. r ,B v ;

[0087] Use the feature fusion module to r ,B v Fusion to obtain B f ;

[0088] Create a blank bird's-eye view for visualization. First, fReduce the dimension to 3 dimensions, traverse each dimension separately, compare each element with the set probability threshold, and fill the corresponding pixel in the blank bird's-eye view with the color corresponding to the target.

[0089] This embodiment provides a bird's-eye view perception method based on a lightweight pure vision model. The model adopts a Transformer architecture, using the lightweight EfficientNet-B4 as a feature extraction network. This reduces model size while maintaining accuracy, making it easier to deploy the network in vehicle-mounted devices and controlling network size. Two different bird's-eye view queries are used for static road and vehicle recognition, addressing the problem that small model network features cannot simultaneously carry multiple target information. Finally, the features are fused, enabling the model to leverage the similarities between static road features and vehicle features to improve model accuracy. This effectively represents the vehicle's surrounding environment with good generalization and robustness.

[0090] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus the necessary general hardware platform, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the embodiment of the present disclosure is essentially or the part that contributes to the prior art can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, disk, CD-ROM), including a number of instructions for enabling a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the method described in each embodiment of the embodiment of the present disclosure.

[0091] This embodiment also provides a device for generating a bird's-eye view of a vehicle, which is used to implement the above-mentioned embodiments and preferred embodiments. Details already described will not be repeated here. As used below, the term "module" may refer to a combination of software and / or hardware that implements a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation using hardware, or a combination of software and hardware, is also possible and contemplated.

[0092] FIG6 is a structural block diagram of a device for generating a bird's-eye view of a vehicle according to an embodiment of the present disclosure. As shown in FIG6 , the device includes:

[0093] An acquisition module 60 is configured to acquire a set of road scene images in multiple directions captured by a vehicle-mounted camera of a target vehicle;

[0094] a processing module 62, configured to preprocess the road scene image set into a target image set of a target size, and obtain camera parameters of the vehicle-mounted camera;

[0095] An extraction module 64 is configured to extract multi-scale image features of the target image set using a convolutional network and create a three-dimensional coordinate grid with the center point of the target vehicle as the origin;

[0096] The generating module 66 is configured to generate a bird's-eye view of the environment of the target vehicle using the multi-scale image features, the camera parameters, and the three-dimensional coordinate grid.

[0097] Optionally, the processing module includes: a reading unit, used to read the scaling matrix and translation matrix from the pixel coordinate system to the camera coordinate system from the preset calibration parameters of the vehicle-mounted camera, and determine the scaling matrix and the translation matrix as camera intrinsic parameters; a calculation unit, used to obtain the first rotation matrix and the first translation matrix of the camera relative to the target vehicle according to the camera installation position of the vehicle-mounted camera, and calculate the third rotation matrix and the third translation matrix of the vehicle yaw angle according to the second rotation matrix and the second translation matrix of the vehicle position at the current moment relative to the vehicle position in the initial state, and calculate the camera extrinsic parameters through the first rotation matrix, the first translation matrix, the second rotation matrix, the second translation matrix, the third rotation matrix and the third translation matrix; wherein, the camera parameters include the camera intrinsic parameters and the camera extrinsic parameters.

[0098] Optionally, the extraction module includes: an input unit for inputting the target image set into a pre-trained convolutional network, wherein the convolutional network includes N layers of network structures connected in series, and N is greater than 5; an output unit for outputting first-scale image features and second-scale image features from the M-th layer and P-th layer network structures of the convolutional network, respectively, and determining the first-scale image features and the second-scale image features as multi-scale image features of the target image set, wherein M and P are both less than N.

[0099] Optionally, the generation module includes: a first generation unit, used to generate a first position code of the drivable road and a second position code of the target vehicle using the three-dimensional coordinate grid; a projection unit, used to project each element in the multi-scale image feature from a pixel coordinate system to a feature vector of a world coordinate system based on the camera parameters; a second generation unit, used to generate a third position code of the road scene image set using the feature vector; and a third generation unit, used to generate a bird's-eye view of the environment of the target vehicle using the first position code, the second position code, and the third position code.

[0100] Optionally, the third generation unit includes: a reading subunit, used to read the first model parameters and the second model parameters of the convolutional network; an embedding subunit, used to embed the first model parameters and the second model parameters into the first position code and the second position code respectively to obtain a first query feature and a second query feature, wherein the first query feature is used to characterize the shape feature and position feature of the static road, and the second query feature is used to characterize the shape feature and position feature of the vehicle; a processing subunit, used to use a projection network to map the multi-scale image feature into a first intermediate feature, embed the intermediate feature into the third position code to obtain a second intermediate feature, and configure the first intermediate feature as a value and the second intermediate feature as a key; a first generation subunit, used to use the first intermediate feature and the second intermediate feature to obtain a feature map of the bird's-eye view to be generated; a second generation subunit, used to use the feature map to generate an environmental bird's-eye view of the target vehicle.

[0101] Optionally, the first generation sub-unit is also used to: input the first query feature and the second query feature into the multi-head cross attention module with the key and value respectively to obtain a third query feature and a fourth query feature; upsample and decode the third query feature and the fourth query feature respectively to obtain a third intermediate feature and a fourth intermediate feature corresponding to the static road and the target vehicle respectively; and fuse the third intermediate feature and the fourth intermediate feature into a feature map of the bird's-eye view to be generated.

[0102] Optionally, the second generating sub-unit is also used to: create a blank bird's-eye view; reduce the dimension of the feature map to three dimensions to obtain target elements of three dimensions, wherein the three dimensions include a static road dimension, other vehicle dimension, and own vehicle dimension; traverse each target element in each dimension respectively to filter out designated targets whose probability of occurrence is greater than a set probability threshold; find a target color that matches the dimension of the designated target; fill the target color in the corresponding element position of the blank bird's-eye view to obtain an environmental bird's-eye view of the target vehicle.

[0103] It should be noted that the above modules can be implemented through software or hardware. For the latter, it can be implemented in the following ways, but not limited to: the above modules are all located in the same processor; or the above modules are located in different processors in any combination.

[0104] An embodiment of the present disclosure further provides a storage medium storing a computer program, wherein the computer program is configured to execute the steps of any one of the above method embodiments when running.

[0105] Optionally, in this embodiment, the storage medium may be configured to store a computer program for performing the following steps:

[0106] S1, obtaining a set of road scene images in multiple directions captured by the onboard camera of the target vehicle;

[0107] S2, preprocessing the road scene image set into a target image set of a target size, and obtaining camera parameters of the vehicle-mounted camera;

[0108] S3, extracting multi-scale image features of the target image set using a convolutional network, and creating a three-dimensional coordinate grid with the center point of the target vehicle as the origin;

[0109] S4, generating a bird's-eye view of the environment of the target vehicle using the multi-scale image features, the camera parameters, and the three-dimensional coordinate grid.

[0110] Optionally, in this embodiment, the above-mentioned storage medium may include but is not limited to: a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk, and other media that can store computer programs.

[0111] An embodiment of the present disclosure further provides an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to perform the steps in any one of the above method embodiments.

[0112] Optionally, the electronic device may further include a transmission device and an input / output device, wherein the transmission device is connected to the processor, and the input / output device is connected to the processor.

[0113] Optionally, in this embodiment, the processor may be configured to execute the following steps through a computer program:

[0114] S1, obtaining a set of road scene images in multiple directions captured by the onboard camera of the target vehicle;

[0115] S2, preprocessing the road scene image set into a target image set of a target size, and obtaining camera parameters of the vehicle-mounted camera;

[0116] S3, extracting multi-scale image features of the target image set using a convolutional network, and creating a three-dimensional coordinate grid with the center point of the target vehicle as the origin;

[0117] S4, generating a bird's-eye view of the environment of the target vehicle using the multi-scale image features, the camera parameters, and the three-dimensional coordinate grid.

[0118] In yet another embodiment provided by the present disclosure, a computer program product including instructions is further provided. When the computer program product is executed on a computer, the computer is enabled to execute the method for generating a bird's-eye view of a vehicle in the above embodiment.

[0119] An embodiment of the present disclosure also provides a vehicle, comprising the electronic device as described above.

[0120] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function according to the embodiment of the present disclosure is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center via a wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) method. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more available media integrated. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive (SSD)).

[0121] Each embodiment in this specification is described in a related manner. Similar portions between the various embodiments can be referenced to each other. Each embodiment focuses on the differences from other embodiments. The embodiments of the apparatus, electronic device, computer-readable storage medium, and computer program product containing instructions thereof are generally similar to the method embodiments, so their description is relatively simple. For related portions, reference can be made to the description of the method embodiments.

[0122] The above are only preferred embodiments of the present disclosure and are not intended to limit the scope of protection of the present disclosure. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present disclosure are included in the scope of protection of the present disclosure.

Claims

1. A method for generating a bird's-eye view of a vehicle, wherein, Including: Obtain a set of road scene images in multiple directions collected by an in-vehicle camera of a target vehicle; Preprocess the set of road scene images into a set of target images of a target size, and obtain the camera parameters of the in-vehicle camera; Extract multi-scale image features of the set of target images using a convolutional network, and create a three-dimensional coordinate grid with the center point of the target vehicle as the origin; Generate an environmental bird's-eye view of the target vehicle using the multi-scale image features, the camera parameters, and the three-dimensional coordinate grid.

2. The method for generating a bird's-eye view of a vehicle according to claim 1, wherein, Obtaining the camera parameters of the in-vehicle camera includes: Read the scaling matrix and translation matrix from the pixel coordinate system to the camera coordinate system from the preset calibration parameters of the in-vehicle camera, and determine the scaling matrix and the translation matrix as the camera internal parameters; Obtain the first rotation matrix and the first translation matrix of the camera relative to the target vehicle according to the camera installation position of the in-vehicle camera, calculate the third rotation matrix and the third translation matrix of the vehicle yaw angle according to the second rotation matrix and the second translation matrix of the vehicle position at the current moment relative to the vehicle position in the initial state, and calculate the camera external parameters through the first rotation matrix, the first translation matrix, the second rotation matrix, the second translation matrix, the third rotation matrix and the third translation matrix; Wherein, the camera parameters include the camera internal parameters and the camera external parameters.

3. The method for generating a bird's-eye view of a vehicle according to claim 1, wherein, Extracting multi-scale image features of the set of target images using a convolutional network includes: Input the set of target images into a pre-trained convolutional network, wherein the convolutional network includes N network structures connected in series in sequence, and N is greater than 5; Output the first-scale image features and the second-scale image features from the Mth and Pth network structures of the convolutional network respectively, and determine the first-scale image features and the second-scale image features as the multi-scale image features of the set of target images, wherein both M and P are less than N.

4. The method for generating a bird's-eye view of a vehicle according to claim 1, wherein, Generating an environmental bird's-eye view of the target vehicle using the multi-scale image features, the camera parameters, and the three-dimensional coordinate grid includes: Generate a first position encoding of the drivable road and a second position encoding of the target vehicle using the three-dimensional coordinate grid; Project each element in the multi-scale image features from the pixel coordinate system into a feature vector in the world coordinate system based on the camera parameters; Generate a third position encoding of the set of road scene images using the feature vector; Generate an environmental bird's-eye view of the target vehicle using the first position encoding, the second position encoding, and the third position encoding.

5. The method for generating a bird's-eye view of a vehicle according to claim 4, wherein, Generating an environmental bird's-eye view of the target vehicle using the first position encoding, the second position encoding, and the third position encoding includes: Read the first model parameter and the second model parameter of the convolutional network; Embed the first model parameter and the second model parameter into the first position encoding and the second position encoding respectively to obtain a first query feature and a second query feature, wherein the first query feature is used to characterize the shape feature and position feature of the static road, and the second query feature is used to characterize the shape feature and position feature of the vehicle; Using a projection network to map the multi-scale image features to first intermediate features, embedding the intermediate features into the third positional encoding to obtain second intermediate features, configuring the first intermediate features as values, and configuring the second intermediate features as keys; Using the first intermediate features and the second intermediate features to obtain a feature map of the bird's-eye view to be generated; Using the feature map to generate an environmental bird's-eye view of the target vehicle.

6. The method for generating a bird's-eye view of a vehicle according to claim 5, wherein, Using the first intermediate features and the second intermediate features to obtain a feature map of the bird's-eye view to be generated includes: Inputting the first query feature and the second query feature, the key, and the value into a multi-head cross-attention module respectively to obtain a third query feature and a fourth query feature; Performing upsampling decoding on the third query feature and the fourth query feature respectively to obtain a third intermediate feature corresponding to the static road and a fourth intermediate feature corresponding to the target vehicle respectively; Fusing the third intermediate feature and the fourth intermediate feature into a feature map of the bird's-eye view to be generated.

7. The method for generating a bird's-eye view of a vehicle according to claim 5, wherein, Using the feature map to generate an environmental bird's-eye view of the target vehicle includes: Creating a blank bird's-eye view; Reducing the dimension of the feature map to three dimensions to obtain target elements in three dimensions, where the three dimensions include a static road dimension, an other vehicle dimension, and a self-vehicle dimension; Traversing each target element in each dimension respectively to screen out specified targets with an occurrence probability greater than a set probability threshold; Searching for a target color that matches the dimension of the specified target; Filling the target color at the corresponding element position in the blank bird's-eye view to obtain an environmental bird's-eye view of the target vehicle.

8. An apparatus for generating a bird's-eye view of a vehicle, wherein, Includes: An acquisition module for acquiring a set of road scene images in multiple directions collected by an on-vehicle camera of a target vehicle; A processing module for preprocessing the set of road scene images into a set of target images with a target size and acquiring camera parameters of the on-vehicle camera; An extraction module for extracting multi-scale image features of the set of target images by using a convolutional network and creating a three-dimensional coordinate grid with the center point of the target vehicle as the origin; A generation module for generating an environmental bird's-eye view of the target vehicle by using the multi-scale image features, the camera parameters, and the three-dimensional coordinate grid.

9. The vehicle bird's-eye view generation device according to claim 8, wherein, The processing module includes: A reading unit for reading a scaling matrix and a translation matrix from pixel coordinates to camera coordinates from preset calibration parameters of the on-vehicle camera and determining the scaling matrix and the translation matrix as camera internal parameters; A calculation unit for obtaining a first rotation matrix and a first translation matrix of the camera relative to the target vehicle according to the camera installation position of the on-vehicle camera, calculating a third rotation matrix and a third translation matrix of the vehicle yaw angle according to a second rotation matrix and a second translation matrix of the vehicle position at the current moment relative to the vehicle position in the initial state, and calculating camera external parameters through the first rotation matrix, the first translation matrix, the second rotation matrix, the second translation matrix, the third rotation matrix, and the third translation matrix; where the camera parameters include the camera internal parameters and the camera external parameters.

10. The generating device for the bird's-eye view of a vehicle according to claim 8, wherein, The extraction module includes: An input unit for inputting the target image set into a pre-trained convolutional network, where the convolutional network includes N network structures connected in series in sequence, and N is greater than 5; An output unit for respectively outputting a first-scale image feature and a second-scale image feature from the M-th and P-th network structures of the convolutional network, and determining the first-scale image feature and the second-scale image feature as the multi-scale image features of the target image set, where both M and P are less than N.

11. The generating device for the bird's-eye view of a vehicle according to claim 8, wherein, The generation module includes: A first generation unit for generating a first position encoding of the drivable road and a second position encoding of the target vehicle by using the three-dimensional coordinate grid; A projection unit for projecting each element in the multi-scale image features from the pixel coordinate system into a feature vector in the world coordinate system based on the camera parameters; A second generation unit for generating a third position encoding of the road scene image set by using the feature vector; A third generation unit for generating an environmental bird's-eye view of the target vehicle by using the first position encoding, the second position encoding, and the third position encoding.

12. The vehicle bird's-eye view generation device according to claim 11, wherein, The third generation unit includes: A reading subunit for reading a first model parameter and a second model parameter of the convolutional network; An embedding subunit for respectively embedding the first model parameter and the second model parameter into the first position encoding and the second position encoding to obtain a first query feature and a second query feature, where the first query feature is used to characterize the shape feature and position feature of the static road, and the second query feature is used to characterize the shape feature and position feature of the vehicle; A processing subunit for mapping the multi-scale image features into a first intermediate feature by using a projection network, embedding the intermediate feature into the third position encoding to obtain a second intermediate feature, configuring the first intermediate feature as value, and configuring the second intermediate feature as key; A first generation subunit for obtaining a feature map of the bird's-eye view to be generated by using the first intermediate feature and the second intermediate feature; A second generation subunit for generating an environmental bird's-eye view of the target vehicle by using the feature map.

13. The vehicle bird's-eye view generation device according to claim 12, wherein, The first generation subunit is further used for: Inputting the first query feature and the second query feature, the key and the value into a multi-head cross-attention module respectively to obtain a third query feature and a fourth query feature; Performing upsampling decoding on the third query feature and the fourth query feature respectively to obtain a third intermediate feature corresponding to the static road and a fourth intermediate feature corresponding to the target vehicle; Fusing the third intermediate feature and the fourth intermediate feature into a feature map of the bird's-eye view to be generated.

14. The generating device for the bird's-eye view of the vehicle according to claim 12, wherein, The second generation subunit is further used for: Creating a blank bird's-eye view; Reducing the dimension of the feature map to three dimensions to obtain target elements in three dimensions, where the three dimensions include a static road dimension, an other vehicle dimension, and a self-vehicle dimension; Traversing each target element in each dimension respectively, and screening out a specified target whose occurrence probability is greater than a set probability threshold. Find a target color that matches the dimension of the specified target; fill the target color at the corresponding element position in the blank bird's-eye view to obtain the environmental bird's-eye view of the target vehicle.

15. The vehicle bird's-eye view generation device according to claim 8, wherein, The generating device of the vehicle bird's-eye view is applied to a vehicle.

16. A storage medium, wherein, A computer program is stored in the storage medium, wherein the computer program is configured to execute the method described in any one of claims 1 to 7 when running.

17. An electronic device, comprising a memory and a processor, wherein, A computer program is stored in the memory, and the processor is configured to run the computer program to execute the method described in any one of claims 1 to 7.

18. A vehicle, wherein, It includes the electronic device described in claim 17.

Citation Information

Patent Citations

  • Image processing method, device and equipment and computer readable storage medium

    CN114723955A

  • Bird-eye view feature generation method based on vehicle-mounted look-around image

    CN115588175A

  • Automatic driving image recognition method and device and recognition equipment

    CN116229394A

  • Object recognition method and device, vehicle and storage medium

    CN117079231A

  • Vehicle aerial view generation method and device, storage medium and electronic device

    CN117830526A

Cited By

  • Multi-view target detection tracking method and device based on adaptive fusion and time sequence association

    CN120747168A

  • Model training method and device, electronic equipment, medium, product and vehicle

    CN121505396A