Vehicle control method and device, vehicle, storage medium and program product

By extracting the top view features in multiple sensor data and determining road environment traffic information, the problem of low computing efficiency of environmental perception of autonomous vehicles is solved, and more efficient and accurate environmental perception and vehicle control are achieved.

CN119975342APending Publication Date: 2025-05-13BYD CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510169154.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-14
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The prior art methods used for the perception of the surrounding environment of vehicles in the field of autonomous driving have problems such as high computational complexity, low computational efficiency and insufficient robustness, which leads to the inability to obtain the surrounding environment information of the vehicle in real time.

Method used

By acquiring environmental perception data around the vehicle collected by multiple sensors, the top-view features are extracted, and the road environment traffic information of the vehicle is determined based on these features, thereby determining the vehicle's control scheme.

Benefits of technology

The calculation efficiency and accuracy of the vehicle control method are improved, the road environment traffic information around the vehicle can be determined faster, and the vehicle's environmental perception ability and driving safety are enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119975342A_ABST
    Figure CN119975342A_ABST
Patent Text Reader

Abstract

The invention discloses a vehicle control method and device, a vehicle, a storage medium and a program product, relates to the technical field of automobiles, and aims to solve the problem that the calculation efficiency is low when the surrounding environment of the vehicle is recognized. The vehicle control method comprises the steps that environment sensing data, collected by various sensors, around a vehicle is obtained; extracting overlook features from the environmental perception data; determining road environment traffic information of the vehicle based on the overlook characteristics; and determining a control scheme of the vehicle based on the road environment traffic information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of automobile technology, and in particular to a vehicle control method, a device, a vehicle, a storage medium and a program product. Background Art

[0002] Object detection plays an important role in the field of autonomous driving. It can accurately identify and analyze the vehicle's surrounding environment information (other vehicles, pedestrians, obstacles, etc.), thereby providing users with a safe and reliable autonomous driving experience.

[0003] Currently, a combination of point cloud and image is used to extract the vehicle's surrounding environment information, but it has problems such as high computational complexity, low computational efficiency, and lack of robustness, resulting in the inability to obtain the vehicle's surrounding environment information in real time. Summary of the invention

[0004] The object of the present invention is to provide a vehicle control method, device, vehicle, storage medium and program product, aiming to solve the problem of low calculation efficiency when identifying the vehicle's surrounding environment.

[0005] In order to achieve the above object, the present invention adopts the following technical scheme:

[0006] In a first aspect, the present invention provides a vehicle control method, comprising: acquiring environmental perception data of the vehicle's surroundings collected by multiple sensors; extracting bird's-eye view features from the environmental perception data; determining the vehicle's road environment traffic information based on the bird's-eye view features; and determining the vehicle's control scheme based on the road environment traffic information.

[0007] The vehicle control method provided by the embodiment of the present application uses environmental perception data around the vehicle collected by multiple sensors, which is not limited by the line of sight of a single sensor, and helps the vehicle to perceive the surrounding environment more comprehensively, thereby improving the reliability of the vehicle control method. The vehicle control method provided by the present application extracts overhead features from environmental perception data, and generates it from a bird's-eye view from above, so that the road environment traffic information is displayed in a more intuitive way, and the road environment traffic information around the vehicle can be determined more quickly, thereby improving the calculation efficiency, and thus improving the calculation efficiency and accuracy of the vehicle control method.

[0008] In some embodiments, environmental perception data includes multi-view images; extracting bird's-eye view features from the environmental perception data includes: performing feature extraction on the multi-view images to obtain multi-view image features; for image features of each perspective in the multi-view image features, calculating the position encoding of the image features of each perspective; determining width feature information of the image features of each perspective; and determining the bird's-eye view features based at least on the position encoding of the image features of each perspective and the width feature information of the image features of each perspective.

[0009] In some embodiments, determining width feature information of image features of each viewing angle includes: performing height feature compression on the image features of each viewing angle to obtain compressed image features of each viewing angle; and extracting width feature information from the compressed image features of each viewing angle.

[0010] In some embodiments, width feature information is extracted from the compressed image features of each perspective, including: concatenating the height feature information of the compressed image features of each perspective to obtain initial width feature information; and using an attention mechanism model to refine the initial width feature information to obtain width feature information.

[0011] In some embodiments, the attention mechanism model is a lightweight Transformer model.

[0012] In some embodiments, calculating the position coding of the image features of each viewing angle includes: calculating the position coding of the image features of each viewing angle based on a reference position coding method.

[0013] In some embodiments, the position encoding of the image feature at each viewing angle includes at least one of the following: the distance of the image feature relative to the vehicle on the top-down plane, the rotation angle of the image feature relative to the vehicle on the top-down plane, and the height of the image feature relative to the ground.

[0014] In some embodiments, when the position encoding of the image feature at each perspective includes the height of the image feature relative to the ground, the overhead view feature is determined based on the position encoding of the image feature at each perspective and the width feature information of the image feature at each perspective, including: weighting the height of the image feature at each perspective relative to the ground and the width feature information of the image feature at each perspective to obtain weighted height and width feature information; and inputting the weighted height and width feature information into a Transformer decoder to obtain the position information of the overhead view feature.

[0015] In some embodiments, the overhead view feature is determined based at least on the position coding of the image feature of each viewing angle and the width feature information of the image feature of each viewing angle, including: determining the position information of the image feature of each viewing angle at the overhead view; determining the overhead view feature based on the position coding of the image feature of each viewing angle, the width feature information of the image feature of each viewing angle, and the position information of the image feature of each viewing angle at the overhead view.

[0016] In some embodiments, when the position encoding of the image feature at each perspective includes the height of the image feature relative to the ground, the overhead feature is determined based on the position encoding of the image feature at each perspective, width feature information of the image feature at each perspective, and position information of the image feature at each perspective at a overhead perspective, including: weighting the height of the image feature at each perspective relative to the ground and the width feature information of the image feature at each perspective to obtain weighted height and width feature information; inputting the weighted height and width feature information, the width feature information of the image feature at each perspective, and the position information of the image feature at the overhead perspective into a Transformer decoder to obtain the position information of the overhead feature.

[0017] In some embodiments, the width feature information of the image features of each perspective is used as the key vector of the Transformer decoder, the weighted height and width feature information is used as the value vector of the Transformer decoder, and the position information of the image features of each perspective in the top-down perspective is used as the query vector of the Transformer decoder.

[0018] In some embodiments, determining the position information of the image features of each perspective in the overhead perspective includes: taking the position of the vehicle as the origin, dividing the surrounding environment of the vehicle into a plurality of regular grid units from a overhead perspective above the vehicle, and obtaining the center coordinates of each grid unit; based on the center coordinates of each grid unit, determining the position information of the image features of each perspective in the overhead perspective.

[0019] In some embodiments, based on the overhead features, the road environment traffic information of the vehicle is determined, including: constructing a bird's-eye view environment plan view based on the overhead features; extracting features of each target object from the bird's-eye view environment plan view; and using a target detection algorithm to identify the vehicle's road traffic information based on the features of each target object.

[0020] In some embodiments, the road environment traffic information of the vehicle includes at least one of the following: lane lines, traffic signs, pedestrians, and vehicles.

[0021] In some embodiments, a control scheme for a vehicle is determined based on road environment traffic information, including: planning a vehicle's driving route based on road traffic information; determining a control scheme for the vehicle based on the vehicle's driving route; the control scheme is used to adjust the vehicle's handling parameters so that the vehicle travels along the driving route.

[0022] In a second aspect, the present application provides an electronic device, comprising: a processor and a memory; the memory stores instructions executable by the processor; when the processor is configured to execute the instructions, the electronic device implements the method of the first aspect above.

[0023] In a third aspect, the present application provides a vehicle, which includes the electronic device according to the second aspect.

[0024] In a fourth aspect, the present application provides a computer-readable storage medium, which includes: computer software instructions; when the computer software instructions are executed in an electronic device, the electronic device implements the method of the first aspect above.

[0025] In a fifth aspect, the present application provides a computer program product, which includes a computer program; when the computer program runs in an electronic device, the electronic device implements the method of the first aspect.

[0026] The beneficial effects of the second to fifth aspects mentioned above refer to the corresponding description of the first aspect and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0028] Figure 1 A schematic diagram of the composition of the vehicle control system provided for this application;

[0029] Figure 2 A schematic diagram of a vehicle control method provided in an embodiment of the present application;

[0030] Figure 3 A flow chart of another vehicle control method provided in an embodiment of the present application;

[0031] Figure 4 A flow chart of another vehicle control method provided in an embodiment of the present application;

[0032] Figure 5 A schematic diagram of a process for extracting top-view features provided in an embodiment of the present application;

[0033] Figure 6 A schematic diagram of the composition of a vehicle control device provided in an embodiment of the present application;

[0034] Figure 7 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application.

[0035] Reference numerals: sensor 100, processor 200 DETAILED DESCRIPTION

[0036] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0037] In the description of the present invention, it should be understood that the terms "upper", "lower", "left", "right", "front", "back", "inside", "outside" and the like indicate directions or positional relationships based on the directions or relative positional relationships shown in the accompanying drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific direction, be constructed and operated in a specific direction, and therefore cannot be understood as a limitation on the present invention. Unless otherwise specified, the above-mentioned directional description can be flexibly set in the process of actual application under the condition that the relative positional relationship shown in the accompanying drawings is satisfied.

[0038] In the description of the present invention, it should be noted that, unless otherwise clearly specified and limited, the terms "installed", "connected", "connected", and "connected" should be understood in a broad sense, for example, it can be a fixed connection, a detachable connection, or an integral connection. It can be directly connected, or indirectly connected through an intermediate medium, or it can be the internal communication of two components. For ordinary technicians in this field, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.

[0039] In the embodiments of the present invention, the terms "comprises", "comprising" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, article or device including a series of elements includes not only those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such process, article or device. In the absence of further restrictions, an element defined by the sentence "comprising a ..." does not exclude the presence of other identical elements in the process, article or device including the element.

[0040] In the embodiments of the present invention, words such as "exemplary" or "for example" are used to indicate examples, illustrations or explanations. Any embodiment or design described as "exemplary" or "for example" in the embodiments of the present invention should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Specifically, the use of words such as "exemplary" or "for example" is intended to present related concepts in a specific way.

[0041] In the description of this specification, specific features, structures, materials or characteristics may be combined in an appropriate manner in any one or more embodiments or examples.

[0042] Object detection plays an important role in the field of autonomous driving. It can accurately identify and analyze the vehicle's surrounding environment information (other vehicles, pedestrians, obstacles, etc.), thereby providing users with a safe and reliable autonomous driving experience.

[0043] At present, the mainstream multimodal detection methods (based on point clouds and images) in the field of autonomous driving rely on deep learning. Although the effect is acceptable, their high computational complexity, large model size and huge training data requirements limit their application in real-time and resource-constrained scenarios. In addition, image features are extracted through attention mechanism and multi-scale feature fusion, and target detection and tracking are performed. Although this method reduces the risk of missed detection, it is insufficient to extract deep semantic information due to the use of a lightweight convolutional neural network model (Mobilenetv2) to extract features. Therefore, a multimodal hybrid unified 3D detection and tracking method is proposed, which fuses different sensor data into Bird's Eye View (BEV) features and integrates 3D target detection and tracking. This multimodal hybrid unified 3D detection and tracking method aims to improve the real-time, accuracy and robustness of target detection, but its Transformer method uses a multi-layer Transformer decoder, which is slow, and the deformable attention operation causes a large amount of random memory reads, which becomes a bottleneck for edge computing.

[0044] The target detection and environment perception in the above methods have the problems of high computational complexity, low computational efficiency and insufficient robustness.

[0045] In response to the above technical problems, the present application provides a vehicle control method, including: obtaining environmental perception data around the vehicle collected by multiple sensors, which is not limited by the line of sight of a single sensor, helping the vehicle to perceive the surrounding environment more comprehensively and improve the reliability of the vehicle control method. Extracting bird's-eye view features from the environmental perception data, the bird's-eye view features are generated from a bird's-eye view, so that the road environment traffic information is presented in a more intuitive way. Based on this, the road environment traffic information of the vehicle is determined more quickly and accurately. Based on the road environment traffic information, the vehicle control plan is determined, which improves the calculation efficiency and accuracy of the vehicle control method.

[0046] The embodiments provided in this application are described in detail below in conjunction with the accompanying drawings.

[0047] The vehicle control method provided in this application can be applied to Figure 1 In the vehicle control system shown. Figure 1 As shown, the vehicle control system includes: a sensor 100 and a processor 200 , wherein various sensors 100 are communicatively connected to the processor 200 .

[0048] In some embodiments, multiple sensors 100 are used to collect environmental perception data around the vehicle and send it to the processor 200.

[0049] Exemplarily, the various sensors 100 may include cameras, lidars, millimeter-wave radars, ultrasonic radars, and the like.

[0050] In some embodiments, various sensors 100 are mounted on a vehicle.

[0051] In some embodiments, the processor 200 is used to detect the surrounding environment of the vehicle. Exemplarily, the processor 200 responds to the environmental perception data sent by the various sensors 100, extracts the top-view features from the environmental perception data, thereby determining the road environment traffic information of the vehicle, and then determining the control scheme of the vehicle.

[0052] In some embodiments, the processor 200 may be a processor in a vehicle, such as a central processing unit (CPU), etc., which is not limited in the embodiments of the present application.

[0053] In some embodiments, the processor 200 may be a server, for example, a single server, or a server cluster consisting of multiple servers, or a cloud server, etc., which is not limited in the embodiments of the present application.

[0054] In some embodiments, the processor 200 may be communicatively connected to a vehicle controller, such that the processor 200 may send a control scheme for the vehicle determined based on the road environment traffic information of the vehicle to the vehicle controller so that the vehicle controller executes the control scheme.

[0055] It should be noted that the system architecture described in the embodiments of the present application is for more clearly illustrating the technical solutions of the embodiments of the present application, and does not constitute a limitation on the technical solutions provided in the embodiments of the present application. A person of ordinary skill in the art can know that with the evolution of the system architecture, the technical solutions provided in the embodiments of the present application are also applicable to similar technical problems.

[0056] See also Figure 2 , is a flow chart of a vehicle control method provided in an embodiment of the present application. Figure 2 As shown, the vehicle control method provided by the present application can be applied to Figure 1 The processor shown in the figure can be implemented by the above vehicle control system, which specifically includes the following steps S201 to S204:

[0057] S201. Acquire environmental perception data around the vehicle collected by multiple sensors.

[0058] Exemplarily, environmental perception data refers to perception data obtained by synchronizing and fusing information about the vehicle's surrounding environment collected by multiple sensors.

[0059] Exemplarily, data synchronization is used to align sensor data with different acquisition frequencies and timestamps. For example, data synchronization can use global positioning system timestamps to align multiple sensor data.

[0060] Exemplarily, data fusion is used to integrate data from multiple sensors. For example, data fusion can be implemented using a Kalman filter algorithm, a particle filter algorithm, or a deep learning model.

[0061] S202: Extract bird's-eye view features from environmental perception data.

[0062] Exemplarily, the bird's-eye view feature refers to environmental data obtained by mapping environmental perception data in a three-dimensional space to a two-dimensional plane. The two-dimensional plane is usually a plane from a bird's-eye view, that is, a plane looking down on the vehicle.

[0063] Exemplarily, the specific implementation process of step S202 refers to the following steps S2021 to S2024, which will not be described in detail here.

[0064] S203: Determine the road environment traffic information of the vehicle based on the bird's-eye view features.

[0065] Exemplarily, the road environment traffic information of the vehicle includes at least one of the following: lane lines, traffic signs, pedestrians, and vehicles.

[0066] For example, lane lines are used to indicate the information of each lane of a road and guide vehicles to drive within the specified range. For example, lane information includes: the number of lanes, lane width, lane type (driving, lane change, parking strip, etc.), and lane line status (clear, damaged, etc.).

[0067] Exemplarily, traffic signs are used to provide information such as roads and traffic rules to assist vehicles in safe driving. For example, traffic signs include speed limit signs, stop signs, yield signs, etc.

[0068] For example, pedestrians can help vehicles determine the trajectory of pedestrian activity and assist vehicles in safe driving. For example, the trajectory of pedestrian activity includes: the pedestrian's position, speed, direction, whether to cross the road, etc.

[0069] Exemplarily, the vehicle is another vehicle traveling around the vehicle, and the operating status of the other vehicle is determined to assist in determining the control scheme of the vehicle. For example, the operating status of the other vehicle includes: position, speed, acceleration, type, etc.

[0070] S204: Determine a vehicle control plan based on road environment traffic information.

[0071] Exemplarily, assuming that the road environment traffic information includes: the vehicle in front is decelerating, and there are oncoming vehicles in the left lane and / or right lane behind, based on traffic rules (type of lane lines and traffic signs) and vehicle behavior judgment (driving status of other vehicles), it is determined that the vehicle continues to drive in the current lane, that is, the vehicle's control scheme is to reduce acceleration and increase braking (that is, vehicle deceleration), and the steering angle remains unchanged (that is, maintaining the driving direction).

[0072] It is understandable that the vehicle control method provided by the embodiment of the present application, through the environmental perception data around the vehicle collected by multiple sensors, is not limited by the line of sight of a single sensor, which helps the vehicle to perceive the surrounding environment more comprehensively and improves the reliability of the vehicle control method. The vehicle control method provided by the present application extracts the top-down features from the environmental perception data and generates it from a top-down perspective, so that the road environment traffic information is displayed in a more intuitive way, and the road environment traffic information around the vehicle can be determined more quickly, and the calculation efficiency is improved, thereby improving the calculation efficiency and accuracy of the vehicle control method.

[0073] In some embodiments, the environmental perception data includes multi-view images, such as Figure 3 As shown, the above step S202 can be specifically implemented as the following steps S2021 to S2024:

[0074] S2021. Perform feature extraction on the multi-view image to obtain multi-view image features.

[0075] In some embodiments, the multi-view images are environment perception image data at different viewing angles, and can provide environment perception data at different angles, depths, and details.

[0076] In some embodiments, extracting features from multi-view images includes: performing image preprocessing on the multi-view images to obtain processed multi-view images; and performing feature detection on the processed multi-view images to obtain multi-view image features.

[0077] Exemplarily, the image preprocessing operation includes a denoising operation and an image enhancement operation.

[0078] Exemplarily, the denoising operation includes using a filtering technique (Gaussian filtering or mean filtering, etc.) to remove noise from the multi-view images.

[0079] Exemplarily, the image enhancement operation includes adjusting the contrast and brightness of the multi-view image to highlight the features of the multi-view image.

[0080] Exemplarily, feature detection includes feature point detection, feature description, feature matching and feature fusion.

[0081] Exemplarily, feature point detection uses a Scale-Invariant Feature Transform (SIFT) method or a Speeded Up Robust Features (SURF) method to detect feature points of multi-view images, and further, an edge detection algorithm is used to extract edge information of feature points of multi-view images.

[0082] Exemplarily, the feature description is based on edge information of the multi-view image feature points to describe local information of the multi-view image feature points, so as to find similar multi-view image features in the multi-view images.

[0083] Exemplarily, feature matching is used to match the same features of environmental perception data at different viewing angles in multi-view images.

[0084] For example, feature fusion is to fuse the features extracted from each view to form a multi-view image feature (F I ).

[0085] Exemplarily, the multi-view image feature (F I ) can satisfy the following formula (1):

[0086]

[0087] Among them, N C Used to indicate the number of views, H I Used to represent the height of the multi-view image, W I It is used to represent the width of the multi-view image, and C is used to represent the number of feature channels (edge, texture, color information, etc.).

[0088] S2022. For the image features of each viewing angle in the multi-view image features, calculate the position code of the image features of each viewing angle.

[0089] In some embodiments, the image feature of each viewing angle satisfies the following formula (2):

[0090] p i,j =[u i,j ,v i,j ] T Formula (2)

[0091] Among them, P i,j Used to represent the image features of each view, U i,j ,v i,j The coordinates of the image features in each view angle.

[0092] In some embodiments, the position encoding of the image feature at each viewing angle includes at least one of the following: the distance of the image feature relative to the vehicle on the top-down plane, the rotation angle of the image feature relative to the vehicle on the top-down plane, and the height of the image feature relative to the ground.

[0093] In some embodiments, the position encoding is used to obtain specific position information of image features in 3D space.

[0094] Exemplarily, the determination of the position code may refer to the following description and will not be elaborated here.

[0095] S2023. Determine width feature information of image features of each viewing angle.

[0096] Exemplarily, the width feature information is used to indicate the size of the area covered by the image feature in the horizontal direction.

[0097] Exemplarily, the width and feature information are determined by referring to the following steps Sc1 to Sc2, which will not be described in detail here.

[0098] S2024: Determine a bird's-eye view feature based at least on the position code of the image feature of each viewing angle and the width feature information of the image feature of each viewing angle.

[0099] As a possible implementation method, when the position encoding of the image feature of each perspective includes the height of the image feature relative to the ground, the height of the image feature of each perspective relative to the ground and the width feature information of the image feature of each perspective are weighted to obtain weighted height and width feature information; the weighted height and width feature information are input into the Transformer decoder to obtain the position information of the top-view feature. The specific implementation method refers to the following steps Sd1 to Sd2, which will not be repeated here.

[0100] As another possible implementation, determine the position information of the image feature of each perspective in the overhead perspective. Determine the overhead feature based on the position coding of the image feature of each perspective, the width feature information of the image feature of each perspective, and the position information of the image feature of each perspective in the overhead perspective. For the specific implementation method, refer to the following steps Se1 to Se2, which will not be described in detail here.

[0101] It is understandable that feature extraction of multi-view images can avoid large blind spots in a single image due to occlusion, improve the vehicle's environmental perception ability, and calculate position coding, which helps the vehicle understand the relative position of image features. At the same time, position coding can process image features under different viewing angles and reduce the impact of viewing angle changes on the vehicle's environmental perception ability. In addition, the use of width information features can more accurately define the shape of image features, so that the bird's-eye view features can more comprehensively and accurately reflect the information of the road traffic environment.

[0102] In some embodiments, the step S2022 of calculating the position coding of the image features of each viewing angle in the multi-view image features may be implemented as follows: calculating the position coding of the image features of each viewing angle based on a reference position coding method.

[0103] Exemplarily, the reference position encoding method includes: performing deep discretization and 3D projection on the image features of each perspective in the multi-view image features to obtain the 3D Cartesian coordinates of the image features of each perspective (with the vehicle as the origin); calculating the polar coordinates of the image features of each perspective based on the 3D Cartesian coordinates of the image features of each perspective; encoding the polar coordinates of the image features of each perspective based on Fourier position coding to obtain the position coding of each reference point of the image features; aggregating the position coding of multiple reference points of the image features to determine the position coding of the image features.

[0104] Exemplarily, depth discretization is used to associate the image features of each view with D discrete depth intervals to generate D reference points, and the homogeneous coordinates of the D reference points satisfy the following formula (3):

[0105] Where k∈|D| Formula (3)

[0106] Among them, d k Used to represent depth values. It is used to represent the image features promoted to the kth reference point through homogeneous coordinate transformation, u i,j ×d k ×d k The three-dimensional vector representing the k-th reference point.

[0107] Exemplarily, the 3D projection is used to project the reference point into a unified 3D space to obtain the 3D Cartesian coordinates (c i,j,k ) is expressed by the following formula (4):

[0108] C i,j,k =[x i,j,k ,y i,j,k , Z i,j,k ] T Formula (4)

[0109] Among them, x i,j,k ,y i,j,k , z i,j,k It is used to represent the coordinate position of the kth reference point in 3D space. T is used to represent the transpose operation, which is used to convert a row matrix into a column matrix.

[0110] For example, the 3D Cartesian coordinates of the reference point (c i,j,k ) can satisfy the following formula (5):

[0111]

[0112] in, Used to represent the intrinsic parameter matrix of multi-view images, The rotation matrix used to represent the nth view, The translation vector used to represent the nth view.

[0113] Exemplarily, the polar coordinates of the image features at each viewing angle satisfy the following formula (6):

[0114]

[0115] Among them, d i,j,k Used to represent the distance of the image feature relative to the vehicle on the top-view plane, θ i,j,k Used to represent the rotation angle of the image feature relative to the vehicle on the top-view plane, x i,j,k ,y i,j,k Used to represent the coordinate position of image features on the top-down plane.

[0116] Exemplarily, the Fourier position encoding (ξ) encodes the polar coordinates of each reference point of the image feature, and can satisfy the following formula (7):

[0117]

[0118] Among them, z i,j,k Used to represent the height of image features relative to the ground. Used to represent the position information of the image features of each view, Concat is used to combine multiple encoded information (the distance of the image feature relative to the vehicle on the top-down plane, the rotation angle of the image feature relative to the vehicle on the top-down plane, and the height of the image feature relative to the ground) to obtain the position information of the image features of each viewing angle.

[0119] Exemplarily, the position codes of multiple reference points of the image features are aggregated to obtain the position information of the image features of each viewing angle satisfying the following formula (8):

[0120]

[0121] Among them, s i,j,k Reference coefficients predicted by lightweight convolutional heads, MLP is used to represent multi-layer perceptrons, which are used to further extract image features. The position code used to represent each reference point.

[0122] For example, assume that the 3D Cartesian coordinates of the image feature are [x q ,y q , zq ] T , the position information of the image feature can be calculated based on the above formula (7) and formula (8) to obtain the following formula (9):

[0123] ψ q =MLP(Concat(ξ(d q ),ξ(sinθ q ),ξ(cosθ q ),ξ(z q ))) Formula (9)

[0124] Among them, z q Used to represent the height of the image feature relative to the ground, d q Represents the distance of the image feature relative to the vehicle on the top-view plane, θ q Used to represent the rotation angle of the image feature relative to the vehicle on the top-view plane.

[0125] For example, q 、sinθ q and cosθ q The above formula (6) can be used to calculate formula (10):

[0126]

[0127] Among them, x q and q Used to indicate the location of image features in the top-down plane.

[0128] In some embodiments, the above step S2023 can be specifically implemented as the following steps Sc1-Sc2:

[0129] Sc1. Perform high-level feature compression on the image features of each viewing angle to obtain compressed image features of each viewing angle.

[0130] Exemplarily, height feature compression is used to indicate compressing the height features of the image features of each perspective into a single row, and the height feature compression can be performed by using maximum pooling or average pooling.

[0131] Exemplarily, maximum pooling is used to divide the image features of each perspective into regions in the height direction (ie, the vertical direction), and select the maximum value in each region as the height feature value of the region.

[0132] Exemplarily, average pooling is used to divide the image features of each perspective into regions in the height direction (ie, the vertical direction), and select the average value in each region as the height feature value of the region.

[0133] Sc2. Extract width feature information from the compressed image features of each viewing angle.

[0134] In some embodiments, the above step Sc2 can be specifically implemented as the following steps:

[0135] Step 1: Concatenate the height feature information of the compressed image features of each viewing angle to obtain initial width feature information.

[0136] Exemplarily, the initial width feature information is obtained by concatenating the height feature information of each view, wherein the height feature information is obtained by performing maximum pooling calculation on the height dimension of the compressed image feature.

[0137] Step 2: Use the attention mechanism model to refine the initial width feature information to obtain width feature information.

[0138] Exemplarily, the attention mechanism model is a lightweight Transformer model, namely RefineTransformer.

[0139] Exemplarily, the attention mechanism model refines the initial width feature information by: aligning the width feature information between multiple views based on a self-attention operation; associating the width feature information of multiple perspectives based on a cross-attention operation, and restoring the details of the width feature information lost due to maximum pooling; obtaining processed width feature information, and inputting the processed width feature information into a feedforward network to obtain width feature information.

[0140] Exemplarily, the self-attention operation aligns the width feature information between different views by calculating the similarity between the width feature information of multiple views.

[0141] Exemplarily, the cross-attention operation recovers the width feature information lost by maximum pooling by associating the aligned width feature information with the image features.

[0142] For example, the feedforward network is a basic neural network structure in deep learning and is widely used to process various tasks, including feature extraction and classification.

[0143] For example, the attention mechanism model has a linear complexity with respect to the input image size. For an image feature size of (H, W), the computational complexity of the attention mechanism model satisfies the following formula (11):

[0144] Computational complexity = O(W 2 )+O(WH) Formula (11)

[0145] Among them, O(W 2) is used to represent the complexity of self-attention operation, and O(WH) is used to represent the complexity of cross-attention operation.

[0146] It can be understood that the width information obtained through the attention mechanism model can more accurately represent the width feature information of the image features, thereby improving the vehicle's environmental perception ability.

[0147] In some embodiments, when the position code of the image feature at each viewing angle includes the height of the image feature relative to the ground, the above step S2024 can be specifically implemented as the following steps Sd1 to Sd2:

[0148] Sd1. Weight the height of the image feature of each viewing angle relative to the ground and the width feature information of the image feature of each viewing angle to obtain weighted height and width feature information.

[0149] Exemplarily, the weighted height and width feature information is determined based on a deep learning model, which is trained in advance and can determine the weighted height and width feature information based on the height and width feature information of the image features of each perspective.

[0150] Sd2: Input the weighted height and width feature information into the Transformer decoder to obtain the position information of the bird’s-eye view feature.

[0151] It should be noted that the Transformer decoder can effectively fuse the height and width feature information of image features to obtain the geometric shape and physical size of image features.

[0152] It can be understood that by determining the position information of the bird's-eye view feature based on the weighted height and width feature information, more critical features of the image features can be determined, which facilitates the identification of road traffic information and improves the vehicle's environmental perception capabilities.

[0153] In some embodiments, the above step S2024 can also be implemented as the following steps Se1-Se2:

[0154] Se1. Determine the position information of the image features of each perspective in the bird's-eye view.

[0155] In some embodiments, determining the position information of the bird's-eye view feature includes: taking the position of the vehicle as the origin, dividing the vehicle's surrounding environment into a plurality of regular grid units from a bird's-eye view perspective above the vehicle, and obtaining the center coordinates of each grid unit; based on the center coordinates of each grid unit, determining the position information of the image feature of each perspective at the bird's-eye view perspective.

[0156] For example, assume that the center coordinates of each grid are [x, y] T, the position information of the image features of each perspective on the top-down plane (q B ) can satisfy the following formula (12):

[0157] q B =MLP(Concat(ξ(d),ξ(sinθ),ξ(cosθ))) Formula (12)

[0158] Wherein, d is used to represent the position of the center coordinate on the top-view plane, and θ is used to represent the rotation angle of the center coordinate on the top-view plane relative to the vehicle.

[0159] Exemplarily, d, sinθ and cosθ may satisfy the following formula (13):

[0160]

[0161] Among them, x and y are used to represent the position of the center coordinates.

[0162] Se2. Determine the bird's-eye view feature based on the position coding of the image feature of each viewing angle, the width feature information of the image feature of each viewing angle, and the position information of the image feature of each viewing angle in the bird's-eye view.

[0163] Exemplarily, when the position encoding of the image feature of each perspective includes the height of the image feature relative to the ground, the height of the image feature of each perspective relative to the ground and the width feature information of the image feature of each perspective are weighted to obtain weighted height and width feature information; the weighted height and width feature information, the width feature information of the image feature of each perspective, and the position information of the image feature of each perspective in a bird's-eye view are input into the Transformer decoder to obtain the position information of the bird's-eye view feature.

[0164] Exemplarily, the weighted height and width feature information may be determined based on the above step Se2, which will not be elaborated herein.

[0165] Exemplarily, the width feature information of the image features of each view is used as the key vector of the Transformer decoder.

[0166] For example, in the attention mechanism of Transformer, the key vector is used to compare with the query vector to determine whether the width feature information should be weighted and participate in the calculation of the final position information.

[0167] Exemplarily, the weighted height and width feature information is used as the value vector of the Transformer decoder.

[0168] For example, the value vector contains detailed information associated with the key vector, which will be weighted and used to generate the final output.

[0169] Exemplarily, the position information of the image features of each perspective in the top-down perspective is used as the query vector of the Transformer decoder.

[0170] Illustratively, the query vector represents a location or region of interest in the BEV space and is used to retrieve information from the value vector associated with the key vector.

[0171] Exemplarily, the Transformer decoder learns the relationship between the weighted height and width feature information, the width feature information of the image features of each perspective, and the position information of the image features of each perspective in the overhead view through the attention model, and determines the position information of the overhead feature based on the relationship between the three.

[0172] It can be understood that based on the position encoding of the image features of each perspective, the width feature information of the image features of each perspective, and the position information of the image features of each perspective in the bird's-eye view, the bird's-eye view features facilitate determining the spatial position information of the image features, thereby improving the accuracy of road traffic information and thereby improving the safety of autonomous driving.

[0173] In some embodiments, Figure 4 As shown, the above step S203 can be specifically implemented as the following steps S2031 to S2033:

[0174] S2031. Construct a bird's-eye view environment plan based on the bird's-eye view features.

[0175] Exemplarily, the top view of the environment plan is a BEV feature map, which is a feature map that projects information in a three-dimensional space onto a two-dimensional plane and is represented in the form of a grid. Each grid represents a certain spatial area and contains feature information in the area.

[0176] S2032: Extract features of each target object from the overhead environment plan view.

[0177] Exemplarily, based on the feature extraction model, features of each target object are extracted from the overhead environment plan view.

[0178] Exemplarily, the feature extraction model is a trained model, and similar features are extracted as features of each target object by comparing the overhead environment plan view with feature images in the training samples.

[0179] S2033: Using a target detection algorithm, based on the characteristics of each target object, identify the road traffic information of the vehicle.

[0180] Exemplarily, the target detection algorithm may be a deep learning target detection model (BEV Detection, BEVDet). For example, the features of each target object are input into the BEVDet model, and the BEVDet model recognizes the features and determines the road traffic information of the vehicle.

[0181] It can be understood that the use of a bird's-eye view of the environment plan can more intuitively display the characteristics of the target object, facilitate the identification of the vehicle's road traffic information, and improve computing efficiency and driving safety.

[0182] In some embodiments, the above step S204 can be specifically implemented as the following steps S2041-S2042:

[0183] S2041. Plan a vehicle's driving route based on road traffic information.

[0184] Exemplarily, it is assumed that the road environment traffic information includes: the vehicle ahead is decelerating, there are no vehicles behind the left lane and / or the right lane, and the vehicle's driving route is planned based on traffic rules (types of lane lines and traffic signs).

[0185] For example, assuming that the current lane line allows lane change, and the current vehicle is on a straight road with no intersection ahead, the planned vehicle route is to change lanes to the left or right.

[0186] S2042: Determine a control plan for the vehicle based on the vehicle's driving route.

[0187] Among them, the control scheme is used to adjust the vehicle's handling parameters so that the vehicle can travel according to the route.

[0188] Exemplarily, the operating parameters of the vehicle include at least one of the following: steering angle, acceleration, braking, etc.

[0189] Exemplarily, the steering angle is used to indicate the turning angle of the front wheels of the vehicle, that is, the driving direction of the vehicle.

[0190] For example, acceleration is used to indicate the rate of change of vehicle speed and to control the acceleration or deceleration of the vehicle.

[0191] Exemplarily, braking is used to refer to the process of reducing the speed of a vehicle or stopping the vehicle.

[0192] For example, assuming that the vehicle's driving route is to change lanes to the left or right, that is, the vehicle's control scheme is to control the vehicle to change lanes to the left or right, that is, reducing acceleration and increasing braking (that is, vehicle deceleration), and changing the steering angle (that is, driving direction to the left or right).

[0193] It is understandable that determining the vehicle control scheme based on road traffic information can effectively reduce the occurrence of traffic accidents, reduce the economic losses of users, and improve the safety of vehicle driving.

[0194] See also Figure 5 , which is a specific implementation method for extracting top-view features provided in this application.

[0195] a1. Perform feature extraction on multi-view images to obtain multi-view image features.

[0196] a2. Based on the reference position encoding method, calculate the position encoding of the image features of each perspective.

[0197] Exemplarily, the distance of the image feature relative to the vehicle on the top-view plane, the rotation angle of the image feature relative to the vehicle on the top-view plane, and the height of the image feature relative to the ground.

[0198] a3. Concatenate the height feature information of the compressed image features of each viewing angle to obtain initial width feature information.

[0199] a4. Use the attention mechanism model to refine the initial width feature information to obtain width feature information.

[0200] a5. Weighting the height of the image feature of each viewing angle relative to the ground and the width feature information of the image feature of each viewing angle to obtain weighted height and width feature information.

[0201] a6. Determine the position information of the image features of each perspective in the bird's-eye view.

[0202] a7. Input the weighted height and width feature information, the width feature information of the image features of each perspective, and the position information of the image features of each perspective in the bird's-eye view into the Transformer decoder to obtain the bird's-eye view features.

[0203] For example, Figure 6 A schematic diagram of the composition of a vehicle control device provided in an embodiment of the present application. Figure 6 As shown, the vehicle control device 800 includes: a communication module 801 and a processing module 802. The communication module 801 is used to obtain environmental perception data around the vehicle collected by various sensors. The processing module 802 is used to extract bird's-eye view features from the environmental perception data; based on the bird's-eye view features, determine the road environment traffic information of the vehicle; based on the road environment traffic information, determine the vehicle control plan.

[0204] In some embodiments, the environmental perception data includes multi-view images, and the processing module 802 is specifically used to extract features from the multi-view images to obtain multi-view image features; for the image features of each perspective in the multi-view image features, calculate the position code of the image features of each perspective; determine the width feature information of the image features of each perspective; and determine the bird's-eye view features based at least on the position code of the image features of each perspective and the width feature information of the image features of each perspective.

[0205] In some embodiments, the processing module 802 is specifically used to perform height feature compression on the image features of each viewing angle to obtain compressed image features of each viewing angle; and extract width feature information from the compressed image features of each viewing angle.

[0206] In some embodiments, the processing module 802 is specifically used to concatenate the height feature information of the compressed image features of each perspective to obtain initial width feature information; and use the attention mechanism model to refine the initial width feature information to obtain width feature information.

[0207] In some embodiments, the attention mechanism model is a lightweight Transformer model.

[0208] In some embodiments, the processing module 802 is specifically configured to calculate the position coding of the image features of each viewing angle based on a reference position coding method.

[0209] In some embodiments, the position encoding of the image features of each viewing angle includes at least one of the following:

[0210] The distance of the image feature relative to the vehicle on the top-down plane, the rotation angle of the image feature relative to the vehicle on the top-down plane, and the height of the image feature relative to the ground.

[0211] In some embodiments, when the position encoding of the image feature of each perspective includes the height of the image feature relative to the ground, the processing module 802 is specifically used to weight the height of the image feature of each perspective relative to the ground and the width feature information of the image feature of each perspective to obtain weighted height and width feature information; the weighted height and width feature information is input into the Transformer decoder to obtain the bird's-eye view feature.

[0212] In some embodiments, the processing module 802 is specifically used to determine the position information of the image features of each perspective at the overhead perspective; the overhead features are determined based on the position coding of the image features of each perspective, the width feature information of the image features of each perspective, and the position information of the image features of each perspective at the overhead perspective.

[0213] In some embodiments, when the position encoding of the image feature of each perspective includes the height of the image feature relative to the ground, the processing module 802 is specifically used to weight the height of the image feature of each perspective relative to the ground and the width feature information of the image feature of each perspective to obtain weighted height and width feature information; the weighted height and width feature information, the width feature information of the image feature of each perspective, and the position information of the image feature of each perspective at a bird's-eye view are input into the Transformer decoder to obtain the position information of the bird's-eye view feature.

[0214] In some embodiments, the width feature information of the image features of each perspective is used as the key vector of the Transformer decoder, the weighted height and width feature information is used as the value vector of the Transformer decoder, and the position information of the image features of each perspective in the top-down perspective is used as the query vector of the Transformer decoder.

[0215] In some embodiments, the processing module 802 is specifically used to divide the vehicle's surrounding environment into multiple regular grid units from a bird's-eye view from above the vehicle, taking the vehicle's position as the origin, and obtain the center coordinates of each grid unit; based on the center coordinates of each grid unit, determine the position information of the image features of each perspective in the bird's-eye view.

[0216] In some embodiments, the processing module 802 is specifically used to construct a bird's-eye view environment plan view based on bird's-eye view features; extract features of each target object from the bird's-eye view environment plan view; and use a target detection algorithm to identify the vehicle's road environment traffic information based on the features of each target object.

[0217] In some embodiments, the road environment traffic information of the vehicle includes at least one of the following: lane lines, traffic signs, pedestrians, and vehicles.

[0218] In some embodiments, the processing module 802 is specifically used to plan the vehicle's driving route based on road environment traffic information; determine the vehicle's control scheme based on the vehicle's driving route; and the control scheme is used to adjust the vehicle's operating parameters so that the vehicle travels along the driving route.

[0219] In the case of implementing the functions of the above-mentioned integrated modules in the form of hardware, the embodiment of the present invention provides a possible structural diagram of the electronic device involved in the above-mentioned embodiment. Figure 7 As shown, the electronic device 900 includes: a processor 902 , a communication interface 903 , and a bus 904 . Optionally, the electronic device 900 may further include a memory 901 .

[0220] The processor 902 may be a processor that implements or executes various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure of the present application. The processor 902 may be a central processing unit, a general-purpose processor, a digital signal processor, an application-specific integrated circuit, a field programmable gate array, or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It may implement or execute various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure of the present application. The processor 902 may also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, and the like.

[0221] The communication interface 903 is used to connect with other devices via a communication network, such as Ethernet, wireless access network, wireless local area network (WLAN), etc.

[0222] The memory 901 may be a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM) or other types of dynamic storage devices that can store information and instructions, or an electrically erasable programmable read-only memory (EEPROM), a disk storage medium or other magnetic storage device, or any other medium that can be used to carry or store desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto.

[0223] As a possible implementation, the memory 901 may exist independently of the processor 902, and the memory 901 may be connected to the processor 902 via a bus 904 to store instructions or program codes. When the processor 902 calls and executes the instructions or program codes stored in the memory 901, the vehicle driving method provided in the embodiment of the present invention can be implemented.

[0224] In another possible implementation, the memory 901 may also be integrated with the processor 902 .

[0225] The bus 904 may be an extended industry standard architecture (EISA) bus, etc. The bus 904 may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 7 Only one thick line is used in the diagram, but this does not mean that there is only one bus or only one type of bus.

[0226] Through the description of the above implementation methods, technical personnel in the relevant field can clearly understand that for the convenience and simplicity of description, only the division of the above-mentioned functional modules is used as an example. In actual applications, the above-mentioned functions can be assigned to different functional modules as needed, that is, the internal structure of the service calling device can be divided into different functional modules to complete all or part of the functions described above.

[0227] The embodiment of the present application also provides a computer-readable storage medium. All or part of the processes in the above method embodiments can be completed by computer instructions to instruct the relevant hardware, and the program can be stored in the above computer-readable storage medium. When the program is executed, it may include the processes of the above method embodiments. The computer-readable storage medium can be the memory or memory of any of the above embodiments. The above computer-readable storage medium can also be an external storage device of the above service calling device, such as a plug-in hard disk, a smart memory card (smart media card, SMC), a secure digital (secure digital, SD) card, a flash card (flash card), etc. equipped on the above service calling device. Further, the above computer-readable storage medium can also include both the internal storage unit of the above service calling device and an external storage device. The above computer-readable storage medium is used to store the above computer program and other programs and data required by the above service calling device. The above computer-readable storage medium can also be used to temporarily store data that has been output or is to be output.

[0228] An embodiment of the present application also provides a computer program product, which includes a computer program. When the computer program product is run on a computer, the computer is enabled to execute any one of the vehicle control methods provided in the above embodiments.

[0229] The above are only specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art who is familiar with the technical field can easily think of changes or substitutions within the technical scope disclosed by the present invention, which should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention should be based on the protection scope of the claims.

Claims

1. A vehicle control method, characterized in that: The method comprises: Obtain environmental perception data around the vehicle collected by multiple sensors; Extracting bird's-eye view features from the environmental perception data; Determining the road environment traffic information of the vehicle based on the bird's-eye view features; Based on the road environment traffic information, a control scheme for the vehicle is determined.

2. The method according to claim 1, characterized in that The environmental perception data includes a multi-view image; and extracting the bird's-eye view feature from the environmental perception data includes: Extracting features from the multi-view images to obtain multi-view image features; For the image features of each viewing angle in the multi-view image features, calculating the position code of the image features of each viewing angle; Determining width feature information of the image features of each viewing angle; The bird's-eye view feature is determined based at least on the position code of the image feature of each viewing angle and the width feature information of the image feature of each viewing angle.

3. The method according to claim 2, characterized in that The determining of the width feature information of the image feature of each viewing angle includes: Performing high-feature compression on the image features of each viewing angle to obtain compressed image features of each viewing angle; The width feature information is extracted from the compressed image features of each viewing angle.

4. The method according to claim 3, characterized in that The extracting the width feature information from the compressed image features of each viewing angle includes: Concatenating the height feature information of the compressed image features of each viewing angle to obtain initial width feature information; The initial width feature information is refined by using an attention mechanism model to obtain the width feature information.

5. The method according to claim 4, characterized in that The attention mechanism model is a lightweight Transformer model.

6. The method according to claim 2, characterized in that The calculating the position encoding of the image features of each viewing angle comprises: Based on the reference position encoding method, the position encoding of the image features of each viewing angle is calculated.

7. The method according to claim 6, characterized in that The position encoding of the image features of each viewing angle includes at least one of the following: The distance of the image feature relative to the vehicle on the top-view plane, the rotation angle of the image feature relative to the vehicle on the top-view plane, and the height of the image feature relative to the ground.

8. The method according to claim 2, characterized in that: In a case where the position code of the image feature of each viewing angle includes the height of the image feature relative to the ground, determining the bird's-eye view feature based on the position code of the image feature of each viewing angle and the width feature information of the image feature of each viewing angle includes: Weighting the height of the image feature of each viewing angle relative to the ground and the width feature information of the image feature of each viewing angle to obtain weighted height and width feature information; The weighted height and width feature information is input into the Transformer decoder to obtain the position information of the top-view feature.

9. The method according to claim 2, characterized in that: The determining the top-view feature based at least on the position coding of the image feature of each viewing angle and the width feature information of the image feature of each viewing angle comprises: Determine the position information of the image features of each viewing angle at the top-down viewing angle; The bird's-eye view feature is determined based on the position code of the image feature of each viewing angle, the width feature information of the image feature of each viewing angle, and the position information of the image feature of each viewing angle at the bird's-eye view.

10. The method according to claim 9, characterized in that In a case where the position code of the image feature at each viewing angle includes the height of the image feature relative to the ground, determining the overlooking feature based on the position code of the image feature at each viewing angle, width feature information of the image feature at each viewing angle, and position information of the image feature at each viewing angle at a overlooking viewing angle includes: Weighting the height of the image feature of each viewing angle relative to the ground and the width feature information of the image feature of each viewing angle to obtain weighted height and width feature information; The weighted height and width feature information, the width feature information of the image feature of each viewing angle, and the position information of the image feature of each viewing angle at a bird's-eye view are input into a Transformer decoder to obtain the position information of the bird's-eye view feature.

11. The method according to claim 10, characterized in that The width feature information of the image features of each viewing angle is used as the key vector of the Transformer decoder, the weighted height and width feature information is used as the value vector of the Transformer decoder, and the position information of the image features of each viewing angle in the top-down viewing angle is used as the query vector of the Transformer decoder.

12. The method according to claim 9, characterized in that The determining the position information of the image feature of each viewing angle at the top-down viewing angle includes: Taking the position of the vehicle as the origin, dividing the surrounding environment of the vehicle into a plurality of regular grid units from a bird's-eye view above the vehicle, and obtaining the center coordinates of each grid unit; Based on the central coordinates of each grid unit, position information of the image feature of each viewing angle in the bird's-eye view is determined.

13. The method according to claim 1, characterized in that The determining the road environment traffic information of the vehicle based on the top-view feature includes: Constructing a bird's-eye view environment plan view based on the bird's-eye view features; Extracting features of each target object from the overhead environmental plan view; A target detection algorithm is used to identify the road traffic information of the vehicle based on the features of each target object.

14. The method according to claim 1, characterized in that The road environment traffic information of the vehicle includes at least one of the following: Lane lines, traffic signs, pedestrians, and vehicles.

15. The method according to any one of claims 1 to 14, characterized in that The determining of the control scheme of the vehicle based on the road environment traffic information includes: Planning a driving route of the vehicle based on the road traffic information; Based on the driving route of the vehicle, a control scheme of the vehicle is determined; the control scheme is used to adjust the operating parameters of the vehicle so that the vehicle travels according to the driving route.

16. An electronic device, characterized in that: The invention comprises a processor and a memory, wherein the processor is coupled to the memory; the memory is used to store computer instructions, and the computer instructions are loaded and executed by the processor so that the computer device implements the vehicle control method as described in any one of claims 1 to 15.

17. A vehicle, characterized in that: Comprising the electronic device as claimed in claim 16.

18. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes computer-executable instructions, and when the computer-executable instructions are executed on a computer, the computer is caused to execute the vehicle control method according to any one of claims 1 to 15.

19. A computer program product, characterized in that The computer program product includes a computer program, and when the computer program is run on an electronic device, the electronic device executes the vehicle control method according to any one of claims 1 to 15.