Vehicle control method and device and computer readable storage medium
By combining surround-view cameras and front-view cameras with a deep neural network model, the vehicle can detect objects in the environment at both near and far distances. This solves the problem of uniformity in perception tasks in existing technologies, reduces waste of hardware resources, and improves the environmental perception capabilities of autonomous driving.
Patent Information
- Application Number
- CN202410563877.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-08
- Publication Date
- 2025-11-11
AI Technical Summary
Existing visual perception technologies lack a unified approach to near-field and long-field perception in autonomous driving, resulting in wasted camera hardware resources and a lack of correlation between perception tasks.
By combining surround-view cameras and front-view cameras, the surround-view cameras acquire feature images from a bird's-eye view to achieve close-range detection, while the front-view cameras acquire feature images from a non-bird's-eye view to achieve long-range detection. Image processing and target detection are then performed using a deep neural network model to generate complete environmental perception results.
It enables vehicles to detect environmental objects at both close and long distances, reducing the waste of camera hardware resources, lowering the overall vehicle cost, and improving the comprehensiveness and accuracy of environmental perception.
Smart Images

Figure CN120932189A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of vehicle technology, and more specifically, to a vehicle control method, apparatus, and computer-readable storage medium in the field of vehicle technology. Background Technology
[0002] Autonomous driving technology is a crucial development direction for the automotive industry now and in the future, and visual perception has seen rapid development and attracted increasing attention from experts and scholars in recent years due to its low hardware cost. In autonomous driving environmental perception, lane detection, obstacle detection, and traffic light detection are fundamental and extremely important perception tasks, playing an irreplaceable role in realizing autonomous driving or assisted driving functions.
[0003] Current visual perception technologies mainly output perception tasks in a single-task manner, such as only outputting lane line detection perception tasks, obstacle detection perception tasks, or traffic light detection perception tasks. The perception tasks are not unified and lack correlation. This not only fails to achieve both near and far-field perspective perception, but also results in a significant waste of camera hardware resources. Summary of the Invention
[0004] This application provides a vehicle control method, apparatus, and computer-readable storage medium, which can not only realize the detection of environmental objects at close and long distances, but also avoid the waste of camera hardware resources.
[0005] A first aspect provides a vehicle control method, comprising: acquiring a first feature image from a bird's-eye view and a second feature image from a non-bird's-eye view; wherein the first feature image is obtained by acquiring a surround-view environment image captured by a surround-view camera deployed on the vehicle, and the second feature image is obtained by acquiring a forward-view environment image captured by a forward-view camera deployed on the vehicle, the surround-view camera including the forward-view camera, the detection range of the surround-view camera being smaller than the detection range of the forward-view camera; detecting a first target object in a first target region based on the first feature image to obtain a first detection result; wherein the first target object includes lane lines and / or a first obstacle, and the first target region is an environmental region corresponding to the viewpoint of the surround-view camera; detecting a second target object in a second target region based on the second feature image to obtain a second detection result; wherein the second target object includes a traffic indicator and / or a second obstacle, and the second target region is an environmental region corresponding to the viewpoint of the forward-view camera; and controlling the vehicle according to the first detection result or the second detection result.
[0006] In the above technical solution, the vehicle control method provided in this application embodiment can detect near-range environmental objects by using the surround-view camera to collect surround-view environmental images, and can detect far-range environmental objects by using the forward-view camera to collect forward-view environmental images. This allows the vehicle to have both near-range and far-range environmental object detection capabilities, thereby expanding the vehicle's detection range. Furthermore, since the forward-view camera for collecting forward-view environmental images is included within the surround-view camera, the forward-view environmental images can be collected through the forward-view camera within the surround-view camera, eliminating the need for an additional forward-view camera. This avoids wasting camera hardware resources and helps save on overall vehicle costs.
[0007] In conjunction with the first aspect, in some possible implementations, the surround view environment image includes target region images corresponding to multiple viewpoints; obtaining the first feature image includes: performing a first processing on the target region images corresponding to the multiple viewpoints to obtain preprocessed images corresponding to the multiple viewpoints; wherein, the first processing includes at least one of virtual camera processing, normalization processing, resizing processing, denoising processing, and enhancement processing; performing two second processings on the preprocessed images corresponding to the multiple viewpoints to obtain two frames of first downsampled images at a first scale and two frames of second downsampled images at a second scale corresponding to the multiple preprocessed images; wherein, the second processing includes feature extraction and downsampling, and the first scale is different from the second scale; the two frames of first downsampled images corresponding to the multiple preprocessed images... The process involves fusing the preprocessed images to obtain a first fused image corresponding to each of the multiple preprocessed images; fusing two frames of second downsampled images corresponding to each of the multiple preprocessed images to obtain a second fused image corresponding to each of the multiple preprocessed images; performing local spatial transformation on the first fused images corresponding to each of the multiple preprocessed images to obtain a first local spatial bird's-eye view corresponding to each of the multiple viewpoints from the bird's-eye view perspective; performing local spatial transformation on the second fused images corresponding to each of the multiple preprocessed images to obtain a second local spatial bird's-eye view corresponding to each of the multiple viewpoints from the bird's-eye view perspective; wherein the scale of the first local spatial bird's-eye view is different from the scale of the second local spatial bird's-eye view; and generating the first feature image based on the first local spatial bird's-eye view and the second local spatial bird's-eye view corresponding to each of the multiple viewpoints.
[0008] In combination with the first aspect and the above implementation, in some possible implementations, the step of performing two second processing steps on the preprocessed images corresponding to each of the plurality of viewpoints to obtain two frames of first downsampled images at the first scale and two frames of second downsampled images at the second scale corresponding to each of the plurality of preprocessed images includes:
[0009] For each viewpoint, the preprocessed image is input into two pre-trained shared networks to obtain the output results of the two shared networks respectively.
[0010] The output results of the two shared networks each include a first downsampled image and a second downsampled image. Both shared networks include a deep neural network model for feature extraction and a deep learning model for object detection.
[0011] In combination with the first aspect and the above implementation methods, in some possible implementation methods, generating the first feature image based on the first local spatial bird's-eye view and the second local spatial bird's-eye view corresponding to each of the multiple perspectives includes: stitching the first local spatial bird's-eye view and the second local spatial bird's-eye view corresponding to each of the multiple perspectives along the channel dimension to obtain the stitched local spatial bird's-eye view corresponding to each of the multiple perspectives; summing the stitched local spatial bird's-eye view corresponding to each of the multiple perspectives to obtain the first feature image.
[0012] In combination with the first aspect and the above implementation methods, in some possible implementation methods, obtaining the second feature image includes: determining the target view of the forward-looking camera among the plurality of viewpoints; fusing the first fusion image and the second fusion image corresponding to the target view to obtain a forward-looking fusion image; and upsampling the forward-looking fusion image to obtain the second feature image.
[0013] In combination with the first aspect and the above implementation methods, in some possible implementation methods, the step of detecting the first target object in the first target region based on the first feature image to obtain the first detection result includes: cropping the first feature image to obtain a cropped image containing the first target object; upsampling the cropped image to obtain a first upsampled image; and detecting the first target object in the first target region based on the first upsampled image to obtain the first detection result.
[0014] In conjunction with the first aspect and the above implementations, in some possible implementations, the first upsampled image includes a sampled image of the lane line and / or a sampled image of the first obstacle; the step of detecting the first target object in the first target region based on the first upsampled image to obtain the first detection result includes: inputting the sampled image of the lane line into a pre-trained lane line detection head network to obtain the detection result of the lane line, and determining the detection result of the lane line as the first detection result; wherein, the detection result of the lane line includes at least one of lane line segmentation, lane line embedding, lane line color, lane line type, and lane line function; and / or, inputting the sampled image of the first obstacle into a pre-trained first obstacle detection head network to obtain the detection result of the first obstacle, and determining the detection result of the first obstacle as the first detection result; wherein, the detection result of the first obstacle includes at least one of heatmap, offset, dimension annotation, and rotation map.
[0015] In combination with the first aspect and the above implementation methods, in some possible implementation methods, the step of detecting the second target object in the second target region based on the second feature image to obtain the second detection result includes: upsampling the second feature image to obtain a second upsampled image; and detecting the second target object in the second target region based on the second upsampled image to obtain the second detection result.
[0016] In conjunction with the first aspect and the above implementation methods, in some possible implementation methods, the second upsampled image includes a sampled image of the traffic sign and a sampled image of the second obstacle; the step of detecting the second target object in the second target area based on the second upsampled image to obtain the second detection result includes: inputting the sampled image of the traffic sign into a pre-trained traffic sign detection head network to obtain the detection result of the traffic sign, and determining the detection result of the traffic sign as the second detection result; and / or, inputting the sampled image of the second obstacle into a pre-trained second obstacle detection head network to obtain the detection result of the second obstacle, and determining the detection result of the second obstacle as the second detection result.
[0017] In conjunction with the first aspect and the above implementation methods, in some possible implementation methods, if the first target object includes the lane line and the first obstacle, and the second target object includes the traffic indicator and the second obstacle, controlling the vehicle based on the first detection result or the second detection result includes: if the detection result of the first obstacle in the first detection result is the same as the detection result of the second obstacle in the second detection result, and the second obstacle is located within the detection range of the surround-view camera, controlling the vehicle based on the first detection result; if the detection result of the first obstacle in the first detection result is a part of the detection result of the second obstacle in the second detection result, and the first obstacle is located within the detection range of the surround-view camera, controlling the vehicle based on the second detection result.
[0018] Secondly, a vehicle control device is provided, the vehicle control device comprising:
[0019] An image acquisition module is used to acquire a first feature image from a bird's-eye view and a second feature image from a non-bird's-eye view; wherein, the first feature image is obtained by acquiring a surround-view environment image from a surround-view camera arranged in the vehicle, and the second feature image is obtained by acquiring a forward-view environment image from a forward-view camera arranged in the vehicle, wherein the surround-view camera includes the forward-view camera, and the detection range of the surround-view camera is smaller than the detection range of the forward-view camera.
[0020] The first detection module is used to detect a first target object in the first target region based on the first feature image to obtain a first detection result; wherein the first target object includes lane lines and / or a first obstacle, and the first target region is the environmental region corresponding to the viewpoint of the surround-view camera;
[0021] The second detection module is used to detect a second target object in the second target region based on the second feature image to obtain a second detection result; wherein the second target object includes a traffic sign and / or a second obstacle, and the second target region is the environmental region corresponding to the viewpoint of the forward-looking camera;
[0022] The vehicle control module is used to control the vehicle based on the first detection result or the second detection result.
[0023] In conjunction with the second aspect, in some possible implementations, the image acquisition module described above includes:
[0024] The first acquisition unit is configured to perform a first processing on the target region images corresponding to each of the multiple viewpoints to obtain preprocessed images corresponding to each of the multiple viewpoints; wherein the first processing includes at least one of virtual camera processing, normalization processing, resizing processing, denoising processing, and enhancement processing; perform two second processing on the preprocessed images corresponding to each of the multiple viewpoints to obtain two frames of first downsampled images at a first scale and two frames of second downsampled images at a second scale corresponding to each of the multiple preprocessed images; wherein the second processing includes feature extraction and downsampling, and the first scale is different from the second scale; and fuse the two frames of first downsampled images corresponding to each of the multiple preprocessed images to obtain a first fused image corresponding to each of the multiple preprocessed images. The image is processed as follows: Two frames of second downsampled images corresponding to each of the plurality of preprocessed images are fused to obtain a second fused image corresponding to each of the plurality of preprocessed images; a local spatial transformation is performed on the first fused image corresponding to each of the plurality of preprocessed images to obtain a first local spatial bird's-eye view corresponding to each of the plurality of viewpoints under the bird's-eye view; a local spatial transformation is performed on the second fused image corresponding to each of the plurality of preprocessed images to obtain a second local spatial bird's-eye view corresponding to each of the plurality of viewpoints under the bird's-eye view; wherein the scale of the first local spatial bird's-eye view is different from the scale of the second local spatial bird's-eye view; the first feature image is generated based on the first local spatial bird's-eye view and the second local spatial bird's-eye view corresponding to each of the plurality of viewpoints.
[0025] In conjunction with the second aspect and the above implementation, in some possible implementations, the first acquisition unit, in performing two second processing steps on the preprocessed images corresponding to each of the multiple viewpoints to obtain two frames of first downsampled images at the first scale and two frames of second downsampled images at the second scale corresponding to each of the multiple preprocessed images, is used to input the preprocessed images corresponding to each viewpoint into two pre-trained shared networks respectively to obtain the output results corresponding to each of the two shared networks; wherein, the output results corresponding to each of the two shared networks include one frame of first downsampled image and one frame of second downsampled image, and both shared networks include a deep neural network model for feature extraction and a deep learning model for object detection.
[0026] In combination with the second aspect and the above implementation, in some possible implementations, the first acquisition unit, in generating the first feature image based on the first local spatial bird's-eye view and the second local spatial bird's-eye view corresponding to each of the multiple perspectives, is used to stitch together the first local spatial bird's-eye view and the second local spatial bird's-eye view corresponding to each of the multiple perspectives along the channel dimension to obtain the stitched local spatial bird's-eye view corresponding to each of the multiple perspectives; and to sum up the stitched local spatial bird's-eye view corresponding to each of the multiple perspectives to obtain the first feature image.
[0027] In combination with the second aspect and the above implementation methods, in some possible implementations, the image acquisition module includes:
[0028] The second acquisition unit is used to determine the target view of the forward-looking camera among the multiple viewpoints; fuse the first fusion image and the second fusion image corresponding to the target viewpoint to obtain a forward-looking fusion image; and upsample the forward-looking fusion image to obtain the second feature image.
[0029] In combination with the second aspect and the above implementation methods, in some possible implementation methods, the first detection module includes:
[0030] The cropping unit is used to crop the first feature image to obtain a cropped image containing the first target object;
[0031] The first sampling unit is used to upsample the cropped image to obtain a first upsampled image;
[0032] The first detection unit is configured to detect the first target object in the first target region based on the first upsampled image, and obtain the first detection result.
[0033] In conjunction with the second aspect and the above implementation methods, in some possible implementation methods, the first upsampled image includes a sampled image of the lane line and / or a sampled image of the first obstacle. Specifically, the first detection unit is used to input the sampled image of the lane line into a pre-trained lane line detection head network to obtain the detection result of the lane line, and to determine the detection result of the lane line as the first detection result; wherein the detection result of the lane line includes at least one of lane line segmentation, lane line embedding, lane line color, lane line type, and lane line function; and / or, input the sampled image of the first obstacle into a pre-trained first obstacle detection head network to obtain the detection result of the first obstacle, and to determine the detection result of the first obstacle as the first detection result; wherein the detection result of the first obstacle includes at least one of heatmap, offset, dimension annotation, and rotation map.
[0034] In combination with the second aspect and the above implementation methods, in some possible implementation methods, the second detection module includes:
[0035] The second sampling unit is used to upsample the second feature image to obtain a second upsampled image;
[0036] The second detection unit is used to detect the second target object in the second target region based on the second upsampled image, and obtain the second detection result.
[0037] In combination with the second aspect and the above implementation, in some possible implementations, the second upsampled image includes the sampled image of the traffic indicator and the sampled image of the second obstacle;
[0038] The aforementioned second detection unit is specifically used to input the sampled image of the traffic sign into a pre-trained traffic sign detection head network to obtain the detection result of the traffic sign, and to determine the detection result of the traffic sign as the second detection result; and / or, to input the sampled image of the second obstacle into a pre-trained second obstacle detection head network to obtain the detection result of the second obstacle, and to determine the detection result of the second obstacle as the second detection result.
[0039] In conjunction with the second aspect and the above implementation methods, in some possible implementation methods, the vehicle control module is specifically used to control the vehicle according to the first detection result if the detection result of the first obstacle in the first detection result is the same as the detection result of the second obstacle in the second detection result, and the second obstacle is located within the detection range of the surround view camera; and to control the vehicle according to the second detection result if the detection result of the first obstacle in the first detection result is a part of the detection result of the second obstacle in the second detection result, and the first obstacle is located within the detection range of the surround view camera.
[0040] Thirdly, a vehicle is provided, including a memory and a processor. The memory is used to store executable program code, and the processor is used to call and run the executable program code from the memory, causing the vehicle to perform the vehicle control method of the first aspect or any possible implementation thereof.
[0041] Fourthly, an electronic device is provided, including a memory and a processor. The memory is used to store executable program code, and the processor is used to call and run the executable program code from the memory, causing the vehicle to perform the vehicle control method of the first aspect or any possible implementation thereof.
[0042] Fifthly, a computer program product is provided, comprising: computer program code, which, when run on a computer, causes the computer to execute the vehicle control method of the first aspect or any possible implementation thereof.
[0043] In a sixth aspect, a computer-readable storage medium is provided, which stores computer program code that, when executed on a computer, causes the computer to perform the vehicle control method of the first aspect or any possible implementation thereof. Attached Figure Description
[0044] Figure 1 A schematic flowchart of a vehicle control method provided in an embodiment of this application is shown;
[0045] Figure 2 A schematic diagram of a vehicle with autonomous driving capabilities is shown;
[0046] Figure 3 A schematic diagram is shown showing the fusion of downsampled images of the same scale corresponding to each viewpoint;
[0047] Figure 4 An architectural block diagram of the task processing model provided in this application is shown;
[0048] Figure 5 A schematic diagram of the first feature image is shown;
[0049] Figure 6 This illustration shows a structural schematic of a vehicle control device provided in an embodiment of this application;
[0050] Figure 7 A schematic diagram of the structure of a vehicle provided in an embodiment of this application is shown. Detailed Implementation
[0051] The technical solutions in this application will be clearly and thoroughly described below with reference to the accompanying drawings. In the description of the embodiments of this application, unless otherwise stated, " / " means "or," for example, A / B can mean A or B. "And / or" in the text is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. Furthermore, in the description of the embodiments of this application, "multiple" refers to two or more than two.
[0052] Hereinafter, the terms "first" and "second" are used for descriptive purposes only and should not be construed as implying or suggesting relative importance or implicitly indicating the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature.
[0053] The following is an embodiment of a vehicle control method provided in this application.
[0054] Figure 1A schematic flowchart of a vehicle control method provided in an embodiment of this application is shown, such as... Figure 1 As shown, the vehicle control method provided in this application embodiment is applied to the intelligent driving controller of a vehicle with autonomous driving function, such as... Figure 2 As shown, Figure 2 A schematic diagram of a vehicle with autonomous driving capabilities is shown. The vehicle includes a surround-view camera system comprising a front-view camera C1, a left front-view camera C2, a right front-view camera C3, a left rear-view camera C4, a right rear-view camera C5, and a rear-view camera C6. It is evident that the surround-view camera system includes front-view cameras. The surround-view camera system, composed of multiple cameras positioned around the vehicle, can acquire image data from 360 degrees around the vehicle, thereby achieving comprehensive environmental image monitoring. Because the surround-view camera system consists of multiple cameras positioned around the vehicle, its detection range is smaller than that of the front-view cameras; for example, the detection range of the surround-view camera is 0 to 50 meters, while the detection range of the front-view camera is 0 to 100 meters.
[0055] The above vehicle control methods include the following schemes:
[0056] S110: Obtain the first feature image from a bird's-eye view and the second feature image from a non-bird's-eye view.
[0057] In one exemplary embodiment, a bird's-eye view (BEV) is a two-dimensional projection method for observing and depicting terrain, urban layout, road networks, or other geospatial information from a bird's-eye view perspective directly above the vehicle. In the field of autonomous driving, the bird's-eye view provides a global and intuitive top-down view of the environment surrounding the vehicle. The first feature image under the bird's-eye view is obtained by capturing surround-view images of the surrounding environment from cameras positioned on the vehicle; that is, the first feature image is the bird's-eye view obtained from surround-view images. A non-bird's-eye view is not a bird's-eye view. In this embodiment, the non-bird's-eye view is the forward-looking view. The second feature image under the non-bird's-eye view is obtained by capturing forward-looking images of the surrounding environment from cameras positioned on the vehicle; that is, the second feature image does not belong to the bird's-eye view but to the feature image under the forward-looking view.
[0058] S120: Detect the first target object in the first target region based on the first feature image to obtain a first detection result.
[0059] After obtaining the first feature image, the detection branch of the BEV perspective task is used to detect the first target object in the first target region. The first target object includes lane lines and / or the first obstacle. The first target region is the environmental region corresponding to the perspective of the surround-view camera, such as the environmental region corresponding to the perspectives of C1-C6, i.e., the environmental region corresponding to the 360-degree view around the vehicle. If the first target object exists in the first feature image, the first detection result includes relevant information about the identified first target object, such as its location and distance from the vehicle. In the field of autonomous driving, the camera, as a sensor for vehicle driving control, belongs to the vehicle's vision (which can be understood as the vehicle's eyes). The first detection result of the first target object can also be called the environmental perception result of the first target object.
[0060] S130: Detect the second target object in the second target region based on the second feature image to obtain the second detection result.
[0061] After obtaining the second feature image, the detection branch of the non-BEV viewpoint task is used to detect the second target object in the second target region. The second target object includes traffic signs and / or second obstacles. Traffic signs include traffic lights (such as red and green lights), traffic signs, etc. The second target region is the environmental region corresponding to the viewpoint of the forward-looking camera. The second detection result of the second target object can also be called the environmental perception result of the second target object.
[0062] S140: Control the vehicle based on the first detection result or the second detection result.
[0063] After obtaining the first and second detection results, a target detection result is selected from the first and second detection results, and the target detection result is used for vehicle control. For example, the target detection result is used for vehicle obstacle avoidance control, vehicle turning control, and vehicle parking control. The target detection result can be either the first or the second detection result.
[0064] If the first detection result is used as the target detection result for vehicle control, since the first detection result is obtained through the first feature image, which is obtained through the surround-view camera capturing the surrounding environment, and the detection range of the surround-view camera is smaller than that of the front-view camera, then the vehicle can detect near-field objects around the vehicle, enabling it to detect objects in the vicinity of the vehicle. If the second detection result is used as the target detection result for vehicle control, since the second detection result is obtained through the second feature image, which is obtained through the front-view camera capturing the forward-view environment, and the detection range of the front-view camera is larger than that of the surround-view camera, then the vehicle can detect far-field objects, enabling it to detect both near- and far-field objects.
[0065] The vehicle control method provided in this application embodiment enables the detection of near-range environmental objects through surround-view images acquired by a surround-view camera, and the detection of far-range environmental objects through forward-view images acquired by a forward-view camera. This allows the vehicle to have both near-range and far-range environmental object detection capabilities, thereby expanding the vehicle's detection range. Furthermore, since the forward-view camera for acquiring forward-view environmental images is included within the surround-view camera, the acquisition of forward-view environmental images can be achieved through the forward-view camera within the surround-view camera, eliminating the need for an additional forward-view camera. This avoids wasting camera hardware resources and helps save on overall vehicle costs.
[0066] The following are Figure 1 The specific implementation methods of each step in the illustrated embodiment will be explained below:
[0067] In one possible implementation, the surround view environment image includes target area images corresponding to multiple viewpoints. Each camera in the surround view camera system corresponds to a viewpoint, and each camera is responsible for capturing the target area image corresponding to its viewpoint. In other words, the surround view environment image includes target area images corresponding to multiple viewpoints.
[0068] The following methods are used to obtain the first feature image:
[0069] The target region images corresponding to each of the multiple viewpoints are subjected to a first processing to obtain preprocessed images corresponding to each of the multiple viewpoints; wherein, the first processing includes at least one of virtual camera processing, normalization processing, resizing processing, denoising processing, and enhancement processing;
[0070] The preprocessed images corresponding to each of the multiple viewpoints are subjected to two second processing steps to obtain two frames of first downsampled images at a first scale and two frames of second downsampled images at a second scale, respectively, for each of the multiple preprocessed images; wherein, the second processing includes feature extraction and downsampling, and the first scale is different from the second scale;
[0071] The two frames of first downsampled images corresponding to each of the plurality of preprocessed images are fused to obtain the first fused image corresponding to each of the plurality of preprocessed images;
[0072] The two frames of second downsampled images corresponding to each of the plurality of preprocessed images are fused to obtain the second fused image corresponding to each of the plurality of preprocessed images;
[0073] Local spatial transformation is performed on the first fused image corresponding to each of the multiple preprocessed images to obtain the first local spatial bird's-eye view corresponding to each of the multiple perspectives under the bird's-eye view.
[0074] The second fused image corresponding to each of the multiple preprocessed images is subjected to local spatial transformation to obtain the second local spatial bird's-eye view corresponding to each of the multiple perspectives under the bird's-eye view; wherein, the scale of the first local spatial bird's-eye view is different from the scale of the second local spatial bird's-eye view.
[0075] The first feature image is generated based on the first local spatial bird's-eye view and the second local spatial bird's-eye view corresponding to the multiple perspectives.
[0076] Regarding the generation of the first feature image: First, the target region images corresponding to multiple viewpoints are subjected to a first processing (i.e., preprocessing) to obtain preprocessed images corresponding to multiple viewpoints. The first processing includes at least one of the following operations: virtual camera processing, normalization processing, resizing processing, denoising processing, and enhancement processing.
[0077] Next, the preprocessed images corresponding to each of the multiple viewpoints undergo two second processing steps. These second processing steps include feature extraction and downsampling; feature extraction is performed first, followed by downsampling. After one second processing step on the preprocessed images corresponding to each of the multiple viewpoints, one frame of first-downsampled image at the first scale and one frame of second-downsampled image at the second scale are obtained for each of the multiple preprocessed images. After two second processing steps, two frames of first-downsampled images at the first scale and two frames of second-downsampled images at the second scale are obtained for each of the multiple preprocessed images. The first scale and the second scale are different; the second scale is larger than the first scale and is an even multiple of the first scale. For example, the first scale is 8 times the scale, and the second scale is 16 times the scale, meaning each preprocessed image corresponds to two frames of first-downsampled images at an 8 times scale and two frames of second-downsampled images at a 16 times scale.
[0078] Next, the downsampled images of the same scale corresponding to each viewpoint are fused. This involves fusing two frames of the first downsampled images corresponding to each of the multiple preprocessed images to obtain a first fused image for each of the multiple preprocessed images, and fusing two frames of the second downsampled images corresponding to each of the multiple preprocessed images to obtain a second fused image for each of the multiple preprocessed images. For example... Figure 3 As shown, Figure 3 The diagram illustrates the fusion of downsampled images of the same scale corresponding to each viewpoint. F11 represents a first downsampled image at a first scale in one frame, F12 represents a first downsampled image at a first scale in another frame, and F13 represents a first fused image; F21 represents a second downsampled image at a second scale in one frame, F22 represents a second downsampled image at a second scale in another frame, and F23 represents a second fused image.
[0079] Based on the above fusion processing, the first fused image and the second fused image corresponding to each of the multiple preprocessed images are obtained, which are the first fused image and the second fused image corresponding to each of the multiple viewpoints.
[0080] Next, the image processing for local spatial transformation is implemented through a fully connected layer (FC). Since the FC layer plays a role in feature transformation and dimensionality reduction in the network structure, its working mechanism involves linear transformation and non-linear activation function processing of a set of input feature vectors to generate new feature representations. The FC layer performs local spatial transformation on the first fused images corresponding to each of the multiple preprocessed images, obtaining first local spatial bird's-eye view images corresponding to multiple perspectives. Similarly, the FC layer performs local spatial transformation on the second fused images corresponding to each of the multiple preprocessed images, obtaining second local spatial bird's-eye view images corresponding to multiple perspectives. Thus, the first and second local spatial bird's-eye view images corresponding to multiple perspectives are obtained. Finally, the first feature image is generated using the first and second local spatial bird's-eye view images corresponding to multiple perspectives. Since the first feature image is generated from multiple local spatial bird's-eye views, each corresponding to a first and second local spatial bird's-eye view, and these multiple views are stitched together to form a complete 360-degree view, the first feature image is a 360-degree bird's-eye view frame, which can be understood as a 360-degree spatial bird's-eye view. By performing local spatial transformation on the first and second fused images through a fully connected layer, the number of model parameters and computational cost can be reduced.
[0081] This application provides a task processing model for both BEV (Battery Electric Vehicle) perspective tasks and non-BEV perspective tasks in order to achieve the processing of these tasks. Figure 4An architectural block diagram of the task processing model provided in this application is shown, as follows: Figure 4 As shown, the task processing model includes two shared networks (a first shared network and a second shared network), a fusion module, a fully connected layer, a first decoder, a second decoder, a lane line detection head network, a first obstacle detection head network, a third decoder, a traffic indicator detection head network, and a second obstacle detection head network. The outputs of the first and second shared networks are connected to the input of the fusion module. The output of the fusion module is connected to the input of the fully connected layer and the input of the third decoder. The output of the fully connected layer is connected to the inputs of the first and second decoders. The output of the first decoder is connected to the input of the lane line detection head network. The output of the second decoder is connected to the input of the first obstacle detection head network. The output of the third decoder is connected to the inputs of the traffic indicator detection head network and the second obstacle detection head network.
[0082] Among them, the branch formed by the fully connected layer, the first decoder, the second decoder, the lane line detection head network, and the first obstacle detection head network is the detection branch for the BEV perspective task; the branch formed by the third decoder, the traffic indicator detection head network, and the second obstacle detection head network is the detection branch for the non-BEV perspective task.
[0083] The first and second fused images corresponding to multiple perspectives are output through the fusion module. Both shared networks include a deep neural network model for feature extraction and a deep learning model for object detection. The output of the deep neural network model is connected to the input of the deep learning model.
[0084] Deep neural network models, such as the backbone, specifically FastERNIE-T0, typically refer to the part of a deep neural network used to extract low-level features. It originates from the basic network structure in image classification tasks. In object detection tasks, the backbone is responsible for extracting feature maps of different levels from the input image. These feature maps contain image information ranging from coarse to fine, which is beneficial for identifying objects of different sizes in the image. Another deep learning model is the Feature Pyramid Network (FPN) model. The FPN model addresses the scale diversity problem in object detection. It constructs a feature pyramid, processing the different levels of features extracted by the backbone to generate feature layers of different resolutions. Each level of feature map can be used to predict objects of different scales, thus improving the performance of small object detection. The FPN model can implement downsampling operations. In the FPN model, the downsampling process mainly occurs in the basic convolutional neural network backbone, obtaining feature representations at different levels. Therefore, each shared network can implement a second processing step, namely feature extraction and downsampling. The two Backbone+FPN models used in this application have the same network architecture, which is a lightweight network structure. Compared with the structure of a single Backbone+FPN model, it can effectively improve the feature extraction capability of the model.
[0085] In one possible implementation, the above-mentioned second processing is performed twice on the preprocessed images corresponding to each of the multiple viewpoints to obtain two frames of first downsampled images at the first scale and two frames of second downsampled images at the second scale corresponding to each of the multiple preprocessed images, including the following scheme:
[0086] For each viewpoint, the preprocessed image is input into two pre-trained shared networks to obtain the output results of the two shared networks respectively; wherein, the output results of the two shared networks respectively include a first downsampled image and a second downsampled image.
[0087] like Figure 4 As shown, for the processing of the preprocessed image corresponding to each viewpoint, the preprocessed image corresponding to each viewpoint is input into two shared networks, namely... Figure 4 Input 1 includes the preprocessed image corresponding to each viewpoint. Since the preprocessed image corresponding to each viewpoint is obtained through virtual camera operation, the virtual camera operation transforms the image of each viewpoint to a standard viewpoint, which can increase the model's generalization ability to adapt to different vehicle models.
[0088] The preprocessed image corresponding to each viewpoint is input into two shared networks, that is, the preprocessed image corresponding to each viewpoint is input into the first shared network and the second shared network respectively. The first shared network performs a second processing on the input preprocessed image to obtain the output result of the first shared network, which includes a first downsampled image at a first scale and a second downsampled image at a second scale. Similarly, the second shared network performs a second processing on the input preprocessed image to obtain the output result of the second shared network, which includes a first downsampled image at a first scale and a second downsampled image at a second scale. In this way, two first downsampled images at the first scale and two second downsampled images at the second scale can be obtained for each of the multiple preprocessed images.
[0089] In one possible implementation, generating the first feature image based on the first local spatial bird's-eye view and the second local spatial bird's-eye view corresponding to each of the multiple perspectives includes the following schemes:
[0090] By stitching together the first and second local spatial bird's-eye view images corresponding to the multiple perspectives along the channel dimension, a stitched local spatial bird's-eye view image corresponding to the multiple perspectives is obtained.
[0091] The first feature image is obtained by summing the stitched local spatial bird's-eye view images corresponding to the multiple perspectives.
[0092] After obtaining the first and second local spatial bird's-eye view images corresponding to multiple perspectives, perform a concat operation on images of different scales under the same perspective, and a sum operation on images under different perspectives.
[0093] The concat operation for images at different scales from the same viewpoint involves stitching together the first and second local spatial bird's-eye views corresponding to multiple viewpoints along the image's channel dimension to obtain a stitched local spatial bird's-eye view corresponding to each viewpoint. For example, for the right rear view of a vehicle, the first and second local spatial bird's-eye views corresponding to the right rear view are stitched together along the image's channel dimension to obtain a stitched local spatial bird's-eye view corresponding to the right rear view.
[0094] The summation operation on images from different perspectives includes the following: After setting up the surround-view cameras, the perspectives of two adjacent cameras overlap, meaning the target area images captured by two adjacent cameras intersect. Specifically, the first local spatial bird's-eye view images corresponding to two adjacent perspectives intersect, and similarly, the second local spatial bird's-eye view images corresponding to two adjacent perspectives also intersect. Therefore, based on the intersection between the images, the stitched local spatial bird's-eye view images corresponding to multiple perspectives are summed to obtain the first feature image.
[0095] like Figure 5 As shown, Figure 5 A schematic diagram of the first feature image is shown. Assuming that the multiple viewpoints corresponding to the surround-view camera are the left front view, right front view, left rear view, and right rear view, V1 represents the stitched local spatial bird's-eye view corresponding to the left front view, V2 represents the stitched local spatial bird's-eye view corresponding to the right front view, V3 represents the stitched local spatial bird's-eye view corresponding to the left rear view, V4 represents the stitched local spatial bird's-eye view corresponding to the right rear view, V12 represents the intersection of V1 and V2, V13 represents the intersection of V1 and V32, V34 represents the intersection of V3 and V4, and V24 represents the intersection of V2 and V4. The image finally formed by V1-V4 is the first feature image.
[0096] One possible implementation involves obtaining the second feature image using the following methods:
[0097] Determine the target viewpoint of the forward-facing camera among the multiple viewpoints;
[0098] The first fused image and the second fused image corresponding to the target viewpoint are fused to obtain the forward-looking fused image;
[0099] The forward-looking fused image is upsampled to obtain the second feature image.
[0100] Because the second feature image is obtained from the forward-looking environment image captured by the forward-looking camera, and the first and second fused images corresponding to multiple viewpoints each include the first and second fused images corresponding to the viewpoint of the forward-looking camera, the viewpoint of the forward-looking camera is determined from the multiple viewpoints, called the target viewpoint. The first and second fused images corresponding to the target viewpoint are then determined from the first and second fused images corresponding to the multiple viewpoints. Next, the first and second fused images corresponding to the target viewpoint are fused to obtain the forward-looking fused image. Then, the forward-looking fused image is upsampled to obtain the second feature image. Wherein, as... Figure 4 As shown, upsampling the forward-looking fused image means inputting the forward-looking fused image into the third decoder, and the third decoder outputting the forward-looking fused image.
[0101] In one possible implementation, the above-mentioned detection of the first target object in the first target region based on the first feature image to obtain the first detection result includes the following schemes:
[0102] The first feature image is cropped to obtain a cropped image containing the first target object;
[0103] The cropped image is upsampled to obtain a first upsampled image;
[0104] The first target object in the first target region is detected based on the first upsampled image to obtain the first detection result.
[0105] After obtaining the first feature image, the first target object in the first feature image is cropped to obtain a cropped image containing the first target object. Then, the cropped image is upsampled to obtain a first upsampled image at a third scale. The first target object in the first target region is then detected using the first upsampled image to obtain a first detection result. Specifically, when the first target object includes lane lines and / or a first obstacle, if the cropped image of the first target object includes cropped images of lane lines and / or the first obstacle, the cropped image of the lane lines is input to the first decoder, which upsamples the cropped image of the lane lines to obtain a sampled image of the lane lines; and / or, the cropped image of the first obstacle is input to the second decoder, which upsamples the cropped image of the first obstacle to obtain a sampled image of the first obstacle. That is, the first upsampled image includes sampled images of lane lines and / or the first obstacle.
[0106] In one possible implementation, when the first target object includes lane lines and / or a first obstacle, the first upsampled image includes a sampled image of the lane lines and / or a sampled image of the first obstacle. The detection of the first target object in the first target region based on the first upsampled image to obtain the first detection result includes the following schemes:
[0107] The sampled image of the lane line is input into a pre-trained lane line detection head network to obtain the detection result of the lane line, and the detection result of the lane line is determined as the first detection result;
[0108] And / or,
[0109] The sampled image of the first obstacle is input into a pre-trained first obstacle detection head network to obtain the detection result of the first obstacle, and the detection result of the first obstacle is determined as the first detection result.
[0110] like Figure 4As shown, the sampled image of the lane lines is input into the lane line detection head network. The lane line detection head network detects the lane lines in the first target area and obtains the lane line detection results, i.e. Figure 4 Output 1 in the output, wherein the lane line detection result includes at least one of lane line segmentation, lane line embedding, lane line color, lane line type, lane line function, etc., and the lane line detection result is determined as the first detection result.
[0111] And / or, the sampled image of the first obstacle is input into the first obstacle detection head network, the first obstacle detection head network detects the first obstacle in the first target area, and obtains the detection result of the first obstacle, i.e. Figure 4 Output 2 in the first obstacle detection result includes at least one of the following: heatmap, offset, dimension, rotation, etc. The first obstacle detection result is determined as the first detection result, that is, the first detection result includes the lane line detection result and / or the first obstacle detection result.
[0112] In one possible implementation, the above-mentioned detection of the second target object in the second target region based on the second feature image to obtain the second detection result includes the following schemes:
[0113] The second feature image is upsampled to obtain a second upsampled image;
[0114] The second target object in the second target region is detected based on the second upsampled image to obtain the second detection result.
[0115] like Figure 4 As shown, the second feature image is input into the third decoder, the third decoder upsamples the second feature image to obtain the second upsampled image, and then the second target object in the second target region is detected by the second upsampled image to obtain the second detection result, which is a three-dimensional detection result.
[0116] In one possible implementation, when the first target object includes lane lines and / or a first obstacle, the second upsampled image includes a sampled image of the traffic sign and a sampled image of the second obstacle. The detection of the second target object in the second target area based on the second upsampled image to obtain the second detection result includes the following schemes:
[0117] The sampled image of the traffic sign is input into a pre-trained traffic sign detection head network to obtain the detection result of the traffic sign, and the detection result of the traffic sign is determined as the second detection result;
[0118] And / or,
[0119] The sampled image of the second obstacle is input into the pre-trained second obstacle detection head network to obtain the detection result of the second obstacle, and the detection result of the second obstacle is determined as the second detection result.
[0120] like Figure 4 As shown, the sampled image of the traffic sign is input into the traffic sign detection head network. The traffic sign detection head network detects the traffic signs in the second target area and obtains the detection result of the traffic signs. Figure 4 Output 3 in the middle, the detection result of the traffic sign is the three-dimensional detection result of the traffic sign, and the detection result of the traffic sign is determined as the second detection result.
[0121] And / or, the sampled image of the second obstacle is input into the second obstacle detection head network, the second obstacle detection head network detects the second obstacle in the second target region, and obtains the detection result of the second obstacle, i.e. Figure 4 Output 4 in the middle, the detection result of the second obstacle is the three-dimensional detection result of the second obstacle, and the detection result of the second obstacle is determined as the second detection result, that is, the second detection result includes the detection result of the traffic indicator and / or the detection result of the second obstacle.
[0122] In one possible implementation, if the first target object includes lane lines and a first obstacle, and the second target object includes a traffic indicator and a second obstacle, the first detection result includes the detection result of the lane lines and the detection result of the first obstacle, and the second detection result includes the detection result of the traffic indicator and the detection result of the second obstacle, then controlling the vehicle based on the first detection result or the second detection result includes the following schemes:
[0123] If the detection result of the first obstacle in the first detection result is the same as the detection result of the second obstacle in the second detection result, and the second obstacle is within the detection range of the surround view camera, the vehicle is controlled according to the first detection result;
[0124] If the detection result of the first obstacle in the first detection result is part of the detection result of the second obstacle in the second detection result, and the first obstacle is within the detection range of the surround view camera, the vehicle is controlled according to the second detection result.
[0125] In the case of controlling the vehicle using the first detection result, if the detection result of the first obstacle in the first detection result is the same as the detection result of the second obstacle in the second detection result, and the second obstacle is located within the detection range of the surround view camera, it means that all obstacles identified by the forward-looking environment image collected by the forward-looking camera have also been identified by the surround-view environment image collected by the surround view camera. Since the field of view of the surround view camera is wider than that of the forward-looking camera, the first detection result is used first to control the vehicle, thereby realizing vehicle control using information of detected near-field environmental objects.
[0126] When controlling the vehicle using the second detection result, if the detection result of the first obstacle in the first detection result is a part of the detection result of the second obstacle in the second detection result, and the first obstacle is within the detection range of the surround view camera, it means that both the forward-view image captured by the forward-view camera and the surround-view image captured by the surround view camera have identified the first obstacle. However, the forward-view image also identifies obstacles other than the first obstacle, while the surround-view image does not identify obstacles other than the first obstacle. This means that obstacles other than the first obstacle are not within the detection range of the surround-view camera, but are within the detection range of the forward-view camera. Therefore, the second detection result is used to control the vehicle first, thereby realizing vehicle control using information of detected distant environmental objects. This allows the vehicle to have both the ability to detect near-distance environmental objects and the ability to detect distant environmental objects.
[0127] The following are embodiments of the apparatus described in this application, which can be used to execute the embodiments of the method described in this application. For details not disclosed in the apparatus embodiments of this application, please refer to the embodiments of the method described in this application.
[0128] Figure 6 This application provides a schematic diagram of the structure of a vehicle control device according to an embodiment of the present application. Figure 6 As shown, the vehicle control device 600 includes:
[0129] Image acquisition module 610 is used to acquire a first feature image from a bird's-eye view and a second feature image from a non-bird's-eye view; wherein, the first feature image is obtained by acquiring a surround-view environment image from a surround-view camera arranged in the vehicle, and the second feature image is obtained by acquiring a forward-view environment image from a forward-view camera arranged in the vehicle, wherein the surround-view camera includes the forward-view camera, and the detection range of the surround-view camera is smaller than the detection range of the forward-view camera.
[0130] The first detection module 620 is used to detect a first target object in the first target region based on the first feature image and obtain a first detection result; wherein, the first target object includes lane lines and / or a first obstacle, and the first target region is the environmental region corresponding to the viewpoint of the surround-view camera;
[0131] The second detection module 630 is used to detect a second target object in the second target region based on the second feature image and obtain a second detection result; wherein the second target object includes a traffic sign and / or a second obstacle, and the second target region is the environmental region corresponding to the viewpoint of the forward-looking camera;
[0132] The vehicle control module 640 is used to control the vehicle based on the first detection result or the second detection result.
[0133] In one possible implementation, the image acquisition module 610 includes:
[0134] The first acquisition unit is configured to perform a first processing on the target region images corresponding to each of the multiple viewpoints to obtain preprocessed images corresponding to each of the multiple viewpoints; wherein the first processing includes at least one of virtual camera processing, normalization processing, resizing processing, denoising processing, and enhancement processing; perform two second processing on the preprocessed images corresponding to each of the multiple viewpoints to obtain two frames of first downsampled images at a first scale and two frames of second downsampled images at a second scale corresponding to each of the multiple preprocessed images; wherein the second processing includes feature extraction and downsampling, and the first scale is different from the second scale; and fuse the two frames of first downsampled images corresponding to each of the multiple preprocessed images to obtain a first fused image corresponding to each of the multiple preprocessed images. The image is processed as follows: Two frames of second downsampled images corresponding to each of the plurality of preprocessed images are fused to obtain a second fused image corresponding to each of the plurality of preprocessed images; a local spatial transformation is performed on the first fused image corresponding to each of the plurality of preprocessed images to obtain a first local spatial bird's-eye view corresponding to each of the plurality of viewpoints under the bird's-eye view; a local spatial transformation is performed on the second fused image corresponding to each of the plurality of preprocessed images to obtain a second local spatial bird's-eye view corresponding to each of the plurality of viewpoints under the bird's-eye view; wherein the scale of the first local spatial bird's-eye view is different from the scale of the second local spatial bird's-eye view; the first feature image is generated based on the first local spatial bird's-eye view and the second local spatial bird's-eye view corresponding to each of the plurality of viewpoints.
[0135] In one possible implementation, the first acquisition unit, in performing two second processes on the preprocessed images corresponding to the plurality of viewpoints to obtain two frames of first downsampled images at the first scale and two frames of second downsampled images at the second scale corresponding to the plurality of preprocessed images, is configured to input the preprocessed images corresponding to each viewpoint into two pre-trained shared networks respectively to obtain the output results corresponding to the two shared networks respectively; wherein, the output results corresponding to the two shared networks respectively include one frame of first downsampled image and one frame of second downsampled image, and both shared networks include a deep neural network model for feature extraction and a deep learning model for object detection.
[0136] In one possible implementation, the first acquisition unit, in generating the first feature image based on the first and second local spatial bird's-eye views corresponding to each of the multiple perspectives, is used to stitch the first and second local spatial bird's-eye views along the channel dimension to obtain the stitched local spatial bird's-eye view corresponding to each of the multiple perspectives; and to sum the stitched local spatial bird's-eye views corresponding to each of the multiple perspectives to obtain the first feature image.
[0137] In one possible implementation, the image acquisition module 610 includes:
[0138] The second acquisition unit is used to determine the target view of the forward-looking camera among the multiple viewpoints; fuse the first fusion image and the second fusion image corresponding to the target viewpoint to obtain a forward-looking fusion image; and upsample the forward-looking fusion image to obtain the second feature image.
[0139] In one possible implementation, the first detection module 620 includes:
[0140] The cropping unit is used to crop the first feature image to obtain a cropped image containing the first target object;
[0141] The first sampling unit is used to upsample the cropped image to obtain a first upsampled image;
[0142] The first detection unit is configured to detect the first target object in the first target region based on the first upsampled image, and obtain the first detection result.
[0143] In one possible implementation, the first upsampled image includes a sampled image of the lane line and / or a sampled image of the first obstacle. Specifically, the first detection unit is configured to input the sampled image of the lane line into a pre-trained lane line detection head network to obtain the detection result of the lane line, and determine the detection result of the lane line as the first detection result; wherein the detection result of the lane line includes at least one of lane line segmentation, lane line embedding, lane line color, lane line type, and lane line function; and / or, input the sampled image of the first obstacle into a pre-trained first obstacle detection head network to obtain the detection result of the first obstacle, and determine the detection result of the first obstacle as the first detection result; wherein the detection result of the first obstacle includes at least one of heatmap, offset, dimension annotation, and rotation map.
[0144] In one possible implementation, the second detection module 630 includes:
[0145] The second sampling unit is used to upsample the second feature image to obtain a second upsampled image;
[0146] The second detection unit is used to detect the second target object in the second target region based on the second upsampled image, and obtain the second detection result.
[0147] In one possible implementation, the second upsampled image includes a sampled image of the traffic sign and a sampled image of the second obstacle;
[0148] The aforementioned second detection unit is specifically used to input the sampled image of the traffic sign into a pre-trained traffic sign detection head network to obtain the detection result of the traffic sign, and to determine the detection result of the traffic sign as the second detection result; and / or, to input the sampled image of the second obstacle into a pre-trained second obstacle detection head network to obtain the detection result of the second obstacle, and to determine the detection result of the second obstacle as the second detection result.
[0149] In one possible implementation, the vehicle control module 610 is specifically configured to control the vehicle based on the first detection result if the detection result of the first obstacle in the first detection result is the same as the detection result of the second obstacle in the second detection result, and the second obstacle is within the detection range of the surround-view camera; and to control the vehicle based on the second detection result if the detection result of the first obstacle in the first detection result is a part of the detection result of the second obstacle in the second detection result, and the first obstacle is within the detection range of the surround-view camera.
[0150] It should be noted that the vehicle control device provided in the above embodiments is only illustrated by the division of the above functional modules when executing the vehicle control method. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the vehicle control device and the vehicle control method embodiments provided in the above embodiments belong to the same concept. Therefore, for details not disclosed in the device embodiments of this application, please refer to the embodiments of the vehicle control method of this application, which will not be repeated here.
[0151] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.
[0152] Figure 7 This application provides a schematic diagram of the structure of a vehicle according to an embodiment of the present application. Figure 7 As shown, the vehicle 700 includes a memory 701 and a processor 702. The memory 701 stores executable program code 7011, and the processor 702 is used to call and execute the executable program code 7011 to perform a vehicle control method.
[0153] This embodiment can divide the vehicle into functional modules according to the above method example. For example, each function can be assigned to a separate module, or two or more functions can be integrated into one processing module. The integrated module can be implemented in hardware. It should be noted that the module division in this embodiment is illustrative and only represents one logical functional division. In actual implementation, there may be other division methods.
[0154] When each functional module is divided according to its corresponding function, the vehicle may include: an image acquisition module, a first detection module, a second detection module, a vehicle control module, etc. It should be noted that all relevant content of each step involved in the above method embodiments can be referenced from the functional description of the corresponding functional module, and will not be repeated here.
[0155] The vehicle provided in this embodiment is used to execute the vehicle control method described above, and therefore can achieve the same effect as the above implementation method.
[0156] When using integrated units, the vehicle may include a processing module and a storage module. The processing module is used to control and manage the vehicle's movements. The storage module is used to support the vehicle in executing relevant program code and data.
[0157] The processing module may be a processor or a controller, which can implement or execute the various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure of this application. The processor may also be a combination of functions that implement computing capabilities, such as a combination of one or more microprocessors, a combination of digital signal processing (DSP) and a microprocessor, etc., and the storage module may be a memory.
[0158] This embodiment also provides an electronic device, which includes a memory and a processor, wherein the memory stores executable program code, and the processor is used to call and execute the executable program code to perform a vehicle display control method.
[0159] This embodiment also provides a computer-readable storage medium storing computer program code. When the computer program code is run on a computer, the computer executes the above-described related method steps to implement a vehicle control method in the above embodiment.
[0160] This embodiment also provides a computer program product that, when run on a computer, causes the computer to perform the aforementioned steps to implement a vehicle control method as described in the above embodiment.
[0161] In addition, the vehicle or electronic device provided in the embodiments of this application may specifically be a chip, component or module. The vehicle or electronic device may include a connected processor and a memory. The memory is used to store instructions. When the vehicle or electronic device is running, the processor may call and execute the instructions to make the chip execute a vehicle control method in the above embodiments.
[0162] In this embodiment, the vehicle, electronic device, computer-readable storage medium, computer program product or chip are all used to execute the corresponding vehicle control method provided above. Therefore, the beneficial effects that can be achieved can be referred to the beneficial effects of the corresponding vehicle control method provided above, and will not be repeated here.
[0163] Through the above description of the embodiments, those skilled in the art will understand that, for the sake of convenience and brevity, only the division of the above functional modules is used as an example. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.
[0164] In the embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another device, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.
[0165] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A vehicle control method, characterized in that, The vehicle control method includes: A first feature image from a bird's-eye view and a second feature image from a non-bird's-eye view are acquired; wherein, the first feature image is obtained by capturing a surround-view environment image from a surround-view camera deployed on the vehicle, and the second feature image is obtained by capturing a forward-view environment image from a forward-view camera deployed on the vehicle, wherein the surround-view camera includes the forward-view camera, and the detection range of the surround-view camera is smaller than the detection range of the forward-view camera. Based on the first feature image, a first target object in the first target region is detected to obtain a first detection result; wherein, the first target object includes lane lines and / or a first obstacle, and the first target region is the environmental region corresponding to the viewpoint of the surround-view camera; The second target object in the second target region is detected based on the second feature image to obtain a second detection result; wherein, the second target object includes a traffic sign and / or a second obstacle, and the second target region is the environmental region corresponding to the viewpoint of the forward-looking camera; The vehicle is controlled based on the first or the second detection result.
2. The vehicle control method according to claim 1, characterized in that, The surround view environment image includes target area images corresponding to multiple viewpoints; Obtaining the first feature image includes: The target region images corresponding to each of the multiple viewpoints are subjected to a first processing to obtain preprocessed images corresponding to each of the multiple viewpoints; wherein, the first processing includes at least one of virtual camera processing, normalization processing, resizing processing, noise reduction processing, and enhancement processing; The preprocessed images corresponding to each of the multiple viewpoints are subjected to two second processing steps to obtain two frames of first downsampled images at a first scale and two frames of second downsampled images at a second scale, respectively, for each of the multiple preprocessed images; wherein, the second processing includes feature extraction and downsampling, and the first scale is different from the second scale; The two frames of first downsampled images corresponding to each of the plurality of preprocessed images are fused to obtain the first fused image corresponding to each of the plurality of preprocessed images; The two frames of second downsampled images corresponding to each of the plurality of preprocessed images are fused to obtain the second fused image corresponding to each of the plurality of preprocessed images; Local spatial transformation is performed on the first fused image corresponding to each of the multiple preprocessed images to obtain the first local spatial bird's-eye view corresponding to each of the multiple perspectives under the bird's-eye view. The second fused image corresponding to each of the multiple preprocessed images is subjected to local spatial transformation to obtain the second local spatial bird's-eye view corresponding to each of the multiple perspectives under the bird's-eye view; wherein, the scale of the first local spatial bird's-eye view is different from the scale of the second local spatial bird's-eye view. The first feature image is generated based on the first local spatial bird's-eye view and the second local spatial bird's-eye view corresponding to the multiple perspectives.
3. The vehicle control method according to claim 2, characterized in that, The step of performing two second processing steps on the preprocessed images corresponding to each of the multiple viewpoints to obtain two frames of first downsampled images at the first scale and two frames of second downsampled images at the second scale corresponding to each of the multiple preprocessed images includes: For each viewpoint, the preprocessed image is input into two pre-trained shared networks to obtain the output results of the two shared networks respectively. The output results of the two shared networks each include a first downsampled image and a second downsampled image. Both shared networks include a deep neural network model for feature extraction and a deep learning model for object detection.
4. The vehicle control method according to claim 2, characterized in that, The step of generating the first feature image based on the first local spatial bird's-eye view and the second local spatial bird's-eye view corresponding to each of the multiple perspectives includes: By stitching together the first and second local spatial bird's-eye view images corresponding to the multiple perspectives along the channel dimension, a stitched local spatial bird's-eye view image corresponding to the multiple perspectives is obtained. The first feature image is obtained by summing the stitched local spatial bird's-eye view images corresponding to the multiple perspectives.
5. The vehicle control method according to claim 2, characterized in that, Obtaining the second feature image includes: Determine the target viewpoint of the forward-facing camera among the multiple viewpoints; The first fused image and the second fused image corresponding to the target viewpoint are fused to obtain the forward-looking fused image; The forward-looking fused image is upsampled to obtain the second feature image.
6. The vehicle control method according to claim 1, characterized in that, The step of detecting the first target object in the first target region based on the first feature image to obtain the first detection result includes: The first feature image is cropped to obtain a cropped image containing the first target object; The cropped image is upsampled to obtain a first upsampled image; The first target object in the first target region is detected based on the first upsampled image to obtain the first detection result.
7. The vehicle control method according to claim 6, characterized in that, The first upsampled image includes a sampled image of the lane line and / or a sampled image of the first obstacle; The step of detecting the first target object in the first target region based on the first upsampled image to obtain the first detection result includes: The sampled image of the lane line is input into a pre-trained lane line detection head network to obtain the detection result of the lane line, and the detection result of the lane line is determined as the first detection result; wherein, the detection result of the lane line includes at least one of lane line segmentation, lane line embedding, lane line color, lane line type and lane line function; And / or, The sampled image of the first obstacle is input into a pre-trained first obstacle detection head network to obtain the detection result of the first obstacle, and the detection result of the first obstacle is determined as the first detection result; wherein, the detection result of the first obstacle includes at least one of heat map, offset, dimension annotation and rotation map.
8. The vehicle control method according to claim 1, characterized in that, The second detection result obtained by detecting the second target object in the second target region based on the second feature image includes: The second feature image is upsampled to obtain a second upsampled image; The second target object in the second target region is detected based on the second upsampled image to obtain the second detection result.
9. The vehicle control method according to claim 8, characterized in that, The second upsampled image includes the sampled image of the traffic sign and the sampled image of the second obstacle; The step of detecting the second target object in the second target region based on the second upsampled image to obtain the second detection result includes: The sampled image of the traffic sign is input into a pre-trained traffic sign detection head network to obtain the detection result of the traffic sign, and the detection result of the traffic sign is determined as the second detection result; And / or, The sampled image of the second obstacle is input into a pre-trained second obstacle detection head network to obtain the detection result of the second obstacle, and the detection result of the second obstacle is determined as the second detection result.
10. The vehicle control method according to any one of claims 1 to 9, characterized in that, If the first target object includes the lane line and the first obstacle, and the second target object includes the traffic indicator and the second obstacle, controlling the vehicle based on the first detection result or the second detection result includes: If the detection result of the first obstacle in the first detection result is the same as the detection result of the second obstacle in the second detection result, and the second obstacle is within the detection range of the surround view camera, the vehicle is controlled according to the first detection result; If the detection result of the first obstacle in the first detection result is part of the detection result of the second obstacle in the second detection result, and the first obstacle is within the detection range of the surround view camera, the vehicle is controlled according to the second detection result.
11. A vehicle control device, characterized in that, The vehicle control device includes: An image acquisition module is used to acquire a first feature image from a bird's-eye view and a second feature image from a non-bird's-eye view; wherein, the first feature image is obtained by acquiring a surround-view environment image from a surround-view camera arranged in the vehicle, and the second feature image is obtained by acquiring a forward-view environment image from a forward-view camera arranged in the vehicle, wherein the surround-view camera includes the forward-view camera, and the detection range of the surround-view camera is smaller than the detection range of the forward-view camera. The first detection module is used to detect a first target object in the first target region based on the first feature image to obtain a first detection result; wherein the first target object includes lane lines and / or a first obstacle, and the first target region is the environmental region corresponding to the viewpoint of the surround-view camera; The second detection module is used to detect a second target object in the second target region based on the second feature image to obtain a second detection result; wherein the second target object includes a traffic sign and / or a second obstacle, and the second target region is the environmental region corresponding to the viewpoint of the forward-looking camera; The vehicle control module is used to control the vehicle based on the first detection result or the second detection result.
12. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed, implements the vehicle control method as described in any one of claims 1 to 10.