Target detection method and device, terminal equipment and storage medium

By generating a bird's-eye view feature map in the BEV space using the depth features and parameters of the camera feature map, the problem of low target detection accuracy in the prior art is solved, and higher detection accuracy is achieved.

CN120976512APending Publication Date: 2025-11-18HAOMO TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410614740.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-05-17
Publication Date
2025-11-18

AI Technical Summary

Technical Problem

Existing technologies for target detection in BEV space are insufficient to meet practical needs, resulting in low target detection accuracy.

Method used

By acquiring feature maps of the images to be detected captured by each camera on the vehicle, feature projection processing is performed to obtain depth features, and bird's-eye view feature maps are generated using the intrinsic and extrinsic parameters of the cameras. Target detection is then performed by combining multiple set horizontal planes.

Benefits of technology

It improves the accuracy of target detection, generates rich bird's-eye view feature maps, and enhances the accuracy of detection results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120976512A_ABST
    Figure CN120976512A_ABST
Patent Text Reader

Abstract

The invention is suitable for the technical field of computers, and provides a target detection method and device, terminal equipment and a storage medium, and the method comprises the steps: obtaining a feature map corresponding to a to-be-detected image shot by each camera of a vehicle; performing feature projection processing on each column of features in each feature map to obtain depth features corresponding to each column of features in each feature map under different depths; generating a bird's-eye view feature map corresponding to each set horizontal plane according to the internal and external parameters of each camera and the depth features under different depths corresponding to each column of features in each feature map; and performing target detection processing according to each bird's-eye view feature map to obtain a target detection result. According to the method, feature projection processing is carried out on each column of features in each feature map, abundant depth features can be obtained, and abundant aerial view angle feature maps can be obtained by combining internal and external parameters of each camera and a plurality of depth features, so that the accuracy of target detection is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the technical field of computers, and particularly relates to a target detection method and device, a terminal device, and a storage medium. BACKGROUND

[0002] Bird's Eye View (BEV) is a commonly used representation space in the field of automatic driving. Target detection in the BEV space is a basic task in the automatic driving scenario. However, the existing target detection technology in the BEV space cannot meet the actual needs, and the accuracy of target detection is reduced. SUMMARY

[0003] The embodiments of the present application provide a target detection method and device, a terminal device, and a storage medium, which improve the accuracy of target detection.

[0004] In a first aspect, the embodiments of the present application provide a target detection method, comprising:

[0005] obtaining feature maps corresponding to images to be detected photographed by each camera of a vehicle;

[0006] performing feature projection processing on each column of features in each of the feature maps to obtain depth features of each column of features in each of the feature maps at different depths;

[0007] generating a Bird's Eye View (BEV) feature map corresponding to each of the set horizontal planes according to the intrinsic and extrinsic parameters of each camera and the depth features of each column of features in each of the feature maps at different depths;

[0008] performing target detection processing according to each of the BEV feature maps to obtain a target detection result.

[0009] Optionally, the generating of the BEV feature map corresponding to each of the set horizontal planes according to the intrinsic and extrinsic parameters of each camera and the depth features of each column of features in each of the feature maps at different depths comprises:

[0010] calculating a relationship between a depth and an angle of each column of features in each of the feature maps at each of the set horizontal planes according to the intrinsic and extrinsic parameters of each camera and pixel coordinates of each pixel point in each of the feature maps;

[0011] determining the BEV feature map corresponding to each of the set horizontal planes according to the relationship between the depth and the angle of each column of features in each of the feature maps at each of the set horizontal planes.

[0012] Optionally, the depth-angle relationship of each column of features in each of the feature maps corresponding to each of the setting horizontal planes is calculated according to the internal and external parameters of each of the cameras and the pixel coordinates of each pixel point in each of the feature maps, and the depth-angle relationship of each column of features in each of the feature maps corresponding to each of the setting horizontal planes is calculated according to the internal and external parameters of each of the cameras and the pixel coordinates of each pixel point in each of the feature maps.

[0013] The first coordinate set of each column of features in each of the feature maps on each of the setting horizontal planes is calculated according to the internal and external parameters of each of the cameras and the pixel coordinates of each pixel point in each of the feature maps.

[0014] The first coordinate set of each column of features in each of the feature maps on each of the setting horizontal planes is calculated according to the internal and external parameters of each of the cameras and the pixel coordinates of each pixel point in each of the feature maps.

[0015] Optionally, the bird's-eye view angle feature map corresponding to each of the setting horizontal planes is determined according to the depth-angle relationship of each column of features in each of the feature maps corresponding to each of the setting horizontal planes.

[0016] The three-dimensional coordinate set of the depth feature at different depths corresponding to each column of features in each of the feature maps on each of the setting horizontal planes is determined according to the depth-angle relationship of each column of features in each of the feature maps corresponding to each of the setting horizontal planes.

[0017] The bird's-eye view angle feature map corresponding to each of the setting horizontal planes is generated according to the three-dimensional coordinate set of the depth feature at different depths corresponding to each column of features in each of the feature maps on each of the setting horizontal planes.

[0018] Optionally, the three-dimensional coordinate set of the depth feature at different depths corresponding to each column of features in each of the feature maps on each of the setting horizontal planes is determined according to the depth-angle relationship of each column of features in each of the feature maps corresponding to each of the setting horizontal planes.

[0019] The angle corresponding to each of the depth features at different depths corresponding to each column of features in each of the feature maps on each of the setting horizontal planes is calculated according to the depth-angle relationship of each column of features in each of the feature maps corresponding to each of the setting horizontal planes and the depth feature at different depths corresponding to each column of features in each of the feature maps.

[0020] For each column of features in each feature map under each set horizontal plane, the angles corresponding to the depth features at different depths are subjected to inverse polar coordinate transformation to obtain the three-dimensional coordinate set of the depth features at different depths corresponding to each column of features in each feature map under each set horizontal plane.

[0021] Optionally, generating a bird's-eye view feature map corresponding to each of the set horizontal planes based on the three-dimensional coordinate set of depth features corresponding to each column of features in each of the feature maps at different depths under each of the set horizontal planes includes:

[0022] Based on the three-dimensional coordinate set of the depth features corresponding to each column of features in each feature map at different depths under each set horizontal plane, and the pixel coordinate set corresponding to each column of features in each feature map, a designated matrix corresponding to each set horizontal plane is generated; wherein, the value of the target element in the designated matrix is ​​1, and the value of all other elements in the designated matrix except the target element is 0; the position of the target element is determined by the correspondence between the three-dimensional coordinate set of the depth features corresponding to each column of features in each feature map at different depths and the pixel coordinate set corresponding to each column of features in each feature map;

[0023] The depth features corresponding to each column of features in each of the feature maps at different depths are multiplied by the corresponding specified matrix to obtain each of the bird's-eye view feature maps.

[0024] Optionally, the step of performing target detection processing based on the feature maps of each of the bird's-eye view perspectives to obtain target detection results includes:

[0025] The feature maps of each bird's-eye view are aggregated to obtain voxel features;

[0026] The voxel features are compressed to obtain the target bird's-eye view features;

[0027] The target detection result is obtained by performing target detection processing on the bird's-eye view features of the target.

[0028] Secondly, embodiments of this application provide a target detection device, comprising:

[0029] The acquisition unit is used to acquire feature maps corresponding to the images to be detected captured by each camera of the vehicle;

[0030] The first processing unit is used to perform feature projection processing on each column of features in each feature map to obtain the depth features corresponding to each column of features in each feature map at different depths.

[0031] The first generation unit is used to generate a bird's-eye view feature map corresponding to each set horizontal plane based on the intrinsic and extrinsic parameters of each camera and the depth features corresponding to each column of features in each feature map at different depths.

[0032] The first detection unit is used to perform target detection processing based on the feature maps of each bird's-eye view to obtain target detection results.

[0033] Thirdly, embodiments of this application provide a terminal device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the target detection method as described in any one of the first aspects above.

[0034] Fourthly, embodiments of this application provide a computer-readable storage medium storing a computer program that, when executed by a processor, implements the target detection method as described in any one of the first aspects above.

[0035] Fifthly, embodiments of this application provide a computer program product that, when run on a terminal device, enables the terminal device to execute the target detection method described in any one of the first aspects.

[0036] The beneficial effects of the embodiments in this application compared with the prior art are:

[0037] This application provides a target detection method that acquires feature maps corresponding to images to be detected captured by various cameras on a vehicle; performs feature projection processing on each column of features in each feature map to obtain depth features corresponding to each column of features at different depths; generates bird's-eye view feature maps corresponding to various set horizontal planes based on the intrinsic and extrinsic parameters of each camera and the depth features corresponding to each column of features in each feature map at different depths; and performs target detection processing based on the various bird's-eye view feature maps to obtain target detection results. This method performs feature projection processing on each column of features in each feature map, thereby obtaining rich depth features. Then, by combining the intrinsic and extrinsic parameters of each camera and the aforementioned multiple depth features, it can generate bird's-eye view feature maps corresponding to various set horizontal planes, thus obtaining rich bird's-eye view feature maps and improving the accuracy of target detection. Attached Figure Description

[0038] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0039] Figure 1 This is a flowchart illustrating the implementation of a target detection method provided in an embodiment of this application;

[0040] Figure 2 This is a flowchart illustrating the implementation of a target detection method according to another embodiment of this application;

[0041] Figure 3 This is a flowchart illustrating the implementation of a target detection method provided in another embodiment of this application;

[0042] Figure 4 This is a schematic diagram of the target detection device provided in one embodiment of this application;

[0043] Figure 5 This is a schematic diagram of the structure of a terminal device provided in an embodiment of this application. Detailed Implementation

[0044] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.

[0045] It should be understood that, when used in this application specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or a collection thereof.

[0046] It should also be understood that the term “and / or” as used in this application specification and the appended claims means any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.

[0047] As used in this application specification and the appended claims, the term "if" may be interpreted, depending on the context, as "when," "once," "in response to determination," or "in response to detection." Similarly, the phrase "if determined" or "if detected [the described condition or event]" may be interpreted, depending on the context, as meaning "once determined," "in response to determination," "once detected [the described condition or event]," or "in response to detection [the described condition or event]."

[0048] Furthermore, in the description of this application and the appended claims, the terms "first," "second," "third," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.

[0049] References to "one embodiment" or "some embodiments" as described in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.

[0050] In practical applications, Bird's Eye View (BEV) is a commonly used representation space in the field of autonomous driving, and object detection in the BEV space is a fundamental task in autonomous driving scenarios. However, existing techniques for object detection in the BEV space are insufficient to meet practical needs, reducing the accuracy of object detection. Therefore, embodiments of this application provide an object detection method to improve the accuracy of object detection.

[0051] It should be noted that in all embodiments of this application, in order to achieve target detection in the BEV space, the vehicle is equipped with multiple cameras, and the shooting range of each camera does not completely overlap, so as to achieve complete shooting of the environment around the vehicle.

[0052] In some possible embodiments, the vehicle may be equipped with six cameras: a front-view camera, a left-front-view camera, a right-front-view camera, a rear-view camera, a left-rear-view camera, and a right-rear-view camera. Specifically, the front-view camera is used to capture images of the scene directly in front of the vehicle; the left-front-view camera is used to capture images of the scene to the left front of the vehicle; the right-front-view camera is used to capture images of the scene to the right front of the vehicle; the rear-view camera is used to capture images of the scene directly behind the vehicle; the left-rear-view camera is used to capture images of the scene to the left rear of the vehicle; and the right-rear-view camera is used to capture images of the scene to the right rear of the vehicle.

[0053] Please see Figure 1 , Figure 1 This is a flowchart illustrating the implementation of a target detection method according to an embodiment of this application. In this embodiment, the target detection method is executed by a terminal device. The terminal device includes, but is not limited to, devices such as laptops, desktop computers, and computers.

[0054] like Figure 1 As shown, a target detection method provided in one embodiment of this application may include S101 to S104, which are described in detail below:

[0055] In S101, feature maps corresponding to the images to be detected captured by each camera of the vehicle are obtained.

[0056] In practical applications, when a user needs to perform target detection in the BEV space, they can send a target detection request to the terminal device.

[0057] In this embodiment, the terminal device detecting a target detection request sent by the user can be achieved by detecting a preset operation on the terminal device. The preset operation can be set according to actual needs and is not limited here. For example, the preset operation can be clicking a preset control on the terminal device. Based on this, when the terminal device detects that its preset control has been clicked, it can determine that a preset operation has been detected, i.e., the aforementioned target detection request has been detected.

[0058] After detecting the above target detection request, the terminal device can acquire the images to be detected captured by each camera of the vehicle, and perform feature extraction processing on each image to obtain the feature map corresponding to each image.

[0059] In some possible embodiments, the terminal device can input each image to be detected into a convolutional neural network for processing, that is, perform convolution operations on each image to be detected to obtain the feature map corresponding to each image to be detected.

[0060] It should be noted that the specific format of each feature map is as follows: in, Let B represent the i-th feature map, C represent the batch size of the convolutional neural network, H represent the height of the i-th feature map, and W represent the width of the i-th feature map.

[0061] In S102, feature projection processing is performed on each column of features in each feature map to obtain the depth features at different depths corresponding to each column of features in each feature map.

[0062] In this embodiment, for any given feature map, the terminal device can perform feature projection processing on each column of features in the feature map to obtain the depth features corresponding to each column of features at different depths, i.e., at each depth in a preset depth set. The depth range of the preset depth set can be determined according to actual needs and is not limited here.

[0063] In some possible embodiments, the depth range of the preset depth set can be determined based on the size of the BEV space. For example, assuming the size of the BEV space is 100, the depth range of the preset depth set is [0, 100].

[0064] Based on this, each depth value in the preset depth set can be determined according to the depth range and the preset number of samples. The preset number of samples can be determined according to actual needs and is not limited here. For example, the preset number of samples can be 50.

[0065] Specifically, the i-th depth value = i * (upper limit of the depth range / preset number of samples). For example, assuming the upper limit of the depth range is 100 and the preset number of samples is 50, then the i-th depth value is 2i.

[0066] In this embodiment of the application, for any feature map, the terminal device can use a multilayer perceptron to project the features corresponding to each pixel in each column of the feature map onto each depth in a preset depth set, thereby obtaining the depth features corresponding to each column of features in the feature map at different depths.

[0067] For example, suppose the preset depth set is Feature map A certain column of features is The terminal device can project this column of features onto a preset depth set using a multilayer perceptron to obtain the depth features corresponding to this column of features at different depths: For feature maps After performing the above operations on all columns, the terminal device can obtain the feature map. Each column of features in the image is projected onto the depth features of a preset depth set: F Depth ∈(B,C,D,W). Where D dimension corresponds to the preset depth set. At various depths.

[0068] Based on this, the terminal device can obtain the depth features at different depths corresponding to each column of features in each feature map.

[0069] In S103, based on the intrinsic and extrinsic parameters of each camera and the depth features at different depths corresponding to each column of features in each feature map, a bird's-eye view feature map corresponding to each set horizontal plane is generated.

[0070] In practical applications, the camera's intrinsic parameters, also known as intrinsic parameters or intrinsic parameter matrix, are parameters related to the camera's own characteristics, such as the camera's focal length and pixel size, which can be obtained by calibrating the camera.

[0071] The camera's extrinsic parameters, also known as extrinsic parameters or extrinsic parameter matrix, refer to the camera's parameters in the world coordinate system, such as the camera's position and rotation direction. These can be obtained through camera calibration. The world coordinate system (Xw, Yw, Zw) is the reference system for the target object's position, and the position of the dot can be freely set for computational convenience.

[0072] In this embodiment, since the BEV space is a three-dimensional space, in order to obtain an accurate bird's-eye view feature map, the terminal device can make a planar assumption, that is, set multiple predetermined horizontal planes in the BEV space. The number of predetermined horizontal planes can be determined according to actual needs and is not limited here. For example, the number of predetermined horizontal planes can be four.

[0073] It should be noted that the height (i.e., z) of the horizontal plane varies depending on the setting.

[0074] In this embodiment of the application, after obtaining multiple set vertical planes, the terminal device can generate a bird's-eye view feature map corresponding to each set horizontal plane based on the intrinsic and extrinsic parameters of each camera and the depth features at different depths corresponding to each column of features in each feature map.

[0075] It should be noted that each set horizontal plane corresponds to a bird's-eye view feature map.

[0076] In one embodiment of this application, the terminal device can specifically be configured as follows: Figure 2 Steps S201 to S202, which show how to determine the bird's-eye view feature map corresponding to each set horizontal plane, are described in detail below:

[0077] In S201, based on the intrinsic and extrinsic parameters of each camera and the pixel coordinates of each pixel in each feature map, the relationship between depth and angle corresponding to each column of features in each feature map under each set horizontal plane is calculated.

[0078] In this embodiment, for any given horizontal plane, the terminal device can calculate the relationship between the depth features of each column of features in each feature map at different depths and the given horizontal plane based on the intrinsic and extrinsic parameters of each camera and the pixel coordinates of each pixel in each feature map. That is, the relationship between the depth and angle corresponding to each column of features in each feature map.

[0079] Based on this, the terminal device can obtain the relationship between the depth and angle of each column of features in each feature map under each set horizontal plane.

[0080] In one embodiment of this application, the terminal device may specifically perform step S201 according to the following steps, as detailed below:

[0081] Based on the intrinsic and extrinsic parameters of each camera and the pixel coordinates of each pixel in each feature map, the first coordinate set corresponding to each column of features in each feature map is calculated under each set horizontal plane.

[0082] Polar coordinate transformation is performed on the first coordinate set of each column of features in each of the respective feature maps on each of the respective set horizontal planes to obtain the relationship between the depth and angle of each column of features in each of the respective feature maps under each of the respective set horizontal planes.

[0083] In this embodiment, for any given horizontal plane, the terminal device can calculate the first coordinate set corresponding to each pixel in each column of pixel points in each feature map under each given horizontal plane, based on the intrinsic and extrinsic parameters of each camera and the pixel coordinates of each pixel in each feature map. That is, the first coordinate set corresponding to each column of features in each feature map.

[0084] Based on this, the terminal device can obtain the first coordinate set corresponding to each column of features in each feature map under each set horizontal plane.

[0085] Subsequently, for any given horizontal plane, the terminal device can perform polar coordinate transformation on the first coordinate set of each column of features in each feature map on the given horizontal plane to obtain the relationship between the depth and angle of each column of features in each feature map under the given horizontal plane.

[0086] Based on this, the terminal device can obtain the relationship between the depth and angle of each column of features in each feature map under each set horizontal plane.

[0087] For example, suppose there is a certain horizontal plane Z = z, and the i-th feature map The set of pixel coordinates for a specific column (i.e., the set of pixel coordinates corresponding to a specific column of features) is as follows: Then the terminal device can use the intrinsic and extrinsic parameters of each camera to generate feature maps. The pixel coordinates of each pixel in the image are projected into the BEV space to obtain the feature map. The first coordinate set corresponding to each column of features Because the plane assumption is used, the first coordinate set All points in the coordinate system lie on the same horizontal plane Z = z. At this point, the terminal device can transform the first coordinate set on the Z = z plane to a polar coordinate system, i.e., perform a polar coordinate transformation, thereby obtaining the feature map. The relationship between depth d and angle a for each column of features.

[0088] In this embodiment, after the terminal device obtains the relationship between the depth and angle of each column of features in each feature map under each set horizontal plane, it can fit the relationship between the depth and angle of each column of features in each feature map under each set horizontal plane to obtain the corresponding quadratic polynomial.

[0089] For example, a quadratic polynomial could be: a = m0d i 2 +m1d i +m2. Where, d i Let m0, m1, and m2 represent the i-th depth. m0, m1, and m2 are all constants and can be determined according to actual needs.

[0090] In S202, based on the relationship between the depth and angle of each column of features in each of the feature maps under each of the set horizontal planes, the bird's-eye view feature map corresponding to each of the set horizontal planes is determined.

[0091] In this embodiment, for any given horizontal plane, the terminal device can determine the bird's-eye view feature map corresponding to that horizontal plane based on the relationship between the depth and angle of each column of features in each feature map.

[0092] Specifically, for any given horizontal plane, the terminal device can determine the relationship between the depth features and BEV space coordinates corresponding to each column of features in each feature map at different depths based on the relationship between the depth and angle of each column of features in the given horizontal plane, thereby obtaining the bird's-eye view feature map corresponding to the given horizontal plane.

[0093] Based on this, the terminal device can obtain bird's-eye view feature maps corresponding to each set horizontal plane.

[0094] In one embodiment of this application, the terminal device can specifically be configured as follows: Figure 3 Steps S301 to S302, as shown, yield bird's-eye view feature maps corresponding to each set horizontal plane, detailed below:

[0095] In S301, based on the relationship between depth and angle of each column of features in each feature map under each set horizontal plane, the three-dimensional coordinate set of depth features at different depths corresponding to each column of features in each feature map under each set horizontal plane is determined.

[0096] In this embodiment, for any given horizontal plane, the terminal device can determine the three-dimensional coordinate set of the depth features corresponding to each column of features in each feature map at different depths under the given horizontal plane based on the relationship between the depth and angle of each column of features in each feature map in the given horizontal plane.

[0097] Specifically, for any given horizontal plane, the terminal device can determine the three-dimensional coordinate set of the depth features corresponding to each column of features in each feature map at different depths based on the depth features at different depths in each feature map and the relationship between the depth and angle of each column of features in each feature map in the given horizontal plane.

[0098] Based on this, the terminal device can obtain the three-dimensional coordinate set of the depth features corresponding to each column of features in each feature map at different depths under each set horizontal plane.

[0099] In one embodiment of this application, the terminal device may specifically perform step S301 according to the following steps, as detailed below:

[0100] Based on the relationship between the depth and angle of each column of features in each feature map under each set horizontal plane, and the depth features of each column of features in each feature map at different depths, the angles corresponding to the depth features of each column of features in each feature map at different depths under each set horizontal plane are calculated.

[0101] For each column of features in each feature map under each set horizontal plane, the angles corresponding to the depth features at different depths are subjected to inverse polar coordinate transformation to obtain the three-dimensional coordinate set of the depth features at different depths corresponding to each column of features in each feature map under each set horizontal plane.

[0102] In this embodiment, for any given horizontal plane, the terminal device can calculate the angle corresponding to the depth feature at different depths of each column of features in each feature map under the given horizontal plane, based on the relationship between the depth and angle of each column of features in each feature map under the given horizontal plane, and the depth features of each column of features in each feature map at different depths.

[0103] Based on this, the terminal device can obtain the angles corresponding to the depth features of each column of features in each feature map at different depths under each set horizontal plane.

[0104] Subsequently, for any given horizontal plane, the terminal device can perform polar coordinate inverse transformation on the angles corresponding to the depth features at different depths for each column of features in each feature map under that given horizontal plane, to obtain the three-dimensional coordinate set of the depth features at different depths for each column of features in each feature map under that given horizontal plane.

[0105] Based on this, the terminal device can obtain the three-dimensional coordinate set of the depth features corresponding to each column of features in each feature map at different depths under each set horizontal plane.

[0106] For example, suppose we set the horizontal plane Z = z, and the i-th feature map Projecting a column of features onto a preset depth set The obtained depth features at different depths are: Then the terminal device can determine the i-th feature map based on the set horizontal plane Z=z. The relationship between depth and angle for each column of features, and the i-th feature map. Each column of features corresponds to a depth feature at a different depth. The i-th feature map is calculated under the set horizontal plane. Each column of features corresponds to the set of angles A = {a0, ..., a0} at different depths. D-1 Afterwards, the terminal device can perform an inverse polar coordinate transformation on the above angle set, that is, transform it to the Cartesian coordinate system, to obtain the three-dimensional coordinate set of the depth features at different depths corresponding to each column of features in each feature map under the set horizontal plane: Among them, the three-dimensional coordinate set Representative characteristics The corresponding three-dimensional point coordinates in the D dimension (at different depths).

[0107] In S302, a bird's-eye view feature map corresponding to each set of set horizontal planes is generated based on the three-dimensional coordinate set of the depth features corresponding to each column of features in each feature map at different depths under each set horizontal plane.

[0108] In this embodiment, for any given horizontal plane, the terminal device can generate a bird's-eye view feature map corresponding to that horizontal plane based on the three-dimensional coordinate set of the depth features at different depths corresponding to each column of features in each feature map under that horizontal plane.

[0109] Based on this, the terminal device can obtain bird's-eye view feature maps corresponding to each set horizontal plane.

[0110] In one embodiment of this application, in order to improve generation efficiency, the terminal device can generate bird's-eye view feature maps corresponding to each set horizontal plane according to the following steps, detailed below:

[0111] Based on the three-dimensional coordinate set of the depth features corresponding to each column of features in each feature map at different depths under each set horizontal plane, and the pixel coordinate set corresponding to each column of features in each feature map, a designated matrix corresponding to each set horizontal plane is generated; wherein, the value of the target element in the designated matrix is ​​1, and the value of all other elements in the designated matrix except the target element is 0; the position of the target element is determined by the correspondence between the three-dimensional coordinate set of the depth features corresponding to each column of features in each feature map at different depths and the pixel coordinate set corresponding to each column of features in each feature map;

[0112] The depth features corresponding to each column of features in each of the aforementioned feature maps at different depths are multiplied by the corresponding specified matrix to obtain the respective bird's-eye view feature maps.

[0113] In this embodiment, for any given horizontal plane, the terminal device can generate a specified matrix corresponding to that horizontal plane based on the three-dimensional coordinate set of the depth features at different depths corresponding to each column of features in each feature map under that specified horizontal plane, and the pixel coordinate set corresponding to each column of features in each feature map. The target element in the specified matrix has a value of 1, and all other elements in the specified matrix have a value of 0.

[0114] It should be noted that the position of the target element is determined by the correspondence between the three-dimensional coordinate set of the depth features corresponding to each column of features in each feature map at different depths under the aforementioned set horizontal plane, and the pixel coordinate set corresponding to each column of features in each feature map.

[0115] Based on this, the terminal device can generate the specified matrix corresponding to each set horizontal plane.

[0116] Subsequently, for any given horizontal plane, the terminal device can multiply the specified matrix corresponding to the given horizontal plane with the depth features corresponding to each column of features in each feature map at different depths, thereby indexing the depth features corresponding to each column of features in each feature map at different depths, and thus obtaining the bird's-eye view feature map corresponding to the given horizontal plane.

[0117] Based on this, the terminal device can obtain the bird's-eye view feature map corresponding to each set horizontal plane.

[0118] In S104, target detection processing is performed based on the various bird's-eye view feature maps to obtain target detection results.

[0119] In this embodiment of the application, after obtaining the bird's-eye view feature maps corresponding to each set horizontal plane in the BEV space, the terminal device can integrate the various bird's-eye view feature maps to obtain the final target bird's-eye view feature.

[0120] Afterwards, the terminal device can perform target detection on the above-mentioned bird's-eye view features and obtain the target detection results.

[0121] The target detection results include, but are not limited to, lane lines, obstacles, and pedestrians.

[0122] In one embodiment of this application, the terminal device may obtain the target detection result according to the following steps, detailed below:

[0123] The feature maps of each bird's-eye view are aggregated to obtain voxel features;

[0124] The voxel features are compressed to obtain the target bird's-eye view features;

[0125] The target detection result is obtained by performing target detection processing on the bird's-eye view features of the target.

[0126] In this embodiment, the terminal device can aggregate the feature maps of each bird's-eye view to obtain the initial bird's-eye view features, i.e., voxel features.

[0127] Afterwards, the terminal device can use the Space-to-Channel (S2C) module in the Unified BEV (M2BEV) technology to compress the above voxel features to obtain the target bird's-eye view features.

[0128] Finally, the terminal device can perform target detection on the above-mentioned bird's-eye view features to obtain the target detection results.

[0129] As can be seen from the above, the target detection method provided in this application acquires feature maps corresponding to the images to be detected captured by various cameras on a vehicle; performs feature projection processing on each column of features in each feature map to obtain depth features corresponding to each column of features at different depths; generates bird's-eye view feature maps corresponding to each set horizontal plane based on the intrinsic and extrinsic parameters of each camera and the depth features corresponding to each column of features in each feature map at different depths; and performs target detection processing based on each bird's-eye view feature map to obtain the target detection result. This method performs feature projection processing on each column of features in each feature map, thereby obtaining rich depth features. Then, by combining the intrinsic and extrinsic parameters of each camera and the above-mentioned multiple depth features, bird's-eye view feature maps corresponding to each set horizontal plane can be generated, thus obtaining rich bird's-eye view feature maps and improving the accuracy of target detection.

[0130] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0131] Corresponding to the target detection method described in the above embodiments, Figure 4 A schematic diagram of a target detection device according to an embodiment of this application is shown. For ease of explanation, only the parts relevant to the embodiment of this application are shown. (Refer to...) Figure 4 The target detection device 400 includes: a first acquisition unit 41, a first processing unit 42, a first generation unit 43, and a first detection unit 44. Wherein:

[0132] The acquisition unit 41 is used to acquire feature maps corresponding to the images to be detected captured by each camera of the vehicle.

[0133] The first processing unit 42 is used to perform feature projection processing on each column of features in each feature map to obtain the depth features corresponding to each column of features in each feature map at different depths.

[0134] The first generation unit 43 is used to generate a bird's-eye view feature map corresponding to each set horizontal plane based on the intrinsic and extrinsic parameters of each camera and the depth features corresponding to each column of features in each feature map at different depths.

[0135] The first detection unit 44 is used to perform target detection processing based on the various bird's-eye view feature maps to obtain target detection results.

[0136] In one embodiment of this application, the first generation unit 43 specifically includes: a first calculation unit and a first determination unit. Wherein:

[0137] The first calculation unit is used to calculate the relationship between depth and angle for each column of features in each feature map under each set horizontal plane, based on the intrinsic and extrinsic parameters of each camera and the pixel coordinates of each pixel in each feature map.

[0138] The first determining unit is used to determine the bird's-eye view feature map corresponding to each set horizontal plane based on the relationship between the depth and angle of each column of features in each feature map under each set horizontal plane.

[0139] In one embodiment of this application, the first computing unit specifically includes: a second computing unit and a second processing unit. Wherein:

[0140] The second calculation unit is used to calculate, based on the intrinsic and extrinsic parameters of each camera and the pixel coordinates of each pixel in each feature map, the first coordinate set corresponding to each column of features in each feature map under each set horizontal plane.

[0141] The second processing unit is used to perform polar coordinate transformation on the first coordinate set of each column of features in each of the feature maps in each of the set horizontal planes, so as to obtain the relationship between the depth and angle of each column of features in each of the feature maps under each of the set horizontal planes.

[0142] In one embodiment of this application, the first determining unit specifically includes: a second determining unit and a second generating unit. Wherein:

[0143] The second determining unit is used to determine the three-dimensional coordinate set of the depth features corresponding to each column of features in each feature map at different depths under each set horizontal plane, based on the relationship between depth and angle of each column of features in each feature map under each set horizontal plane.

[0144] The second generation unit is used to generate a bird's-eye view feature map corresponding to each set of set horizontal planes based on the three-dimensional coordinate set of the depth features corresponding to each column of features in each feature map at different depths under each set horizontal plane.

[0145] In one embodiment of this application, the second determining unit specifically includes: a second calculation unit and a third processing unit. Wherein:

[0146] The second calculation unit uses a method based on the relationship between the depth and angle of each column of features in each feature map under each set horizontal plane, and the depth features of each column of features in each feature map at different depths, to calculate the angle corresponding to the depth features of each column of features in each feature map at different depths under each set horizontal plane.

[0147] The third processing unit is used to perform polar coordinate inverse transformation on the angles corresponding to the depth features at different depths for each column of features in each feature map under each set horizontal plane, so as to obtain the three-dimensional coordinate set of the depth features at different depths for each column of features in each feature map under each set horizontal plane.

[0148] In one embodiment of this application, the second generation unit specifically includes: a third generation unit and a third determination unit. Wherein:

[0149] The third generation unit is used to generate a designated matrix corresponding to each of the set horizontal planes based on the three-dimensional coordinate set of the depth features corresponding to each column of features in each of the feature maps at different depths under each set horizontal plane, and the pixel coordinate set corresponding to each column of features in each of the feature maps; wherein, the value of the target element in the designated matrix is ​​1, and the value of the other elements in the designated matrix other than the target element is 0; the position of the target element is determined by the correspondence between the three-dimensional coordinate set of the depth features corresponding to each column of features in each of the feature maps at different depths and the pixel coordinate set corresponding to each column of features in each of the feature maps.

[0150] The third determining unit is used to multiply the depth features corresponding to each column of features in each of the feature maps at different depths with the corresponding specified matrix to obtain each of the bird's-eye view feature maps.

[0151] In one embodiment of this application, the first detection unit 44 specifically includes: a fourth processing unit, a fifth processing unit, and a second detection unit. Wherein:

[0152] The fourth processing unit is used to aggregate the various bird's-eye view feature maps to obtain voxel features.

[0153] The fifth processing unit is used to compress the voxel features to obtain the target bird's-eye view features.

[0154] The second detection unit is used to perform target detection processing on the bird's-eye view features of the target to obtain the target detection result.

[0155] It should be noted that the information interaction and execution process between the above-mentioned devices / units are based on the same concept as the method embodiments of this application. For details on their specific functions and technical effects, please refer to the method embodiments section, and they will not be repeated here.

[0156] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is merely an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this application. The specific working process of the units and modules in the above system can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0157] Figure 5 This is a schematic diagram of the structure of a terminal device provided in an embodiment of this application. Figure 5 As shown, the terminal device 5 in this embodiment includes: at least one processor 50 ( Figure 5 (Only one is shown) a processor, a memory 51, and a computer program 52 stored in the memory 51 and executable on the at least one processor 50, which, when executing the computer program 52, implements the steps in any of the above-described target detection method embodiments.

[0158] The terminal device may include, but is not limited to, a processor 50 and a memory 51. Those skilled in the art will understand that... Figure 5 This is merely an example of terminal device 5 and does not constitute a limitation on terminal device 5. It may include more or fewer components than shown in the figure, or combine certain components, or different components, such as input / output devices, network access devices, etc.

[0159] The processor 50 may be a Central Processing Unit (CPU), or it may be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor.

[0160] In some embodiments, the memory 51 may be an internal storage unit of the terminal device 5, such as the RAM of the terminal device 5. In other embodiments, the memory 51 may be an external storage device of the terminal device 5, such as a plug-in hard drive, Smart Media Card (SMC), Secure Digital (SD) card, or Flash Card equipped on the terminal device 5. Furthermore, the memory 51 may include both internal and external storage units of the terminal device 5. The memory 51 is used to store the operating system, applications, bootloader, data, and other programs, such as the program code of the computer program. The memory 51 can also be used to temporarily store data that has been output or will be output.

[0161] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps described in the various method embodiments above.

[0162] This application provides a computer program product that, when run on a terminal device, enables the terminal device to implement the steps described in the various method embodiments above.

[0163] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments of this application can be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include at least: any entity or device capable of carrying computer program code to a terminal device, a recording medium, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium. Examples include USB flash drives, portable hard drives, magnetic disks, or optical disks. In some jurisdictions, according to legislation and patent practice, computer-readable media cannot be electrical carrier signals or telecommunication signals.

[0164] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0165] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.

Claims

1. A target detection method, characterized in that, include: Obtain feature maps corresponding to the images to be detected captured by each camera of the vehicle; Perform feature projection processing on each column of features in each feature map to obtain the depth features at different depths corresponding to each column of features in each feature map. Based on the intrinsic and extrinsic parameters of each camera and the depth features corresponding to each column of features in each feature map at different depths, a bird's-eye view feature map corresponding to each set horizontal plane is generated. Target detection is performed based on the feature maps of each bird's-eye view to obtain the target detection results.

2. The target detection method as described in claim 1, characterized in that, The step of generating a bird's-eye view feature map corresponding to each set horizontal plane based on the intrinsic and extrinsic parameters of each camera and the depth features corresponding to each column of features in each feature map at different depths includes: Based on the intrinsic and extrinsic parameters of each camera and the pixel coordinates of each pixel in each feature map, the relationship between depth and angle corresponding to each column of features in each feature map is calculated under each set horizontal plane. Based on the relationship between depth and angle of each column of features in each feature map under each set horizontal plane, the bird's-eye view feature map corresponding to each set horizontal plane is determined.

3. The target detection method as described in claim 2, characterized in that, The step of calculating the relationship between depth and angle for each column of features in each feature map under each set horizontal plane, based on the intrinsic and extrinsic parameters of each camera and the pixel coordinates of each pixel in each feature map, includes: Based on the intrinsic and extrinsic parameters of each camera and the pixel coordinates of each pixel in each feature map, the first coordinate set corresponding to each column of features in each feature map is calculated under each set horizontal plane. Polar coordinate transformation is performed on the first coordinate set of each column of features in each of the respective feature maps on each of the respective set horizontal planes to obtain the relationship between the depth and angle of each column of features in each of the respective feature maps under each of the respective set horizontal planes.

4. The target detection method as described in claim 2, characterized in that, The step of determining the bird's-eye view feature map corresponding to each set horizontal plane based on the relationship between depth and angle of each column of features in each feature map under each set horizontal plane includes: Based on the relationship between depth and angle of each column of features in each feature map under each set horizontal plane, determine the three-dimensional coordinate set of depth features at different depths corresponding to each column of features in each feature map under each set horizontal plane; Based on the three-dimensional coordinate set of the depth features corresponding to each column of features in each feature map at different depths under each set horizontal plane, a bird's-eye view feature map corresponding to each set horizontal plane is generated.

5. The target detection method as described in claim 4, characterized in that, The step of determining the three-dimensional coordinate set of depth features at different depths corresponding to each column of features in each feature map under each set horizontal plane, based on the relationship between depth and angle for each column of features in each feature map under each set horizontal plane, includes: Based on the relationship between the depth and angle of each column of features in each feature map under each set horizontal plane, and the depth features of each column of features in each feature map at different depths, the angles corresponding to the depth features of each column of features in each feature map at different depths under each set horizontal plane are calculated. For each column of features in each feature map under each set horizontal plane, the angles corresponding to the depth features at different depths are subjected to inverse polar coordinate transformation to obtain the three-dimensional coordinate set of the depth features at different depths corresponding to each column of features in each feature map under each set horizontal plane.

6. The target detection method as described in claim 4, characterized in that, The step of generating a bird's-eye view feature map corresponding to each of the set horizontal planes based on the three-dimensional coordinate set of the depth features corresponding to each column of features in each of the feature maps at different depths includes: Based on the three-dimensional coordinate set of the depth features corresponding to each column of features in each feature map at different depths under each set horizontal plane, and the pixel coordinate set corresponding to each column of features in each feature map, a designated matrix corresponding to each set horizontal plane is generated; wherein, the value of the target element in the designated matrix is ​​1, and the value of all other elements in the designated matrix except the target element is 0; the position of the target element is determined by the correspondence between the three-dimensional coordinate set of the depth features corresponding to each column of features in each feature map at different depths and the pixel coordinate set corresponding to each column of features in each feature map; The depth features corresponding to each column of features in each of the feature maps at different depths are multiplied by the corresponding specified matrix to obtain each of the bird's-eye view feature maps.

7. The target detection method according to any one of claims 1-6, characterized in that, The target detection process based on the feature maps of each bird's-eye view to obtain the target detection result includes: The feature maps of each bird's-eye view are aggregated to obtain voxel features; The voxel features are compressed to obtain the target bird's-eye view features; The target detection result is obtained by performing target detection processing on the bird's-eye view features of the target.

8. A target detection device, characterized in that, include: The acquisition unit is used to acquire feature maps corresponding to the images to be detected captured by each camera of the vehicle; The first processing unit is used to perform feature projection processing on each column of features in each feature map to obtain the depth features corresponding to each column of features in each feature map at different depths. The first generation unit is used to generate a bird's-eye view feature map corresponding to each set horizontal plane based on the intrinsic and extrinsic parameters of each camera and the depth features corresponding to each column of features in each feature map at different depths. The first detection unit is used to perform target detection processing based on the feature maps of each bird's-eye view to obtain target detection results.

9. A terminal device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the target detection method as described in any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the target detection method as described in any one of claims 1 to 7.