3D target detection method and device based on multi-view fusion

The 3D target detection method based on multi-view fusion addresses the inefficiency of conventional methods by performing feature fusion and 3D target detection in an end-to-end process, resulting in improved detection efficiency and accuracy for autonomous driving applications.

JP2025517403AActive Publication Date: 2025-06-05BEIJING HORIZON ROBOTICS TECH RES & DEV CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2024568637
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-05-18
Filing Date
2023-02-08
Publication Date
2025-06-05
Estimated Expiration
2043-02-08

AI Technical Summary

Technical Problem

Conventional 3D detection methods for autonomous driving require performing 3D detection on each image collected by the autonomous driving carrier and then fusing the results, leading to low detection efficiency.

Method used

A 3D target detection method and apparatus based on multi-view fusion, which involves acquiring images from multiple camera viewpoints, extracting features, mapping them to a bird's-eye viewpoint space, performing feature fusion, and predicting target objects' three-dimensional spatial information.

Benefits of technology

This method improves detection efficiency by performing multi-view feature fusion followed by 3D target detection, completing 3D target detection in an end-to-end manner without the need for post-processing, thereby enhancing the accuracy and speed of 3D target detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025517403000001_ABST
    Figure 2025517403000001_ABST
Patent Text Reader

Abstract

The embodiment of the present disclosure discloses a 3D target detection method and device based on multi-view fusion. In the method, feature extraction is performed on at least one image of multi-camera viewpoints collected by a multi-camera system, and feature data including target object features in the extracted multi-camera viewpoint space are mapped to the same bird's-eye viewpoint space based on the internal parameters of the multi-camera system and the vehicle parameters, and corresponding feature data in the bird's-eye viewpoint space of at least one image are obtained, and bird's-eye viewpoint fusion features are obtained by feature fusion. Target prediction is performed on the target object in the bird's-eye viewpoint fusion feature to obtain the 3D space information of the target object. When performing 3D target detection based on multi-view fusion using the solution of the embodiment of the present disclosure, first perform feature fusion of multi-views, and then perform 3D target detection, thereby completing 3D detection of the scenario object in the bird's-eye viewpoint in an end-to-end manner, and improving the detection efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This disclosure claims priority to a Chinese patent application filed on May 18, 2022, bearing application number 202210544237.0 and titled "3D target detection method and apparatus based on multi-view fusion," the entire contents of which are incorporated herein by reference.

[0002] The present disclosure relates to the field of computer vision, and in particular to a 3D target detection method and apparatus based on multi-view fusion. [Background technology]

[0003] With the development of science and technology, the application of autonomous driving technology in people's lives is becoming more and more widespread. The autonomous driving carrier can perform 3D detection on target objects (vehicles, pedestrians, lidars, etc.) within a certain distance of the surroundings to obtain the three-dimensional spatial information of the target object. Based on the three-dimensional spatial information of the target object, it can perform distance measurement and speed measurement for the target object to achieve better driving control.

[0004] Currently, an autonomous driving carrier can collect multiple images with different perspectives, then perform 3D detection on each image respectively, and finally fuse the 3D detection results of each image to generate three-dimensional spatial information of target objects in the environment surrounding the carrier. Summary of the Invention [Problem to be solved by the invention]

[0005] In the conventional technical solutions, 3D detection needs to be performed on each image collected by the autonomous driving carrier, and then the 3D detection results of each image need to be fused to obtain other vehicle information in the 360-degree environment around the carrier, resulting in low detection efficiency. [Means for solving the problem]

[0006] In order to solve the above technical problems, the present disclosure proposes a 3D target detection method and apparatus based on multi-view fusion.

[0007] According to one aspect of the present disclosure, acquiring at least one image from a collection of multiple camera viewpoints; performing feature extraction on the at least one image to obtain feature data including corresponding target object features in a multi-camera viewpoint space of the at least one image; Mapping the corresponding feature data in the multi-camera viewpoint space of the at least one image into a same bird's-eye viewpoint space according to the internal parameters of the multi-camera system and vehicle parameters, and obtaining the corresponding feature data in the bird's-eye viewpoint space of the at least one image; performing feature fusion on the corresponding feature data in the bird's-eye view space of the at least one image to obtain a bird's-eye view fusion feature; performing target prediction for the target object in the bird's-eye view fusion feature to obtain three-dimensional spatial information of the target object.

[0008] According to another aspect of the present disclosure, an image receiving module for acquiring at least one image from the collected multi-camera viewpoints; a feature extraction module for performing feature extraction on the at least one image acquired by the image receiving module to obtain feature data including corresponding target object features in a multi-camera viewpoint space of the at least one image; an image feature mapping module for mapping corresponding feature data in the multi-camera viewpoint space of the at least one image acquired by the feature extraction module into the same bird's-eye viewpoint space according to internal parameters of the multi-camera system and vehicle parameters, thereby obtaining corresponding feature data in the bird's-eye viewpoint space of the at least one image; an image fusion module for performing feature fusion on the corresponding feature data in the bird's-eye view space of the at least one image obtained by the image feature mapping module to obtain a bird's-eye view fusion feature; and a 3D detection module for performing target prediction on the target object in the bird's-eye view fusion features obtained by the image fusion module, and obtaining three-dimensional spatial information of the target object.

[0009] According to yet another aspect of the present disclosure, there is provided a computer-readable storage medium having a computer program stored therein, the computer program being used to execute the above-mentioned 3D target detection method based on multi-view fusion.

[0010] According to yet another aspect of the present disclosure, there is provided an electronic device, the electronic device comprising: A processor; a memory for storing instructions executable by said processor; The processor is used to read the executable instructions from the memory and execute the instructions to realize the above multi-view fusion based 3D target detection method. Effect of the Invention

[0011] According to the 3D target detection method and device based on multi-view fusion provided in the above embodiment of the present disclosure, feature extraction is performed on at least one image of multi-camera viewpoints collected by a multi-camera system, and feature data including target object features in the extracted multi-camera viewpoint space are mapped to the same bird's-eye viewpoint space based on the internal parameters of the multi-camera system, and corresponding feature data in the bird's-eye viewpoint space of at least one image are obtained, and corresponding feature data in the bird's-eye viewpoint space of at least one image are feature-fused to obtain bird's-eye viewpoint fusion features. Furthermore, target prediction is performed on the target object in the bird's-eye viewpoint fusion feature to obtain three-dimensional space information of the target object. When performing 3D target detection based on multi-view fusion using the solution of the embodiment of the present disclosure, first perform feature fusion of multi-views and then perform 3D target detection, and complete 3D target detection of the scenario object in the bird's-eye viewpoint in an end-to-end manner, avoiding the post-processing step in normal multi-view 3D detection, and improving detection efficiency. [Brief description of the drawings]

[0012] The above and other objects, features and advantages of the present disclosure will become more apparent by describing the embodiments of the present disclosure in more detail with reference to the drawings. The drawings are used to provide a further understanding of the embodiments of the present disclosure, constitute a part of the specification, and are used to interpret the present disclosure together with the embodiments of the present disclosure, and are not intended to constitute limitations on the present disclosure. In the drawings, the same reference numerals generally indicate the same components or steps. [Figure 1] FIG. 1 is a scenario diagram to which the present disclosure is applicable. [Diagram 2] FIG. 1 is a system block diagram of an in-vehicle autonomous driving system provided in an embodiment of the present disclosure. [Diagram 3] 4 is a flowchart of a 3D target detection method based on multi-view fusion provided in an exemplary embodiment of the present disclosure. [Figure 4] FIG. 2 is a block diagram illustrating a multi-camera system provided in one exemplary embodiment of the present disclosure for collecting images. [Diagram 5] 1 is a schematic diagram of images from multiple camera viewpoints provided in one exemplary embodiment of the present disclosure; [Figure 6] FIG. 2 is a block schematic diagram of a feature extraction provided in one exemplary embodiment of the present disclosure. [Figure 7] FIG. 2 is a schematic diagram illustrating generating a bird's-eye view image from images collected by a multi-camera system provided in one exemplary embodiment of the present disclosure. [Figure 8] FIG. 2 is a block schematic diagram of a target detection provided in an exemplary embodiment of the present disclosure. [Figure 9] 11 is a flowchart for determining feature data in a bird's-eye view space provided in one exemplary embodiment of the present disclosure. [Figure 10] FIG. 2 is a block schematic diagram of performing step S303 and step S304 provided in an exemplary embodiment of the present disclosure. [Figure 11] 1 is a flowchart of target detection provided in an exemplary embodiment of the present disclosure. [Figure 12] FIG. 2 is a schematic diagram of an output result of a prediction network provided in an exemplary embodiment of the present disclosure. [Figure 13] 11 is another flowchart of target detection provided in an exemplary embodiment of the present disclosure. [Figure 14] FIG. 2 is a schematic diagram of a Gaussian kernel provided in one exemplary embodiment of the present disclosure; [Figure 15] FIG. 2 is a schematic diagram of a heat map provided in one exemplary embodiment of the present disclosure. [Figure 16] 11 is another flowchart of target detection provided in an exemplary embodiment of the present disclosure. [Figure 17] FIG. 2 is a structural diagram of a 3D target detection device based on multi-view fusion provided in an exemplary embodiment of the present disclosure. [Figure 18] FIG. 2 is another structural diagram of a 3D target detection device based on multi-view fusion provided in an exemplary embodiment of the present disclosure. [Figure 19]FIG. 2 is a block diagram of an electronic device provided in an exemplary embodiment of the present disclosure. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0013] Hereinafter, exemplary embodiments based on the present disclosure will be described in detail with reference to the drawings. Obviously, the described embodiments are only some embodiments of the present disclosure, and are not all embodiments of the present disclosure, and it should be understood that the present disclosure is not limited to the exemplary embodiments described herein. Application Summary

[0014] In order to ensure safety during autonomous driving, the autonomous driving carrier can detect target objects (e.g., vehicles, pedestrians, lidars, etc.) within a certain distance around the carrier in real time to obtain three-dimensional spatial information (e.g., attributes such as position, size, orientation angle and category) of the 3D target objects. Based on the three-dimensional spatial information of the target objects, distance measurement and speed measurement are performed on the target objects to realize better driving control. Here, the autonomous driving carrier may be a vehicle, an airplane, etc.

[0015] The autonomous driving carrier can use a multi-camera system to collect multiple images with different viewpoints, and then perform 3D target detection on each image respectively, for example, filtering, de-duplication and other operations on the target objects on the multiple images collected by the cameras with different viewpoints respectively. Finally, the 3D detection results of each image are fused to generate three-dimensional spatial information of the target object in the environment surrounding the carrier. As can be seen, the conventional technical solution requires 3D detection on each image collected by the autonomous driving carrier respectively, and then fusion of the 3D detection results of each image, resulting in low detection efficiency.

[0016] In view of this, the embodiment of the present disclosure provides a 3D target detection method and device based on multi-view fusion. When 3D target detection is performed by the solution of the present disclosure, the autonomous driving carrier can perform feature extraction on at least one image of multi-camera viewpoints collected by a multi-camera system to obtain feature data in a multi-camera viewpoint space, the feature data including target object features. According to the internal parameters of the multi-camera system and the vehicle parameters, the feature data in the multi-camera viewpoint space is mapped to the same bird's-eye viewpoint space to obtain corresponding feature data in the bird's-eye viewpoint space of at least one image. Then, the corresponding feature data in the bird's-eye viewpoint space of the at least one image is feature-fused to obtain a bird's-eye viewpoint fusion feature. Target prediction is performed on the target object in the bird's-eye viewpoint fusion feature to obtain three-dimensional space information of the target object in the environment surrounding the carrier.

[0017] In the solution of the embodiment of the present disclosure, when performing 3D target detection based on multi-view fusion, the feature data of the multi-camera views of at least one image are simultaneously mapped into the same bird's-eye view space, so that more reasonable and better fusion can be performed. In addition, the fused bird's-eye view fusion features are directly detected in the bird's-eye view space for 3D spatial information of each target object in the surroundings of the vehicle environment. Therefore, when performing 3D target detection based on multi-view fusion using the solution of the embodiment of the present disclosure, multi-view feature fusion is first performed, and then 3D target detection is performed, so that the 3D target detection of the scenario object in the bird's-eye view is completed in an end-to-end manner, which avoids the post-processing step in the normal multi-view 3D target detection, and improves the detection efficiency. Exemplary System

[0018] The embodiments of the present disclosure can be applied to application scenarios where 3D target detection needs to be performed, such as autonomous driving application scenarios.

[0019] For example, in an autonomous driving application scenario, a multi-camera system is deployed on an autonomous driving carrier (hereinafter referred to as "carrier"), and images from different perspectives are collected by the multi-camera system, and then three-dimensional spatial information of target objects in the environment around the carrier is obtained by 3D target detection based on multi-view fusion according to the solution of the embodiment of the present disclosure.

[0020] FIG. 1 is a diagram of a scenario to which the present disclosure is applied.

[0021] As shown in FIG. 1, an embodiment of the present disclosure is applied to a driving assistance or autonomous driving application scenario, in which a driving assistance or autonomous driving carrier 100 is provided with an on-board autonomous driving system 200 and a multi-camera system 300, and the on-board autonomous driving system 200 is electrically connected to the multi-camera system 300. The multi-camera system 300 is used to collect images of the environment surrounding the carrier, and the on-board autonomous driving system 200 is used to obtain the images collected by the multi-camera system 300, perform 3D target detection based on multi-view fusion, and obtain three-dimensional spatial information of the target object in the environment surrounding the carrier.

[0022] FIG. 2 is a system block diagram of an in-vehicle autonomous driving system provided in an embodiment of the present disclosure.

[0023] As shown in FIG. 2, the vehicle-mounted automatic driving system 200 includes an image receiving module 201, a feature extraction module 202, an image feature mapping module 203, an image fusion module 204, and a 3D detection module 205. The image receiving module 201 is used to acquire at least one image collected by the multi-camera system 300. The feature extraction module 202 is used to perform feature extraction on the at least one image acquired by the image receiving module 201 to obtain feature data. The image feature mapping module 203 is used to map the feature data of the at least one image from the multi-camera viewpoint space to the same bird's-eye viewpoint space. The image fusion module 204 is used to perform feature fusion on the corresponding feature data in the bird's-eye viewpoint space of the at least one image to obtain a bird's-eye viewpoint fusion feature. The 3D detection module 205 is used to perform target prediction on the target object in the bird's-eye viewpoint fusion feature acquired by the image fusion module 204 to obtain three-dimensional spatial information of the target object in the surrounding environment of the carrier.

[0024] The multi-camera system 300 includes multiple cameras with different viewpoints, each camera is used to collect an environmental image of one viewpoint, and the multiple cameras cover a 360-degree environmental range around the carrier. Each camera defines its own camera viewpoint coordinate system, which forms a respective camera viewpoint space, and the environmental image collected by each camera is an image in the corresponding camera viewpoint space. Exemplary Methods

[0025] FIG. 3 is a flowchart of a 3D target detection method based on multi-view fusion provided in an exemplary embodiment of the present disclosure.

[0026] This embodiment can be applied to an in-vehicle automatic driving system 200, and as shown in FIG. 3, includes the following steps.

[0027] In step S301, at least one image from multiple camera viewpoints is acquired.

[0028] Here, the at least one image may be collected by at least one camera of the multi-camera system. Exemplarily, the at least one image may be collected in real time by the multi-camera system, or may be collected in advance by the multi-camera system.

[0029] FIG. 4 is a block diagram illustrating a multi-camera system provided in one exemplary embodiment of the present disclosure for collecting images.

[0030] 4, in one embodiment, the multi-camera system can collect multiple images from different viewpoints in real time, for example, images 1, 2...N, and transmit the collected images to the in-vehicle autonomous driving system in real time. In this way, the images acquired by the in-vehicle autonomous driving system can characterize the actual state of the environment around the carrier at the current time.

[0031] FIG. 5 is a schematic diagram of images from multiple camera viewpoints provided in one exemplary embodiment of the present disclosure.

[0032] As shown in (1) to (6) of FIG. 5, in one embodiment, the multi-camera system can include six cameras. The six cameras are installed at the front end, the left front end, the right front end, the rear end, the left rear end, and the right rear end of the carrier, respectively. In this way, at any one time, the multi-camera system can capture six different viewpoint images, e.g., a front view image (I front ), left front view image (I frontleft ), right front view image (I frontright ), backsight image (I rear ), left rear view image (I rearleft ) and right rear view image (I rearright ) can be collected.

[0033] Here, each image includes target objects of each category such as, but not limited to, roads, traffic lights, road signs, vehicles (small cars, buses, trucks, etc.), pedestrians, riders, etc. Depending on the difference in the category position of the target object in the surrounding environment of the carrier, the category and position of the target object included in each image are different.

[0034] In step S302, feature extraction is performed on at least one image to obtain feature data including target object features corresponding to the at least one image in the multi-camera viewpoint space.

[0035] In one embodiment, the vehicle autonomous driving system can respectively extract feature data in a corresponding camera viewpoint space from each image, which may include target object features for indicating a target object in the image, including but not limited to image texture information, edge contour information, semantic information, etc.

[0036] Here, the image texture information is used to characterize the image texture of the target object, the edge contour information is used to characterize the edge contour of the target object, and the semantic information is used to characterize the category of the target object, including but not limited to roads, traffic lights, road signs, vehicles (such as small cars, buses, trucks, etc.), pedestrians, lidars, etc.

[0037] FIG. 6 is a block schematic diagram of a feature extraction provided in one exemplary embodiment of the present disclosure.

[0038] As shown in Figure 6, the in-vehicle autonomous driving system can employ a neural network to perform feature extraction for at least one image (image 1-N) and obtain corresponding feature data 1-N in the multi-camera viewpoint space for each image.

[0039] For example, an in-vehicle autonomous driving system uses forward-looking images (I front ) and extract features from the front view image (I front ) feature data f front The left front view image (I frontleft ) and extract features from the left front view image (I frontleft ) feature data f frontleft The right front view image (I frontright ) and extract features from the right front view image (I frontright ) feature data f frontright The backsight image (I rear ) and extract features from the backsight image (I rear ) feature data f rear The left rear view image (I rearleft ) and extract features from the left rear view image (I rearleft ) feature data f rearleft The right rear view image (I rearright ) and extract features from the right rear view image (I rearright ) feature data f rearright can be obtained.

[0040] In step S303, based on the internal parameters of the multi-camera system and the vehicle parameters, corresponding feature data in the multi-camera viewpoint space of at least one image are mapped to the same bird's-eye viewpoint space, and corresponding feature data in the bird's-eye viewpoint space of the at least one image are obtained.

[0041] Here, the internal parameters of the multi-camera system include the internal camera parameters and the external camera parameters of each camera, the internal camera parameters being parameters related to the characteristics of the camera itself, such as the focal length of the camera, pixel size, etc., and the external camera parameters being parameters in the world coordinate system, such as the position of the camera, the rotation direction, etc. The vehicle parameters are the transformation matrix from the vehicle coordinate system (VCS) to the bird's-eye view coordinate system (BEV), and the vehicle coordinate system is the coordinate system in which the carrier is located.

[0042] For example, an in-vehicle autonomous driving system uses forward-looking images (I front ) feature data f front are mapped to the same bird's-eye view space, and the front view image (I front ) feature data F front A left front view image (I frontleft ) feature data f frontleft are mapped to the same bird's-eye view space, and the left front view image (I frontleft ) feature data F frontleft A right front view image (I frontright ) feature data f frontright are mapped to the same bird's-eye view space, and the right front view image (I frontright ) feature data F frontright The backsight image (I rear ) feature data f rear are mapped to the same bird's-eye view space, and the back-view image (I rear ) feature data F rear The left rear view image (I rearleft ) feature data f rearleft are mapped to the same bird's-eye view space, and the left rear view image (I rearleft ) feature data F rearleft The right rear view image (I rearright ) feature data f rearrightare mapped to the same bird's-eye view space, and the right rear view image (I rearright ) feature data F rearright Get the.

[0043] In step S304, corresponding feature data in the bird's-eye view space of at least one image are feature-fused to obtain a bird's-eye view fusion feature.

[0044] Here, the bird's-eye view fusion feature is used to characterize feature data in the bird's-eye view space of the target object around the carrier, and the feature data in the bird's-eye view space of the target object may include attributes such as the shape, size, category, orientation angle, relative position, etc. of the target object, but are not limited to them.

[0045] In one embodiment, the vehicle-mounted autonomous driving system can perform additive feature fusion on the corresponding feature data in the bird's-eye view space of at least one image to obtain a bird's-eye view fusion feature. Specifically, it can be expressed by the following formula:

number

[0046] Here, F′ denotes the bird's-eye view fusion feature, and Add denotes the additive feature fusion calculation performed on the corresponding feature data in the bird's-eye view space of at least one image.

[0047] It should be noted that the embodiment of step S304 is not limited thereto, and may perform feature fusion on corresponding feature data in the bird's-eye view space of images from different camera viewpoints, for example, by using methods such as multiplication, overlapping, etc.

[0048] FIG. 7 is a schematic diagram illustrating generating a bird's-eye view image from images collected by a multi-camera system provided in one exemplary embodiment of the present disclosure.

[0049] As shown in Fig. 7, for example, the size of the bird's-eye view image may be the same as the size of at least one image collected by the multi-camera system. The bird's-eye view image can represent three-dimensional spatial information of a target object, and the three-dimensional spatial information includes at least one attribute information of the target object, including, but not limited to, 3D position information (i.e., coordinate information of X-axis, Y-axis, and Z-axis), size information (i.e., length, width, and height information), orientation angle information, etc.

[0050] Here, the coordinate information of the X-axis, Y-axis, and Z-axis refers to the coordinate position (x, y, z) of the target object in the bird's-eye view space, and the origin of the coordinate system of the bird's-eye view space is located at one of the positions such as the chassis of the carrier or the center of the carrier, and the X-axis direction is the front to rear direction, the Y-axis direction is the left to right direction, and the Z-axis direction is the vertical up and down direction. The orientation angle is the angle formed by the front direction or traveling direction of the target object in the bird's-eye view space. For example, when the target object is a moving pedestrian, the orientation angle is the angle formed by the traveling direction of the pedestrian in the bird's-eye view space. When the target object is a stationary vehicle, the orientation angle is the angle formed by the head direction of the vehicle in the bird's-eye view space.

[0051] In addition, since at least one image collected by the multi-camera system may contain target objects of different categories, the bird's-eye view image may contain bird's-eye view fusion features of target objects of different categories.

[0052] In step S305, target prediction is performed on the target object in the bird's-eye view fusion feature to obtain three-dimensional space information of the target object.

[0053] Here, the three-dimensional space information may include at least one of attributes such as the position, size, and orientation angle of the target object in the bird's-eye view coordinate system. The position is the coordinate position (x, y, z) of the target object relative to the carrier in the bird's-eye view space, the size is the length, width, and height (Height, Width, Length) of the target object in the bird's-eye view space, and the orientation angle is the orientation angle (rotation yaw) of the target object in the bird's-eye view space.

[0054] FIG. 8 is a block schematic diagram of a target detection provided in one exemplary embodiment of the present disclosure.

[0055] As shown in FIG. 8, in one embodiment, the in-vehicle autonomous driving system can use one or more prediction networks to perform 3D target prediction for target objects in bird's-eye view fusion features, and obtain three-dimensional spatial information of each target object in the environment surrounding the carrier.

[0056] When an in-vehicle autonomous driving system uses multiple prediction networks to perform 3D target prediction, each prediction network can output one or more attributes of the target object, and the attributes output by different prediction networks are also different.

[0057] In the solution of the embodiment of the present disclosure, when performing 3D target detection based on multi-view fusion, multi-view feature fusion is first performed and then 3D target detection is performed, so that 3D target detection of scenario objects in a bird's-eye view is completed in an end-to-end manner, and the post-processing step in conventional multi-view 3D target detection can be avoided and the detection efficiency can be improved.

[0058] FIG. 9 is a flowchart for determining feature data in a bird's-eye view space provided in one exemplary embodiment of the present disclosure.

[0059] As shown in FIG. 9, based on the embodiment shown in FIG. 3 above, step S303 may include the following steps:

[0060] In step S3031, a transformation matrix from the camera coordinate system of the multi-cameras of the multi-camera system to the bird's-eye view coordinate system is determined based on the internal parameters of the multi-camera system and the vehicle parameters.

[0061] Here, the internal parameters of the multi-camera system include the camera internal parameters and camera external parameters of each camera, the camera external parameters are the transformation matrix from the camera coordinate system of the multi-camera to the vehicle coordinate system, the vehicle parameters are the transformation matrix from the vehicle coordinate system (VCS) to the bird's-eye view coordinate system (BEV), and the vehicle coordinate system is the coordinate system in which the carrier is located.

[0062] In one specific embodiment, step S3031 includes the steps of: A step of acquiring camera internal parameters and camera external parameters of each of the multiple cameras in the multi-camera system, and acquiring a transformation matrix from a vehicle coordinate system to a bird's-eye view coordinate system; The method includes determining a transformation matrix from the camera coordinate system of the multi-camera to the bird's-eye view coordinate system based on the camera external parameters, the camera internal parameters, and a transformation matrix from the vehicle coordinate system to the bird's-eye view coordinate system of the multi-camera.

[0063] In one embodiment, the in-vehicle autonomous driving system can determine a transformation matrix H from the multi-camera camera coordinate system to the bird's-eye view coordinate system using the following formula:

number

[0064] where @ denotes matrix multiplication and T camera→vcs denotes the transformation matrix from the camera coordinate system to the vehicle coordinate system, and T camera→vcs characterizes the camera extrinsic parameters, and T vcs→camera denotes a transformation matrix from the vehicle coordinate system to the bird's-eye view coordinate system, and K denotes a camera internal parameter.

[0065] The camera external parameters, i.e., the transformation matrix from the camera coordinate system to the vehicle coordinate system, can be obtained by calibration of the multi-camera system, and usually does not change once the calibration is completed. The transformation matrix from the vehicle coordinate system to the bird's-eye view coordinate system can be calculated from the artificially set bird's-eye view range (for example, a range surrounded by 100 meters from the front, back, left, and right) and the resolution of the bird's-eye view image (for example, 512 x 512).

[0066] In this way, it is possible to determine the transformation matrix corresponding to each camera in the multi-camera system. For example, the in-vehicle autonomous driving system can determine the transformation matrix H from the camera coordinate system of the front end camera to the bird's-eye view coordinate system based on the transformation matrix from the camera coordinate system of the front end camera to the vehicle coordinate system, the transformation matrix from the vehicle coordinate system to the bird's-eye view coordinate system, and the camera internal parameters of the front end camera. front→bev Based on the transformation matrix from the camera coordinate system of the left front end camera to the vehicle coordinate system, the transformation matrix from the vehicle coordinate system to the bird's-eye view coordinate system, and the camera internal parameters of the left front end camera, a transformation matrix H frontleft→bev Based on the transformation matrix from the camera coordinate system of the right front end camera to the vehicle coordinate system, the transformation matrix from the vehicle coordinate system to the bird's-eye view coordinate system, and the camera internal parameters of the right front end camera, a transformation matrix H frontright→bev Based on the transformation matrix from the camera coordinate system of the rear end camera to the vehicle coordinate system, the transformation matrix from the vehicle coordinate system to the bird's-eye view coordinate system, and the camera internal parameters of the rear end camera, a transformation matrix H rear→bev Based on the transformation matrix from the camera coordinate system of the left rear end camera to the vehicle coordinate system, the transformation matrix from the vehicle coordinate system to the bird's-eye view coordinate system, and the camera internal parameters of the left rear end camera, a transformation matrix H rearleft→bevBased on the transformation matrix from the camera coordinate system of the right rear end camera to the vehicle coordinate system, the transformation matrix from the vehicle coordinate system to the bird's-eye view coordinate system, and the camera internal parameters of the right rear end camera, a transformation matrix H rearright→bev Determine.

[0067] In this embodiment, each camera has its own transformation matrix from the camera viewpoint coordinate system to the bird's-eye view coordinate system, so that the prediction network adopted in the embodiments of the present application for 3D target detection can be applied to a multi-camera system, eliminating the need to train the prediction network from scratch and improving detection efficiency.

[0068] In step S3032, based on a transformation matrix from the multi-camera camera coordinate system to the bird's-eye view coordinate system, corresponding feature data in the multi-camera viewpoint space of at least one image is transformed from the multi-camera viewpoint space to the bird's-eye view space, and corresponding feature data in the bird's-eye view space of the at least one image is obtained.

[0069] In one embodiment, the in-vehicle automatic driving system can obtain corresponding feature data in the bird's-eye view space of at least one image by matrix multiplying the transformation matrix of each camera and the feature data in the viewpoint space of each camera. Specifically, it can be expressed by the following formula.

number

[0070] Here, F is the corresponding feature data F in the bird's-eye view space of at least one image. front , F frontleft , F frontright , F rear , F rearleft and F rearright where H is the transformation matrix H corresponding to each camera in the multi-camera system. front , H frontleft , H frontright , H rear , H rearleftand H rearright where f is feature data f in the multi-camera viewpoint space of at least one image. front , f frontleft , f frontright , f rear , f rearleft and f rearright Shows.

[0071] As can be seen, the embodiments of the present disclosure can not only be applied to multi-camera systems of different models, but also perform more rational feature fusion by calculating respective transformation matrices (homographies) for different cameras in a multi-camera system, and then mapping each feature data into bird's-eye view space based on each transformation matrix of each camera, and obtaining corresponding feature data in the bird's-eye view space of each image.

[0072] Note that the two steps of step S302 and step S3031 may be executed synchronously or asynchronously, and may be determined according to the actual application situation.

[0073] FIG. 10 is a block schematic diagram of performing step S303 and step S304 provided in one exemplary embodiment of the present disclosure.

[0074] 10, after the execution of steps S302 and S3031 is completed, the feature space transformation of step S3032 is performed based on the transformation matrix from the camera coordinate system of each camera to the bird's-eye view coordinate system obtained in step S3031 and the feature data in the corresponding camera viewpoint space obtained in step S302, to obtain feature data in the bird's-eye view space. Finally, step S304 is performed to perform feature fusion on the feature data in the bird's-eye view space of the multi-camera viewpoints, to obtain bird's-eye view fusion features.

[0075] FIG. 11 is a flowchart of target detection provided in one exemplary embodiment of the present disclosure.

[0076] As shown in FIG. 11, based on the embodiment shown in FIG. 3 above, step S305 may include the following steps:

[0077] In step S3051, a prediction network is used to obtain a corresponding heat map from the bird's-eye view fusion features for determining a first preset coordinate value in the bird's-eye view coordinate system of the target object, and obtain other attribute maps for determining a second preset coordinate value, size and orientation angle in the bird's-eye view coordinate system of the target object.

[0078] Here, the prediction network may be a neural network for performing target prediction for a target object. Since it is necessary to perform three-dimensional space information prediction with different attributes for a target object, the prediction network may be of multiple types. Different prediction networks are used to predict three-dimensional space information with different attributes.

[0079] For example, if the attribute to be predicted is a first preset coordinate value of the target object, a prediction network corresponding to the first preset coordinate value can be used to process the bird's-eye view fusion feature in the bird's-eye view image to obtain a heat map, and then the heat map can be used to determine the first preset coordinate value in the bird's-eye view coordinate system of the target object. The size of the heat map can be the same as that of the bird's-eye view image.

[0080] In addition, for example, if the attributes that need to be predicted are the second preset coordinate value, size, and orientation angle of the target object, a prediction network corresponding to the second preset coordinate value, size, and orientation angle can be used to process the bird's-eye view fusion features in the bird's-eye view image to obtain another attribute map, thereby determining the second preset coordinate value, size, and orientation angle in the bird's-eye view coordinate system of the target object using the other attribute map.

[0081] Here, the first preset coordinate value is the (x, y) position in the bird's-eye view coordinate system, the second preset coordinate value is the z position in the bird's-eye view coordinate system, the size is the length, width, and height, and the orientation angle is the orientation angle.

[0082] In step S3052, a first preset coordinate value in the bird's-eye view coordinate system of the target object is determined based on peak information in the heat map, and a second preset coordinate value, size and orientation angle in the bird's-eye view coordinate system of the target object are determined from another attribute map based on the first preset coordinate value in the bird's-eye view coordinate system of the target object.

[0083] Here, the peak information is the center value of a Gaussian kernel, that is, the center point of the target object.

[0084] After predicting the first preset coordinate value in the bird's-eye view space of the target object, other attribute maps can output their respective attribute information using the attribute output results of the heat map, so that the second preset coordinate value, size and orientation angle in the bird's-eye view coordinates of the target object can be predicted from the other attribute maps based on the first preset coordinate value in the bird's-eye view coordinate system of the target object.

[0085] In step S3053, three-dimensional spatial information of the target object is determined based on the first preset coordinate value, the second preset coordinate value, the size and the orientation angle of the target object in the bird's-eye view coordinate system.

[0086] In one embodiment, the vehicle-mounted autonomous driving system can determine the first preset coordinate value and the second preset coordinate value as the (x, y, z) position of the target object in the bird's-eye view space, determine the size as the length, width, and height of the target object in the bird's-eye view space, and determine the orientation angle as the orientation angle of the target object in the bird's-eye view space. Finally, determine the three-dimensional spatial information of the target object in the surrounding environment of the carrier based on the (x, y, z) position, length, width, height, and orientation angle.

[0087] Fig. 12 is a schematic diagram of the output result of the prediction network provided in one exemplary embodiment of the present disclosure. In Fig. 12, the center A of the smallest circle is the position of the carrier, and the block position B around the center is the target object around the carrier.

[0088] In addition, the in-vehicle autonomous driving system can further display a three-dimensional spatial projection of the target object on images from multiple camera viewpoints collected by the multi-camera system, to facilitate the user to intuitively understand the three-dimensional spatial information of the target object from the in-vehicle display.

[0089] As can be seen, the embodiment of the present disclosure can process bird's-eye view images based on a prediction network to obtain heat maps and other attribute maps, and input the bird's-eye view fusion features obtained by feature fusion into the heat maps and other attribute maps to directly predict the 3D spatial information of the target object, thereby improving the efficiency of 3D target detection.

[0090] FIG. 13 is another flowchart of target detection provided in an exemplary embodiment of the present disclosure.

[0091] As shown in FIG. 13, based on the embodiment shown in FIG. 11 above, step S305 may further include the following steps:

[0092] In step S3054, during the training phase of the prediction network, a first loss function is constructed between the heat map output by the prediction network and the truth heat map, and a second loss function is constructed between other attribute maps predicted by the prediction network and other truth attribute maps.

[0093] In one embodiment, the in-vehicle autonomous driving system can construct a Gaussian kernel for each target object based on the position of each target object in the bird's-eye view fusion features.

[0094] Fig. 14 is a schematic diagram of a Gaussian kernel provided in an exemplary embodiment of the present disclosure. As shown in Fig. 14, when constructing a Gaussian kernel, a single N x N Gaussian kernel can be generated centered on the position (i, j) of the target object. Here, the value of the center of the Gaussian kernel is 1, the surrounding values ​​are attenuated downward to 0, and the color from white to black indicates that the value is attenuated from 1 to 0.

[0095] Fig. 15 is a schematic diagram of a heat map provided in an exemplary embodiment of the present disclosure. As shown in Fig. 15, a truth heat map can be obtained by placing the Gaussian kernel of each target object on the heat map. In Fig. 15, each white area represents one Gaussian kernel, i.e., one target object, for example, target objects 1 to 6.

[0096] For generating other truth attribute diagrams, the method for generating a truth heat map can be referred to, and the description thereof will be omitted here.

[0097] After determining the truth heat map, a first loss function can be constructed based on the truth heat map and the heat map output by the prediction network, where the first loss function can determine the distribution of the difference between the output prediction value of the prediction network and the truth value, and is used to monitor the training process of the prediction network.

[0098] In one embodiment, the first loss function Lcls Specifically, it may be constructed as follows:

number

[0099] Here, y' i,j denotes the first preset coordinate value in the truth heat map at the (i,j) position, 1 denotes a peak in the heat map, y i,j denotes the first preset coordinate value in the heat map predicted by the prediction network for position (i,j), α and β are adjustable hyperparameters, and the ranges of α and β are both between 0 and 1, N denotes the total number of target objects in the bird's-eye view fusion feature, and h,w denote the size of the bird's-eye view fusion feature.

[0100] In one embodiment, the second loss function L reg Specifically, it may be constructed as follows:

number

[0101] Here, B' is the true value of the second preset coordinate value, size and orientation angle in the bird's-eye view coordinate system of the target object, B is the predicted value of the second preset coordinate value, size and orientation angle in the bird's-eye view coordinate system of the target object predicted by the prediction network, and N represents the number of target objects in the bird's-eye view fusion feature.

[0102] In step S3055, a total loss function in the training stage of the prediction network is determined based on the first loss function and the second loss function, and the training process of the prediction network is monitored.

[0103] In one embodiment, the total loss function during the training phase of the prediction network is: Obtaining weight values ​​of a first loss function and weight values ​​of a second loss function; determining a total loss function in the training phase of the prediction network based on the first loss function, a weight value of the first loss function, the second loss function, and a weight value of the second loss function.

[0104] In this case, when predicting the 3D spatial information of a target object using a prediction network, different attributes have different importance during the training process, so the importance of the corresponding loss functions is also different. Therefore, different weights are assigned to the loss functions corresponding to different attributes according to the importance of each attribute during the training process.

[0105] Here, the total loss function L during the training phase of the prediction network is 3d may be determined by the following formula:

number

[0106] Here, L cls is the first loss function, L reg is the second loss function, λ 1 is the weight value of the first loss function, and λ 2 is the weight value of the second loss function. 1 and λ 2 are all between 0 and 1, and λ 1 >λ 2 , λ 1 +λ 2 =1.

[0107] As can be seen, when the embodiments of the present disclosure train a prediction network, a total loss function is constructed to monitor the overall training process, thereby ensuring that the outputs of various attributes of the prediction network are more accurate, and further ensuring that the efficiency of 3D target detection is higher.

[0108] FIG. 16 is another flowchart of target detection provided in an exemplary embodiment of the present disclosure.

[0109] As shown in FIG. 16, based on the embodiment shown in FIG. 3 above, step S305 may further include the following steps:

[0110] In step S3056, feature extraction is performed on the bird's-eye view fusion feature using a neural network to obtain bird's-eye view fusion feature data including the target object feature.

[0111] In one embodiment, the vehicle-mounted autonomous driving system can use a neural network to perform calculations such as convolution on the bird's-eye view fusion features to realize feature extraction and obtain bird's-eye view fusion feature data. The bird's-eye view fusion feature data includes target object features for characterizing different dimensions of the target object, i.e., scenario information from different dimensions in the bird's-eye view space of the target object.

[0112] Here, the neural network may be pre-trained and may be a neural network for feature extraction. Optionally, the neural network for feature extraction is not limited to a specific network result, such as resnet, densenet, mobilenet, etc.

[0113] In step S3057, target prediction is performed on the target object in the bird's-eye view fusion feature data including the target object features using the prediction network, and three-dimensional spatial information of the target object is obtained.

[0114] As can be seen, in the embodiment of the present disclosure, before training the bird's-eye view fusion feature by a prediction network, feature extraction is performed on the bird's-eye view fusion feature to obtain bird's-eye view fusion feature data, and then the prediction network is used to predict the bird's-eye view fusion feature data including the target object features, so that the prediction result is more accurate, i.e., the 3D spatial information of the determined target object is more accurate.

[0115] Based on the embodiment shown in FIG. 3, step S302: The step may include a step of using a deep neural network to perform a convolution calculation on an image corresponding to each viewpoint, and acquiring feature data of a plurality of different resolutions in a multi-camera viewpoint space of the image corresponding to each viewpoint, the feature data corresponding to each image corresponding to each viewpoint, respectively, and including target object features.

[0116] Here, the deep neural network may be a pre-trained neural network for feature extraction. Optionally, the neural network for feature extraction is not limited to a specific network result, for example, resnet, densenet, mobilenet, etc. By using the deep neural network to perform calculations such as convolution and pooling on the image of the target viewpoint, feature data of multiple different resolutions (scales) corresponding to the image of the target viewpoint can be obtained.

[0117] For example, the size of image A at a certain viewpoint is H×W×3, where H is the height of image A, W is the width of image A, and 3 indicates that there are three channels. For example, in the case of an RGB image, 3 indicates three channels of RGB (R red, G green, B blue), and in the case of a YUV image, 3 indicates three channels of YUV (Y luminance signal, U blue component signal, V red component signal). Image A is input to a deep neural network, and after performing calculations such as convolution using the deep neural network, a feature matrix of H1×W1×N dimensions is output, where H1 and W1 are the height and width of the feature (generally smaller than H and W, N is the number of channels, and N is larger than 3). By fitting training to the input data by the neural network, it is possible to obtain feature data of the input image with multiple different resolutions including target object features, such as low-level image textures, edge contour information, and high-level semantic information corresponding to different resolutions. After obtaining the feature data of the images from each viewpoint, the subsequent spatial transformation, multi-view feature fusion and target prediction steps can be performed to obtain the 3D spatial information of the target object.

[0118] As can be seen, the embodiment of the present disclosure uses a deep neural network to perform calculations such as convolution and pooling on the images corresponding to each viewpoint to obtain feature data of multiple different resolutions for each viewpoint image, which can better reflect the image features collected by the corresponding viewpoint camera and improve the efficiency of subsequent 3D target detection. Exemplary Apparatus

[0119] 17 is a structural diagram of a 3D target detection device based on multi-view fusion provided in an exemplary embodiment of the present disclosure. The 3D target detection device based on multi-view fusion may be installed in an electronic device such as a terminal device, a server, or a carrier for driving assistance or automatic driving, and may be installed in an in-vehicle automatic driving system to execute the 3D target detection method based on multi-view fusion of any one of the above embodiments of the present disclosure. As shown in FIG. 17, the 3D target detection device based on multi-view fusion of the embodiment includes an image receiving module 201, a feature extraction module 202, an image feature mapping module 203, an image fusion module 204, and a 3D detection module 205.

[0120] Here, the image receiving module 201 is used to obtain at least one image from the collected multi-camera viewpoints.

[0121] The feature extraction module 202 is used to perform feature extraction on the at least one image acquired by the image receiving module, and obtain corresponding feature data in a multi-camera viewpoint space of the at least one image, where the feature data includes target object features.

[0122] The image feature mapping module 203 is used to map corresponding feature data in the multi-camera viewpoint space of the at least one image acquired by the feature extraction module to the same bird's-eye viewpoint space based on the internal parameters of the multi-camera system and vehicle parameters, and obtain corresponding feature data in the bird's-eye viewpoint space of the at least one image.

[0123] The image fusion module 204 is used for feature fusing the corresponding feature data in the bird's-eye view space of the at least one image obtained by the image feature mapping module, to obtain a bird's-eye view fusion feature.

[0124] The 3D detection module 205 is used to perform target prediction on the target object in the bird's-eye view fusion features obtained by the image fusion module, and obtain three-dimensional spatial information of the target object.

[0125] As can be seen, when performing 3D target detection based on multi-view fusion, the device of the embodiment of the present disclosure can perform more reasonable and better fusion by simultaneously mapping feature data from the multi-camera viewpoints of at least one image into the same bird's-eye view space through middle fusion. In addition, the fused bird's-eye view fusion features directly detect the 3D spatial information of each target object in the surroundings of the vehicle environment in the bird's-eye view space. Therefore, when performing 3D target detection based on multi-view fusion using the device of the embodiment of the present disclosure, the 3D detection of the scenario object in the bird's-eye view is completed in an end-to-end manner, and the post-processing step in the normal multi-view 3D target detection is avoided, thereby improving the detection efficiency.

[0126] FIG. 18 is another structural diagram of a 3D target detection device based on multi-view fusion provided in an exemplary embodiment of the present disclosure.

[0127] Further, in the structural diagram shown in FIG. 18, the image feature mapping module 203 includes: a transformation matrix determination unit 2031 for determining a transformation matrix from a camera coordinate system of the multi-camera system to a bird's-eye view coordinate system based on the internal parameters of the multi-camera system and vehicle parameters; and a spatial transformation unit 2032 for transforming corresponding feature data in the multi-camera viewpoint space of the at least one image from the multi-camera viewpoint space to the bird's-eye viewpoint space based on a transformation matrix from the multi-camera camera coordinate system to the bird's-eye viewpoint coordinate system determined by the transformation matrix determination unit 2031, thereby obtaining corresponding feature data in the bird's-eye viewpoint space of the at least one image.

[0128] In one possible embodiment, the transformation matrix determination unit 2031 is a transformation matrix acquisition subunit for acquiring camera internal parameters and camera external parameters of the multiple cameras in the multi-camera system, and acquiring a transformation matrix from a vehicle coordinate system to a bird's-eye view coordinate system; and a transformation matrix determination subunit for determining a transformation matrix from the multi-camera camera coordinate system to the bird's-eye view coordinate system based on the multi-camera camera external parameters, camera internal parameters, and a transformation matrix from the vehicle coordinate system to the bird's-eye view coordinate system acquired by the transformation matrix acquisition subunit.

[0129] Furthermore, the 3D detection module 205 a detection network acquisition unit 2051 for using a prediction network to obtain a corresponding heat map from the bird's-eye view fusion feature for determining a first preset coordinate value in the bird's-eye view coordinate system of the target object, and obtaining other attribute maps for determining a second preset coordinate value, size and orientation angle in the bird's-eye view coordinate system of the target object; an information detection unit 2052 for determining a first preset coordinate value in the bird's-eye view coordinate system of the target object based on peak information in the heat map acquired by the detection network acquisition unit 2051, and for determining a second preset coordinate value, size and orientation angle in the bird's-eye view coordinate system of the target object from the other attribute map based on the first preset coordinate value in the bird's-eye view coordinate system of the target object; and an information determination unit 2053 for determining three-dimensional spatial information of the target object based on a first preset coordinate value, a second preset coordinate value, a size and an orientation angle in the bird's-eye view coordinate system of the target object detected by the information detection unit 2052.

[0130] In one possible embodiment, the 3D detection module 205 comprises: a loss function construction unit 2054 for constructing a first loss function between the heat map predicted by the prediction network and the truth heat map during the training phase of the prediction network, and constructing a second loss function between the other attribute map predicted by the prediction network and the other truth attribute map; and a total loss function determining unit 2055 for determining a total loss function in the training phase of the prediction network based on the first loss function constructed by the loss function constructing unit 2054 and the second loss function, and monitoring the training process of the prediction network.

[0131] In one possible embodiment, the total loss function determination unit 2055 determines a weight value obtaining subunit for obtaining a weight value of a first loss function and a weight value of a second loss function; and a total loss function determination subunit for determining a total loss function in a training stage of a prediction network based on the first loss function, the second loss function constructed by the loss function construction unit 2054, and the weight values ​​of the first loss function and the weight values ​​of the second loss function obtained by the weight value acquisition subunit.

[0132] In one possible embodiment, the 3D detection module 205 comprises: a fusion feature extraction unit 2056 for performing feature extraction on the bird's-eye view fusion feature by using a neural network to obtain bird's-eye view fusion feature data including target object features; The apparatus further includes a target prediction unit 2057 for performing target prediction on the target object in the bird's-eye view fusion feature data including the target object features obtained by the feature extraction unit 2056 using a prediction network, and obtaining three-dimensional spatial information of the target object.

[0133] Furthermore, the feature extraction module 202 The system includes a feature extraction unit 2021 for performing convolution calculations on images corresponding to each viewpoint using a deep neural network, and obtaining feature data of a plurality of different resolutions including target object features that correspond to the images corresponding to each viewpoint in a multi-camera viewpoint space. Exemplary Electronic Devices

[0134] Hereinafter, an electronic device according to an embodiment of the present disclosure will be described with reference to FIG.

[0135] FIG. 19 is a block diagram of an electronic device provided in an exemplary embodiment of the present disclosure.

[0136] As shown in FIG. 19, the electronic device 11 includes one or more processors 111 and a memory 112.

[0137] The processor 111 may be a central processing unit (CPU) or other type of processing unit having data processing and / or instruction execution capabilities and may control other components in the electronic device 11 to perform desired functions.

[0138] The memory 112 may include one or more computer program products, which may include various types of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may include, for example, random access memory (RAM) and / or high-speed cache memory (cache). The non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc. One or more computer program instructions may be stored in the computer-readable storage medium, and the processor 111 may execute the program instructions to realize the 3D target detection method based on multi-view fusion according to the embodiments of the present disclosure described above and / or other desired functions. The computer-readable storage medium may further store various contents, such as an input signal, a signal component, a noise component, etc.

[0139] In one example, electronic device 11 may further include input devices 113 and output devices 114, with these components being connected together by a bus system and / or other type of connection mechanism (not shown).

[0140] Furthermore, the input device 113 may further include, for example, a keyboard and a mouse.

[0141] The output device 114 can output various information including the determined distance information, direction information, etc. The output device 114 may include, for example, a display, a speaker, a printer, a communication network, and a remote output device connected thereto.

[0142] Of course, for the sake of simplicity, Fig. 20 shows only some of the components related to the present disclosure in the electronic device 11, and omits components such as buses, input / output interfaces, etc. In addition, the electronic device 11 may further include any other suitable components according to specific application situations. Exemplary Computer Program Products and Computer-Readable Storage Media

[0143] In addition to the above methods and apparatus, an embodiment of the present disclosure may also be a computer program product including computer program instructions that, when executed by a processor, cause the processor to perform steps in a 3D target detection method based on multi-view fusion according to various embodiments of the present disclosure described in the "Exemplary Method" section above of this specification.

[0144] Additionally, an embodiment of the present disclosure may be a computer-readable storage medium having computer program instructions stored thereon, which, when executed by a processor, cause the processor to perform steps in a 3D target detection method based on multi-view fusion according to various embodiments of the present disclosure described in the "Exemplary Method" section above of this specification.

[0145] Although the basic principle of the present disclosure has been described above with reference to specific embodiments, it should be noted that the advantages, advantages, effects, etc. mentioned in the present disclosure are merely examples and are not limiting, and these advantages, advantages, effects, etc. are not considered essential to each embodiment of the present disclosure. In addition, the specific details disclosed above are merely for illustrative purposes and to facilitate understanding, but are not limiting, and the above details do not limit that the present disclosure must be realized using the above specific details.

[0146] Block diagrams of devices, apparatus, equipment, and systems according to the present disclosure are not intended to require or imply that they must be connected, arranged, or configured in the manner shown in the block diagrams, as illustrative examples only. As one skilled in the art will recognize, these devices, apparatus, equipment, and systems can be connected, arranged, or configured in any manner. For example, the terms "including," "containing," and "having" are open-ended terms and can be used interchangeably to mean "including, but not limited to." Unless the context expressly indicates otherwise, "or" and "and" as used herein mean and can be used interchangeably to mean "and / or." As used herein, "for example" means and can be used interchangeably to mean "for example, but not limited to."

[0147] It should be further noted that in the devices, apparatuses and methods of the present disclosure, each component or each step can be disassembled and / or recombined, which disassembly and / or recombination should be considered as equivalent solutions of the present disclosure.

Claims

1. acquiring at least one image from a collection of multiple camera viewpoints; performing feature extraction on the at least one image to obtain feature data including corresponding target object features in a multi-camera viewpoint space of the at least one image; Mapping the corresponding feature data in the multi-camera viewpoint space of the at least one image into a same bird's-eye viewpoint space based on the internal parameters of the multi-camera system and vehicle parameters, and obtaining the corresponding feature data in the bird's-eye viewpoint space of the at least one image; performing feature fusion on the corresponding feature data in the bird's-eye view space of the at least one image to obtain a bird's-eye view fusion feature; and performing target prediction on the target object in the bird's-eye view fusion features to obtain three-dimensional spatial information of the target object.

2. The step of mapping the corresponding feature data in the multi-camera viewpoint space of the at least one image to a same bird's-eye viewpoint space based on the internal parameters of the multi-camera system and vehicle parameters, and obtaining the corresponding feature data in the bird's-eye viewpoint space of the at least one image, includes: determining a transformation matrix from a camera coordinate system of the multi-camera system to a bird's-eye view coordinate system based on internal parameters of the multi-camera system and vehicle parameters; The method of claim 1, further comprising a step of transforming corresponding feature data in the multi-camera viewpoint space of the at least one image from the multi-camera viewpoint space to the bird's-eye viewpoint space based on a transformation matrix from the camera coordinate system of the multi-camera to the bird's-eye viewpoint coordinate system, thereby obtaining corresponding feature data in the bird's-eye viewpoint space of the at least one image.

3. The step of determining a transformation matrix from a camera coordinate system of the multi-camera system to a bird's-eye view coordinate system based on the internal parameters of the multi-camera system and vehicle parameters includes: Acquiring camera internal parameters and camera external parameters of each of the cameras in the multi-camera system, and acquiring a transformation matrix from a vehicle coordinate system to the bird's-eye view coordinate system; The method of claim 2, further comprising: determining a transformation matrix from the camera coordinate system of the multi-camera to the bird's-eye view coordinate system based on the camera external parameters of the multi-camera, the camera internal parameters, and a transformation matrix from the vehicle coordinate system to the bird's-eye view coordinate system.

4. The step of performing target prediction on the target object in the bird's-eye view fusion feature to obtain three-dimensional spatial information of the target object includes: obtaining a corresponding heat map from the bird's-eye view fusion feature using a prediction network, for determining a first preset coordinate value in the bird's-eye view coordinate system of the target object, and obtaining other attribute maps for determining a second preset coordinate value, a size and an orientation angle of the target object in the bird's-eye view coordinate system; determining the first preset coordinate value in the bird's-eye view coordinate system of the target object based on peak information in the heat map, and determining the second preset coordinate value, size and orientation angle in the bird's-eye view coordinate system of the target object from the other attribute map based on the first preset coordinate value in the bird's-eye view coordinate system of the target object; and determining three-dimensional spatial information of the target object based on the first preset coordinate value, the second preset coordinate value, the size and the orientation angle in a bird's-eye view coordinate system of the target object.

5. In the training phase of the prediction network, a first loss function is constructed between the heat map predicted by the prediction network and the truth heat map, and a second loss function is constructed between the other attribute map predicted by the prediction network and the other truth attribute map; 5. The method of claim 4, further comprising: determining a total loss function during a training phase of the predictive network based on the first loss function and the second loss function to monitor the training process of the predictive network.

6. Determining a total loss function in a training phase of the prediction network based on the first loss function and the second loss function includes: obtaining weight values ​​of the first loss function and weight values ​​of the second loss function; and determining a total loss function in a training phase of the predictive network based on the first loss function, a weight value of the first loss function, the second loss function, and a weight value of the second loss function.

7. The step of performing target prediction on the target object in the bird's-eye view fusion feature to obtain three-dimensional spatial information of the target object includes: performing feature extraction on the bird's-eye view fusion feature using a neural network to obtain bird's-eye view fusion feature data including the target object feature; The method according to claim 1 or 4, further comprising a step of performing target prediction for the target object in the bird's-eye view fusion feature data including the target object features using a prediction network to obtain three-dimensional spatial information of the target object.

8. The step of performing feature extraction on the at least one image to obtain feature data including corresponding target object features in a multi-camera viewpoint space of the at least one image includes: The method of claim 1 , comprising a step of using a deep neural network to convolute images corresponding to each viewpoint to obtain feature data of a plurality of different resolutions that correspond to the images corresponding to each viewpoint in a multi-camera viewpoint space and include the target object features.

9. an image receiving module for acquiring at least one image from the collected multi-camera viewpoints; a feature extraction module for performing feature extraction on the at least one image acquired by the image receiving module to obtain feature data including corresponding target object features in a multi-camera viewpoint space of the at least one image; an image feature mapping module for mapping corresponding feature data in the multi-camera viewpoint space of the at least one image acquired by the feature extraction module into a same bird's-eye viewpoint space based on internal parameters of a multi-camera system and vehicle parameters, thereby obtaining corresponding feature data in the bird's-eye viewpoint space of the at least one image; an image fusion module for performing feature fusion on the corresponding feature data in the bird's-eye view space of the at least one image obtained by the image feature mapping module to obtain a bird's-eye view fusion feature; and a 3D detection module for performing target prediction on the target object in the bird's-eye view fusion features obtained by the image fusion module, and obtaining three-dimensional spatial information of the target object.

10. A computer-readable storage medium storing a computer program for executing the 3D target detection method based on multi-view fusion according to any one of claims 1 to 8.

11. A processor; a memory for storing instructions executable by said processor; The processor is used to read and execute the executable instructions from the memory to realize the 3D target detection method based on multi-view fusion according to any one of claims 1 to 8. An electronic device.

Citation Information

Patent Citations

  • Intersection multi-view target detection method and system based on angular point pooling

    CN113673444A

  • Section line recognition device

    JP2018097782A