3D target detection method and device based on multi-view fusion

The 3D target detection method for autonomous vehicles uses multi-view fusion to enhance detection efficiency by mapping and fusing features from multiple camera viewpoints into a bird's-eye viewpoint space, improving the accuracy of 3D spatial information generation.

JP7778252B2Active Publication Date: 2025-12-01BEIJING HORIZON ROBOTICS TECH RES & DEV CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2024568637
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2022-05-18
Filing Date
2023-02-08
Publication Date
2025-12-01
Estimated Expiration
2043-02-08

AI Technical Summary

Technical Problem

Conventional 3D detection methods for autonomous driving vehicles require separate detection and fusion of images from multiple viewpoints, leading to low detection efficiency.

Method used

A 3D target detection method based on multi-view fusion that performs feature extraction and mapping of images from multiple camera viewpoints into a bird's-eye viewpoint space, followed by feature fusion and target prediction to obtain three-dimensional spatial information.

Benefits of technology

This method improves detection efficiency by performing multi-view feature fusion and 3D target detection in an end-to-end manner, avoiding post-processing steps and enhancing the accuracy of 3D spatial information generation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007778252000007
    Figure 0007778252000007
  • Figure 0007778252000008
    Figure 0007778252000008
  • Figure 0007778252000009
    Figure 0007778252000009
Patent Text Reader

Abstract

The embodiment of the present disclosure discloses a 3D target detection method and device based on multi-view fusion. In the method, feature extraction is performed on at least one image of multi-camera viewpoints collected by a multi-camera system, and feature data including target object features in the extracted multi-camera viewpoint space are mapped to the same bird's-eye viewpoint space based on the internal parameters of the multi-camera system and the vehicle parameters, and corresponding feature data in the bird's-eye viewpoint space of at least one image are obtained, and bird's-eye viewpoint fusion features are obtained by feature fusion. Target prediction is performed on the target object in the bird's-eye viewpoint fusion feature to obtain the 3D space information of the target object. When performing 3D target detection based on multi-view fusion using the solution of the embodiment of the present disclosure, first perform feature fusion of multi-views, and then perform 3D target detection, thereby completing 3D detection of the scenario object in the bird's-eye viewpoint in an end-to-end manner, and improving the detection efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This disclosure claims priority to a Chinese patent application filed on May 18, 2022, bearing application number 202210544237.0 and entitled "3D target detection method and apparatus based on multi-view fusion," the entire contents of which are incorporated herein by reference.

[0002] The present disclosure relates to the field of computer vision, and in particular to a 3D target detection method and apparatus based on multi-view fusion. [Background technology]

[0003] With the development of science and technology, the application of autonomous driving technology in people's lives is becoming more and more widespread. Autonomous driving vehicles can perform 3D detection of target objects (vehicles, pedestrians, lidars, etc.) within a certain distance from the vehicle to obtain the 3D spatial information of the target object. Based on the 3D spatial information of the target object, distance and speed measurements can be performed to achieve better driving control.

[0004] Currently, an autonomous driving carrier can collect multiple images with different viewpoints, then perform 3D detection on each image separately, and finally fuse the 3D detection results of each image to generate three-dimensional spatial information of target objects in the environment surrounding the carrier. Summary of the Invention [Problem to be solved by the invention]

[0005] Conventional technical solutions require 3D detection for each image collected by the autonomous driving carrier, and then fuse the 3D detection results of each image to obtain information about other vehicles in the 360-degree environment around the carrier, resulting in low detection efficiency. [Means for solving the problem]

[0006] In order to solve the above technical problems, the present disclosure proposes a 3D target detection method and apparatus based on multi-view fusion.

[0007] According to one aspect of the present disclosure, acquiring at least one image from the collected multi-camera viewpoints; performing feature extraction on the at least one image to obtain feature data including corresponding target object features in a multi-camera viewpoint space of the at least one image; Mapping the corresponding feature data in the multi-camera viewpoint space of the at least one image into the same bird's-eye viewpoint space based on the internal parameters of the multi-camera system and vehicle parameters, and obtaining the corresponding feature data in the bird's-eye viewpoint space of the at least one image; performing feature fusion on the corresponding feature data in the bird's-eye view space of the at least one image to obtain bird's-eye view fusion features; and performing target prediction on the target object in the bird's-eye view fusion features to obtain three-dimensional spatial information of the target object.

[0008] According to another aspect of the present disclosure, an image receiving module for acquiring at least one image from the collected multi-camera viewpoints; a feature extraction module for performing feature extraction on the at least one image acquired by the image receiving module to obtain feature data including corresponding target object features in a multi-camera viewpoint space of the at least one image; an image feature mapping module for mapping corresponding feature data in the multi-camera viewpoint space of the at least one image acquired by the feature extraction module to the same bird's-eye viewpoint space based on internal parameters of the multi-camera system and vehicle parameters, thereby obtaining corresponding feature data in the bird's-eye viewpoint space of the at least one image; an image fusion module for performing feature fusion on the corresponding feature data in the bird's-eye view space of the at least one image obtained by the image feature mapping module to obtain bird's-eye view fusion features; and a 3D detection module for performing target prediction on the target object in the bird's-eye view fusion features obtained by the image fusion module, and obtaining three-dimensional spatial information of the target object.

[0009] According to yet another aspect of the present disclosure, there is provided a computer-readable storage medium having a computer program stored therein, the computer program being used to execute the above-described 3D target detection method based on multi-view fusion.

[0010] According to yet another aspect of the present disclosure, there is provided an electronic device, the electronic device comprising: a processor; a memory for storing instructions executable by said processor; The processor is used to read the executable instructions from the memory and execute the instructions to implement the above multi-view fusion based 3D target detection method. [Effects of the Invention]

[0011] The 3D target detection method and apparatus based on multi-view fusion provided in the above embodiments of the present disclosure include: performing feature extraction on at least one image from multiple camera viewpoints collected by a multi-camera system; mapping the extracted feature data, including target object features in the multi-camera viewpoint space, to the same bird's-eye viewpoint space based on the internal parameters of the multi-camera system; obtaining corresponding feature data in the bird's-eye viewpoint space of at least one image; and performing feature fusion on the corresponding feature data in the bird's-eye viewpoint space of the at least one image to obtain bird's-eye viewpoint fusion features. Furthermore, target prediction is performed on the target object in the bird's-eye viewpoint fusion features to obtain 3D spatial information of the target object. When performing 3D target detection based on multi-view fusion using the solution of the embodiments of the present disclosure, multi-view viewpoint feature fusion is first performed, followed by 3D target detection, completing 3D target detection of scenario objects in the bird's-eye viewpoint in an end-to-end manner, avoiding the post-processing step required in conventional multi-view 3D detection, and improving detection efficiency. [Brief explanation of the drawings]

[0012] The above and other objects, features, and advantages of the present disclosure will become more apparent by describing the embodiments of the present disclosure in more detail with reference to the drawings. The drawings are used to provide a further understanding of the embodiments of the present disclosure, constitute a part of the specification, and are used to interpret the present disclosure together with the embodiments of the present disclosure, and are not intended to constitute limitations on the present disclosure. In the drawings, the same reference numerals generally indicate the same components or steps. [Figure 1] FIG. 1 is a scenario diagram to which the present disclosure is applied. [Figure 2] FIG. 1 is a system block diagram of an in-vehicle autonomous driving system provided in an embodiment of the present disclosure. [Figure 3] 1 is a flowchart of a 3D target detection method based on multi-view fusion provided in an exemplary embodiment of the present disclosure. [Figure 4] FIG. 1 is a block diagram illustrating a multi-camera system for collecting images provided in one exemplary embodiment of the present disclosure. [Figure 5] 1 is a schematic diagram of images from multiple camera viewpoints provided in one exemplary embodiment of the present disclosure; [Figure 6] FIG. 2 is a block schematic diagram of feature extraction provided in one exemplary embodiment of the present disclosure. [Figure 7] FIG. 2 is a schematic diagram illustrating generating a bird's-eye view image from images collected by a multi-camera system provided in one exemplary embodiment of the present disclosure. [Figure 8] FIG. 2 is a block schematic diagram of a target detection provided in an exemplary embodiment of the present disclosure. [Figure 9] 10 is a flowchart for determining feature data in a bird's-eye view space provided in an exemplary embodiment of the present disclosure. [Figure 10] FIG. 10 is a block schematic diagram of performing step S303 and step S304 provided in one exemplary embodiment of the present disclosure. [Figure 11] 1 is a flowchart of target detection provided in an exemplary embodiment of the present disclosure. [Figure 12] FIG. 1 is a schematic diagram of an output result of a prediction network provided in an exemplary embodiment of the present disclosure. [Figure 13] 10 is another flowchart of target detection provided in an exemplary embodiment of the present disclosure. [Figure 14] FIG. 2 is a schematic diagram of a Gaussian kernel provided in one exemplary embodiment of the present disclosure. [Figure 15] FIG. 2 is a schematic diagram of a heat map provided in one exemplary embodiment of the present disclosure. [Figure 16] 10 is another flowchart of target detection provided in an exemplary embodiment of the present disclosure. [Figure 17] 1 is a structural diagram of a 3D target detection device based on multi-view fusion provided in an exemplary embodiment of the present disclosure; FIG. [Figure 18] FIG. 10 is another structural diagram of a 3D target detection device based on multi-view fusion provided in an exemplary embodiment of the present disclosure. [Figure 19]FIG. 2 is a block diagram of an electronic device provided in an exemplary embodiment of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION

[0013] Hereinafter, exemplary embodiments based on the present disclosure will be described in detail with reference to the drawings. Obviously, the described embodiments are only some embodiments of the present disclosure, but not all embodiments of the present disclosure, and it should be understood that the present disclosure is not limited to the exemplary embodiments described herein. Application Overview

[0014] To ensure safety during autonomous driving, the autonomous vehicle can detect target objects (e.g., vehicles, pedestrians, lidars, etc.) within a certain distance around the vehicle in real time and obtain three-dimensional spatial information (e.g., attributes such as position, size, orientation angle, and category) of the 3D target objects. Based on the three-dimensional spatial information of the target objects, distance and speed measurements can be performed on the target objects to achieve better driving control. Here, the autonomous vehicle may be a vehicle, an airplane, etc.

[0015] The autonomous driving carrier can use a multi-camera system to collect multiple images with different viewpoints, and then perform 3D target detection on each image separately, for example, filtering target objects, deduplicating, and the like on the multiple images collected by the cameras with different viewpoints. Finally, the 3D detection results of each image are fused to generate three-dimensional spatial information of the target object in the environment surrounding the carrier. As can be seen, the conventional technical solution requires 3D detection on each image collected by the autonomous driving carrier separately, and then fusing the 3D detection results of each image, resulting in low detection efficiency.

[0016] In view of this, an embodiment of the present disclosure provides a 3D target detection method and apparatus based on multi-view fusion. When performing 3D target detection using the solution of the present disclosure, an autonomous driving vehicle can perform feature extraction on at least one image from multiple camera viewpoints collected by a multi-camera system to obtain feature data in a multi-camera viewpoint space, where the feature data includes target object features. Based on internal parameters of the multi-camera system and vehicle parameters, the feature data in the multi-camera viewpoint space is mapped to a bird's-eye viewpoint space to obtain corresponding feature data in the bird's-eye viewpoint space of at least one image. The corresponding feature data in the bird's-eye viewpoint space of the at least one image are then subjected to feature fusion to obtain bird's-eye viewpoint fusion features. Target prediction is performed on the target object in the bird's-eye viewpoint fusion features to obtain 3D spatial information of the target object in the environment surrounding the vehicle.

[0017] In the solution of the embodiment of the present disclosure, when performing 3D target detection based on multi-view fusion, the feature data of the multi-camera views of at least one image can be simultaneously mapped into the same bird's-eye view space, so that more reasonable and better fusion can be achieved. In addition, the fused bird's-eye view fusion features can be used to generate a bird's-eye view fusion feature in the bird's-eye view space. The vehicle is located The 3D spatial information of each target object in the surrounding environment is directly detected. Therefore, when performing 3D target detection based on multi-view fusion using the solution of the embodiment of the present disclosure, multi-view feature fusion is first performed, and then 3D target detection is performed, thereby completing 3D target detection of scenario objects in a bird's-eye view in an end-to-end manner, avoiding the post-processing step in conventional multi-view 3D target detection, and improving detection efficiency. Exemplary System

[0018] The embodiments of the present disclosure can be applied to application scenarios where 3D target detection needs to be performed, such as autonomous driving application scenarios.

[0019] For example, in an autonomous driving application scenario, a multi-camera system is installed on an autonomous driving carrier (hereinafter referred to as "carrier"), and images from different perspectives are collected by the multi-camera system, and then 3D target detection based on multi-view fusion according to the solution of the embodiment of the present disclosure is used to obtain three-dimensional spatial information of target objects in the environment surrounding the carrier.

[0020] FIG. 1 is a scenario diagram to which the present disclosure is applied.

[0021] As shown in FIG. 1 , an embodiment of the present disclosure is applied to a driver assistance or autonomous driving application scenario, in which an onboard autonomous driving system 200 and a multi-camera system 300 are disposed on a driver assistance or autonomous driving carrier 100, and the onboard autonomous driving system 200 is electrically connected to the multi-camera system 300. The multi-camera system 300 is used to collect images of the environment around the carrier, and the onboard autonomous driving system 200 is used to acquire the images collected by the multi-camera system 300 and perform 3D target detection based on multi-view fusion to obtain three-dimensional spatial information of target objects in the environment around the carrier.

[0022] FIG. 2 is a system block diagram of an in-vehicle autonomous driving system provided in an embodiment of the present disclosure.

[0023] As shown in FIG. 2, the in-vehicle autonomous driving system 200 includes an image receiving module 201, a feature extraction module 202, an image feature mapping module 203, an image fusion module 204, and a 3D detection module 205. The image receiving module 201 is used to acquire at least one image collected by a multi-camera system 300. The feature extraction module 202 is used to perform feature extraction on the at least one image acquired by the image receiving module 201 to obtain feature data. The image feature mapping module 203 is used to map the feature data of the at least one image from the multi-camera viewpoint space to the same bird's-eye viewpoint space. The image fusion module 204 is used to perform feature fusion on corresponding feature data in the bird's-eye viewpoint space of the at least one image to obtain bird's-eye viewpoint fusion features. The 3D detection module 205 is used to perform target prediction on a target object in the bird's-eye viewpoint fusion features acquired by the image fusion module 204 to obtain three-dimensional spatial information of the target object in the environment surrounding the carrier.

[0024] The multi-camera system 300 includes multiple cameras with different viewpoints, each of which is used to collect an environmental image from one viewpoint, and the multiple cameras cover a 360-degree environmental range around the carrier. Each camera defines its own camera viewpoint coordinate system, which forms a respective camera viewpoint space, and the environmental image collected by each camera is an image in the corresponding camera viewpoint space. Exemplary Methods

[0025] FIG. 3 is a flowchart of a 3D target detection method based on multi-view fusion provided in an exemplary embodiment of the present disclosure.

[0026] This embodiment can be applied to an in-vehicle automatic driving system 200, and as shown in FIG. 3, includes the following steps.

[0027] In step S301, at least one image from a collection of multiple camera viewpoints is acquired.

[0028] Here, the at least one image may be collected by at least one camera of a multi-camera system. Exemplarily, the at least one image may be an image collected in real time by the multi-camera system, or an image collected in advance by the multi-camera system.

[0029] FIG. 4 is a block diagram illustrating a multi-camera system for collecting images provided in one exemplary embodiment of the present disclosure.

[0030] 4, in one embodiment, the multi-camera system can collect multiple images from different viewpoints in real time, for example, images 1, 2...N, and transmit the collected images to the in-vehicle autonomous driving system in real time. In this way, the images acquired by the in-vehicle autonomous driving system can characterize the current state of the environment around the carrier.

[0031] FIG. 5 is a schematic diagram of images from multiple camera viewpoints provided in one exemplary embodiment of the present disclosure.

[0032] As shown in (1) to (6) of Figure 5, in one embodiment, the multi-camera system can include six cameras. The six cameras are installed at the front end, left front end, right front end, rear end, left rear end, and right rear end of the carrier, respectively. In this way, at any given time, the multi-camera system can capture six different perspective images, for example, a front view image (I front ), left anterior view image (I frontleft ), right front view image (I frontright ), backsight image (I rear ), left rear view image (I rearleft ) and right rearview image (I rearright ) can be collected.

[0033] Here, each image includes target objects representing categories such as, but not limited to, roads, traffic lights, road signs, vehicles (small cars, buses, trucks, etc.), pedestrians, riders, etc. Categories of target objects in the environment surrounding the carrier. or Due to differences in position, the categories and positions of target objects included in each image also differ.

[0034] In step S302, feature extraction is performed on at least one image to obtain feature data including target object features corresponding to each of the at least one image in the multi-camera viewpoint space.

[0035] In one embodiment, the autonomous driving system can extract feature data in a corresponding camera viewpoint space from each image, which may include target object features for indicating a target object in the image, including but not limited to image texture information, edge contour information, semantic information, etc.

[0036] Here, the image texture information is used to characterize the image texture of the target object, the edge contour information is used to characterize the edge contour of the target object, and the semantic information is used to characterize the category of the target object, including but not limited to roads, traffic lights, road signs, vehicles (small cars, buses, trucks, etc.), pedestrians, riders, etc.

[0037] FIG. 6 is a block schematic diagram of feature extraction provided in one exemplary embodiment of the present disclosure.

[0038] As shown in Figure 6, the in-vehicle autonomous driving system employs a neural network to perform feature extraction for at least one image (images 1-N), and can obtain corresponding feature data 1-N in the multi-camera viewpoint space for each image.

[0039] For example, an in-vehicle autonomous driving system uses forward-looking images (I front ) feature extraction is performed on the foreground image (I front ) feature data f in the front camera viewpoint space front The left front view image (I frontleft ) feature extraction is performed on the left front view image (I frontleft ) feature data f in the viewpoint space of the left front camera frontleft The right front view image (I frontright ) and extract features from the right front view image (I frontright ) feature data f in the viewpoint space of the right front camera frontright The backsight image (I rear ) and extract features from the backsight image (I rear ) feature data f in the rear camera viewpoint space rear The left rear view image (I rearleft ) and extracting features from the left rear view image (I rearleft ) feature data f in the left rear camera viewpoint space rearleft The right rearview image (I rearright ) and extract features from the right rear view image (I rearright ) feature data f in the viewpoint space of the right rear camera rearright can be obtained.

[0040] In step S303, based on the internal parameters of the multi-camera system and the vehicle parameters, the corresponding feature data in the multi-camera viewpoint space of at least one image is mapped to the same bird's-eye viewpoint space, and the corresponding feature data in the bird's-eye viewpoint space of the at least one image is obtained.

[0041] Here, the internal parameters of the multi-camera system include the internal camera parameters and external camera parameters of each camera, where the internal camera parameters are parameters related to the characteristics of the camera itself, such as the focal length and pixel size of the camera, and the external camera parameters are parameters in the world coordinate system, such as the position and rotation direction of the camera. The vehicle parameters are the transformation matrix from the vehicle coordinate system (VCS) to the bird's-eye view coordinate system (BEV), and the vehicle coordinate system is the coordinate system in which the carrier is located.

[0042] For example, an in-vehicle autonomous driving system uses forward-looking images (I front ) feature data f in the front camera viewpoint space front are mapped to the same bird's-eye view space, and the foreground image (I front ) feature data F in the bird's-eye view space front A left front view image (I frontleft ) feature data f in the viewpoint space of the left front camera frontleft are mapped to the same bird's-eye view space, and the left front view image (I frontleft ) feature data F in the bird's-eye view space frontleft A right front view image (I frontright ) feature data f in the viewpoint space of the right front camera frontright are mapped to the same bird's-eye view space, and the right front view image (I frontright ) feature data F in the bird's-eye view space frontright and obtain the backsight image (I rear ) feature data f in the rear camera viewpoint space rear are mapped to the same bird's-eye view space, and the back-view image (I rear ) feature data F in the bird's-eye view space rear The left rear view image (I rearleft ) feature data f in the left rear camera viewpoint space rearleft are mapped to the same bird's-eye view space, and the left rear view image (I rearleft ) feature data F in the bird's-eye view space rearleft The right rearview image (I rearright ) feature data f in the viewpoint space of the right rear camera rearrightare mapped to the same bird's-eye view space, and the right rear view image (I rearright ) feature data F in the bird's-eye view space rearright Get.

[0043] In step S304, the corresponding feature data in the bird's-eye view space of at least one image are subjected to feature fusion to obtain bird's-eye view fusion features.

[0044] Here, the bird's-eye view fusion feature is used to characterize the feature data in the bird's-eye view space of the target object around the carrier, and the feature data in the bird's-eye view space of the target object may include, but is not limited to, attributes such as the shape, size, category, orientation angle, and relative position of the target object.

[0045] In one embodiment, the in-vehicle autonomous driving system can perform additive feature fusion on corresponding feature data in the bird's-eye view space of at least one image to obtain a bird's-eye view fusion feature, which can be expressed by the following formula:

number

[0046] where F' denotes the bird's-eye view fusion feature, and Add denotes the additive feature fusion calculation performed on the corresponding feature data in the bird's-eye view space of at least one image.

[0047] It should be noted that the embodiment of step S304 is not limited thereto, and may perform feature fusion on the corresponding feature data in the bird's-eye view space of images from different camera viewpoints, for example, by multiplication, overlapping, etc.

[0048] FIG. 7 is a schematic diagram illustrating generating a bird's-eye view image from images collected by a multi-camera system provided in one exemplary embodiment of the present disclosure.

[0049] 7, the size of the bird's-eye view image may be the same as the size of at least one image collected by the multi-camera system. The bird's-eye view image may represent three-dimensional spatial information of the target object, and the three-dimensional spatial information includes at least one attribute information of the target object, including, but not limited to, 3D position information (i.e., coordinate information of X-axis, Y-axis, and Z-axis), size information (i.e., length, width, and height information), orientation angle information, etc.

[0050] Here, the coordinate information of the X-axis, Y-axis, and Z-axis refers to the coordinate position (x, y, z) of the target object in the bird's-eye view space, and the origin of the coordinate system of the bird's-eye view space is located at one of the positions such as the carrier chassis or the carrier center, with the X-axis direction being the front-to-rear direction, the Y-axis direction being the left-to-right direction, and the Z-axis direction being the vertical up-down direction. The heading angle is the angle formed by the front direction or traveling direction of the target object in the bird's-eye view space. For example, if the target object is a moving pedestrian, the heading angle is the angle formed by the pedestrian's traveling direction in the bird's-eye view space. If the target object is a stationary vehicle, the heading angle is the angle formed by the vehicle's head direction in the bird's-eye view space.

[0051] In addition, since at least one image collected by the multi-camera system may contain target objects of different categories, the bird's-eye view image may contain bird's-eye view fusion features of target objects of different categories.

[0052] In step S305, target prediction is performed on the target object in the bird's-eye view fusion feature to obtain three-dimensional space information of the target object.

[0053] Here, the three-dimensional spatial information may include at least one of attributes such as the position, size, and orientation angle of the target object in the bird's-eye view coordinate system. The position is the coordinate position (x, y, z) of the target object relative to the carrier in the bird's-eye view space, the size is the length, width, and height (Height, Width, Length) of the target object in the bird's-eye view space, and the orientation angle is the orientation angle (rotation yaw) of the target object in the bird's-eye view space.

[0054] FIG. 8 is a block schematic diagram of a target detection provided in one exemplary embodiment of the present disclosure.

[0055] As shown in FIG. 8, in one embodiment, the in-vehicle autonomous driving system can use one or more prediction networks to perform 3D target prediction for target objects in bird's-eye view fusion features, and obtain three-dimensional spatial information of each target object in the environment surrounding the carrier.

[0056] When an in-vehicle autonomous driving system uses multiple prediction networks to perform 3D target prediction, each prediction network can output one or more attributes of the target object, and the attributes output by different prediction networks will also be different.

[0057] In the solution of the embodiments of the present disclosure, when 3D target detection based on multi-view fusion is performed, multi-view feature fusion is first performed and then 3D target detection is performed, so that 3D target detection of scenario objects in a bird's-eye view can be completed end-to-end, which avoids the post-processing step in conventional multi-view 3D target detection and improves detection efficiency.

[0058] FIG. 9 is a flowchart for determining feature data in a bird's-eye view space provided in one exemplary embodiment of the present disclosure.

[0059] As shown in FIG. 9, based on the embodiment shown in above FIG. 3, step S303 may include the following steps:

[0060] In step S3031, a transformation matrix from the camera coordinate system of the multi-cameras of the multi-camera system to the bird's-eye view coordinate system is determined based on the internal parameters of the multi-camera system and the vehicle parameters.

[0061] Here, the internal parameters of the multi-camera system include the internal camera parameters and external camera parameters of each camera, where the external camera parameters are the transformation matrix from the camera coordinate system of the multi-camera to the vehicle coordinate system, and the vehicle parameters are the transformation matrix from the vehicle coordinate system (VCS) to the bird's-eye view coordinate system (BEV), and the vehicle coordinate system is the coordinate system in which the carrier is located.

[0062] In one specific embodiment, step S3031 includes: A step of acquiring intrinsic camera parameters and extrinsic camera parameters of the multiple cameras in the multi-camera system, and acquiring a transformation matrix from a vehicle coordinate system to a bird's-eye view coordinate system; The method includes determining a transformation matrix from the camera coordinate system of the multi-camera to the bird's-eye view coordinate system based on the external camera parameters of the multi-camera, the internal camera parameters, and the transformation matrix from the vehicle coordinate system to the bird's-eye view coordinate system.

[0063] In one embodiment, the in-vehicle autonomous driving system can determine a transformation matrix H from the multi-camera camera coordinate system to the bird's-eye view coordinate system using the following equation:

number

[0064] where @ denotes matrix multiplication and T camera→vcs denotes the transformation matrix from the camera coordinate system to the vehicle coordinate system, and T camera→vcs characterizes the camera extrinsic parameters, and T vcs→camera denotes a transformation matrix from the vehicle coordinate system to the bird's-eye view coordinate system, and K denotes a camera internal parameter.

[0065] The camera external parameters, i.e., the transformation matrix from the camera coordinate system to the vehicle coordinate system, can be obtained by calibration of the multi-camera system and usually does not change once calibration is complete. The transformation matrix from the vehicle coordinate system to the bird's-eye view coordinate system can be calculated from an artificially set bird's-eye view range (e.g., a range surrounded by 100 meters from the front, rear, left, and right) and the resolution of the bird's-eye view image (e.g., 512 × 512).

[0066] In this way, it is possible to determine the transformation matrix corresponding to each camera in the multi-camera system. For example, the in-vehicle autonomous driving system may determine the transformation matrix H from the camera coordinate system of the front-end camera to the bird's-eye view coordinate system based on the transformation matrix from the camera coordinate system of the front-end camera to the vehicle coordinate system, the transformation matrix from the vehicle coordinate system to the bird's-eye view coordinate system, and the camera internal parameters of the front-end camera. front→bev and determine a transformation matrix H from the camera coordinate system of the left front end camera to the bird's-eye view coordinate system based on the transformation matrix from the camera coordinate system of the left front end camera to the vehicle coordinate system, the transformation matrix from the vehicle coordinate system to the bird's-eye view coordinate system, and the camera internal parameters of the left front end camera. frontleft→bev and determine a transformation matrix H from the camera coordinate system of the right front camera to the bird's-eye view coordinate system based on the transformation matrix from the camera coordinate system of the right front camera to the vehicle coordinate system, the transformation matrix from the vehicle coordinate system to the bird's-eye view coordinate system, and the camera internal parameters of the right front camera. frontright→bev and determine a transformation matrix H from the camera coordinate system of the rear-end camera to the bird's-eye view coordinate system based on the transformation matrix from the camera coordinate system of the rear-end camera to the vehicle coordinate system, the transformation matrix from the vehicle coordinate system to the bird's-eye view coordinate system, and the internal camera parameters of the rear-end camera. rear→bev and determine a transformation matrix H from the camera coordinate system of the left rear end camera to the bird's-eye view coordinate system based on the transformation matrix from the camera coordinate system of the left rear end camera to the vehicle coordinate system, the transformation matrix from the vehicle coordinate system to the bird's-eye view coordinate system, and the camera internal parameters of the left rear end camera. rearleft→bevand determine a transformation matrix H from the camera coordinate system of the right rear end camera to the bird's-eye view coordinate system based on the transformation matrix from the camera coordinate system of the right rear end camera to the vehicle coordinate system, the transformation matrix from the vehicle coordinate system to the bird's-eye view coordinate system, and the camera internal parameters of the right rear end camera. rearright→bev Determine.

[0067] In this embodiment, each camera has its own transformation matrix from the camera viewpoint coordinate system to the bird's-eye view coordinate system, so the prediction network adopted in the embodiments of this application for 3D target detection can be applied to a multi-camera system, eliminating the need to train the prediction network from scratch and improving detection efficiency.

[0068] In step S3032, based on a transformation matrix from the multi-camera camera coordinate system to the bird's-eye view coordinate system, the corresponding feature data in the multi-camera view space of at least one image is transformed from the multi-camera view space to the bird's-eye view space, and the corresponding feature data in the bird's-eye view space of at least one image is obtained.

[0069] In one embodiment, the in-vehicle autonomous driving system calculates the transformation matrix of each camera and Each camera By matrix multiplying the feature data in the viewpoint space, corresponding feature data in the bird's-eye viewpoint space of at least one image can be obtained. Specifically, this can be expressed by the following equation.

number

[0070] Here, F is the corresponding feature data F in the bird's-eye view space of at least one image. front , F frontleft , F frontright , F rear , F rearleft and F rearright where H is the transformation matrix H corresponding to each camera in the multi-camera system. front , H frontleft , H frontright , H rear , Hrearleft and H rearright where f is the feature data f in the multi-camera viewpoint space of at least one image. front , f frontleft , f frontright , f rear , f rearleft and f rearright Shows.

[0071] As can be seen, the embodiments of the present disclosure not only apply to multi-camera systems of different models, but also enable more rational feature fusion by calculating respective transformation matrices (homographies) for different cameras in a multi-camera system, and then mapping each feature data to a bird's-eye view space based on each transformation matrix of each camera, thereby obtaining corresponding feature data in the bird's-eye view space for each image.

[0072] The two steps, step S302 and step S3031, may be executed synchronously or asynchronously, and may be determined according to the actual application situation.

[0073] FIG. 10 is a block schematic diagram of performing step S303 and step S304 provided in one exemplary embodiment of the present disclosure.

[0074] 10, after the execution of steps S302 and S3031 is completed, the feature space transformation of step S3032 is performed based on the transformation matrix from the camera coordinate system of each camera to the bird's-eye view coordinate system obtained in step S3031 and the feature data in the corresponding camera viewpoint space obtained in step S302, to obtain feature data in the bird's-eye view space. Finally, step S304 is performed to perform feature fusion on the feature data in the bird's-eye view space of the multi-camera viewpoints, to obtain bird's-eye view fusion features.

[0075] FIG. 11 is a flowchart of target detection provided in one exemplary embodiment of the present disclosure.

[0076] As shown in FIG. 11, based on the embodiment shown in above FIG. 3, step S305 may include the following steps:

[0077] In step S3051, a prediction network is used to obtain a corresponding heat map from the bird's-eye view fusion features for determining a first preset coordinate value in the bird's-eye view coordinate system of the target object, and obtain other attribute maps for determining a second preset coordinate value, size and orientation angle in the bird's-eye view coordinate system of the target object.

[0078] Here, the prediction network may be a neural network for performing target prediction for a target object. Since it is necessary to perform 3D spatial information prediction with different attributes for the target object, the prediction network may be of multiple types. Different prediction networks are used to predict 3D spatial information with different attributes.

[0079] For example, if the attribute to be predicted is a first preset coordinate value of the target object, the prediction network corresponding to the first preset coordinate value can be used to process the bird's-eye view fusion feature in the bird's-eye view image to obtain a heat map, and then the heat map can be used to determine the first preset coordinate value in the bird's-eye view coordinate system of the target object, where the size of the heat map can be the same as the size of the bird's-eye view image.

[0080] Also, for example, if the attributes that need to be predicted are the second preset coordinate value, size, and orientation angle of the target object, a prediction network corresponding to the second preset coordinate value, size, and orientation angle can be used to process the bird's-eye view fusion features in the bird's-eye view image to obtain another attribute map, and thereby the second preset coordinate value, size, and orientation angle of the target object in the bird's-eye view coordinate system can be determined using the other attribute map.

[0081] Here, the first preset coordinate value is the (x, y) position in the bird's-eye view coordinate system, the second preset coordinate value is the z position in the bird's-eye view coordinate system, the size is the length, width, and height, and the orientation angle is the orientation angle.

[0082] In step S3052, a first preset coordinate value in the bird's-eye view coordinate system of the target object is determined based on peak information in the heat map, and a second preset coordinate value, size and orientation angle in the bird's-eye view coordinate system of the target object are determined from another attribute map based on the first preset coordinate value in the bird's-eye view coordinate system of the target object.

[0083] Here, the peak information is the center value of the Gaussian kernel, that is, the center point of the target object.

[0084] After predicting the first preset coordinate value of the target object in the bird's-eye view space, other attribute maps can output their respective attribute information using the attribute output results of the heat map, so that the second preset coordinate value, size, and orientation angle of the target object in the bird's-eye view coordinate system can be predicted from the other attribute maps based on the first preset coordinate value of the target object in the bird's-eye view coordinate system.

[0085] In step S3053, three-dimensional spatial information of the target object is determined based on the first preset coordinate value, the second preset coordinate value, the size and the orientation angle of the target object in the bird's-eye view coordinate system.

[0086] In one embodiment, the in-vehicle autonomous driving system can determine the first and second preset coordinate values ​​as the (x, y, z) position of the target object in the bird's-eye view space, determine the size as the length, width, and height of the target object in the bird's-eye view space, and determine the orientation angle as the orientation angle of the target object in the bird's-eye view space. Finally, the system determines three-dimensional spatial information of the target object in the environment surrounding the carrier based on the (x, y, z) position, length, width, height, and orientation angle.

[0087] 12 is a schematic diagram of the output result of the prediction network provided in one exemplary embodiment of the present disclosure. In FIG. 12, the center A of the smallest circle is the position of the carrier, and the block position B around the center is the target object around the carrier.

[0088] In addition, the in-vehicle autonomous driving system can further display a three-dimensional spatial projection of the target object on images from multiple camera viewpoints collected by the multi-camera system, making it easier for the user to intuitively understand the three-dimensional spatial information of the target object from the in-vehicle display.

[0089] As can be seen, the embodiments of the present disclosure can process bird's-eye view images based on a prediction network to obtain heat maps and other attribute maps, and input bird's-eye view fusion features obtained by feature fusion into the heat maps and other attribute maps to directly predict the 3D spatial information of target objects and improve the efficiency of 3D target detection.

[0090] FIG. 13 is another flowchart of target detection provided in an exemplary embodiment of the present disclosure.

[0091] As shown in FIG. 13, based on the embodiment shown in FIG. 11 above, step S305 can further include the following steps:

[0092] In step S3054, during the training phase of the prediction network, predictionA first loss function is constructed between the generated heat map and the truth-value heat map, and a second loss function is constructed between the other attribute map predicted by the prediction network and the other truth-value attribute map.

[0093] In one embodiment, the in-vehicle autonomous driving system can construct a Gaussian kernel for each target object based on the position of each target object in the bird's-eye view fusion features.

[0094] 14 is a schematic diagram of a Gaussian kernel provided in one exemplary embodiment of the present disclosure. As shown in FIG. 14, when constructing a Gaussian kernel, one Gaussian kernel with a size of N×N can be generated with the center at the position (i, j) of the target object. Here, the value at the center of the Gaussian kernel is 1, and the surrounding values ​​are attenuated downward to 0, and the color changing from white to black indicates that the value is attenuated from 1 to 0.

[0095] 15 is a schematic diagram of a heat map provided in an exemplary embodiment of the present disclosure. As shown in FIG. 15, a truth heat map can be obtained by placing the Gaussian kernel of each target object on the heat map. In FIG. 15, each white area represents one Gaussian kernel, i.e., one target object, for example, target objects 1 to 6.

[0096] Note that the method for generating other truth attribute diagrams can refer to the method for generating truth heat maps, and the description thereof will be omitted here.

[0097] After determining the truth heat map, a first loss function can be constructed based on the truth heat map and the heat map output by the prediction network, where the first loss function can determine the distribution of the difference between the output prediction value of the prediction network and the truth value, and is used to monitor the training process of the prediction network.

[0098] In one embodiment, the first loss function L cls Specifically, may be constructed as follows:

number

[0099] where y' i,j indicates the first preset coordinate value in the truth heat map at the (i,j) position, 1 indicates the peak in the heat map, and y i,j denotes the first preset coordinate value in the heat map predicted by the prediction network for position (i, j), α and β are adjustable hyperparameters, and the ranges of α and β are both between 0 and 1, N denotes the total number of target objects in the bird's-eye view fusion feature, and h, w denote the size of the bird's-eye view fusion feature.

[0100] In one embodiment, the second loss function L reg Specifically, may be constructed as follows:

number

[0101] Here, B' is the true value of the second preset coordinate value, size, and orientation angle of the target object in the bird's-eye view coordinate system, B is the predicted value of the second preset coordinate value, size, and orientation angle of the target object in the bird's-eye view coordinate system predicted by the prediction network, and N represents the number of target objects in the bird's-eye view fusion feature.

[0102] In step S3055, a total loss function in the training stage of the prediction network is determined based on the first loss function and the second loss function, and the training process of the prediction network is monitored.

[0103] In one embodiment, the total loss function during the training phase of the prediction network is: obtaining weight values ​​of a first loss function and weight values ​​of a second loss function; and determining a total loss function in the training phase of the prediction network based on the first loss function, the weight value of the first loss function, the second loss function, and the weight value of the second loss function.

[0104] In this study, when predicting the 3D spatial information of a target object using a prediction network, different attributes have different importance during the training process, and therefore the importance of the corresponding loss functions also differs. Therefore, different weights are assigned to the loss functions corresponding to different attributes based on the importance of each attribute during the training process.

[0105] Here, the total loss function L during the training phase of the predictive network is 3d may be determined by the following formula:

number

[0106] where L cls is the first loss function, L reg is the second loss function, λ1 is the weight value of the first loss function, and λ2 is the weight value of the second loss function. λ1 and λ2 are both between 0 and 1, λ1>λ2, λ1+λ2=1.

[0107] As can be seen, when embodiments of the present disclosure train a prediction network, they construct a total loss function training The process is monitored to ensure that the output of various attributes of the prediction network is more accurate, and that the efficiency of 3D target detection is higher.

[0108] FIG. 16 is another flowchart of target detection provided in an exemplary embodiment of the present disclosure.

[0109] As shown in FIG. 16, based on the embodiment shown in FIG. 3 above, step S305 may further include the following steps:

[0110] In step S3056, feature extraction is performed on the bird's-eye view fusion features using a neural network to obtain bird's-eye view fusion feature data including the target object features.

[0111] In one embodiment, the in-vehicle autonomous driving system can use a neural network to perform calculations such as convolution on the bird's-eye view fusion features to realize feature extraction and obtain bird's-eye view fusion feature data, which includes target object features for characterizing different dimensions of the target object, i.e., scenario information from different dimensions in the bird's-eye view space of the target object.

[0112] Here, the neural network may be a pre-trained neural network for feature extraction. Optionally, the neural network for feature extraction may be a specific network. structure , for example, but not limited to, resnet, densenet, mobilenet, etc.

[0113] In step S3057, target prediction is performed on the target object in the bird's-eye view fusion feature data including the target object features using the prediction network, and three-dimensional space information of the target object is obtained.

[0114] As can be seen, in the embodiment of the present disclosure, before training the bird's-eye view fusion features with a prediction network, feature extraction is performed on the bird's-eye view fusion features to obtain bird's-eye view fusion feature data, and then the bird's-eye view fusion feature data including the target object features is predicted using the prediction network, so that the prediction result is more accurate, i.e., the 3D spatial information of the determined target object is more accurate.

[0115] Based on the embodiment shown in FIG. 3 above, step S302: The method may include a step of using a deep neural network to perform convolution calculations on images corresponding to each viewpoint, and obtaining feature data of multiple different resolutions in a multi-camera viewpoint space for the images corresponding to each viewpoint, wherein the feature data respectively correspond to the images corresponding to each viewpoint and include target object features.

[0116] Here, the deep neural network may be a pre-trained neural network for feature extraction. Optionally, the neural network for feature extraction may be a specific network. structure For example, but not limited to, resnet, densenet, mobilenet, etc. By using a deep neural network to perform calculations such as convolution and pooling on the image of the target viewpoint, feature data at multiple different resolutions (scales) corresponding to the image of the target viewpoint can be obtained.

[0117] For example, the size of image A from a certain viewpoint is H×W×3, where H is the height of image A, W is the width of image A, and 3 indicates that there are three channels. For example, in the case of an RGB image, 3 indicates the three channels of RGB (R red, G green, B blue), and in the case of a YUV image, 3 indicates the three channels of YUV (Y luminance signal, U blue component signal, V red component signal). Image A is input to a deep neural network, and after the deep neural network performs calculations such as convolution, an H1×W1×N dimensional feature matrix is ​​output, where H1 and W1 are the height and width of the feature (generally smaller than H and W, and N is the number of channels, N is greater than 3). By fitting training to the input data using the neural network, feature data of multiple different resolutions including target object features of the input image can be obtained, such as low-level image texture, edge contour information, and high-level semantic information corresponding to different resolutions. After obtaining the feature data of each viewpoint image, the subsequent spatial transformation, multi-view feature fusion and target prediction steps can be performed to obtain the 3D spatial information of the target object.

[0118] As can be seen, the embodiments of the present disclosure use a deep neural network to perform calculations such as convolution and pooling on the images corresponding to each viewpoint to obtain feature data of multiple different resolutions for each viewpoint image, which can better reflect the image features collected by the corresponding viewpoint camera and improve the efficiency of subsequent 3D target detection. Exemplary Apparatus

[0119] 17 is a structural diagram of a 3D target detection device based on multi-view fusion provided in an exemplary embodiment of the present disclosure. The 3D target detection device based on multi-view fusion may be installed in an electronic device such as a terminal device or a server, or in a driver assistance or autonomous driving carrier. For example, the 3D target detection device may be installed in an in-vehicle autonomous driving system to execute the 3D target detection method based on multi-view fusion of any one of the above embodiments of the present disclosure. As shown in FIG. 17 , the 3D target detection device based on multi-view fusion of this embodiment includes an image receiving module 201, a feature extraction module 202, an image feature mapping module 203, an image fusion module 204, and a 3D detection module 205.

[0120] Here, the image receiving module 201 is used to obtain at least one image from the collected multi-camera viewpoints.

[0121] The feature extraction module 202 is used to perform feature extraction on the at least one image acquired by the image receiving module, and obtain corresponding feature data in the multi-camera viewpoint space of the at least one image, where the feature data includes target object features.

[0122] The image feature mapping module 203 is used to map corresponding feature data in the multi-camera viewpoint space of the at least one image acquired by the feature extraction module to the same bird's-eye viewpoint space based on the internal parameters of the multi-camera system and vehicle parameters, and obtain corresponding feature data in the bird's-eye viewpoint space of the at least one image.

[0123] The image fusion module 204 is used to perform feature fusion on the corresponding feature data in the bird's-eye view space of the at least one image obtained by the image feature mapping module, to obtain bird's-eye view fusion features.

[0124] The 3D detection module 205 is used to perform target prediction on the target object in the bird's-eye view fusion features obtained by the image fusion module, and obtain three-dimensional spatial information of the target object.

[0125] As can be seen, when performing 3D target detection based on multi-view fusion, the device according to the embodiment of the present disclosure can perform more reasonable and better fusion by simultaneously mapping feature data from multiple camera views of at least one image into the same bird's-eye view space through middle fusion. In addition, the fused bird's-eye view fusion features can be used to identify the vehicle in the bird's-eye view space. Both are located The 3D spatial information of each target object within the surrounding environment is directly detected. Therefore, when performing 3D target detection based on multi-view fusion using the device of the embodiment of the present disclosure, the 3D detection of scenario objects in a bird's-eye view is completed in an end-to-end manner, which avoids the post-processing step required in conventional multi-view 3D target detection and improves detection efficiency.

[0126] FIG. 18 is another structural diagram of a 3D target detection device based on multi-view fusion provided in an exemplary embodiment of the present disclosure.

[0127] Furthermore, in the structural diagram shown in FIG. 18, the image feature mapping module 203: a transformation matrix determination unit 2031 for determining a transformation matrix from a camera coordinate system of the multi-cameras of the multi-camera system to a bird's-eye view coordinate system based on the internal parameters of the multi-camera system and vehicle parameters; and a space transformation unit 2032 for transforming corresponding feature data in the multi-camera viewpoint space of the at least one image from the multi-camera viewpoint space to the bird's-eye viewpoint space based on the transformation matrix from the multi-camera camera coordinate system to the bird's-eye viewpoint coordinate system determined by the transformation matrix determination unit 2031, thereby obtaining corresponding feature data in the bird's-eye viewpoint space of the at least one image.

[0128] In one possible embodiment, the transformation matrix determination unit 2031 is a transformation matrix acquisition subunit for acquiring internal camera parameters and external camera parameters of the multiple cameras in the multi-camera system, and acquiring a transformation matrix from a vehicle coordinate system to a bird's-eye view coordinate system; and a transformation matrix determination subunit for determining a transformation matrix from the camera coordinate system of the multi-camera to the bird's-eye view coordinate system based on the external camera parameters, internal camera parameters, and the transformation matrix from the vehicle coordinate system to the bird's-eye view coordinate system of the multi-camera acquired by the transformation matrix acquisition subunit.

[0129] Furthermore, the 3D detection module 205 a detection network obtaining unit 2051 for obtaining a corresponding heat map from the bird's-eye view fusion feature using a prediction network to determine a first preset coordinate value in the bird's-eye view coordinate system of the target object, and obtaining other attribute maps to determine a second preset coordinate value, size, and orientation angle in the bird's-eye view coordinate system of the target object; an information detection unit 2052 for determining a first preset coordinate value in the bird's-eye view coordinate system of the target object based on peak information in the heat map acquired by the detection network acquisition unit 2051, and for determining a second preset coordinate value, size and orientation angle in the bird's-eye view coordinate system of the target object from the other attribute map based on the first preset coordinate value in the bird's-eye view coordinate system of the target object; and an information determination unit 2053 for determining three-dimensional spatial information of the target object based on the first preset coordinate value, the second preset coordinate value, the size and the orientation angle of the target object in the bird's-eye view coordinate system detected by the information detection unit 2052.

[0130] In one possible embodiment, the 3D detection module 205: a loss function construction unit 2054 for constructing a first loss function between the heat map predicted by the prediction network and the truth heat map during the training phase of the prediction network, and constructing a second loss function between the other attribute map predicted by the prediction network and the other truth attribute map; The prediction network further includes a total loss function determination unit 2055 for determining a total loss function in the training stage of the prediction network based on the first loss function constructed by the loss function construction unit 2054 and the second loss function, and monitoring the training process of the prediction network.

[0131] In one possible embodiment, the total loss function determination unit 2055 determines a weight value obtaining subunit for obtaining a weight value of a first loss function and a weight value of a second loss function; and an overall loss function determination subunit for determining an overall loss function in the training stage of the prediction network based on the first loss function and the second loss function constructed by the loss function construction unit 2054, and the weight values ​​of the first loss function and the weight values ​​of the second loss function obtained by the weight value acquisition subunit.

[0132] In one possible embodiment, the 3D detection module 205: a fusion feature extraction unit 2056 for performing feature extraction on the bird's-eye view fusion feature by using a neural network to obtain bird's-eye view fusion feature data including target object features; Using a prediction network fusion The system further includes a target prediction unit 2057 for performing target prediction on the target object in the bird's-eye view fusion feature data including the target object features obtained by the feature extraction unit 2056, and obtaining three-dimensional spatial information of the target object.

[0133] Furthermore, the feature extraction module 202 The image processing device includes a feature extraction unit 2021 for performing convolution calculations on images corresponding to each viewpoint using a deep neural network, and obtaining feature data of a plurality of different resolutions that correspond to the images corresponding to each viewpoint in the multi-camera viewpoint space and include target object features. Exemplary Electronic Devices

[0134] Hereinafter, an electronic device according to an embodiment of the present disclosure will be described with reference to FIG.

[0135] FIG. 19 is a block diagram of an electronic device provided in one exemplary embodiment of the present disclosure.

[0136] As shown in FIG. 19, the electronic device 11 includes one or more processors 111 and a memory 112.

[0137] The processor 111 may be a central processing unit (CPU) or other type of processing unit having data processing and / or instruction execution capabilities, and may control other components in the electronic device 11 to perform desired functions.

[0138] The memory 112 may include one or more computer program products, which may include various types of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may include, for example, random access memory (RAM) and / or high-speed cache memory. The non-volatile memory may include, for example, read-only memory (ROM), a hard disk, or flash memory. One or more computer program instructions may be stored in the computer-readable storage medium, and the processor 111 may execute the program instructions to implement the 3D target detection method based on multi-view fusion according to the embodiments of the present disclosure and / or other desired functions. The computer-readable storage medium may also store various contents, such as an input signal, a signal component, and a noise component.

[0139] In one example, electronic device 11 may further include input devices 113 and output devices 114, these components being connected to each other by a bus system and / or other type of connection mechanism (not shown).

[0140] The input device 113 may further include, for example, a keyboard and a mouse.

[0141] The output device 114 can output various information including determined distance information, direction information, etc. The output device 114 may include, for example, a display, a speaker, a printer, a communication network, and a remote output device connected thereto.

[0142] Naturally, for simplicity, Fig. 19 In the figure, only some of the components related to the present disclosure in the electronic device 11 are shown, and components such as buses, input / output interfaces, etc. are omitted. In addition, the electronic device 11 may further include any other appropriate components according to specific application situations. Exemplary Computer Program Products and Computer-Readable Storage Media

[0143] In addition to the above methods and apparatus, embodiments of the present disclosure may also be a computer program product including computer program instructions that, when executed by a processor, cause the processor to perform steps in a 3D target detection method based on multi-view fusion according to various embodiments of the present disclosure, as described in the "Exemplary Method" section above of this specification.

[0144] Additionally, an embodiment of the present disclosure may be a computer-readable storage medium having stored thereon computer program instructions that, when executed by a processor, cause the processor to perform steps in a 3D target detection method based on multi-view fusion according to various embodiments of the present disclosure, as described in the "Exemplary Method" section above of this specification.

[0145] Although the basic principles of the present disclosure have been described above with reference to specific embodiments, it should be noted that the benefits, advantages, effects, etc. mentioned in the present disclosure are merely examples and are not limiting, and these benefits, advantages, effects, etc. are not considered essential to each embodiment of the present disclosure. Furthermore, the specific details disclosed above are merely for illustrative purposes and to facilitate understanding, but are not limiting, and the details do not limit the scope of the present disclosure that must be realized using the specific details.

[0146] Block diagrams of devices, apparatus, instruments, and systems according to the present disclosure are intended to be illustrative examples only and are not intended to require or imply that they must be connected, arranged, or configured in the manner shown in the block diagrams. As will be appreciated by those skilled in the art, these devices, apparatus, instruments, and systems can be connected, arranged, or configured in any manner. For example, terms such as "comprise," "contain," and "have" are open-ended terms and can be used interchangeably to mean "including but not limited to." Unless the context clearly indicates otherwise, the terms "or" and "and" used herein mean and can be used interchangeably with "and / or." The term "for example" used herein means and can be used interchangeably with "for example, but not limited to."

[0147] It should be further noted that in the devices, apparatuses, and methods of the present disclosure, each component or each step can be disassembled and / or reassembled, and such disassembly and / or reassembly should be considered as an equivalent solution of the present disclosure.

Claims

1. A step of acquiring at least two images from multiple camera viewpoints collected by a multi-camera system; performing feature extraction on the at least two images to obtain feature data including corresponding target object features in a multi-camera viewpoint space of the at least two images; mapping corresponding feature data of the at least two images in the multi-camera viewpoint space to the same bird's-eye viewpoint space based on internal parameters of the multi-camera system and vehicle parameters, and obtaining corresponding feature data of the at least two images in the bird's-eye viewpoint space; performing feature fusion on corresponding feature data in the bird's-eye view space of the at least two images to obtain bird's-eye view fusion features; and performing target prediction on the target object in the bird's-eye view fusion features to obtain three-dimensional spatial information of the target object, wherein each included step is performed by an in-vehicle autonomous driving system.

2. The step of mapping corresponding feature data in the multi-camera viewpoint space of the at least two images to the same bird's-eye viewpoint space based on the internal parameters of the multi-camera system and vehicle parameters, and obtaining corresponding feature data in the bird's-eye viewpoint space of the at least two images, comprises: determining a transformation matrix from a camera coordinate system of the multi-cameras of the multi-camera system to a bird's-eye view coordinate system based on internal parameters of the multi-camera system and vehicle parameters; 2. The method of claim 1, further comprising: a step of transforming corresponding feature data in the multi-camera viewpoint space of the at least two images from the multi-camera viewpoint space to the bird's-eye viewpoint space based on a transformation matrix from the camera coordinate system of the multi-camera to the bird's-eye viewpoint coordinate system, thereby obtaining corresponding feature data in the bird's-eye viewpoint space of the at least two images.

3. The step of determining a transformation matrix from a camera coordinate system of the multi-cameras of the multi-camera system to a bird's-eye view coordinate system based on the internal parameters of the multi-camera system and vehicle parameters includes: acquiring internal camera parameters and external camera parameters of the multiple cameras in the multi-camera system, and acquiring a transformation matrix from a vehicle coordinate system to the bird's-eye view coordinate system; and determining a transformation matrix from the camera coordinate systems of the multi-cameras to the bird's-eye view coordinate system based on the camera extrinsic parameters, the camera intrinsic parameters, and a transformation matrix from the vehicle coordinate system to the bird's-eye view coordinate system of the multi-cameras.

4. The step of performing target prediction on the target object in the bird's-eye view fusion feature to obtain three-dimensional space information of the target object includes: Using a prediction network to obtain a corresponding heat map from the bird's-eye view fusion features, for determining a first preset coordinate value in the bird's-eye view coordinate system of the target object, and obtain other attribute maps for determining a second preset coordinate value, size, and orientation angle of the target object in the bird's-eye view coordinate system; determining the first preset coordinate value of the target object in the bird's-eye view coordinate system based on peak information in the heat map, and determining the second preset coordinate value, size, and orientation angle of the target object in the bird's-eye view coordinate system from the other attribute map based on the first preset coordinate value of the target object in the bird's-eye view coordinate system; and determining three-dimensional spatial information of the target object based on the first preset coordinate value, the second preset coordinate value, the size, and the orientation angle of the target object in the bird's-eye view coordinate system.

5. In the training phase of the prediction network, a first loss function is constructed between the heat map predicted by the prediction network and the truth-value heat map, and a second loss function is constructed between the other attribute map predicted by the prediction network and the other truth-value attribute map; 5. The method of claim 4, further comprising: determining a total loss function during a training phase of the prediction network based on the first loss function and the second loss function, and monitoring the training process of the prediction network.

6. determining a total loss function in a training phase of the prediction network based on the first loss function and the second loss function, obtaining weight values ​​of the first loss function and weight values ​​of the second loss function; and determining a total loss function in a training phase of the predictive network based on the first loss function, a weight value of the first loss function, the second loss function, and a weight value of the second loss function.

7. The step of performing target prediction on the target object in the bird's-eye view fusion feature to obtain three-dimensional space information of the target object includes: performing feature extraction on the bird's-eye view fusion features using a neural network to obtain bird's-eye view fusion feature data including the target object features; The method of claim 1 , further comprising: using a prediction network to perform target prediction on the target object in the bird's-eye view fusion feature data including the target object features, and obtaining three-dimensional spatial information of the target object.

8. The step of performing feature extraction on the at least two images to obtain feature data including corresponding target object features in a multi-camera viewpoint space of the at least two images includes: The method of claim 1 , comprising: using a deep neural network to perform a convolution calculation on an image corresponding to each viewpoint, and acquiring feature data of a plurality of different resolutions corresponding to the image corresponding to each viewpoint in the multi-camera viewpoint space and including the target object feature.

9. An image receiving module for acquiring at least two images from multiple camera viewpoints collected by a multi-camera system; a feature extraction module for performing feature extraction on the at least two images acquired by the image receiving module to obtain feature data including corresponding target object features in a multi-camera viewpoint space of the at least two images; an image feature mapping module for mapping corresponding feature data in the multi-camera viewpoint space of the at least two images acquired by the feature extraction module to the same bird's-eye viewpoint space based on internal parameters of the multi-camera system and vehicle parameters, thereby obtaining corresponding feature data in the bird's-eye viewpoint space of the at least two images; an image fusion module for fusing corresponding feature data in the bird's-eye view space of the at least two images obtained by the image feature mapping module to obtain bird's-eye view fusion features; a 3D detection module for performing target prediction on the target object in the bird's-eye view fusion features obtained by the image fusion module, and obtaining three-dimensional spatial information of the target object.

10. A computer-readable storage medium storing a computer program for executing the 3D target detection method based on multi-view fusion according to any one of claims 1 to 8.

11. a processor; a memory for storing instructions executable by said processor; The processor is used to read and execute the executable instructions from the memory to realize the 3D target detection method based on multi-view fusion according to any one of claims 1 to 8. Electronic device.

Citation Information

Patent Citations

  • Intersection multi-view target detection method and system based on angular point pooling

    CN113673444A

  • Section line recognition device

    JP2018097782A