3D Object Detection Method, Device, Electronic Device and Medium

By combining two-dimensional images and point cloud data, the target point cloud data set is determined and correlated, and the problems of low detection accuracy of three-dimensional objects and high false positive rates are solved, achieving high-precision three-dimensional object positioning.

CN112700552BActive Publication Date: 2025-06-27HUAWEI TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202011641585.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-12-31
Publication Date
2025-06-27
Estimated Expiration
2040-12-31

AI Technical Summary

Technical Problem

In some observation angles, the device cannot obtain enough point cloud data, resulting in low detection accuracy of three-dimensional objects and high false positive rate.

Method used

By acquiring the two-dimensional image and point cloud dataset, the target object image in the two-dimensional image is used to determine the target point cloud dataset from the point cloud dataset and correlate it with the target object image to obtain the estimated position of the target object in three-dimensional space.

Benefits of technology

It improves the accuracy of three-dimensional object detection, reduces the false positive rate, and does not need to obtain a large amount of three-dimensional training data, avoiding the problem of poor generalization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112700552B_ABST
    Figure CN112700552B_ABST
Patent Text Reader

Abstract

The present application provides a three-dimensional object detection method, apparatus, electronic device, and medium, which relate to the field of artificial intelligence and can improve the accuracy of three-dimensional object detection. The method includes: a three-dimensional object detection apparatus obtains a two-dimensional image and at least one point cloud data set. Among them, the two-dimensional image includes images of at least one object, and the point cloud data set includes a plurality of point cloud data, and the point cloud data is used to describe the candidate regions of at least one object in three-dimensional space. Then, the three-dimensional object detection apparatus determines a target point cloud data set from at least one point cloud data set according to the target object image in the two-dimensional image. Among them, the point cloud data in the target point cloud data set is used to describe the candidate regions of the target object in three-dimensional space. The three-dimensional object detection apparatus associates the target point cloud data set with the target object image to obtain a detection result. Among them, the detection result indicates the estimated position of the target object in three-dimensional space.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of artificial intelligence (AI), and particularly to a three-dimensional object detection method, apparatus, electronic device, and medium. Background Art

[0002] A robot has the ability to identify objects in the environment, thereby realizing functions such as path planning and obstacle avoidance. Among them, the three-dimensional (3D) spatial dimensions of an object are particularly important for the robot to understand the environment. Exemplarily, after the device obtains the point cloud of the scene, it determines the candidate object region based on the point cloud of the scene, then selects the target points located in the candidate object region from the point cloud, and uses the position information of the target points to adjust the candidate object region, thereby locating the three-dimensional spatial position of the object.

[0003] However, in some observation perspectives, the device cannot obtain enough point clouds, resulting in the inability to identify objects, leading to low three-dimensional object detection accuracy and high false positive rates. Summary of the Invention

[0004] Embodiments of this application provide a three-dimensional object detection method, apparatus, electronic device, and medium, which can improve the accuracy of three-dimensional object detection.

[0005] To achieve the above objective, the embodiments of this application adopt the following technical solutions:

[0006] In a first aspect, embodiments of this application provide a three-dimensional object detection method. The execution subject of this method may be a three-dimensional object detection apparatus. The method includes: obtaining a two-dimensional image and at least one point cloud data set, where the two-dimensional image includes images of at least one object, the point cloud data set includes a plurality of point cloud data, the point cloud data is used to describe the candidate region of the at least one object in the three-dimensional space, the two-dimensional image is information collected by an image sensor, and the point cloud data is information collected by a depth sensor; determining a target point cloud data set from the at least one point cloud data set according to the target object image in the two-dimensional image, where the target object image includes the image of the target object among the at least one object, and the point cloud data in the target point cloud data set is used to describe the candidate region of the target object in the three-dimensional space; associating the target point cloud data set with the target object image to obtain a detection result, where the detection result indicates the estimated position of the target object in the three-dimensional space.

[0007] In this method, a two-dimensional image and at least one point cloud data set are obtained by a three-dimensional object detection device. Among them, the two-dimensional image includes images of at least one object. The point cloud data set includes a plurality of point cloud data, and the point cloud data is used to describe the candidate regions of at least one object in the three-dimensional space. The two-dimensional image is the information collected by an image sensor, and the point cloud data is the information collected by a depth sensor. Then, the three-dimensional object detection device determines a target point cloud data set from at least one point cloud data set according to the target object image in the two-dimensional image. Among them, the target object image includes the image of the target object among at least one object, and the point cloud data in the target point cloud data set is used to describe the candidate region of the target object in the three-dimensional space. Then, the three-dimensional object detection device associates the target point cloud data set with the target object image to obtain a detection result. Among them, the detection result indicates the estimated position of the target object in the three-dimensional space.

[0008] In the three-dimensional object detection method provided in the embodiment of the present application, due to the high processing accuracy of the two-dimensional image, the target object image can accurately present the region of the target object in the two-dimensional image. Using the target object image to screen the target point cloud data set can realize geometric segmentation and clustering of the point cloud data set without obtaining a large amount of three-dimensional training data. Even if the object is occluded, the target point cloud data set can be obtained, which improves the accuracy of the target point cloud data set corresponding to the target object to a certain extent. Moreover, the three-dimensional object detection device associates the target point cloud data set with the target object image to obtain a detection result. Due to the high processing accuracy of the two-dimensional image, even if the point cloud data of the target object is insufficient, the estimated position of the target object in the three-dimensional space can be accurately determined, avoiding the problem of a high false positive rate. The three-dimensional object detection method in the embodiment of the present application does not need to obtain three-dimensional training data, avoiding the problem of "poor generalization" caused by "training a model based on three-dimensional training data".

[0009] In a possible design, the determining the target point cloud data set from the at least one point cloud data set according to the target object image in the two-dimensional image includes: determining a first projection region of the first point cloud data set in the two-dimensional image, where the first point cloud data set is a set among the at least one point cloud data set; and determining the first point cloud data set as the target point cloud data set according to the first projection region and the target image region, where the target image region is the region of the target object image in the two-dimensional image.

[0010] In this method, the three-dimensional object detection device determines a target point cloud dataset from at least one point cloud dataset according to the target object image in the two-dimensional image, including: the three-dimensional object detection device determines a first projection area of the first point cloud dataset in the two-dimensional image. Wherein, the first point cloud dataset is a set among at least one point cloud dataset. Then, the three-dimensional object detection device determines the first point cloud dataset as the target point cloud dataset according to the first projection area and the target image area. Wherein, the target image area is the area of the target object image in the two-dimensional image.

[0011] That is to say, the three-dimensional object detection device determines whether a point cloud dataset is the target point cloud dataset based on two areas (i.e., the target image area and the projection area of a point cloud dataset on the two-dimensional image). Since the target object image belongs to the two-dimensional image, the three-dimensional object detection device has high detection and recognition accuracy for the two-dimensional image. By jointly recognizing the target point cloud dataset with the target object image, the recognition accuracy of the target point cloud dataset can be correspondingly improved.

[0012] In a possible design, the determining the first projection area of the first point cloud dataset in the two-dimensional image includes: determining a first feature point from the feature points represented by the first point cloud dataset according to the depth range of the point clouds in the first point cloud dataset; determining a first projection point of the first feature point in the two-dimensional image according to the conversion parameters between the point cloud data and the two-dimensional image; and taking the area marked by the two-dimensional annotation box corresponding to the first projection point as the first projection area.

[0013] In this method, the three-dimensional object detection device determines the first projection area of the first point cloud dataset in the two-dimensional image, including: the three-dimensional object detection device determines a first feature point from the feature points represented by the first point cloud dataset according to the depth range of the point clouds in the first point cloud dataset, such as the farthest point and the nearest point. Then, the three-dimensional object detection device determines a first projection point of the first feature point in the two-dimensional image according to the conversion parameters between the point cloud data and the two-dimensional image, such as the internal parameters of the depth sensor, the rotation matrix, or the translation matrix. The three-dimensional object detection device takes the area marked by the two-dimensional annotation box corresponding to the first projection point as the first projection area. Exemplarily, the two-dimensional annotation box is a box with the first projection point as the diagonal point.

[0014] That is to say, in the case where the three-dimensional object detection device determines the first feature point in a point cloud dataset, first, the projection point of the first feature point on the two-dimensional image, that is, the first projection point, is determined. Since the first projection point is the projection of the farthest point and the nearest point in the first point cloud dataset on the two-dimensional image. Therefore, the first projection area is the area between the first projection points, that is, the area marked by the two-dimensional annotation box corresponding to the first projection point, thus realizing the accurate projection of the first point cloud dataset on the two-dimensional image.

[0015] In a possible design, determining the first point cloud data set as the target point cloud data set according to the first projection area and the target image area includes: determining the first point cloud data set as the target point cloud data set according to the degree of coincidence between the first projection area and the target image area and the size of the first projection area.

[0016] In this method, the three-dimensional object detection device determines the first point cloud data set as the target point cloud data set according to the first projection area and the target image area, including: the three-dimensional object detection device determines the first point cloud data set as the target point cloud data set according to the degree of coincidence between the first projection area and the target image area and the size of the first projection area.

[0017] That is to say, even if the first projection area coincides with the target image area, if the area of the "first projection area" is too small, the feature points represented by the point cloud data in the first point cloud data set may be part of the target object. Since a part of the target object cannot accurately represent the estimated position of the entire target object in the three-dimensional space, such a point cloud data set is not used as the target point cloud data set. Thus, in the process of determining the target point cloud data set, the three-dimensional object detection device needs to combine two factors, namely, "the degree of coincidence between the first projection area and the target image area" and "the size of the first projection area", to more accurately determine the target point cloud data set.

[0018] In a possible design, the target projection area of the feature points represented by the target point cloud data set in the two-dimensional image satisfies:

[0019]

[0020] where S represents the similarity between the target projection area and the target image area, IOU represents the intersection over union between the target projection area and the target image area, S ∩ represents the overlapping area between the target projection area and the target image area, S ∪ represents the sum of the overlapping area and the non-overlapping area, the non-overlapping area is the area that is not overlapped between the target projection area and the target image area, Lj represents the projection point spacing of the target projection area, the projection point spacing is the distance between the projection points of the target feature points in the two-dimensional image, the target feature points belong to the feature points represented by the target point cloud data set and indicate the depth range of the points in the target point cloud data set, Dij represents the distance between the reference point of the target projection area and the reference point of the target image area, and T represents the similarity threshold.

[0021] In a possible design, the association of the target point cloud dataset and the target object image to obtain a detection result includes: according to the depth range of the points in the target point cloud dataset, inverse mapping some pixel points in the target object image to the three-dimensional space to obtain target inverse mapping points; taking the area marked by the three-dimensional annotation box corresponding to the target inverse mapping points as the detection result.

[0022] In this method, the three-dimensional object detection device associates the target point cloud dataset and the target object image to obtain a detection result, including: the three-dimensional object detection device inverse maps some pixel points in the target object image to the three-dimensional space according to the depth range of the points in the target point cloud dataset to obtain target inverse mapping points. For example, when the target object image is a rectangular area in the two-dimensional image, the pixel points located at the diagonal points are inverse mapped to the three-dimensional space to obtain target inverse mapping points. Then, the three-dimensional object detection device takes the area marked by the three-dimensional annotation box corresponding to the target inverse mapping points as the detection result.

[0023] That is to say, the three-dimensional object detection device uses the target point cloud dataset and the target object image to determine the detection result to avoid the problem of "high false positive rate" caused by the occlusion of the target object and incomplete viewing angle.

[0024] In a possible design, the three-dimensional object detection method of the embodiment of the present application further includes: adjusting the estimated position indicated by the detection result according to a preset adjustment factor, where the adjustment factor indicates the difference between the true position and the estimated position of the target object in the three-dimensional space.

[0025] This method further includes: the three-dimensional object detection device adjusts the estimated position indicated by the detection result according to a preset adjustment factor. The adjustment factor indicates the difference between the true position and the estimated position of the target object in the three-dimensional space, so that the estimated position determined by the three-dimensional object detection device is more in line with the actual object size, improving the accuracy of object detection.

[0026] In a possible design, the number of feature points represented by the point cloud dataset is less than a number threshold to eliminate the point cloud dataset describing the background object, which helps to reduce the computational load of the three-dimensional object detection device.

[0027] Second aspect, an embodiment of the present application provides a three-dimensional object detection device, which may be the device in the above first aspect or any possible design of the first aspect, or a chip that implements the above functions; the three-dimensional object detection device includes corresponding modules, units, or means for implementing the above method, and the module, unit, or means may be implemented by hardware, software, or by hardware executing corresponding software. The hardware or software includes one or more modules or units corresponding to the above functions.

[0028] The three-dimensional object detection device includes an acquisition unit and a processing unit. Among them, the acquisition unit is used to acquire a two-dimensional image and at least one point cloud data set. The two-dimensional image includes an image of at least one object, the point cloud data set includes a plurality of point cloud data, and the point cloud data is used to describe a candidate region of the at least one object in three-dimensional space. The two-dimensional image is information collected by an image sensor, and the point cloud data is information collected by a depth sensor;

[0029] The processing unit is used to determine a target point cloud data set from the at least one point cloud data set according to the target object image in the two-dimensional image, where the target object image includes an image of a target object among the at least one object, and the point cloud data in the target point cloud data set is used to describe the candidate region of the target object in the three-dimensional space;

[0030] The processing unit is further used to associate the target point cloud data set with the target object image to obtain a detection result, where the detection result indicates the estimated position of the target object in the three-dimensional space.

[0031] In a possible design, the processing unit is used to determine a target point cloud data set from the at least one point cloud data set according to the target object image in the two-dimensional image, specifically including:

[0032] Determine a first projection region of the first point cloud data set in the two-dimensional image, where the first point cloud data set is a set among the at least one point cloud data set;

[0033] According to the first projection region and the target image region, determine the first point cloud data set as the target point cloud data set, where the target image region is the region of the target object image in the two-dimensional image.

[0034] In a possible design, the processing unit is used to determine a first projection region of the first point cloud data set in the two-dimensional image, specifically including:

[0035] Determine a first feature point from the feature points represented by the first point cloud dataset according to the depth range of the point clouds in the first point cloud dataset;

[0036] Determine a first projection point of the first feature point in the two-dimensional image according to the conversion parameters between the point cloud data and the two-dimensional image;

[0037] Use the area marked by the two-dimensional bounding box corresponding to the first projection point as the first projection area.

[0038] In a possible design, the processing unit is configured to determine the first point cloud dataset as the target point cloud dataset according to the first projection area and the target image area, specifically including:

[0039] Determine the first point cloud dataset as the target point cloud dataset according to the degree of overlap between the first projection area and the target image area, and the size of the first projection area.

[0040] In a possible design, the target projection area of the feature points represented by the target point cloud dataset in the two-dimensional image satisfies:

[0041]

[0042] where S represents the similarity between the target projection area and the target image area, IOU represents the intersection over union between the target projection area and the target image area, S ∩ represents the overlapping area between the target projection area and the target image area, S ∪ represents the sum of the overlapping area and the non-overlapping area, the non-overlapping area is the area that is not overlapped between the target projection area and the target image area, Lj represents the projection point spacing of the target projection area, the projection point spacing is the distance between the projection points of the target feature points in the two-dimensional image, the target feature points belong to the feature points represented by the target point cloud dataset, and indicate the depth range of the point clouds in the target point cloud dataset, Dij represents the distance between the reference point of the target projection area and the reference point of the target image area, and T represents the similarity threshold.

[0043] In a possible design, the processing unit is configured to associate the target point cloud dataset and the target object image to obtain a detection result, specifically including:

[0044] Inverse map some pixel points in the target object image to the three-dimensional space according to the depth range of the point clouds in the target point cloud dataset to obtain target inverse mapping points;

[0045] Use the area marked by the three-dimensional annotation box corresponding to the target inverse mapping point as the detection result.

[0046] In a possible design, the processing unit is further configured to:

[0047] Adjust the estimated position indicated by the detection result according to a preset adjustment factor, where the adjustment factor indicates the difference between the true position and the estimated position of the target object in the three-dimensional space.

[0048] In a possible design, the number of feature points represented by the point cloud data set is less than a quantity threshold.

[0049] In a third aspect, an embodiment of the present application provides an electronic device, which includes a processor and a memory. The processor and the memory communicate with each other. The processor is configured to execute instructions stored in the memory so that the electronic device performs the three-dimensional object detection method in the first aspect or any one of the designs in the first aspect.

[0050] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, in which instructions are stored, and the instructions are used to instruct a device to perform the three-dimensional object detection method in the first aspect or any one of the designs in the first aspect.

[0051] In a fifth aspect, the present application provides a computer program product containing instructions, which, when running on a device, causes the device to perform the three-dimensional object detection method in the first aspect or any one of the designs in the first aspect.

[0052] In a sixth aspect, an embodiment of the present application provides a chip, including a logic circuit and an input / output interface. The input / output interface is used to communicate with modules outside the chip. For example, the chip can be a chip that implements the function of the three-dimensional object detection device in the first aspect or any one of the possible designs in the first aspect. The input / output interface inputs a two-dimensional image and at least one point cloud data set, and the input / output interface outputs a detection result. The logic circuit is used to run a computer program or instructions to implement the three-dimensional object detection method in the first aspect or any one of the possible designs in the first aspect.

[0053] In a seventh aspect, an embodiment of the present application provides a robot, including: an image sensor, a depth sensor, a processor, and a memory for storing processor-executable instructions. The image sensor is used to collect a two-dimensional image, the depth sensor is used to collect at least one point cloud data set, and the processor is configured with executable instructions to implement the three-dimensional object detection method in the first aspect or any one of the possible designs in the first aspect.

[0054] In an eighth aspect, an embodiment of the present application provides a server, including: a processor and a memory for storing processor-executable instructions. The processor is configured to execute the executable instructions to implement the three-dimensional object detection method in the first aspect or any possible design of the first aspect as described above.

[0055] Among them, the technical effects brought by any design in the second aspect to the eighth aspect can refer to the beneficial effects in the corresponding methods provided above, and will not be elaborated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] Figure 1 FIG. [ID] is a schematic diagram of a system architecture provided by an embodiment of the present application;

[0057] Figure 2 FIG. [ID] is another schematic diagram of a system architecture provided by an embodiment of the present application;

[0058] Figure 3 FIG. [ID] is a schematic flowchart of a three-dimensional object detection method provided by an embodiment of the present application;

[0059] Figure 4a FIG. [ID] is a checkerboard image provided by an embodiment of the present application;

[0060] Figure 4b FIG. [ID] is another schematic flowchart of a three-dimensional object detection method provided by an embodiment of the present application;

[0061] Figure 5 FIG. [ID] is another schematic flowchart of a three-dimensional object detection method provided by an embodiment of the present application;

[0062] Figure 6a FIG. [ID] is a schematic flowchart of a model training stage provided by an embodiment of the present application;

[0063] Figure 6b FIG. [ID] is a schematic flowchart of a model application stage provided by an embodiment of the present application;

[0064] Figure 6c FIG. [ID] is a schematic diagram of a 2D detection frame provided by an embodiment of the present application;

[0065] Figure 7a FIG. [ID] is another schematic flowchart of a three-dimensional object detection method provided by an embodiment of the present application;

[0066] Figure 7b FIG. [ID] is a schematic diagram of normal vector estimation provided by an embodiment of the present application;

[0067] Figure 8 FIG. [ID] is another schematic flowchart of a three-dimensional object detection method provided by an embodiment of the present application;

[0068] Figure 9aA schematic diagram of the positions of the farthest point and the nearest point provided by an embodiment of the present application;

[0069] Figure 9b A schematic diagram of the position of a projection area provided by an embodiment of the present application;

[0070] Figure 9c A schematic diagram of the positions of a target projection area and a target image area provided by an embodiment of the present application;

[0071] Figure 10 A schematic flowchart of another three-dimensional object detection method provided by an embodiment of the present application;

[0072] Figure 11 A schematic diagram of the structure of another device provided by an embodiment of the present application. Detailed implementation manners

[0073] In the specification and drawings of the present application, terms such as "first" and "second" are used to distinguish different objects or different treatments of the same object, rather than to describe a specific order of the objects. In addition, the terms "including" and "having" and any variations thereof mentioned in the description of the present application are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes other unlisted steps or units, or optionally further includes other steps or units inherent to these processes, methods, products, or devices. It should be noted that in the embodiments of the present application, words such as "exemplarily" or "for example" are used to indicate examples, illustrations, or explanations. Any embodiment or design solution described as "exemplarily" or "for example" in the embodiments of the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Rather, the use of words such as "exemplarily" or "for example" is intended to present related concepts in a specific manner.

[0074] To make the present application clearer, some concepts and processing flows mentioned in the present application are briefly introduced first.

[0075] 1. Robustness

[0076] Robustness refers to the ability of a system to survive in abnormal and dangerous situations, or the characteristic of a control system to maintain certain other performances under a certain (structure, size) parameter perturbation.

[0077] 2. False positive rate

[0078] The false positive rate refers to the probability that "the result obtained through the deep learning model is a wrong positive class", that is, the probability that the deep learning model determines a non-target sample as a correct target sample.

[0079] 3. Two - dimensional (2D) images

[0080] A two - dimensional image refers to a planar image that does not contain depth information. Two - dimensional images can include red - green - blue (RGB) images, grayscale images, etc.

[0081] 4. Depth images

[0082] A depth image, also known as a range image, is an image that uses the distance (or depth) from a depth sensor to each point in space as the pixel value. A depth image directly reflects the geometric shape of the visible surface of objects in space.

[0083] 5. Point cloud data

[0084] A point cloud refers to a set of points that express the spatial distribution and surface characteristics of a target object under a certain spatial reference system. In the embodiments of the present application, point cloud data is used to represent the three - dimensional coordinate values of each point in the point cloud under the spatial reference coordinate system. The spatial reference coordinate system can be the coordinate system corresponding to the depth sensor.

[0085] 6. Point cloud clusters

[0086] A point cloud cluster refers to the points represented by a part of the point cloud data that satisfies the preset partitioning rules after a series of calculations (such as geometric segmentation, clustering processing, etc.) on the point cloud data. Among them, the calculation methods can be clustering methods based on point cloud data density, nearest - neighbor methods based on kdtree, k - means methods, and deep - learning methods, etc.

[0087] In the embodiments of the present application, the point cloud data corresponding to a point cloud cluster is described as a "point cloud data set".

[0088] 7. Three - dimensional (3D) object detection

[0089] Three - dimensional object detection can provide an object map to enable the robot to better position itself. Since objects are the basis for the robot to understand the environment, objects can be used as a kind of semantics to improve the navigation intelligence of the robot. Three - dimensional object detection can extend the object from the image plane to the real world and better realize human - robot interaction. Below, the implementation process of a three - dimensional object detection method based on deep learning is given:

[0090] After the device obtains the point cloud of the scene, it determines the candidate object region based on the point cloud of the scene, then selects the target points located in the candidate object region from the point cloud, and uses the position information of the target points to adjust the candidate object region, so as to locate the three-dimensional spatial position of the object. However, in some observation perspectives, the device cannot obtain enough point clouds, resulting in the inability to identify the object, thus leading to low three-dimensional object detection accuracy and high false positive rate.

[0091] In view of this, an embodiment of the present application provides a three-dimensional object detection method. The three-dimensional object detection method provided by the embodiment of the present application can be applied to a device as shown in Figure 1 The device includes a first device 101 and a second device 102. The first device 01 is an image acquisition device, and the image acquisition device includes an image sensor and a depth sensor. Among them, the image sensor is used to acquire two-dimensional images, such as RGB images, grayscale images, etc. The image sensor can be, for example but not limited to, the following introductions: RBG camera, digital single-lens reflex (DSLR) camera, point-and-shoot camera, video camera, wearable device, augmented reality (AR) device, virtual reality (VR) device, vehicle-mounted device, smart screen, etc. The depth sensor is used to acquire depth images. The depth sensor can be, for example but not limited to, the following introductions: depth camera, time of flight (TOF) camera, or lidar, photogrammetric scanner, or light detection and ranging (LiDAR) sensor. The second device 102 is a processing device, and the processing device has a central processing unit (CPU) and / or a graphics processing unit (GPU), and is used to process the images acquired by the image acquisition device, so as to realize three-dimensional object detection.

[0092] It should be noted that the first device 101 and the second device 102 can be set on the robot body, as shown in Figure 1 For example, the first device 101 and the second device 102 can be set on the head of the robot ( Figure 1 not shown), or can be set on the body part of the robot, as shown in Figure 1 Of course, the first device 101 and the second device 102 can also be set on other parts of the robot body, and the embodiment of the present application does not limit this.

[0093] In addition, the first device 101 and the second device 102 can be independent devices or integrated together. For example, the first device 101 is a part of the second device 102. In this case, the first device 101 and the second device 102 are connected through a bus. Exemplarily, the bus can be implemented as a bidirectional synchronous serial bus, and the bidirectional synchronous serial bus includes a serial data line (SDA) and a serial clock line (SCL). In this case, the first device 101 and the second device 102 include an inter-integrated circuit (I2C) interface. The first device 101 and the second device 102 communicate through the bidirectional synchronous serial bus connected by the I2C interface. Alternatively, the first device 101 and the second device 102 include a mobile industry processor interface (MIPI) interface. The first device 101 and the second device 102 communicate through the bidirectional synchronous serial bus connected by the MIPI interface. Alternatively, the first device 101 and the second device 102 include a general-purpose input / output (GPIO) interface. The first device 101 and the second device 102 communicate through the bidirectional synchronous serial bus connected by the GPIO interface.

[0094] In the embodiments of the present application, the case where "the first device 101 and the second device 102 are independent devices" is taken as an example for description. In the case where "the first device 101 and the second device 102 are independent devices", the first device 101 and the second device 102 can be arranged at different positions. For example, the first device 101 is arranged on the body part of the robot, and the second device 102 is arranged outside the robot body, such as Figure 2As shown. In this case, the second device 102 can be a physical device or a cluster of physical devices, such as a terminal, a server, or a cluster of servers. The second device 102 can also be a virtualized cloud device, such as at least one cloud computing device in a cloud computing cluster. Both the first device 101 and the second device 102 can include devices or chips that support wireless communication technologies. Among them, wireless communication technologies can be, for example but not limited to, the following introductions: near field communication (NFC) technology, infrared (IR) technology, global system for mobile communications (GSM), general packet radio service (GPRS), code division multiple access (CDMA), wideband code division multiple access (WCDMA), time-division code division multiple access (TD-SCDMA), long term evolution (LTE), bluetooth (BT), global navigation satellite system (GNSS), or frequency modulation (FM), etc. GNSS can include global positioning system (GPS), global navigation satellite system (GLONASS), beidou navigation satellite system (BDS), quasi-zenith satellite system (QZSS), and / or satellite based augmentation systems (SBAS).

[0095] Among them, Figure 1 and Figure 2 the robots in can be service robots, such as floor cleaning robots in a home environment, robots for door-to-door delivery, children's education robots, etc. Figure 1 and Figure 2 the robots in can also be mechanical robots, such as robots for transporting goods in a factory. Additionally, in Figure 1 andFigure 2 In this example, only a robot is described. The robot can be replaced by intelligent household appliances such as smart speakers, smart TVs, etc., to locate the estimated position of a human body in a three-dimensional space, so as to switch its own working state. For example, when the estimated position of the human body located by the smart speaker in the three-dimensional space is greater than a certain threshold, the audio playback stops. Conversely, when the estimated position of the human body located by the smart speaker in the three-dimensional space is less than a certain threshold, the audio playback starts. Figure 1 and Figure 2 The robot in can also be replaced by a drone, such as a drone for door-to-door delivery, a drone for monitoring forest fire risks, a drone for spraying pesticides and fertilizers, etc.

[0096] To make the technical solution of the present application clearer and easier to understand, the first three-dimensional object detection method provided by the embodiments of the present application will be introduced in the following two stages:

[0097] The embodiments of the present application also provide a second three-dimensional object detection method, which includes two stages, and the specific description is as follows:

[0098] The first stage is the acquisition stage. In this stage, the three-dimensional object detection device acquires the point cloud data corresponding to the two-dimensional image and the depth image. Refer to Figure 3 , and the steps of this stage are introduced as follows:

[0099] S301a. The first device acquires a two-dimensional image.

[0100] Among them, the two-dimensional image includes the planar image of at least one object in the first scene. The first scene is the scene within the scanning range of the first device. For example, when the first device is in the living room, the first scene can be the scene within the range scanned by the first device in the living room, and the objects in the first scene can be, for example but not limited to, people, televisions, tables, chairs, or sofas, etc. When the first device is in the bedroom, the first scene can be the scene within the range scanned by the first device in the bedroom, and the objects in the first scene can be, for example but not limited to, beds or wardrobes, etc. When the first device is in the kitchen, the first scene can be the scene within the range scanned by the first device in the kitchen, and the objects in the first scene can be, for example but not limited to, refrigerators, wine glasses, or plates, etc. When the first device is on the transportation lane, the first scene can be the lane scene scanned by the first device, and the objects in the first scene can be, for example but not limited to, vehicles or tracks, etc. When the first device is monitoring forest fire risks, the first scene can be the forest scene scanned by the first device, and the objects in the first scene can be, for example but not limited to, trees or obstacles, etc.

[0101] Exemplarily, the first device may include an image sensor, and the image sensor can be, for example but not limited to Figure 1The example in . A two-dimensional image is collected by an image sensor.

[0102] S302a. The first device sends the two-dimensional image to the three-dimensional object detection device. Correspondingly, the three-dimensional object detection device receives the two-dimensional image from the first device.

[0103] Exemplarily, when the three-dimensional object detection devices in the first device and the second device are connected by a wired connection, the first device sends the two-dimensional image to the three-dimensional object detection device in the second device through a bus. Correspondingly, the three-dimensional object detection device in the second device receives the two-dimensional image from the first device through the bus. For the introduction of "bus", reference can be made to the relevant description in Figure 1 and will not be elaborated here. When the three-dimensional object detection devices in the first device and the second device communicate through wireless communication technology, the first device sends the two-dimensional image to the three-dimensional object detection device in the second device through the Internet. Correspondingly, the three-dimensional object detection device in the second device receives the two-dimensional image from the first device through the Internet. For the introduction of "wireless communication technology", reference can be made to the relevant description in Figure 1 and will not be elaborated here.

[0104] S301b. The first device collects a depth image.

[0105] Among them, the depth image includes an image composed of the depth values of at least one object in the first scene. For the introduction of "the first scene" and "object", reference can be made to the relevant description in S301a and will not be elaborated here.

[0106] Exemplarily, the first device may include a depth sensor, and the depth sensor may be, for example but not limited to Figure 1 the example in . The depth image is collected by the depth sensor.

[0107] S302b. The first device sends the depth image to the three-dimensional object detection device. Correspondingly, the three-dimensional object detection device receives the depth image from the first device.

[0108] Exemplarily, when the three-dimensional object detection devices in the first device and the second device are connected by a wired connection, the first device sends the depth image to the three-dimensional object detection device in the second device through a bus. Correspondingly, the three-dimensional object detection device in the second device receives the depth image from the first device through the bus. For the introduction of "bus", reference can be made to the relevant description in Figure 1For the relevant descriptions in [reference], they will not be elaborated here. When the 3D object detection devices in the first device and the second device communicate through wireless communication technology, the first device sends a depth image to the 3D object detection device in the second device via the Internet. Correspondingly, the 3D object detection device in the second device receives the depth image from the first device via the Internet. For the introduction of "wireless communication technology", reference can be made to Figure 1 the relevant descriptions in [reference], which will not be elaborated here.

[0109] S303b. The 3D object detection device back-projects the pixel points in the depth image into the coordinate system of the depth sensor to obtain point cloud data in 3D space.

[0110] Exemplarily, the 3D object detection device uses the internal parameters of the depth sensor to back-project the pixel point coordinates (u′, v′) of the depth image into the coordinate system of the depth sensor to obtain point cloud data in 3D space. Among them, the relationship between the point cloud data in 3D space and the pixel point coordinates of the depth image satisfies the following formula:

[0111]

[0112] where u′ represents the abscissa of the pixel point in the depth image, and v′ represents the ordinate of the pixel point in the depth image. x represents the coordinate of the pixel point on the x-axis in the coordinate system of the depth sensor (or the coordinate of the point cloud data in 3D space on the x-axis), y represents the coordinate of the pixel point on the y-axis in the coordinate system of the depth sensor (or the coordinate of the point cloud data in 3D space on the y-axis), and z represents the coordinate of the pixel point on the z-axis in the coordinate system of the depth sensor (or the coordinate of the point cloud data in 3D space on the z-axis). K1 -1 represents the inverse matrix of the internal parameters of the depth sensor.

[0113] It should be noted that the internal parameters K1 of the depth sensor and the internal parameters K2 of the image sensor are pre-calibrated parameters. The calibration process can be, for example but not limited to, the following introduction:

[0114] First, the 3D object detection device acquires multiple sets of checkerboard images at different angles.

[0115] Among them, each set of checkerboard images in the above-mentioned "multiple sets of checkerboard images at different angles" can include a 2D image and a depth image, and they are images collected by the image sensor and the depth sensor at the same time. The checkerboard is a checkerboard of A4 paper size with black and white squares, and the grid distribution can be 10 rows and 8 columns, as Figure 4a shown. In Figure 4a the checkerboard, the squares filled with slashes represent black squares, and the squares without slashes represent white squares.

[0116] Then, the 3D object detection device calculates the coordinates of the diagonal points of the checkerboard in the checkerboard image by the Gauss-Newton method to obtain the internal parameters of the camera, that is, the internal parameter K1 of the depth sensor and the internal parameter K2 of the image sensor.

[0117] In addition, the 3D object detection device can also determine the external parameters of the camera based on the 2D image and the point cloud data in the 3D space.

[0118] Among them, the external parameters of the camera include a rotation matrix and a translation matrix. Exemplarily, referring to Figure 4b , the process of "the 3D object detection device determines the external parameters of the camera" is introduced:

[0119] S3041. The 3D object detection device transforms the point cloud data in the 3D space into the coordinate system of the image sensor to obtain the first coordinate.

[0120] Among them, the first coordinate refers to the coordinate of the point cloud data in the coordinate system of the image sensor.

[0121] Exemplarily, the following formula is satisfied between the first coordinate and the point cloud data in the 3D space:

[0122]

[0123] Among them, x represents the coordinate of the point cloud data in the 3D space on the x-axis, y represents the coordinate of the point cloud data in the 3D space on the y-axis, and z represents the coordinate of the point cloud data in the 3D space on the z-axis. x' represents the coordinate of the point represented by the point cloud data in the 3D space on the x-axis in the coordinate system of the image sensor, y' represents the coordinate of the point represented by the point cloud data in the 3D space on the y-axis in the coordinate system of the image sensor, and z' represents the coordinate of the point represented by the point cloud data in the 3D space on the z-axis in the coordinate system of the image sensor. r represents a 3*3 rotation matrix, and t represents a 3*1 translation matrix.

[0124] S3042. The 3D object detection device transforms the first coordinate into the 2D image coordinate system to obtain the second coordinate.

[0125] Among them, the second coordinate is the coordinate of the point cloud data in the 3D space (that is, the point cloud data determined by S303b) in the 2D image coordinate system.

[0126] Exemplarily, the first coordinate and the second coordinate satisfy the following formula:

[0127]

[0128] Wherein, x′ represents the coordinate of the point represented by the point cloud data in the 3D space on the x-axis in the coordinate system of the image sensor, y′ represents the coordinate of the point represented by the point cloud data in the 3D space on the y-axis in the coordinate system of the image sensor, and z′ represents the coordinate of the point represented by the point cloud data in the 3D space on the z-axis in the coordinate system of the image sensor. u represents the abscissa of the point represented by the point cloud data in the 3D space in the two-dimensional image coordinate system, and v represents the ordinate of the point represented by the point cloud data in the 3D space in the two-dimensional image coordinate system. K2 represents the internal parameters of the image sensor.

[0129] S3043. The three-dimensional object detection device determines the external parameters of the camera according to the pixel point coordinates and the second coordinates in the depth image.

[0130] Exemplarily, the three-dimensional object detection device determines the error (u - u′, v - v′) between the pixel point coordinates (u, v) and the second coordinates (u′, v′) in the depth image, and adjusts the rotation matrix r and the translation matrix t based on the error. The three-dimensional object detection device repeats the above S3041 to S3043 to determine the rotation matrix R and the translation matrix T corresponding to the minimum error.

[0131] The second stage is the detection stage. Refer to Figure 5 , in this stage, the three-dimensional object detection device detects the point cloud data corresponding to the two-dimensional image and the depth image to determine the estimated position of the target object in the 3D space. Wherein, the target object is one of at least one object. The specific steps of the second stage are introduced as follows:

[0132] First, the processing process of the "two-dimensional image" is described:

[0133] S501a. The three-dimensional object detection device detects the two-dimensional image to obtain the detection result of the two-dimensional image.

[0134] Wherein, the detection result of the two-dimensional image includes at least the target object image. The target object image is the image of the target object among at least one object.

[0135] Exemplarily, the implementation process of S501a is as follows: The three-dimensional object detection device inputs the two-dimensional image into the 2D object detection model, and uses the 2D object detection model to detect the two-dimensional image to obtain the detection result of the two-dimensional image. Among them, the 2D object detection model can be, for example but not limited to, the following introduction: SSD (single shot multibox detector) model, DSSD (deconvolution single shot multibox detector) model, YoloV4 or other self-developed models, etc. Exemplarily, the 2D object detection model can be a pre-trained model. Refer to Figure 6a, the steps in the model training phase are described as follows:

[0136] Step a1, Image data annotation. In this step, the two-dimensional images in the pre-acquired sample set are annotated.

[0137] Step a2, Data augmentation. In this step, data augmentation processing is performed on the annotated two-dimensional images, such as brightness transformation, to obtain the images after data augmentation.

[0138] Step a3, Input into the neural network. In this step, the images after data augmentation are input into a neural network, such as a convolutional neural network.

[0139] Step a4, Calculate the loss function. In this step, a convolutional neural network is used to calculate the feature vectors between the input images after data augmentation and the annotated information. This process is called "calculating the loss function".

[0140] Step a5, Save the training weights. In this step, after the above training process, the three-dimensional object detection device saves the weights calculated by the convolutional neural network.

[0141] In this way, after steps a1 to a5, the three-dimensional object detection device can obtain a 2D object detection model.

[0142] In the model application phase, which is the implementation process of S502a. Refer to Figure 6b , the steps in the model application phase are described as follows:

[0143] Step a6, Determine the two-dimensional image. In this step, the three-dimensional object detection device determines the two-dimensional image to be processed, that is, the two-dimensional image obtained in S501a

[0144] in the process.

[0145] Step a7, Load the training weights and network model. In this step, the three-dimensional object detection device loads the training weights and network model to construct a 2D object detection model and inputs the two-dimensional image into the 2D object detection model.

[0146] Step a8, Forward propagation. In this step, the 2D object detection model is used to calculate the input two-dimensional image. This process can be described as "forward propagation".

[0147] Step a9: Predict 2D detection boxes. In this step, the 3D object detection device uses a 2D object detection model to detect the 2D image to obtain the target object image. Exemplarily, the 3D object detection device uses 2D detection boxes to identify the target object image. Among them, the 2D detection box can be a rectangular box, including the pixel point coordinates (x, y) of the upper left corner, the width parameter, and the height parameter. Exemplarily, the number of target objects is denoted as N, N≥1. The detection result of the 2D image of the i-th target object is denoted as DR = {Oi}. Among them, DR represents the detection result of the 2D image, and Oi represents the 2D detection box parameters of the i-th target object. 1≤i≤N. Exemplarily, participating Figure 6c , Figure 6c shows two target objects, and the 2D detection boxes respectively identify the images of a person and a chair, as Figure 6c shown by the thick solid lines in

[0148] Optionally, the detection result of the 2D image further includes at least one of the following:

[0149] First, the category of the target object. Among them, the category of the target object can be, for example but not limited to, a person, a table, a chair, etc.

[0150] Second, the confidence level. Among them, the confidence level indicates the credibility of the detection result of the 2D image. The value of the confidence level is not greater than 1. The higher the value of the confidence level, the higher the credibility of the detection result of the 2D image. Exemplarily, when the confidence level is greater than the confidence level threshold, the 3D object detection device executes S502. Conversely, when the confidence level is less than or equal to the confidence level threshold, the 3D object detection device re-executes step a8 and step a9 until the confidence level exceeds the confidence level threshold, or the number of repetitions of the 2D image in step a6 reaches the first preset value. Since the 3D object detection result is determined based on the target object image, and the target object image is an image that meets the confidence level requirement, the 3D object detection method in the embodiments of the present application can accurately screen out the target point cloud data set, which helps to improve the accuracy of the 3D object detection result.

[0151] Then, the processing process of "the point cloud data corresponding to the depth image" will be described:

[0152] S501b: The 3D object detection device clusters the point cloud data corresponding to the depth image to obtain at least one point cloud data set.

[0153] Among them, the "point cloud data in a point cloud data set" is a part of the "point cloud data obtained in S303b" above. The point cloud data in the point cloud data set is used to describe the candidate area of the object in the first scene. Among them, the points represented by a "point cloud data set" can also be described as a "point cloud cluster".

[0154] Exemplarily, such as Figure 7a shown, the implementation process of S502b can be introduced as follows, for example:

[0155] Step b1, filtering. In this step, the three-dimensional object detection device performs downsampling on the "point cloud data obtained in S303b" to improve the calculation efficiency.

[0156] Step b2, normal estimation. In this step, the three-dimensional object detection device performs normal estimation on the "point cloud data after downsampling in step b1" to determine the surface normal.

[0157] Exemplarily, referring to Figure 7b , taking a sampling point Pi as an example, from the points represented by the "point cloud data after downsampling in step b1", determine the points that meet the first preset condition. For example, the first preset condition can be implemented as: points within a circular area with a radius of 3 cm. Taking the "K points that meet the first preset condition" as an example, create a covariance matrix C according to the coordinates of the K points. Then, decompose the eigenvalues and eigenvectors of the covariance matrix C. Among them, the covariance matrix C satisfies the following formula:

[0158]

[0159] where C represents the covariance matrix, K represents the number of points that meet the first preset condition, Pi represents the i-th sampling point among the K points, represents the average value of the coordinates of the K points, λ i is the i-th eigenvalue of the covariance matrix C, is the j-th eigenvector. The eigenvector with the smallest eigenvalue and in the same direction as the sensing direction of the depth sensor is used as the normal.

[0160] Step b3, plane detection. First, perform clustering based on the normal direction, that is, cluster the normals that meet the Euclidean distance constraint, and find the point cloud data set S composed of points with similar normal directions. Then, perform clustering based on the spatial position, that is, cluster the points in the point cloud data set S, and find the points that meet the Euclidean distance. Finally, substitute the points that meet the Euclidean distance into the plane equation to calculate the least squares solution in the form of AX = B. Among them, the plane equation AX = B satisfies the following formula:

[0161]

[0162] where x1 represents the x-axis coordinate of the first point in the "points that meet the Euclidean distance" in the depth sensor coordinate system, y1 represents the y-axis coordinate of the first point in the "points that meet the Euclidean distance" in the depth sensor coordinate system, z1 represents the z-axis coordinate of the first point in the "points that meet the Euclidean distance" in the depth sensor coordinate system. x mrepresents the x-axis coordinate of the m-th point among the "points satisfying the Euclidean distance" in the depth sensor coordinate system, y m represents the y-axis coordinate of the m-th point among the "points satisfying the Euclidean distance" in the depth sensor coordinate system, z m represents the z-axis coordinate of the m-th point among the "points satisfying the Euclidean distance" in the depth sensor coordinate system. The analytical solution is X = (AA T ) -1 A T B, which is the normal vector to be found. a represents the x-axis coordinate of the normal vector, b represents the y-axis coordinate of the normal vector, and c represents the z-axis coordinate of the normal vector. In this way, the three-dimensional object detection device solves for the values of a, b, and c, thereby obtaining the fitting plane.

[0163] Step b4, Euclidean clustering.

[0164] First, determine the number of points in the fitting plane in step b3, and eliminate the fitting planes with the number of points greater than the number threshold. Since the depth image includes a large amount of background images, such as the image of the ground, there are a large number of pixel points of background objects in the depth image. If the number of points in a certain fitting plane is greater than the number threshold, the probability that this fitting plane belongs to the image area of the background object is relatively large. Correspondingly, the probability that this fitting plane belongs to the image area of the target object is relatively small, and it needs to be eliminated to improve the calculation efficiency.

[0165] Then, cluster the points in the remaining fitting planes, and form a point cloud data set with the coordinates of the points that satisfy the Euclidean distance condition as the point cloud data set of the depth image. Among them, the Euclidean distance condition can be, for example but not limited to, the following description: the Euclidean distance between two points in the fitting plane is less than the distance threshold. The distance threshold can be 2CM or other values, which can be determined according to debugging experience or experimental tests.

[0166] Exemplarily, the depth image includes images of N objects, and the point cloud data set corresponding to the depth image is denoted as S = {Ci}. Among them, Ci represents the point cloud data set of the i-th object.

[0167] In this way, through the above steps b1 to b4, the three-dimensional object detection device can obtain at least one point cloud data set of the depth image.

[0168] It should be noted that the three-dimensional object detection device can first execute the processing steps of the two-dimensional image (i.e., S501a), and then execute the processing steps of the point cloud data (i.e., S501b), or first execute the processing steps of the point cloud data, and then execute the processing steps of the two-dimensional image, or can also execute the processing steps of the two-dimensional image and the processing steps of the point cloud data simultaneously. The embodiments of the present application do not limit this.

[0169] Finally, the processing procedures for the target object image and the point cloud data set will be described again:

[0170] S502. The three-dimensional object detection device determines a target point cloud data set from at least one point cloud data set according to the target object image.

[0171] Among them, the point cloud data in the target point cloud data set is used to describe the estimated area where the target object exists in the first scenario. In the embodiments of the present application, the points represented by the "target point cloud data set" can also be described as "target point cloud clusters".

[0172] Exemplarily, referring to Figure 8 , one of the at least one point cloud data sets is described as the "first point cloud data set". Among them, the points represented by the "first point cloud data set" can also be described as "first point cloud clusters". Taking the first point cloud data set as an example, the "determination process of the target point cloud data set" will be introduced in the case of "projecting the first point cloud data set onto a two-dimensional image":

[0173] S5021. The three-dimensional object detection device determines a first projection area of the first point cloud data set in the two-dimensional image.

[0174] Among them, the two-dimensional image is the image obtained in S501a. Exemplarily, the implementation process of S5021 is as follows:

[0175] Step 1. The three-dimensional object detection device determines first feature points from the feature points represented by the first point cloud data set according to the depth range of the point cloud in the first point cloud data set.

[0176] Exemplarily, the first feature points can be at least one of the following: the farthest point among the feature points represented by the first point cloud data set, the nearest point among the feature points represented by the first point cloud data set.

[0177] Exemplarily, the first point cloud data set is denoted as the point cloud data set Ci. The three-dimensional object detection device searches for the farthest point Pmax and the nearest point Pmin in the point cloud data set Ci as the first feature points.

[0178] Step 2. The three-dimensional object detection device determines a first projection point of the first feature point in the two-dimensional image according to the conversion parameters between the point cloud data and the two-dimensional image.

[0179] Exemplarily, the conversion parameters between the point cloud data and the two-dimensional image can be at least one of the following: the internal parameter K1 of the depth sensor, the rotation matrix R, the translation matrix T.

[0180] Exemplarily, taking the farthest point Pmax as an example, first use formula (6) to determine the coordinates of the farthest point Pmax in the coordinate system of the image sensor.

[0181]

[0182] Among them, x max represents the coordinate of the farthest point Pmax on the x-axis in 3D space, and y max represents the coordinate of the farthest point Pmax on the y-axis in 3D space, and z max represents the coordinate of the farthest point Pmax on the z-axis in 3D space. x′ max represents the coordinate of the farthest point Pmax on the x-axis in the coordinate system of the image sensor, and y′ max represents the coordinate of the farthest point Pmax on the y-axis in the coordinate system of the image sensor, and z′ max represents the coordinate of the farthest point Pmax on the z-axis in the coordinate system of the image sensor. R represents a 3×3 rotation matrix, and T represents a 3×1 translation matrix.

[0183] Then, use formula (7) to determine the coordinates of the farthest point Pmax in the coordinate system of the two-dimensional image.

[0184]

[0185] Among them, x′ max represents the coordinate of the farthest point Pmax on the x-axis in the coordinate system of the image sensor, and y′ max represents the coordinate of the farthest point Pmax on the y-axis in the coordinate system of the image sensor, and z′ max represents the coordinate of the farthest point Pmax on the z-axis in the coordinate system of the image sensor. u max represents the abscissa of the farthest point Pmax in the two-dimensional image coordinate system, and v max represents the ordinate of the farthest point Pmax in the two-dimensional image coordinate system. K2 represents the internal parameters of the image sensor.

[0186] Taking the nearest point Pmin as an example, first use formula (8) to determine the coordinates of the nearest point Pmin in the coordinate system of the image sensor.

[0187]

[0188] Among them, x min represents the coordinate of the nearest point Pmin on the x-axis in 3D space, and y min represents the coordinate of the nearest point Pmin on the y-axis in 3D space, and z min represents the coordinate of the nearest point Pmin on the z-axis in 3D space. x′ min represents the coordinate of the nearest point Pmin on the x-axis in the coordinate system of the image sensor, and y′ min represents the coordinate of the nearest point Pmin on the y-axis in the coordinate system of the image sensor, and z′ minIndicates the z - axis coordinate of the nearest point Pmin in the coordinate system of the image sensor. R represents a 3×3 rotation matrix, and T represents a 3×1 translation matrix.

[0189] Then, use formula (9) to determine the coordinates of the nearest point Pmin in the coordinate system of the two - dimensional image.

[0190]

[0191] Among them, x′ min Indicates the x - axis coordinate of the nearest point Pmin in the coordinate system of the image sensor, y′ min Indicates the y - axis coordinate of the nearest point Pmin in the coordinate system of the image sensor, z′ min Indicates the z - axis coordinate of the nearest point Pmin in the coordinate system of the image sensor. u min Indicates the abscissa of the nearest point Pmin in the two - dimensional image coordinate system, v min Indicates the ordinate of the nearest point Pmin in the two - dimensional image coordinate system. K2 represents the internal parameters of the image sensor.

[0192] Step 3: The three - dimensional object detection device uses the area marked by the two - dimensional annotation box corresponding to the first projection point as the first projection area.

[0193] That is to say, the area marked by the two - dimensional annotation box on the two - dimensional image is the first projection area.

[0194] Exemplarily, the two - dimensional annotation box can be a rectangular box, as Figure 9b shown. The two - dimensional annotation box can be an annotation box with the first projection point as the diagonal point.

[0195] In this way, the three - dimensional object detection device can determine the first projection area of the first point cloud dataset in the two - dimensional image, and then determine whether the first point cloud dataset is the target point cloud dataset.

[0196] S5022: The three - dimensional object detection device determines the target image area of the target object image in the two - dimensional image.

[0197] Exemplarily, the target image area can be the area indicated by the 2D detection box parameters in S501a. For specific reference, see the introduction in S501a, which will not be elaborated here.

[0198] S5023: The three - dimensional object detection device determines that the first point cloud dataset is the target point cloud dataset according to the first projection area and the target image area.

[0199] Among them, there are various implementation methods for S5023, which can be, for example but not limited to, the following introduction:

[0200] The three-dimensional object detection device determines the first point cloud data set as the target point cloud data set according to the degree of overlap between the first projection area and the target image area, and the size of the first projection area.

[0201] That is to say, when determining "whether the first point cloud data set is the target point cloud data set", in addition to considering "the degree of overlap between the first projection area and the target image area", the three-dimensional object detection device also refers to the index of "the size of the first projection area". If the area of the "first projection area" is small, the feature points represented by the point cloud data in the first point cloud data set may be part of the target object. For example, when the target object is a "chair", the feature points represented by the point cloud data in the first point cloud data set may belong to the "backrest" part or the "armrest" part. In this case, there is still overlap between the first projection area and the target image area, but a part of the target object cannot accurately represent the estimated position of the entire target object in the three-dimensional space. Therefore, such a point cloud data set is not used as the target point cloud data set. Considering the above two indexes helps to improve the accuracy of screening the target point cloud data set.

[0202] Exemplarily, the implementation process of S5023 is described through two examples:

[0203] Example 1. The target projection area of the feature points represented by the target point cloud data set in the two-dimensional image satisfies:

[0204]

[0205] Among them, S s represents the similarity between the target projection area and the target image area. IOU s represents the intersection over union between the target projection area and the target image area. S ∩ represents the area of the intersection (or overlapping area) between the target projection area and the target image area, and S ∪ represents the area of the union (or the sum of the overlapping area and the non-overlapping area) between the target projection area and the target image area. Lj1 represents the projection point spacing of the target projection area, and the projection point spacing is the distance between the projection points of the target feature points in the two-dimensional image. The target feature points belong to the feature points represented by the target point cloud data set and indicate the depth range of the feature points represented by the target point cloud data set. Dij1 represents the distance between the reference point of the target projection area and the reference point of the target image area. Among them, the reference point can be the center point, the vertex of the upper left corner, the center point of the side, etc. For example, the reference point of the target projection area can be the center point of the target projection area, the vertex of the upper left corner, the center point of the left side, etc. Similarly, the reference point of the target image area can also be the center point of the target image area, the vertex of the upper left corner, the center point of the left side, etc. T sIndicates the similarity threshold.

[0206] Take Figure 9c as an example. The target projection area is denoted as Ri, and the target image area is denoted as Oi. The overlapping area between the two is as shown in the area filled with diagonal lines in Figure 9c , and the non-overlapping area between the two is as shown in the area without diagonal lines filled in Figure 9c . S ∩ represents the above overlapping area, and S ∪ represents the sum of the above overlapping area and the non-overlapping area. Lj1 represents the projection point spacing of the target projection area, as shown by the diagonal of Ri in Figure 9c . Dij1 represents the distance between the center point of the target projection area and the center point of the target image area, as shown by the thick solid line in Figure 9c . In this way, the three-dimensional object detection device determines whether the first point cloud data set satisfies the above formula (10). If it satisfies, the first point cloud data set is used as the target point cloud data set. Otherwise, if it does not satisfy, the first point cloud data set is not the target point cloud data set.

[0207] Example 2: When the three-dimensional object detection device determines that the IOU in formula (10) s is greater than a second preset value (such as 0.5), the three-dimensional object detection device further combines formula (10) to determine whether the first point cloud data set is the target point cloud data set. For specific details, refer to the relevant description in "Example 1 of S5023", which will not be elaborated here.

[0208] S503: The three-dimensional object detection device combines the target point cloud data set and the target object image to obtain the detection result of the target object.

[0209] Among them, the detection result of the target object indicates the estimated position of the target object in the three-dimensional space.

[0210] Exemplarily, the implementation steps of S503 are as follows in Step 1 and Step 2:

[0211] Step 1: The three-dimensional object detection device inverse maps some pixel points in the target object image to the three-dimensional space according to the depth range of the point cloud in the target point cloud data set to obtain the target inverse mapping points.

[0212] Among them, some pixel points in the target object image can be the diagonal points of the target object image. Take Figure 9c as an example. The diagonal point Pi1(u1, v1) of the 2D detection box Oi is back-projected into the 3D space to obtain PPi1(x1, y1, z1). The coordinates between PPi1 and Pi1 satisfy the following formula:

[0213]

[0214] Among them, z min_i represents the minimum value of the depth range of the target point cloud dataset, u1 represents the abscissa of the diagonal point Pi1, v1 represents the ordinate of the diagonal point Pi1, and K2 -1 represents the inverse matrix of the internal parameters of the image sensor, x1 represents the coordinate of PPi1 on the x-axis, y1 represents the coordinate of PPi1 on the y-axis, and z1 represents the coordinate of PPi1 on the z-axis.

[0215] Back-project the diagonal point Pi2(u2, v2) of the 2D detection box Oi into the 3D space to obtain PPi2(x2, y2, z2). Among them, the coordinates between PPi2 and Pi2 satisfy the following formula:

[0216]

[0217] Among them, z max_i represents the maximum value of the depth range of the target point cloud dataset, u2 represents the abscissa of the diagonal point Pi2, v2 represents the ordinate of the diagonal point Pi2, and K2 -1 represents the inverse matrix of the internal parameters of the image sensor, x2 represents the coordinate of PPi2 on the x-axis, y2 represents the coordinate of PPi2 on the y-axis, and z2 represents the coordinate of PPi2 on the z-axis. In Figure 9c the scenario shown, the target inverse mapping points are PPi1 and PPi2.

[0218] Step 2: The three-dimensional object detection device uses the area marked by the three-dimensional annotation box corresponding to the target inverse mapping point as the detection result, that is, the detection result of the target object, to indicate the estimated position of the target object in the three-dimensional space.

[0219] Exemplarily, taking Figure 9a as an example, the diagonal points of the three-dimensional annotation box are PPi1 and PPi2, and the area marked by the three-dimensional annotation box is the estimated position of the target object in the three-dimensional space.

[0220] Optionally, in some embodiments, the three-dimensional object detection device also executes S504:

[0221] S504: The three-dimensional object detection device adjusts the estimated position of the target object in the three-dimensional space according to a preset adjustment factor.

[0222] Among them, the adjustment factor indicates the difference between the true position and the estimated position of the target object in the first scenario. Exemplarily, based on a large amount of data statistics, the "true position of the object in the three-dimensional space" is usually less than the "estimated position of the object in the three-dimensional space", and the adjustment factor can be a coefficient less than 1. Multiply the respective vertex coordinates of the 3D annotation box identifying the three-dimensional object detection result by the adjustment factor to obtain the adjusted three-dimensional space estimated position to fit the actual position of the target object.

[0223] See Figure 10 , the second three-dimensional object detection method provided by the embodiments of this application includes the following steps:

[0224] S1001. The three-dimensional object detection device acquires a two-dimensional image and at least one point cloud data set.

[0225] Among them, the two-dimensional image is the information collected by the image sensor. The two-dimensional image includes the images of at least one object. For the introduction of the "two-dimensional image", please refer to the relevant descriptions in S301a and S302a for details.

[0226] Among them, the point cloud data is the information collected by the depth sensor. The point cloud data set includes a plurality of point cloud data, and the point cloud data is used to describe the candidate regions of at least one object in the three-dimensional space. For the introduction of the "point cloud data set", please refer to the relevant descriptions in S501b, S301b, S302b and S303b for details.

[0227] S1002. The three-dimensional object detection device determines a target point cloud data set from at least one point cloud data set according to the target object image in the two-dimensional image.

[0228] Among them, the target object image includes the image of the target object among at least one object. Specifically, please refer to the relevant description in S501a. The point cloud data in the target point cloud data set is used to describe the candidate regions of the target object in the three-dimensional space.

[0229] Exemplarily, one set in the "at least one point cloud data set" is described as the "first point cloud data set". Taking the first point cloud data set as an example, the implementation process of S1002 can be referred to the relevant descriptions in S5021 to S5023. Among them, the target projection area of the feature points represented by the target point cloud data set in the two-dimensional image satisfies:

[0230]

[0231] Among them, S represents the similarity between the target projection area and the target image area. IOU represents the intersection-over-union (IOU) between the target projection area and the target image area. S ∩ represents the overlapping area between the target projection area and the target image area. S ∪represents the sum of the overlapping area and the non - overlapping area. The non - overlapping area is the area that does not overlap between the target projection area and the target image area. Lj represents the projection point spacing of the target projection area. The projection point spacing is the distance between the projection points of the target feature points in the two - dimensional image. The target feature points belong to the feature points represented by the target point cloud dataset and are determined based on the end values of the depth range of the target point cloud dataset. Dij represents the distance between the reference point of the target projection area and the reference point of the target image area. T represents the similarity threshold. When the three - dimensional object detection device executes S5023, the above formula (13) can be implemented as formula (10).

[0232] S1003. The three - dimensional object detection device associates the target point cloud dataset and the target object image to obtain a detection result.

[0233] Among them, the detection result indicates the estimated position of the target object in the three - dimensional space. Exemplarily, when S1003 is specifically implemented as S503, the "detection result" in S1003 is implemented as the "detection result of the target object" in S503. For details, refer to the relevant description of S503.

[0234] Due to the high processing accuracy of the two - dimensional image, the target object image can accurately present the area of the target object in the two - dimensional image. Using the target object image to screen the target point cloud dataset can achieve geometric segmentation and clustering of the point cloud dataset without obtaining a large amount of three - dimensional training data. Even if the object is occluded, the target point cloud dataset can still be obtained, which improves the accuracy of the target point cloud dataset corresponding to the target object to a certain extent. Moreover, the three - dimensional object detection device associates the target point cloud dataset and the target object image to obtain a detection result. Due to the high processing accuracy of the two - dimensional image, even if the point cloud data of the target object is insufficient, the estimated position of the target object in the three - dimensional space can be accurately determined, avoiding the problem of a high false positive rate. The three - dimensional object detection method in the embodiments of the present application does not need to obtain three - dimensional training data, avoiding the problem of "poor generalization" caused by "training a model based on three - dimensional training data".

[0235] In some embodiments, the three - dimensional object detection device also executes S1004:

[0236] S1004. The three - dimensional object detection device adjusts the estimated position indicated by the detection result according to a preset adjustment factor, so that the estimated position determined by the three - dimensional object detection device is more in line with the actual object size, thereby improving the accuracy of object detection.

[0237] Among them, the adjustment factor indicates the difference between the true position and the estimated position of the target object in the three-dimensional space. For specific descriptions, refer to the relevant descriptions in S504. Exemplarily, when S1004 is specifically implemented as S504, the "detection result" in S1004 is implemented as the "estimated position of the target object in the three-dimensional space" in S504. For details, refer to the relevant descriptions in S504.

[0238] The above mainly introduces the solutions provided by the embodiments of the present application from the perspective of methods. Next, the three-dimensional object detection device 1020 and the second device 102 provided according to the present application will be described with reference to the accompanying drawings.

[0239] See Figure 1 the schematic structural diagram of the three-dimensional object detection device 1020 in the system architecture diagram shown in Figure 1 As shown, the three-dimensional object detection device 1020 includes: an acquisition unit 1121 and a processing unit 1122.

[0240] The acquisition unit 1121 is configured to acquire a two-dimensional image and at least one point cloud data set. Among them, the two-dimensional image is the information collected by the image sensor, and the two-dimensional image includes the images of at least one object. The point cloud data is the information collected by the depth sensor, and the point cloud data set includes a plurality of point cloud data, and the point cloud data is used to describe the candidate area of at least one object in the three-dimensional space.

[0241] The processing unit 1122 is configured to determine a target point cloud data set from at least one point cloud data set according to the target object image in the two-dimensional image. Among them, the target object image includes the image of the target object among at least one object. The point cloud data in the target point cloud data set is used to describe the candidate area of the target object in the three-dimensional space.

[0242] The processing unit 1122 is further configured to associate the target point cloud data set with the target object image to obtain a detection result. Among them, the detection result indicates the estimated position of the target object in the three-dimensional space.

[0243] Among them, for the specific implementation of the acquisition unit 1121, reference can be made to Figure 3 the relevant descriptions of S302a, S302b and S303b in the embodiments shown, and for the specific implementation of the processing unit 1122, reference can be made to Figure 5 the relevant descriptions of S501a, S501b, S502 and S503 in the embodiments shown, which will not be elaborated here.

[0244] In a possible design, when the processing unit 1122 is used to determine a target point cloud dataset from at least one point cloud dataset according to the target object image in a two-dimensional image, it specifically includes: The processing unit 1122 is used to determine a first projection area of the first point cloud dataset in the two-dimensional image. Wherein, the first point cloud dataset is a set among at least one point cloud dataset. The processing unit 1122 is used to determine the first point cloud dataset as the target point cloud dataset according to the first projection area and the target image area. Wherein, the target image area is the area of the target object image in the two-dimensional image.

[0245] Among them, for the specific implementation of the processing unit 1122, reference can be made to Figure 8 the relevant content descriptions of S5021, S5022, and S5023 in the embodiments shown, which will not be elaborated here.

[0246] In a possible design, when the processing unit 1122 is used to determine a first projection area of the first point cloud dataset in the two-dimensional image, it specifically includes: The processing unit 1122 is used to determine a first feature point from the feature points represented by the first point cloud dataset according to the depth range of the point clouds in the first point cloud dataset. The processing unit 1122 is used to determine a first projection point of the first feature point in the two-dimensional image according to the conversion parameters between the point cloud data and the two-dimensional image. The processing unit 1122 is used to use the area marked by the two-dimensional annotation box corresponding to the first projection point as the first projection area.

[0247] Among them, for the specific implementation of the processing unit 1122, reference can be made to the relevant content descriptions of Step 1, Step 2, and Step 3 in S5021, which will not be elaborated here.

[0248] In a possible design, when the processing unit 1122 is used to determine the first point cloud dataset as the target point cloud dataset according to the first projection area and the target image area, it specifically includes: The processing unit 1122 is used to determine the first point cloud dataset as the target point cloud dataset according to the degree of overlap between the first projection area and the target image area, and the size of the first projection area.

[0249] Among them, for the specific implementation of the processing unit 1122, reference can be made to the relevant content descriptions in S5023, which will not be elaborated here.

[0250] In a possible design, when the processing unit 1122 is used to associate the target point cloud dataset with the target object image to obtain a detection result, it specifically includes: The processing unit 1122 is used to inverse-map some pixel points in the target object image to the three-dimensional space according to the depth range of the point clouds in the target point cloud dataset to obtain target inverse-mapped points. The processing unit 1122 is used to use the area marked by the three-dimensional annotation box corresponding to the target inverse-mapped points as the detection result.

[0251] Among them, for the specific implementation of the processing unit 1122, reference may be made to the relevant content descriptions of steps 1 and 2 in S503, which will not be elaborated here.

[0252] In a possible design, the processing unit 1122 is further configured to adjust the estimated position indicated by the detection result according to a preset adjustment factor. The adjustment factor indicates the difference between the true position and the estimated position of the target object in the three-dimensional space.

[0253] Among them, for the specific implementation of the processing unit 1122, reference may be made to Figure 8 the relevant content description in S504, which will not be elaborated here.

[0254] The three-dimensional object detection device 1020 according to the embodiment of the present application may correspond to executing the method described in the embodiment of the present application. Moreover, the above and other operations and / or functions of each module in the three-dimensional object detection device 1020 are respectively for implementing Figure 2 , Figure 3 , Figure 4b , Figure 5 , Figure 6a , Figure 6b , Figure 7a , Figure 8 the corresponding processes of each method in, and for the sake of brevity, will not be elaborated here.

[0255] In addition, it should be noted that the above-described embodiments are merely illustrative. The modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical modules, that is, they may be located in one place or distributed to multiple network modules. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. In addition, in the drawings of the device embodiments provided in the present application, the connection relationship between the modules indicates that they have a communication connection, which can be specifically implemented as one or more communication buses or signal lines.

[0256] The embodiment of the present application further provides a second device 102 for implementing the function of the three-dimensional object detection device 1020 in the system architecture diagram shown above. Figure 1 Among them, the second device 102 may be a physical device or a physical device cluster, or a virtualized cloud device, such as at least one cloud computing device in a cloud computing cluster. For the sake of easy understanding, the structure of the second device 102 is illustrated by taking the second device 102 as an independent physical device in the present application.

[0257] Figure 11 A schematic structural diagram of the second device 102 is provided, as shown in Figure 11As shown, the second device 102 includes a bus 1101, a processor 1102, a communication interface 1103, and a memory 1104. The processor 1102, the memory 1104, and the communication interface 1103 communicate with each other via the bus 1101. The bus 1101 can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of simplicity, Figure 11 it is only represented by a thick line in the figure, but it does not mean that there is only one bus or one type of bus. The communication interface 1103 is used for external communication. For example, acquiring two-dimensional images and depth images, etc.

[0258] Among them, the processor 1102 can be a Central Processing Unit (CPU). The memory 1104 can include volatile memory, such as Random Access Memory (RAM). The memory 1104 can also include non-volatile memory, such as Read-Only Memory (ROM), flash memory, a Hard Disk Drive (HDD), or a Solid-State Disk (SSD).

[0259] The memory 1104 stores executable code, and the processor 1102 executes the executable code to perform the aforementioned three-dimensional object detection method.

[0260] Specifically, when implementing Figure 1 the embodiments shown, and Figure 1 when each module of the three-dimensional object detection device 1020 described in the embodiments is implemented by software, the software or program code required to execute the functions of the acquisition unit 1121 and the processing unit 1122 in Figure 1 is stored in the memory 1104. The processor 1102 executes the program code corresponding to each module stored in the memory 1104, such as the program code corresponding to the acquisition unit 1121 and the processing unit 1122, to extract the target object image and the target point cloud data set, and then obtain the detection result of the target object. In this way, by associating the target object image and the target point cloud data set, three-dimensional object detection is achieved.

[0261] An embodiment of the present application also provides an electronic device, which includes a processor and a memory. The processor and the memory communicate with each other. The processor is configured to execute instructions stored in the memory, so that the electronic device executes the above-mentioned three-dimensional object detection method.

[0262] An embodiment of the present application also provides a computer-readable storage medium, which includes instructions for instructing a second device 102 to execute the above-mentioned three-dimensional object detection method applied to the three-dimensional object detection device 1020.

[0263] An embodiment of the present application also provides a computer program product. When the computer program product is executed by a computer, the computer executes any one of the above-mentioned three-dimensional object detection methods. The computer program product can be a software installation package. In the case where any one of the above-mentioned three-dimensional object detection methods needs to be used, the computer program product can be downloaded and executed on the computer.

[0264] An embodiment of the present application also provides a chip, which includes a logic circuit and an input / output interface. The input / output interface is used to communicate with modules outside the chip. For example, the chip can be a chip that implements the functions of the above-mentioned three-dimensional object detection device. The input / output interface inputs a two-dimensional image and at least one point cloud data set, and the input / output interface outputs a detection result. The logic circuit is used to run a computer program or instructions to implement the above-mentioned three-dimensional object detection method.

[0265] An embodiment of the present application also provides a robot, which includes an image sensor, a depth sensor, a processor, and a memory for storing processor-executable instructions. The image sensor is used to collect a two-dimensional image, the depth sensor is used to collect at least one point cloud data set, and the processor is configured with executable instructions to implement the above-mentioned three-dimensional object detection method.

[0266] An embodiment of the present application also provides a server, which includes a processor and a memory for storing processor-executable instructions. The processor is configured with executable instructions to implement the above-mentioned three-dimensional object detection method.

[0267] Through the description of the above embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general hardware. Of course, it can also be implemented by dedicated hardware including application-specific integrated circuits, dedicated CPUs, dedicated memories, dedicated components, etc. Generally, functions completed by computer programs can be easily implemented by corresponding hardware, and the specific hardware structures for implementing the same function can also be diverse, such as analog circuits, digital circuits, or dedicated circuits, etc. However, for this application, in more cases, software program implementation is a better implementation method. Based on such an understanding, the technical solution of this application, in essence, or the part that makes a contribution to the prior art, can be embodied in the form of a software product. This computer software product is stored in a readable storage medium, such as a floppy disk, USB flash drive, mobile hard disk, ROM, RAM, magnetic disk, or optical disc of a computer, etc., and includes several instructions to enable a computer device (which can be a personal computer, training device, or network device, etc.) to execute the methods described in various embodiments of this application.

[0268] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product.

[0269] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of this application are generated in whole or in part. The computer can be a general-purpose computer, a dedicated computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website, computer, training device, or data center to another website, computer, training device, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that a computer can store, or a data storage device such as a training device or data center that includes one or more integrated available media. The available medium can be a magnetic medium (for example, a floppy disk, hard disk, magnetic tape), an optical medium (for example, a DVD), or a semiconductor medium, etc.

Claims

1. A three-dimensional object detection method, characterized in that, Including: Obtain a two-dimensional image and at least one point cloud data set. The two-dimensional image includes an image of at least one object, and the point cloud data set includes a plurality of point cloud data, which are used to describe candidate regions of the at least one object in three-dimensional space. The two-dimensional image is information collected by an image sensor, and the point cloud data is information collected by a depth sensor; According to the target object image in the two-dimensional image, determine a target point cloud data set from the at least one point cloud data set. Among them, the target object image includes an image of a target object among the at least one object, and the point cloud data in the target point cloud data set is used to describe the candidate region of the target object in the three-dimensional space. The similarity between the target projection region of the feature points represented by the target point cloud data set in the two-dimensional image and the target image region is greater than a similarity threshold. The target image region is the region of the target object image in the two-dimensional image. The similarity between the target projection region and the target image region is equal to the ratio of the product of the intersection-over-union between the target projection region and the target image region and the projection point spacing of the target projection region to the distance between the reference point of the target projection region and the reference point of the target image region. The projection point spacing is the distance between the projection points of the target feature points in the two-dimensional image; According to the depth range of the points in the target point cloud data set, inverse-map some pixel points in the target object image to the three-dimensional space to obtain target inverse-mapped points; Use the region marked by the three-dimensional bounding box corresponding to the target inverse-mapped points as the detection result, where the detection result indicates the estimated position of the target object in the three-dimensional space.

2. The method according to claim 1, wherein The step of determining a target point cloud data set from the at least one point cloud data set according to the target object image in the two-dimensional image includes: Determine a first projection region of the first point cloud data set in the two-dimensional image, where the first point cloud data set is a set among the at least one point cloud data set; Determine the first point cloud data set as the target point cloud data set according to the first projection region and the target image region.

3. The method according to claim 2, characterized in that, The step of determining the first projection region of the first point cloud data set in the two-dimensional image includes: Determine first feature points from the feature points represented by the first point cloud data set according to the depth range of the points in the first point cloud data set; Determine the first projection points of the first feature points in the two-dimensional image according to the conversion parameters between the point cloud data and the two-dimensional image; Use the region marked by the two-dimensional bounding box corresponding to the first projection points as the first projection region.

4. The method according to claim 2 or 3, characterized in that, The step of determining the first point cloud data set as the target point cloud data set according to the first projection region and the target image region includes: Determine the first point cloud data set as the target point cloud data set according to the degree of overlap between the first projection region and the target image region and the size of the first projection region.

5. The method according to claim 4, characterized in that, The target projection region of the feature points represented by the target point cloud dataset in the two-dimensional image satisfies: Among them, S represents the similarity between the target projection area and the target image area, IOU represents the intersection over union between the target projection area and the target image area, S ∩ represents the overlapping area between the target projection area and the target image area, S ∪ represents the sum of the overlapping area and the non-overlapping area, the non-overlapping area is the area where the target projection area and the target image area do not overlap, Lj represents the projection point spacing of the target projection area, the projection point spacing is the distance between the projection points of the target feature points in the two-dimensional image, the target feature points belong to the feature points represented by the target point cloud dataset, and indicate the depth range of the point cloud in the target point cloud dataset, Dij represents the distance between the reference point of the target projection area and the reference point of the target image area, and T represents the similarity threshold.

6. The method according to claim 1, characterized in that, The method further includes: Adjusting the estimated position indicated by the detection result according to a preset adjustment factor, where the adjustment factor indicates the difference between the true position and the estimated position of the target object in the three-dimensional space.

7. The method according to claim 1, wherein The number of feature points represented by the point cloud dataset is less than a number threshold.

8. A three-dimensional object detection device, characterized in that, Includes: An acquisition unit for acquiring a two-dimensional image and at least one point cloud dataset, where the two-dimensional image includes an image of at least one object, the point cloud dataset includes a plurality of point cloud data, the point cloud data is used to describe the candidate region of the at least one object in the three-dimensional space, the two-dimensional image is information collected by an image sensor, and the point cloud data is information collected by a depth sensor; A processing unit for determining a target point cloud dataset from the at least one point cloud dataset according to the target object image in the two-dimensional image, where the target object image includes an image of a target object among the at least one object, the point cloud data in the target point cloud dataset is used to describe the candidate region of the target object in the three-dimensional space, the similarity between the target projection region of the feature points represented by the target point cloud dataset in the two-dimensional image and the target image region is greater than a similarity threshold, the target image region is the region of the target object image in the two-dimensional image, the similarity between the target projection region and the target image region is equal to the ratio of the product of the intersection-over-union between the target projection region and the target image region and the projection point spacing of the target projection region to the distance between the reference point of the target projection region and the reference point of the target image region, and the projection point spacing is the distance between the projection points of the target feature points in the two-dimensional image; The processing unit is further configured to inverse-map some pixel points in the target object image to the three-dimensional space according to the depth range of the points in the target point cloud dataset to obtain target inverse-mapped points; The processing unit is further configured to use the region annotated by the three-dimensional annotation box corresponding to the target inverse-mapped points as the detection result, where the detection result indicates the estimated position of the target object in the three-dimensional space.

9. The device according to claim 8, characterized in that, The processing unit is configured to determine a target point cloud dataset from the at least one point cloud dataset according to the target object image in the two-dimensional image, specifically including: Determining a first projection region of the first point cloud dataset in the two-dimensional image, where the first point cloud dataset is a set among the at least one point cloud dataset; Determining the first point cloud dataset as the target point cloud dataset according to the first projection region and the target image region.

10. The device according to claim 9, characterized in that, The processing unit is configured to determine a first projection region of the first point cloud dataset in the two-dimensional image, specifically including: Determining first feature points from the feature points represented by the first point cloud dataset according to the depth range of the points in the first point cloud dataset; Determine a first projection point of the first feature point in the two-dimensional image according to the conversion parameters between the point cloud data and the two-dimensional image; Use the area marked by the two-dimensional bounding box corresponding to the first projection point as the first projection area.

11. The device according to claim 9 or 10, characterized in that, The processing unit is configured to determine the first point cloud data set as the target point cloud data set according to the first projection area and the target image area, specifically including: Determine the first point cloud data set as the target point cloud data set according to the degree of overlap between the first projection area and the target image area, and the size of the first projection area.

12. The device according to claim 11, characterized in that, The target projection area in the two-dimensional image of the feature points represented by the target point cloud data set satisfies: Among them, S represents the similarity between the target projection area and the target image area, IOU represents the intersection over union between the target projection area and the target image area, S ∩ represents the overlapping area between the target projection area and the target image area, S ∪ represents the sum of the overlapping area and the non-overlapping area, the non-overlapping area is the area that is not overlapped between the target projection area and the target image area, Lj represents the projection point spacing of the target projection area, the projection point spacing is the distance between the projection points of the target feature points in the two-dimensional image, the target feature points belong to the feature points represented by the target point cloud dataset, and indicate the depth range of the point cloud in the target point cloud dataset, Dij represents the distance between the reference point of the target projection area and the reference point of the target image area, and T represents the similarity threshold.

13. The device according to claim 8, characterized in that, The processing unit is further configured to: Adjust the estimated position indicated by the detection result according to a preset adjustment factor, where the adjustment factor indicates the difference between the true position and the estimated position of the target object in the three-dimensional space.

14. The device according to claim 8, characterized in that, The number of feature points represented by the point cloud data set is less than a number threshold.

15. An electronic device, characterized in that, Including: A processor and a memory, the processor and the memory are coupled, and the memory stores program instructions. When the program instructions stored in the memory are executed by the processor, the three-dimensional object detection method according to any one of claims 1 to 7 is executed.

16. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a program, and when the program is called by the processor, the three-dimensional object detection method according to any one of claims 1 to 7 is executed.

17. A computer program product, characterized in that, The computer program product includes computer instructions, and when the computer instructions run on a computer, the computer is caused to execute the three-dimensional object detection method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Target object recognition method and device applied to unmanned vehicle

    CN106407947A

  • Obstacle detection method and device applied to automatic driving system, and storage medium

    CN110286387A

  • Target detection method and device, equipment and storage medium

    CN112102409A