Object detection method, object detection device, terminal device, and medium

By using the method of combining image sub-regions by three-dimensional point clouds, the problem of waste of computing resources caused by redundancy of candidate areas in the prior art is solved, and more efficient object detection is achieved.

CN114766039BActive Publication Date: 2025-05-09GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202080084709.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-12-12
Filing Date
2020-09-08
Publication Date
2025-05-09
Estimated Expiration
2040-09-08

AI Technical Summary

Technical Problem

Existing object detection methods tend to generate redundant areas when generating candidate areas, resulting in wasting of computing resources and time.

Method used

By obtaining the 3D point cloud of the scene, the scene image is segmented into multiple sub-regions, and the sub-regions are merged according to the 3D point cloud to generate more accurate candidate areas.

Benefits of technology

The number of candidate areas is reduced, computing time and resources are saved, and the efficiency of object detection is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114766039B_ABST
    Figure CN114766039B_ABST
Patent Text Reader

Abstract

The present disclosure provides an object detection method. The method includes: (101) acquiring a scene image of a scene; (102) acquiring a three-dimensional point cloud of the scene; (103) segmenting the scene image into multiple sub-regions; (104) merging the multiple sub-regions according to the three-dimensional point cloud to generate multiple candidate regions; (105) performing object detection on the multiple candidate regions to determine a target object to be detected in the scene image. In addition, the present disclosure also provides an object detection device, a terminal device, and a medium.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application claims priority to and the benefit of U.S. Patent Application No. 62 / 947,372 filed in the U.S. Patent and Trademark Office on December 12, 2019, the entire contents of which are incorporated herein by reference. Technical Field

[0003] The present disclosure relates to the field of image recognition technology, and in particular to an object detection method, an object detection device, a terminal device, and a medium. Background Art

[0004] Object detection can detect objects such as faces or cars in images, which has been widely used in the field of image recognition technology.

[0005] Currently, the mainstream object detection method includes two stages. The first stage is to extract multiple regions that may include objects (i.e., candidate regions) based on the image by using a candidate region generation method. The second stage is to perform feature extraction on the extracted candidate regions, and then identify the category of the object in the candidate region through a classifier.

[0006] In the related art, during object detection, the first stage usually uses selective search, deep learning and other methods to generate candidate regions, which may generate unreasonable redundant candidate regions. Therefore, in the subsequent feature extraction process of the candidate regions, the existence of redundant candidate regions easily leads to a waste of computing resources and computing time. Summary of the invention

[0007] The embodiments of the present disclosure provide an object detection method, an object detection device, a terminal device, and a computer-readable storage medium, which are used to solve the following technical problems in the related art: The object detection method in the related art may generate some unreasonable redundant candidate regions, which may lead to a waste of computing resources and computing time in the subsequent feature extraction process of the candidate regions.

[0008] To this end, an embodiment of the first aspect provides an object detection method. The method includes: acquiring a scene image of a scene; acquiring a three-dimensional point cloud of the scene; segmenting the scene image into multiple sub-regions; merging the multiple sub-regions according to the three-dimensional point cloud to generate multiple candidate regions; and performing object detection on the multiple candidate regions to determine a target object to be detected in the scene image.

[0009] The embodiment of the second aspect provides an object detection device. The device includes: a first acquisition module configured to acquire a scene image of a scene; a second acquisition module configured to acquire a three-dimensional point cloud of the scene; a segmentation module configured to segment the scene image into multiple sub-regions; a merging module configured to merge the multiple sub-regions according to the three-dimensional point cloud to generate multiple candidate regions; and a detection module configured to perform object detection on the multiple candidate regions to determine the target object to be detected in the scene image.

[0010] The third aspect of the embodiment provides a terminal device, including: a memory, a processor, and a computer program stored in the memory and executable by the processor. When the processor executes the computer program, the object detection method according to the first aspect of the embodiment is implemented.

[0011] The embodiment of the fourth aspect provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the object detection method according to the embodiment of the first aspect is implemented.

[0012] The technical solution disclosed in the present disclosure has the following beneficial effects.

[0013] The object detection method of the embodiment of the present disclosure, after segmenting the scene image into multiple sub-regions, merges the multiple sub-regions according to the three-dimensional point cloud of the scene to generate multiple candidate regions, so that the generated candidate regions are more accurate and the number of generated candidate regions is greatly reduced. Since the number of generated candidate regions is reduced, the calculation time is reduced, and the subsequent feature extraction of the candidate regions consumes less computing resources, thereby saving the computing time and computing resources of the object detection and improving the efficiency of the object detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] The above and / or other aspects and advantages of the embodiments of the present disclosure will become apparent and more easily understood from the following description with reference to the accompanying drawings, in which:

[0015] Figure 1 is a flow chart of an object detection method according to an embodiment of the present disclosure.

[0016] Figure 2 is a flow chart of an object detection method according to an embodiment of the present disclosure.

[0017] Figure 3 is a schematic diagram of a method for merging sub-regions according to an embodiment of the present disclosure.

[0018] Figure 4 is a block diagram of an object detection apparatus according to an embodiment of the present disclosure.

[0019] Figure 5is a block diagram of an object detection apparatus according to an embodiment of the present disclosure.

[0020] Figure 6 is a block diagram of a terminal device according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0021] Embodiments of the present disclosure will be described in detail, and examples of the embodiments are shown in the accompanying drawings. Throughout the specification, the same or similar elements and elements having the same or similar functions are represented by the same reference numerals. The embodiments described herein with reference to the accompanying drawings are illustrative and are used to explain the present disclosure, and are not to be construed as limiting the embodiments of the present disclosure.

[0022] Currently, the mainstream object detection method includes two stages. The first stage is to extract multiple regions that may include objects (i.e., candidate regions) based on the image by using the candidate region generation method. The second stage is to extract features from the extracted candidate regions and then identify the category of the object in the candidate region through a classifier.

[0023] In the related art, in the object detection process, the first stage usually uses methods such as selective search and deep learning to generate candidate regions, which may generate unreasonable redundant candidate regions. Therefore, in the subsequent feature extraction process of the candidate regions, the existence of redundant candidate regions easily leads to a waste of computing resources and computing time.

[0024] The embodiment of the present disclosure provides an object detection method for the above technical problem. After acquiring a scene image of a scene, a three-dimensional point cloud of the scene is acquired. The scene image is divided into multiple sub-regions. The multiple sub-regions are merged according to the three-dimensional point cloud to generate multiple candidate regions. Object detection is performed on the multiple candidate regions to determine the target object to be detected in the scene image.

[0025] The object detection method of the embodiment of the present disclosure, when performing object detection, after segmenting the scene image into multiple sub-regions, merges the multiple sub-regions according to the sparse three-dimensional point cloud to generate multiple candidate regions, so that the generated candidate regions are more accurate, and the number of generated candidate regions is greatly reduced. Since the number of generated candidate regions is reduced, the calculation time is reduced, and the subsequent feature extraction of the candidate regions consumes less computing resources, thereby saving the computing time and computing resources of the object detection and improving the efficiency of the object detection.

[0026] The object detection method, the object detection device, the terminal device, and the computer-readable storage medium are described below with reference to the accompanying drawings.

[0027] The following combination Figure 1 The object detection method according to the embodiment of the present disclosure is described in detail. Figure 1The present invention is a flowchart of an object detection method according to an embodiment of the present invention.

[0028] like Figure 1 As shown, the object detection method according to the present disclosure may include the following actions.

[0029] At block 101 , a scene image of a scene is acquired.

[0030] In detail, the object detection method according to the present disclosure can be performed by the object detection device according to the present disclosure. The object detection device can be configured in a terminal device to perform object detection on a scene image of a scene. The terminal device according to an embodiment of the present disclosure can be any hardware device capable of data processing, such as a mobile phone, a tablet computer, a robot, and a wearable device such as a head-mounted mobile device.

[0031] It should be understood that a camera may be configured in the terminal device to capture a scene image of a scene.

[0032] The scene may be a physical scene or a virtual scene, which is not limited here. The scene image may be static or dynamic, which is not limited here.

[0033] At frame 102 , a three-dimensional point cloud of a scene is obtained.

[0034] In detail, a three-dimensional point cloud of a scene may be generated by scanning the scene using a simultaneous localization and mapping (SLAM) system, or a dense three-dimensional point cloud of the scene may be acquired using a depth camera, or a three-dimensional point cloud of the scene may be acquired using other methods, without limitation herein.

[0035] At block 103 , the scene image is segmented into a plurality of sub-regions.

[0036] Each sub-region belongs to at most one object. The sub-region may include a single pixel or multiple pixels, and the size and shape of the sub-region may be set as required and are not limited here.

[0037] In detail, the scene image can be segmented using any method, such as a watershed segmentation algorithm, a pyramid segmentation algorithm, a mean shift segmentation algorithm, etc., which is not limited here.

[0038] In addition, the actions at the blocks 102 and 103 may be performed simultaneously, or the actions at the block 102 may be performed first and then the actions at the block 103, or the actions at the block 103 may be performed first and then the actions at the block 102, without limitation. In other words, the actions at the blocks 102 and 103 need to be performed before the actions at the block 104.

[0039] At block 104 , a plurality of sub-regions are merged according to the three-dimensional point cloud to generate a plurality of candidate regions.

[0040] In detail, the actions at block 104 may be implemented by following the actions at blocks 104a and 104b.

[0041] In frame 104a, first to n-th sub-images and first to n-th sub-3D point clouds corresponding to the first to n-th sub-regions are acquired, where n is a positive integer greater than 1.

[0042] In an embodiment, if a scene image is divided into n sub-regions and marked as the first sub-region to the nth sub-region, the first sub-image to the nth sub-image corresponding to the first sub-region to the nth sub-region can be obtained according to the scene image. According to the three-dimensional point cloud of the scene, the first sub-three-dimensional point cloud to the nth sub-three-dimensional point cloud corresponding to the first sub-region to the nth sub-region are obtained. The first sub-image to the nth sub-image corresponds to the first sub-three-dimensional point cloud to the nth sub-three-dimensional point cloud.

[0043] At frame 104 b , n sub-regions are merged to form a plurality of candidate regions according to the first to n-th sub-images and the first to n-th sub-3D point clouds.

[0044] In detail, the image similarity between the sub-images from the first sub-image to the n-th sub-image can be obtained, and the 3D point similarity between the sub-3D point clouds from the first sub-3D point cloud to the n-th sub-3D point cloud can be obtained. The n sub-regions are merged according to the image similarity and the 3D point similarity, which will not be described here.

[0045] At block 105 , object detection is performed on the plurality of candidate regions to determine a target object to be detected in the scene image.

[0046] In detail, after forming multiple candidate regions, a neural network can be used to extract feature maps of the multiple candidate regions, and then a classification method is used to identify the category of the object in each candidate region, and then the bounding box of each object is regressed to determine the size of each object to achieve object detection in multiple candidate regions, thereby determining the target object to be detected in the scene image.

[0047] The neural network used to extract the feature map of the candidate region can be any feature extraction network. Any image classification neural network can be used to determine the object category. When regressing the bounding box of the object, any regression neural network can be used without limitation.

[0048] It can be understood that when performing object detection of the embodiments of the present disclosure, after the scene image is divided into multiple sub-regions, the sub-regions are merged based on the image similarity between the sub-regions and the three-dimensional point similarity between the sub-regions determined by the sparse three-dimensional point cloud of the scene, so that the generated candidate regions are more accurate and fewer in number.

[0049] It can be understood that the object detection method of the embodiment of the present disclosure can be applied to an AR software development kit (SDK) to provide an object detection function, and developers can use the object detection function in the AR SDK to realize the recognition of objects in the scene, and then realize various functions such as product recommendation in the e-commerce field.

[0050] According to the object detection method of the embodiment of the present disclosure, a scene image of a scene is obtained, and a three-dimensional point cloud of the scene is obtained. The scene image is divided into multiple sub-regions, and the multiple sub-regions are merged according to the three-dimensional point cloud to generate multiple candidate regions. Finally, object detection is performed on the multiple candidate regions to determine the target object to be detected in the scene image. Therefore, in the object detection process, after the scene image is divided into multiple sub-regions, the multiple sub-regions are merged using the three-dimensional point cloud of the scene to generate multiple candidate regions, so that the generated candidate regions are more accurate, and the number of generated candidate regions is greatly reduced. Since the number of generated candidate regions is reduced, the calculation time is reduced, and the subsequent feature extraction of the candidate regions consumes less computing resources, thereby saving the computing time and computing resources of object detection and improving the efficiency of object detection.

[0051] Reference below Figure 2 The object detection method according to the embodiment of the present disclosure is further described. Figure 2 The present invention is a flowchart of an object detection method according to another embodiment of the present invention.

[0052] like Figure 2 As shown, the object detection method according to the embodiment of the present disclosure may include the following actions.

[0053] At block 201 , a scene image of a scene is acquired.

[0054] At block 202 , a scene is scanned by a simultaneous localization and mapping (SLAM) system to generate a three-dimensional point cloud of the scene.

[0055] The SLAM system used in the embodiments of the present disclosure will be briefly described below.

[0056] SLAM systems, as the name implies, are capable of both positioning and mapping. When a user holds or wears a terminal device and starts from an unknown location in an unknown environment, the SLAM system in the terminal device estimates the position and posture of the camera at each moment through the feature points observed by the camera during the movement, and merges the image frames collected by the camera at different moments to reconstruct a complete three-dimensional map of the scene around the user. SLAM systems are widely used in robot positioning and navigation, virtual reality (VR), augmented reality (AR), drones, and unmanned driving. The position and posture of the camera at each moment can be represented by a matrix or vector containing rotation and translation information.

[0057] The SLAM system can usually be divided into a visual front-end module and an optimized back-end module.

[0058] The main task of the visual front end is to solve the camera pose transformation between adjacent frames by using the image frames collected by the camera at different times during the movement and through feature matching, and to complete the fusion of image frames to reconstruct the map.

[0059] The visual front end relies on sensors installed in terminal devices such as robots or mobile phones. Common sensors include cameras (such as monocular cameras, binocular cameras, TOF cameras), inertial measurement units (IMUs), and lidars, and are configured to collect various types of raw data in the actual environment, including laser scanning data, video image data, and point cloud data.

[0060] The optimization backend of the SLAM system is mainly used to optimize and fine-tune the inaccurate camera pose and reconstructed map obtained by the visual front end, which can be separated from the visual front end as an offline operation or integrated into the visual front end.

[0061] In an embodiment, a SLAM system may be used to obtain a three-dimensional point cloud of a scene by using the following method.

[0062] In detail, a camera included in the terminal device may be calibrated in advance to determine the internal parameters of the camera, and then a scene may be scanned using the calibrated camera, and a three-dimensional point cloud of the scene may be generated using the SLAM system.

[0063] To calibrate the camera, you can first print a 7*9 black and white calibration plate on A4 paper. The checkerboard size on the calibration plate is 29.1mm. Stick the calibration plate on a neat and flat wall, and use the camera to be calibrated to shoot a video of the calibration plate. When shooting, move the camera continuously to shoot the calibration plate from different angles and distances. The calibration program is written using algorithms and functions that encapsulate OpenCV. Finally, convert the video into images, and select 50 of the images as calibration images, together with the basic parameters of the calibration plate, and then enter the calibration program to calculate the internal parameters of the camera.

[0064] Points in the world coordinate system are measured in terms of physical size, while points in the image plane are measured in pixels. Intrinsic parameters are used to linearly transform between the two coordinate systems. A point Q(X, Y, Z) in space can be transformed by the intrinsic parameter matrix to obtain the corresponding point Q(u, v) projected by the ray onto the pixel coordinate system on the image plane, where:

[0065]

[0066] K is the intrinsic parameter matrix of the camera.

[0067]

[0068] Where f is the focal length of the camera in millimeters, dx and dy are the length and width of each pixel in millimeters, and u0, v0 are the coordinates of the center of the image, usually in pixels.

[0069] According to the internal parameters of the camera and the height and width of the scene image obtained when the camera shoots the scene, the camera parameter file is written in the format required by the DSO program, and the DSO program is started with the camera parameter file as input. In other words, when the camera is used to scan the scene, the three-dimensional point cloud of the scene can be constructed in real time.

[0070] It should be noted that the above method is only one implementation method of scanning a scene by a SLAM system to generate a three-dimensional point cloud of the scene. In practical applications, the SLAM system can be used to generate a three-dimensional point cloud of the scene by any other method, which is not limited here.

[0071] In addition, in the aforementioned embodiment, the SLAM system is configured to scan a scene to generate a three-dimensional point cloud of the scene. In practical applications, a dense three-dimensional point cloud of the scene can be obtained by a depth camera, or the three-dimensional point cloud of the scene can be obtained by other methods, which are not limited here.

[0072] At block 203 , the scene image is segmented into a plurality of sub-regions.

[0073] At frame 204 , first to n-th sub-images and first to n-th sub-3D point clouds corresponding to the first to n-th sub-regions are acquired, where n is a positive integer greater than 1.

[0074] At frame 205 , an i-th sub-region and a j-th sub-region among n sub-regions are obtained, where i and j are positive integers less than or equal to n.

[0075] The implementation process and principle of the actions in frames 201 to 205 refer to the description of the above embodiment and will not be repeated here.

[0076] At frame 206 , the image similarity between the ith sub-region and the jth sub-region is generated according to the ith sub-image of the ith sub-region and the jth sub-image of the jth sub-region.

[0077] In detail, the structural similarity (SSIM), cosine similarity, mutual information, color similarity, or histogram similarity between the i-th sub-image of the i-th sub-region and the j-th sub-image of the j-th sub-region can be calculated to generate the image similarity between the i-th sub-region and the j-th sub-region.

[0078] At frame 207 , the 3D point similarity between the ith sub-region and the jth sub-region is generated according to the ith sub-3D point cloud of the ith sub-region and the jth sub-3D point cloud of the jth sub-region.

[0079] In detail, the 3D point similarity between the i-th sub-region and the j-th sub-region may be determined by calculating the distance between each 3D point in the i-th sub-3D point cloud of the i-th sub-region and each 3D point in the j-th sub-3D point cloud of the j-th sub-region.

[0080] It can be understood that there may be multiple 3D points in the i-th sub-3D point cloud and in the j-th sub-3D point cloud. Correspondingly, there are multiple distance parameters between the 3D points in the i-th sub-3D point cloud and the 3D points in the j-th sub-3D point cloud. In the embodiment of the present disclosure, the 3D point similarity between the i-th sub-region and the j-th sub-region is generated according to the maximum distance between the 3D points in the i-th sub-3D point cloud and the 3D points in the j-th sub-3D point cloud, or the average value of the distance between the 3D points in the i-th sub-3D point cloud and the 3D points in the j-th sub-3D point cloud.

[0081] In this embodiment, the correspondence between the distance between the 3D points and the 3D point similarity can be preset, so that the 3D point similarity between the ith sub-region and the jth sub-region is generated according to the maximum distance between the 3D points in the ith sub-3D point cloud and the 3D points in the jth sub-3D point cloud, or the average value of the distance between the 3D points in the ith sub-3D point cloud and the 3D points in the jth sub-3D point cloud. The distance between the 3D points and the 3D point similarity can be inversely proportional, that is, the greater the distance, the smaller the 3D point similarity.

[0082] Alternatively, each 3D point in the i-th sub-3D point cloud of the i-th sub-region and each 3D point in the j-th sub-3D point cloud of the j-th sub-region are fitted to a preset model to calculate the distance between the preset models fitted to each 3D point in different sub-3D point clouds, and generate the 3D point similarity between the i-th sub-region and the j-th sub-region.

[0083] In detail, the correspondence between the distance between the preset models and the 3D point similarity can be preset. Each 3D point in the i-th sub-3D point cloud of the i-th sub-region is fitted with the first preset model, and each 3D point in the j-th sub-3D point cloud of the j-th sub-region is fitted with the second preset model, and then the distance between the first preset model and the second preset model is determined. According to the correspondence between the distance between the preset models and the 3D point similarity, the 3D point similarity between the i-th sub-region and the j-th sub-region is determined.

[0084] The preset model may be a preset basic geometric model, such as a sphere, a cylinder, a plane, an ellipsoid, or a complex geometric model composed of basic geometric models, or other preset models, which are not limited here.

[0085] In addition, the method of fitting the 3D points in the sub-3D point cloud to the preset model may be the least square method or any other method, which is not limited here.

[0086] For example, by fitting the 3D points in the i-th sub-3D point cloud to a cylinder, the cylinder is parameterized. For example, the cylinder in space can be represented by parameters such as the center coordinates (X, Y, Z) in the three-dimensional space, the bottom radius, the height, and the orientation. Then, some 3D points are randomly selected from the sub-3D point cloud each time by the random sampling consensus (RANSAC) algorithm. Assuming that these 3D points are in the cylinder, the parameters of the cylinder are calculated, and then the number of 3D points in the sub-3D point cloud on the cylinder is counted, and it is determined whether the number exceeds a preset number threshold. If not, some other 3D points are randomly selected to repeat the above process. Otherwise, it can be determined that the 3D points on the cylinder in the sub-3D point cloud can be fitted with a cylinder, thereby obtaining the cylinder fitted by each 3D point in the i-th sub-3D point cloud. Then, through a similar algorithm, a second preset model fitted by each 3D point in the j-th sub-3D point cloud is obtained, and it is assumed that the second preset model is an ellipse. Then, the distance between the cylinder and the ellipse is calculated, and the 3D point similarity between the i-th sub-region and the j-th sub-region is determined according to the corresponding relationship between the distance and the 3D point similarity.

[0087] The quantity threshold can be set as needed and is not limited here.

[0088] In addition, a distance threshold may be set, and the distances from all 3D points in the sub-3D point cloud to the cylinder may be calculated, thereby determining 3D points whose distances are less than the distance threshold as 3D points on the cylinder.

[0089] At block 208 , the i th sub-region and the j th sub-region are merged according to the image similarity and the 3D point similarity.

[0090] In detail, the image similarity threshold and the 3D point similarity threshold may be pre-set, so that after determining the image similarity and the 3D point similarity between the i-th sub-region and the j-th sub-region, the image similarity and the image similarity threshold are compared, and the 3D point similarity and the 3D point similarity threshold are compared. If the image similarity is greater than the image similarity threshold, and the 3D point similarity is greater than the 3D point similarity threshold, it may be determined that the i-th sub-image of the i-th sub-region and the j-th sub-image of the j-th sub-region are images of different parts of the same object and may be used as a candidate region, so that the i-th sub-region and the j-th sub-region may be merged.

[0091] In addition, if the image similarity between the i-th sub-region and the j-th sub-region is less than or equal to the image similarity threshold, or the three-dimensional point similarity between the i-th sub-region and the j-th sub-region is less than or equal to the three-dimensional point similarity threshold, it can be determined that the i-th sub-image of the i-th sub-region and the j-th sub-image of the j-th sub-region are images of different objects, and thus the i-th sub-region and the j-th sub-region cannot be merged.

[0092] For example, assume that the image similarity threshold is 80% and the three-dimensional point similarity threshold is 70%. Figure 3 It is a partial sub-region after the scene image is divided into multiple sub-regions. The image similarity between region 1 and region 2 is 90%, and the 3D point similarity between region 1 and region 2 is 80%. The image similarity between region 1 and other regions is less than 80%, and the 3D point similarity between region 1 and other regions is less than 70%. The image similarity between region 2 and other regions is less than 80%, and the 3D point similarity between region 2 and other regions is less than 70%. The image similarity between region 3 and other regions is less than 80%, and the 3D point similarity between region 3 and other regions is less than 70%. Region 1 and region 2 are merged to form a candidate region. Region 2 and region 3 are determined as separated candidate regions.

[0093] At block 209 , object detection is performed on the plurality of candidate regions to determine a target object to be detected in the scene image.

[0094] In detail, after forming multiple candidate regions, feature maps of the multiple candidate regions can be extracted by using a neural network, and then a classification method is used to identify the category of the object in each candidate region, and then a bounding box of each object is regressed. The size of each object is determined to achieve object detection of the multiple candidate regions, thereby determining the target object to be detected in the scene image.

[0095] The neural network used to extract the feature map of the candidate region can be any feature extraction network. Any image classification neural network can be used to determine the object category. When regressing the bounding box of the object, any regression neural network can be used without limitation.

[0096] When performing object detection of the embodiment of the present disclosure, after the scene image is divided into multiple sub-regions, the three-dimensional point similarity between the sub-regions and the image similarity between the sub-regions can be generated based on the sparse three-dimensional point cloud, and the sub-regions can be merged based on the similarity to generate multiple candidate regions, thereby determining the target object to be detected in the scene image, making the generated candidate regions more accurate, and greatly reducing the number of generated candidate regions. Since the number of generated candidate regions is reduced, the calculation time is reduced, and the subsequent feature extraction of the candidate regions consumes less computing resources, thereby saving the computing time and computing resources of object detection and improving the efficiency of object detection.

[0097] Combine the following Figure 4 An object detection apparatus according to an embodiment of the present disclosure is described. Figure 4 is a block diagram of an object detection apparatus according to an embodiment of the present disclosure.

[0098] like Figure 4 As shown, the object detection device includes a first acquisition module 11, a second acquisition module 12, a segmentation module 13, a merging module 14 and a detection module 15.

[0099] The first acquisition module 11 is configured to acquire a scene image of a scene.

[0100] The second acquisition module 12 is configured to acquire a three-dimensional point cloud of the scene.

[0101] The segmentation module 13 is configured to segment the scene image into a plurality of sub-regions.

[0102] The merging module 14 is configured to merge multiple sub-regions according to the three-dimensional point cloud to generate multiple candidate regions.

[0103] The detection module 15 is configured to perform object detection on the plurality of candidate regions to determine the target object to be detected in the scene image.

[0104] In an exemplary embodiment, the second acquisition module 12 is configured to scan the scene through a simultaneous localization and mapping (SLAM) system to generate a three-dimensional point cloud of the scene.

[0105] In detail, the object detection device can perform the object detection method described in the aforementioned embodiment. The device can be configured in a terminal device to perform object detection on a scene image of a scene. The terminal device in the embodiments of the present disclosure can be any hardware device capable of data processing, such as a mobile phone, a tablet computer, a robot, a wearable device such as a head-mounted mobile device, etc.

[0106] It should be noted that the implementation process and technical principles of the object detection device of this embodiment can be referred to the description of the object detection method of the embodiment of the first aspect above, and will not be repeated here.

[0107] According to the object detection device of the embodiment of the present disclosure, a scene image of a scene is acquired, and a three-dimensional point cloud of the scene is acquired. The scene image is divided into multiple sub-regions, and the multiple sub-regions are merged according to the three-dimensional point cloud to generate multiple candidate regions. Finally, object detection is performed on the multiple candidate regions to determine the target object to be detected in the scene image. Therefore, in the object detection process, after the scene image is divided into multiple sub-regions, the multiple sub-regions are merged using the three-dimensional point cloud of the scene to generate multiple candidate regions, so that the generated candidate regions are more accurate, and the number of generated candidate regions is greatly reduced. Since the number of generated candidate regions is reduced, the calculation time is reduced, and the subsequent feature extraction of the candidate regions consumes less computing resources, thereby saving the computing time and computing resources of object detection and improving the efficiency of object detection.

[0108] Combine the following Figure 5 An object detection device according to an embodiment of the present disclosure is further described. Figure 5 is a block diagram of an object detection device according to another embodiment of the present disclosure.

[0109] like Figure 5 As shown, based on Figure 4 The merging module 14 includes: an acquiring unit 141 and a merging unit 142 .

[0110] The acquisition unit 141 is configured to acquire first to n-th sub-images and first to n-th sub-3D point clouds corresponding to the first to n-th sub-regions, where n is a positive integer greater than 1.

[0111] The merging unit 142 is configured to merge n sub-regions according to the first to n-th sub-images and the first to n-th sub-3D point clouds to form a plurality of candidate regions.

[0112] In an exemplary embodiment, the merging unit 142 is configured to: obtain the i-th sub-region and the j-th sub-region among n sub-regions, where i and j are positive integers less than or equal to n; generate the image similarity between the i-th sub-region and the j-th sub-region based on the i-th sub-image of the i-th sub-region and the j-th sub-image of the j-th sub-region; generate the three-dimensional point similarity between the i-th sub-region and the j-th sub-region based on the i-th sub-three-dimensional point cloud of the i-th sub-region and the j-th sub-three-dimensional point cloud of the j-th sub-region; and merge the i-th sub-region and the j-th sub-region based on the image similarity and the three-dimensional point similarity.

[0113] In an exemplary embodiment, the merging unit 142 is configured to merge the i-th sub-region and the j-th sub-region when the image similarity is greater than the image similarity threshold and the 3D point similarity is greater than the 3D point similarity threshold.

[0114] It should be noted that the implementation process and technical principles of the object detection device of this embodiment can be referred to the description of the object detection method of the embodiment of the first aspect above, and will not be repeated here.

[0115] Using an object detection device according to an embodiment of the present disclosure, a scene image of a scene is acquired, and a three-dimensional point cloud of the scene is acquired. The scene image is divided into multiple sub-regions, and the multiple sub-regions are merged according to the three-dimensional point cloud to generate multiple candidate regions. Finally, object detection is performed on the multiple candidate regions to determine the target object to be detected in the scene image. Therefore, in the target detection process, after the scene image is divided into multiple sub-regions, the multiple sub-regions are merged using the three-dimensional point cloud of the scene to generate multiple candidate regions, so that the generated candidate regions are more accurate, and the number of generated candidate regions is greatly reduced. Since the number of generated candidate regions is reduced, the calculation time is reduced, and the subsequent feature extraction of the candidate regions consumes less computing resources, thereby saving the computing time and computing resources of object detection and improving the efficiency of object detection.

[0116] In order to implement the above embodiments, the present disclosure also provides a terminal device.

[0117] Figure 6 is a block diagram of a terminal device according to an embodiment of the present disclosure.

[0118] like Figure 6 As shown, the terminal device includes: a memory, a processor, and a computer program stored in the memory and executable by the processor. When the processor executes the computer program, the object detection method according to the embodiment of the first aspect is implemented.

[0119] It should be noted that the implementation process and technical principles of the terminal device of this embodiment can be referred to the description of the object detection method of the embodiment of the first aspect above, and will not be repeated here.

[0120] Using a terminal device according to an embodiment of the present disclosure, a scene image of a scene is acquired, and a three-dimensional point cloud of the scene is acquired. The scene image is divided into multiple sub-regions, and the multiple sub-regions are merged according to the three-dimensional point cloud to generate multiple candidate regions. Finally, object detection is performed on the multiple candidate regions to determine the target object to be detected in the scene image. Therefore, in the object detection process, after the scene image is divided into multiple sub-regions, the multiple sub-regions are merged by using the three-dimensional point cloud of the scene to generate multiple candidate regions, so that the generated candidate regions are more accurate, and the number of generated candidate regions is greatly reduced. Since the number of generated candidate regions is reduced, the calculation time is reduced, and the subsequent feature extraction of the candidate regions consumes less computing resources, thereby saving the computing time and computing resources of object detection and improving the efficiency of object detection.

[0121] In order to implement the above embodiments, the present disclosure further provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the object detection method according to the embodiment of the first aspect is implemented.

[0122] In order to implement the above embodiments, the present disclosure also provides a computer program. When the computer program is executed by a processor, the object detection method according to the embodiment is implemented.

[0123] Throughout the specification, references to “an embodiment,” “some embodiments,” “an example,” “a specific example,” or “some examples” mean that a particular feature, structure, material, or characteristic described in conjunction with the embodiment or example is included in at least one embodiment or example of the present disclosure.

[0124] In addition, terms such as "first" and "second" are used herein for descriptive purposes and are not intended to indicate or imply relative importance or significance. Therefore, features defined with "first" and "second" may include one or more such features.

[0125] Any process or method described in the flowchart or otherwise described herein may be understood as including one or more modules, segments or portions of code of executable instructions for implementing specific logical functions or steps in the process, and the scope of the preferred embodiments of the present disclosure includes other implementations, which should be understood by those skilled in the art.

[0126] It should be understood that each part of the present disclosure can be implemented by hardware, software, firmware or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by an appropriate instruction execution system. For example, if it is implemented by hardware, in another embodiment, the steps or methods can be implemented by one or a combination of the following technologies known in the art: a discrete logic circuit with a logic gate circuit for implementing the logic function of a data signal, a dedicated integrated circuit with an appropriate combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.

[0127] It should be understood by those skilled in the art that all or part of the steps performed by the method in the above embodiments may be completed by the relevant hardware indicated by the program. The program may be stored in a computer-readable storage medium. When the program is executed, one or a combination of the steps of the method in the above embodiments may be completed.

[0128] The storage medium may be a read-only memory, a disk or a CD, etc. Although illustrative embodiments have been shown and described, it should be understood by those skilled in the art that the above embodiments cannot be construed as limiting the present disclosure, and that changes, substitutions and modifications may be made to the embodiments without departing from the scope of the present disclosure.

Claims

1. An object detection method, comprising: Obtaining a scene image of a scene; Obtaining a three-dimensional point cloud of the scene; dividing the scene image into a plurality of sub-regions; Merging the plurality of sub-regions according to image similarities between the sub-regions and 3D point similarities between the sub-regions to generate a plurality of candidate regions, wherein the 3D point similarity between any two sub-regions is determined by calculating the distance between each 3D point in a sub-3D point cloud of one of the two sub-regions and each 3D point in a sub-3D point cloud of the other of the two sub-regions, and the distance between the 3D points is inversely proportional to the 3D point similarity; as well as Object detection is performed on the multiple candidate areas to determine a target object to be detected in the scene image.

2. The method according to claim 1, wherein: Acquiring the three-dimensional point cloud of the scene includes: The scene is scanned by a simultaneous localization and mapping (SLAM) system to generate the three-dimensional point cloud of the scene.

3. The method according to claim 1, wherein: Merging the multiple sub-regions according to the three-dimensional point cloud to generate the multiple candidate regions comprises: Acquire the first to n-th sub-images and the first to n-th sub-3D point clouds corresponding to the first to n-th sub-regions, where n is a positive integer greater than 1; and The n sub-regions are merged according to the first sub-image to the n-th sub-image and the first sub-3D point cloud to the n-th sub-3D point cloud to form the plurality of candidate regions.

4. The method according to claim 3, wherein: Merging the n sub-regions according to the first sub-image to the n-th sub-image and the first sub-3D point cloud to the n-th sub-3D point cloud to form the plurality of candidate regions comprises: Obtain an i-th sub-region and a j-th sub-region among the n sub-regions, wherein i and j are positive integers less than or equal to n; generating an image similarity between the i-th sub-region and the j-th sub-region according to the i-th sub-image of the i-th sub-region and the j-th sub-image of the j-th sub-region; generating a three-dimensional point similarity between the i-th sub-region and the j-th sub-region according to the i-th sub-three-dimensional point cloud of the i-th sub-region and the j-th sub-three-dimensional point cloud of the j-th sub-region; and The i-th sub-region and the j-th sub-region are merged according to the image similarity and the three-dimensional point similarity.

5. The method according to claim 4, wherein: Merging the i-th sub-region and the j-th sub-region according to the image similarity and the three-dimensional point similarity comprises: When the image similarity is greater than an image similarity threshold and the three-dimensional point similarity is greater than a three-dimensional point similarity threshold, the i-th sub-region and the j-th sub-region are merged.

6. An object detection device comprising: A first acquisition module is configured to acquire a scene image of a scene; A second acquisition module is configured to acquire a three-dimensional point cloud of the scene; A segmentation module, configured to segment the scene image into a plurality of sub-regions; a merging module configured to merge the plurality of sub-regions to generate a plurality of candidate regions according to image similarities between the sub-regions and 3D point similarities between the sub-regions, wherein the 3D point similarity between any two sub-regions is determined by calculating the distance between each 3D point in a sub-3D point cloud of one of the two sub-regions and each 3D point in a sub-3D point cloud of the other of the two sub-regions, and the distance between the 3D points is inversely proportional to the 3D point similarity; as well as The detection module is configured to perform object detection on the multiple candidate areas to determine the target object to be detected in the scene image.

7. The device according to claim 6, wherein: The second acquisition module is configured to scan the scene through a simultaneous localization and mapping (SLAM) system to generate the three-dimensional point cloud of the scene.

8. The device according to claim 6, wherein: The merging module comprises: an acquisition unit configured to acquire first to n-th sub-images and first to n-th sub-3D point clouds corresponding to the first to n-th sub-regions, wherein n is a positive integer greater than 1; and A merging unit is configured to merge n sub-regions according to the first sub-image to the n-th sub-image and the first sub-3D point cloud to the n-th sub-3D point cloud to form the plurality of candidate regions.

9. The device according to claim 8, wherein: The merging unit is configured as: Obtain an i-th sub-region and a j-th sub-region among the n sub-regions, wherein i and j are positive integers less than or equal to n; generating an image similarity between the i-th sub-region and the j-th sub-region according to the i-th sub-image of the i-th sub-region and the j-th sub-image of the j-th sub-region; generating a three-dimensional point similarity between the i-th sub-region and the j-th sub-region according to the i-th sub-three-dimensional point cloud of the i-th sub-region and the j-th sub-three-dimensional point cloud of the j-th sub-region; and The i-th sub-region and the j-th sub-region are merged according to the image similarity and the three-dimensional point similarity.

10. The device according to claim 9, wherein: The merging unit is configured to merge the i-th sub-region and the j-th sub-region when the image similarity is greater than an image similarity threshold and the 3D point similarity is greater than a 3D point similarity threshold.

11. A terminal device, comprising: A memory, a processor, and a computer program stored in the memory and executable by the processor, wherein when the processor executes the computer program, the object detection method according to any one of claims 1 to 5 is implemented.

12. A computer-readable storage medium having a computer program stored therein, wherein: When the computer program is executed by a processor, the object detection method according to any one of claims 1 to 5 is implemented.