Target object detection method, processor and machine readable storage medium
By obtaining the key point set and bird's-eye view of the target object in the point cloud data, matching the target feature model in the pre-constructed model feature library, the problem of low target detection accuracy in the prior art is solved, and high-accuracy target detection in complex environments is achieved.
Patent Information
- Application Number
- CN202411875680.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-19
- Publication Date
- 2025-05-16
AI Technical Summary
The methods used for object detection in the prior art are under the influence of complex environments and data noise, resulting in low accuracy of detection results and easy identification of non-target objects.
By obtaining the keypoint set of target objects in point cloud data, a bird's-eye view containing the target objects is determined, and the target feature model in the pre-constructed model feature library is matched based on the keypoint set and bird's-eye view, and the detection results of the target object are finally determined.
In complex work scenarios, the target object can be accurately identified and the accuracy and real-timeness of the detection results can be improved.
Smart Images

Figure CN120014228A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of target detection, and in particular to a method, a processor and a machine-readable storage medium for detecting a target object. Background Art
[0002] At present, in some specific application scenarios, target detection technology is affected by various factors, such as complex environment (bad weather causes the quality of sensor data to decrease), data noise and occlusion, real-time detection, etc., which leads to low accuracy of the detection results of the detected objects, and it is easy to identify non-target objects, thus affecting subsequent operations. Therefore, the target detection methods used in the prior art have the problem of low accuracy. Summary of the invention
[0003] The purpose of the embodiments of the present application is to provide a method, a processor, and a machine-readable storage medium for target object detection, so as to solve the problem of low accuracy of the target detection method used in the prior art.
[0004] In order to achieve the above-mentioned object, a first aspect of an embodiment of the present application provides a method for detecting a target object, the method comprising:
[0005] Acquire point cloud data, where the point cloud data includes a key point set of the target object;
[0006] Determine a bird's-eye view containing a target object based on the point cloud data;
[0007] According to the key point set and the bird's-eye view, a target feature model corresponding to the target object in a pre-built model feature library is determined, and the model feature library contains feature models of multiple different objects;
[0008] The detection result of the target object is determined based on the target key points in the key point set that match the target feature model.
[0009] In an embodiment of the present application, a method for determining a key point set includes: acquiring image data, the image data containing an image of a target object; performing target recognition on the image data to obtain a two-dimensional target frame of the target object in the image data; based on a predetermined calibration relationship between the image data and the point cloud data, determining a point cloud corresponding to a pixel point in the two-dimensional target frame from the point cloud data; and determining a key point set based on the point cloud corresponding to a pixel point in the two-dimensional target frame.
[0010] In an embodiment of the present application, a target feature model corresponding to a target object in a pre-constructed model feature library is determined based on a key point set and a bird's-eye view, including: performing feature extraction on the key point set to obtain a first feature, the first feature including a multi-scale semantic feature and position information; performing feature extraction on the bird's-eye view to obtain a second feature, the second feature including a bird's-eye view feature; splicing the first feature and the second feature to obtain a key point feature; matching each feature model in the model feature library according to the key point feature to determine the number of matching points between the key point set and each feature model; determining the feature model whose matching point number is the target maximum as the target feature model corresponding to the target object.
[0011] In an embodiment of the present application, the feature model includes multiple feature points, each feature point is associated with a feature point feature, and each feature model in the feature library is matched according to the key point feature to determine the number of matching points between the key point set and each feature model, including: determining the Euclidean distance between each key point in the key point set and each feature point of each feature model according to the key point feature and the feature point feature; determining the key point and feature point whose Euclidean distances are the minimum as matching points; determining the number of matching points between the key point set and each feature model to obtain the number of matching points between the key point set and each feature model.
[0012] In an embodiment of the present application, a method for constructing a model feature library includes: obtaining sample point cloud data of multiple sample models, the multiple sample models are models of multiple different objects, and the sample point cloud data includes a feature point set of the sample models; determining regular voxels of multi-scale semantic features around each feature point in the feature point set through a 3D sparse convolutional network; determining voxel features of each feature point at different levels based on the regular voxels of multi-scale semantic features around each feature point; aggregating and connecting the voxel features at different levels to obtain multi-scale semantic features and position information of each feature point; determining a sample bird's-eye view of each sample model based on the sample point cloud data; determining the bird's-eye view features of each feature point based on the bird's-eye view; splicing the multi-scale semantic features, position information and bird's-eye view features into feature point features of each feature point; generating a corresponding feature model based on the feature point features of each feature point in the feature point set of each sample model to form a model feature library.
[0013] In an embodiment of the present application, a bird's-eye view containing a target object is determined based on point cloud data, including: voxelizing the point cloud data to obtain voxelized point cloud data; extracting features from the voxelized point cloud data according to multiple preset downsampling scales to obtain a multi-scale feature map; and performing planar projection processing on the multi-scale feature map to generate a bird's-eye view.
[0014] In an embodiment of the present application, a detection result of a target object is determined based on target key points in a key point set that match a target feature model, including: determining the number of target key points in a key point set that match a target feature model to obtain the number of target key points; when the number of target key points is greater than a preset number of matching points, performing point cloud registration processing on the target key points and the target feature model through a preset algorithm to obtain a position transformation matrix; determining the posture data of the target object based on the position transformation matrix and the target key points; the detection result of the target object includes the posture data of the target object.
[0015] In an embodiment of the present application, the method also includes: determining the number of target key points in the key point set that match the target feature model to obtain the number of target key points; when the number of target key points is less than or equal to the preset number of matching points, determining the candidate area containing the target object in the bird's-eye view according to a preset candidate area generation algorithm; optimizing the candidate area according to the target key points to obtain the target area, wherein the optimization processing includes pooling processing; and determining the detection result of the target object according to the target area.
[0016] A second aspect of an embodiment of the present application provides a processor configured to execute the above-mentioned method for target object detection.
[0017] A third aspect of an embodiment of the present application provides a machine-readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, the above-mentioned method for detecting a target object is implemented.
[0018] The above technical solution obtains point cloud data including a key point set of the target object, and then determines a bird's-eye view containing the target object based on the point cloud data, and then determines a target feature model corresponding to the target object in a pre-constructed model feature library based on the key point set and the bird's-eye view, wherein the model feature library contains feature models of multiple different objects, and finally determines the detection result of the target object based on the target key points in the key point set that match the target feature model. The present application can accurately identify the target object in operation in a complex working scene based on the target model in the model feature library by pre-constructing a model feature library corresponding to the target object, thereby improving the accuracy and real-time performance of the detection results.
[0019] Other features and advantages of the embodiments of the present application will be described in detail in the subsequent specific implementation section. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] The accompanying drawings are used to provide a further understanding of the embodiments of the present application and constitute a part of the specification. Together with the following specific implementations, they are used to explain the embodiments of the present application, but do not constitute a limitation on the embodiments of the present application. In the accompanying drawings:
[0021] Figure 1 A flowchart of a method for detecting a target object provided in an embodiment of the present application;
[0022] Figure 2 A structural block diagram of a computing device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0023] In order to make the purpose, technical scheme and advantages of the embodiments of the present application clearer, the technical scheme in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. It should be understood that the specific implementation methods described herein are only used to illustrate and explain the embodiments of the present application, and are not used to limit the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.
[0024] It should be noted that if the embodiments of the present application involve directional indications (such as up, down, left, right, front, back...), the directional indications are only used to explain the relative position relationship, movement status, etc. between the components under a certain specific posture (as shown in the accompanying drawings). If the specific posture changes, the directional indication will also change accordingly.
[0025] In addition, if there are descriptions involving "first", "second", etc. in the embodiments of the present application, the descriptions of "first", "second", etc. are only used for descriptive purposes and cannot be understood as indicating or suggesting their relative importance or implicitly indicating the number of technical features indicated. Therefore, the features defined as "first" and "second" may explicitly or implicitly include at least one of the features. In addition, the technical solutions between the various embodiments can be combined with each other, but they must be based on the ability of ordinary technicians in the field to implement them. When the combination of technical solutions is contradictory or cannot be implemented, it should be deemed that such combination of technical solutions does not exist and is not within the scope of protection required by this application.
[0026] Figure 1 A flow chart of a method for detecting a target object provided in an embodiment of the present application is shown in FIG. Figure 1 As shown, an embodiment of the present application provides a method for detecting a target object, which is described by taking the method applied to a processor as an example. The method may include the following steps.
[0027] Step S101, obtaining point cloud data, where the point cloud data includes a key point set of a target object.
[0028] Specifically, point cloud data refers to the point cloud data of the overall environment of a certain area to be inspected, and the area contains a target object, which is the object to be inspected, usually a specific object in the environment, such as a vehicle, pedestrian, building, construction equipment and parts, etc. In some examples, the target object may be a hook, a sling or other work accessories, and in other examples, the target object may also be a pedestrian, a vehicle or various reference objects or objects to be worked in the work scene, etc., and examples are not given here one by one.
[0029] In some examples, point cloud data can be collected by laser radar, depth camera or other three-dimensional sensors. In the embodiment of the present application, the point cloud data includes a key point set of the target object, which refers to those points that are particularly important for identifying or describing the target object. The key points usually include high curvature points, edge points and corner points, etc. These points have important feature information and can be used to describe the local shape and structure of the point cloud. These points may have unique geometric features, colors, textures or other properties, so that they can be distinguished from the surrounding points. In one example, a feature-based method can be used to determine the key point set of the target object in the point cloud data, and the key points can be extracted using features such as curvature and surface normal vectors in the point cloud. For example, the curvature of each point can be calculated, and the point with larger curvature can be selected as the key point. In another example, a density-based method can be used to determine the key point set of the target object in the point cloud data, and the key points can be extracted by calculating the density around each point in the point cloud. For example, points with larger density changes can be selected as key points. In this way, by determining the key point set of the target object in the point cloud data, strong data support can be provided for subsequent target inspection tasks.
[0030] Step S102: determining a bird's-eye view containing a target object according to the point cloud data.
[0031] Specifically, a bird's-eye view is a view of a target object and its surroundings viewed vertically from above. This view helps capture the overall layout and shape of the target object, especially in three-dimensional space. In one example, a bird's-eye view can be generated by projecting or transforming point cloud data.
[0032] Step S103, determining a target feature model corresponding to the target object in a pre-constructed model feature library according to the key point set and the bird's-eye view, wherein the model feature library contains feature models of multiple different objects.
[0033] Specifically, in the embodiment of the present application, a model feature library containing feature models of multiple different objects is constructed in advance. These feature models can be constructed based on geometric shapes, statistical features, or deep learning features. It can be understood that the above-mentioned different objects and the target object are different models of objects of the same type. For example, if the target object is a hook, the model feature library includes feature models of multiple different models of hooks. Furthermore, after obtaining the key point set and the bird's-eye view of the target object, the key point set and the bird's-eye view can be matched with each feature model in the model feature library to determine the target feature model of the target object. In one example, the matching process can be based on methods such as shape similarity, feature point matching, or deep learning.
[0034] Step S104, determining the detection result of the target object according to the target key points in the key point set that match the target feature model.
[0035] Specifically, the processor can determine the target key points in the key point set that match the target feature model through a feature matching algorithm. The feature matching algorithm may include neighbor search and Euclidean distance calculation. Then, based on the position, number and distribution of the target key points, the inspection result of the target object is determined. The inspection result may include the shape, size and position of the target object. In this way, the position information of the target object in the three-dimensional coordinate system and the size and shape of the target object are given to provide perception support for downstream tasks.
[0036] In summary, the embodiment of the present application achieves accurate detection of the target object by acquiring point cloud data, determining a bird's-eye view, finding a target feature model, and determining a detection result. The method for target object inspection provided by the embodiment of the present application has broad application prospects in the fields of autonomous driving, robot navigation, three-dimensional reconstruction, etc.
[0037] The above technical solution obtains point cloud data including a key point set of the target object, and then determines a bird's-eye view containing the target object based on the point cloud data, and then determines a target feature model corresponding to the target object in a pre-constructed model feature library based on the key point set and the bird's-eye view, wherein the model feature library contains feature models of multiple different objects, and finally determines the detection result of the target object based on the target key points in the key point set that match the target feature model. The present application can accurately identify the target object in operation in a complex outdoor work scene by pre-constructing a model feature library corresponding to the target object, based on the target model in the model feature library, and improves the accuracy and real-time performance of the detection results.
[0038] In an embodiment of the present application, a method for determining a key point set includes: acquiring image data, the image data containing an image of a target object; performing target recognition on the image data to obtain a two-dimensional target frame of the target object in the image data; based on a predetermined calibration relationship between the image data and the point cloud data, determining a point cloud corresponding to a pixel point in the two-dimensional target frame from the point cloud data; and determining a key point set based on the point cloud corresponding to a pixel point in the two-dimensional target frame.
[0039] Specifically, the processor can receive an image containing a target object captured by a camera or an image sensor, and preprocess the image data, including denoising, graying, and edge detection. Then, target detection is performed on the image data according to the deep learning model to obtain a two-dimensional target frame of the target object. In one example, YOLOX can be used to perform target detection and recognition on image data, and for target objects within the field of view, the YOLOX target detection network is used to perform target detection on the target object. In addition, taking into account the complex environment of the detection scene of some target objects, the embodiment of the application itself inserts a channel attention module into the CSPDarknet backbone network of the YOLOX model to help the network better capture features at different levels, and finally obtain the two-dimensional target frame information, label information, and confidence information of the target object.
[0040] Next, considering that the two-dimensional target box does not belong to detailed pixel-level semantic segmentation, the box may contain other data points besides the target object. Therefore, it is necessary to determine the key point set based on the point cloud data and the two-dimensional target box based on the predetermined calibration relationship between the image data and the point cloud data. Among them, the calibration relationship is the calibration relationship between the image acquisition device and the point cloud acquisition device. In an example, the pixel points in the two-dimensional target box can be converted into three-dimensional space according to the predetermined calibration relationship to obtain the corresponding three-dimensional coordinate range, and then the corresponding point cloud can be filtered out from the point cloud data. Based on the filtered point cloud, the key points are extracted through the point cloud processing algorithm (such as the feature extraction method in the PCL library) to obtain the key point set corresponding to the target object. In another example, the pixel points in the target frame can be reverse mapped to the radar coordinate system and the point set can be screened using downsampling technology to obtain an appropriate amount of point sets as key points. Here, voxel downsampling is considered to be used to screen the point set to ensure that the selected points are distributed more evenly, meet the quantity requirements and cover the key features of the original point cloud. At the same time, because the point set is reverse mapped from the two-dimensional target recognition frame, it is guaranteed that most of the point clouds are components of the target hook, thereby improving the accuracy and efficiency of subsequent processing.
[0041] In this way, the embodiment of the present application combines image data with point cloud data, and utilizes deep learning, machine learning and other technologies to perform target recognition and point cloud processing, so as to accurately determine the key point set of the target object.
[0042] In an embodiment of the present application, a target feature model corresponding to a target object in a pre-constructed model feature library is determined based on a key point set and a bird's-eye view, including: performing feature extraction on the key point set to obtain a first feature, the first feature including a multi-scale semantic feature and position information; performing feature extraction on the bird's-eye view to obtain a second feature, the second feature including a bird's-eye view feature; splicing the first feature and the second feature to obtain a key point feature; matching each feature model in the model feature library according to the key point feature to determine the number of matching points between the key point set and each feature model; determining the feature model whose matching point number is the target maximum as the target feature model corresponding to the target object.
[0043] It can be understood that in order to determine the target feature model corresponding to the target object in the pre-built model feature library according to the key point set and the bird's-eye view, it is necessary to determine the key point features corresponding to the key point set, and determine the target feature model of the target object in the model feature library by the feature matching method. First, the key point set is feature extracted to obtain the first feature, and the first feature includes multi-scale semantic features and position information. Specifically, there are regular voxels of multi-scale semantic features around the key point. For each key point, the non-empty voxels of the minimum resolution level are first identified within the radius, and the voxels in the adjacent voxel set are transformed by PointNet to generate the key point features at this level. By performing the same operation at different levels and aggregating and connecting the features of different levels, the multi-scale semantic features of this key point are generated, and the accurate position information is retained. Further, the bird's-eye view is feature extracted to obtain the second feature, that is, the bird's-eye view feature of the key point. In one example, the bird's-eye view feature of the key point can be obtained by bilinear difference of the bird's-eye view mapping. In another example, a convolutional neural network or other image feature extraction method can be used to obtain the bird's-eye view feature. After obtaining the first and second features of the key point, the first feature and the second feature can be spliced to form a joint feature vector. This feature vector contains both the local information of the key point and the global information of the scene to obtain the key point features of each key point. The three-dimensional structure of the key point features is richer, which is conducive to improving the accuracy of the feature matching results.
[0044] Furthermore, the number of matching points between the key point set and each feature model is determined according to each feature model in the feature library of the key point feature matching model. In one example, the Euclidean distance or cosine similarity between feature vectors can be used as a metric, and a threshold of similarity or distance can be set to determine the matching situation. Only when the similarity between feature points is higher than the threshold or the distance is lower than the threshold, they are considered to match. Then, for each feature model, the number of matching points between each feature model and the key point set is counted. Finally, the feature model with the target maximum number of matching points is determined as the target feature model, that is, among all feature models, the feature model with the most matching points is selected as the target feature model corresponding to the target object. It can be understood that if there are multiple feature models with the same maximum number of matching points, other factors can be further considered to make a decision, such as the quality of matching, the diversity of features, etc.
[0045] In this way, the embodiment of the present application can accurately determine the target feature model corresponding to the target object by combining the key point set and the bird's-eye view to perform feature extraction and matching.
[0046] In an embodiment of the present application, the feature model includes multiple feature points, each feature point is associated with a feature point feature, and each feature model in the feature library is matched according to the key point feature to determine the number of matching points between the key point set and each feature model, including: determining the Euclidean distance between each key point in the key point set and each feature point of each feature model according to the key point feature and the feature point feature; determining the key point and feature point whose Euclidean distances are the minimum as matching points; determining the number of matching points between the key point set and each feature model to obtain the number of matching points between the key point set and each feature model.
[0047] It can be understood that each feature model in the model feature library includes multiple feature points, each feature point is associated with a feature point feature, that is, a feature vector, and the feature point feature is also formed by the fusion of multi-scale semantic features, location information, and bird's-eye view features. In order to determine the number of matching points between the key point set and each feature model in the model feature library, the embodiment of the present application can use the Euclidean distance to measure the similarity between the two feature vectors, and at the same time combine the nearest neighbor search method to quickly detect the target hook.
[0048] In one example, the number of matching points between the key point set and each feature model in the model feature library can be determined based on the Euclidean distance. Specifically, first traverse each feature model in the model feature library, assuming that the first feature model M1 is taken, its feature descriptor set, that is, the feature point set is D_M1 = {d_A1, d_A2, ..., d_Am}, and the key point feature set corresponding to the key point set N is D_N = {d_B1, d_B2, ..., d_Bm}, and the distance matrix J(i, j) = dist(d_Ai, dBj) between M1 and the key point set N is constructed using the Euclidean distance. Furthermore, insert the distance matrix into the KD tree data structure, traverse each row of data J(i,:), find the matching point d_Bj in the nearest key point set corresponding to the model point (feature point of the feature model) d_Ai, and then traverse the j-column data J(:,j) to obtain the nearest matching point corresponding to the point cloud feature point d_Bj; if d_Ai and d_Bj are the closest points to each other, then add this point to the matching point count of the feature model M1. In this way, the number of matching points between each feature model and the key point set in the model feature library can be determined.
[0049] In this way, by calculating the Euclidean distance and finding the matching points with the minimum value, the number of matching points between the key point set and each feature model in the model feature library can be determined. This method has broad application prospects in the fields of target recognition, image registration and 3D reconstruction.
[0050] In an embodiment of the present application, a method for constructing a model feature library includes: obtaining sample point cloud data of multiple sample models, the multiple sample models are models of multiple different objects, and the sample point cloud data includes a feature point set of the sample models; determining regular voxels of multi-scale semantic features around each feature point in the feature point set through a 3D sparse convolutional network; determining voxel features of each feature point at different levels based on the regular voxels of multi-scale semantic features around each feature point; aggregating and connecting the voxel features at different levels to obtain multi-scale semantic features and position information of each feature point; determining a sample bird's-eye view of each sample model based on the sample point cloud data; determining the bird's-eye view features of each feature point based on the bird's-eye view; splicing the multi-scale semantic features, position information and bird's-eye view features into feature point features of each feature point; generating a corresponding feature model based on the feature point features of each feature point in the feature point set of each sample model to form a model feature library.
[0051] It can be understood that the multiple sample models mainly refer to the multiple different models of similar products corresponding to the target object being detected. For example, if the target object is a hook, the multiple sample models refer to the multiple different models of hooks currently on the market. Specifically, in order to construct a model feature library corresponding to the target object, it is first necessary to obtain 3D models of multiple different objects as sample models, and then extract sample point cloud data of each sample model. The sample point cloud data includes a feature point set of the sample model, which includes a point set on the surface of the sample model and possible additional information, such as color and normal.
[0052] Further, a 3D sparse convolutional network is used to process the sample point cloud data. The 3D sparse convolutional network refers to a 3D sparse convolutional network that has been trained using a large amount of sample data, and can be used to extract effective multi-scale semantic features from the point cloud data. It can be understood that since the sparse convolutional network can efficiently process non-empty voxels, that is, voxels containing points, it is particularly suitable for processing sparse 3D data. Preferably, in order to improve the data quality, the sample point cloud data can be pre-processed before inputting the sample point cloud data into the 3D sparse convolutional network, such as normalization, denoising, downsampling, etc. The sample point cloud data is divided into a regular 3D voxel grid, where each voxel can be regarded as a small 3D cube, which may contain zero, one or more points. Then, for each feature point in the feature point set of the sample model, a 3D sparse convolutional network is used to determine the semantic features at multiple scales around it, and convolution operations are applied to voxels of different scales to capture semantic information from local to global. After extracting voxel features at different levels (i.e., different scales) from the 3D sparse convolutional network, these voxel features at different levels are aggregated and connected to form multi-scale semantic features for each feature point. In one example, this can be achieved by splicing, weighted averaging, or other more complex feature fusion methods. Preferably, the position information of each feature point and its three-dimensional coordinates in 3D space are retained at the same time.
[0053] At the same time, a bird's-eye view of each sample model can be generated based on the sample point cloud data. The bird's-eye view is usually an image that overlooks the scene from above and is projected onto a 2D plane. Then, image processing techniques (such as convolutional neural networks) are used to extract features on the bird's-eye view, namely, bird's-eye view features. It can be understood that the bird's-eye view features may contain information about the layout, direction, and relative position of objects in the scene.
[0054] Finally, the multi-scale semantic features, location information, and bird's-eye view features are spliced to form the complete feature point features of each feature point. For each sample model, the corresponding feature model is generated according to the feature point features of all feature points in its feature point set. The feature models of all sample models are combined to form a model feature library.
[0055] In this way, the embodiment of the present application generates a model feature library with rich information by extracting multi-scale semantic features, aggregating connected voxel features, introducing bird's-eye view features, etc., providing strong support for subsequent feature matching and recognition tasks.
[0056] In an embodiment of the present application, a bird's-eye view containing a target object is determined based on point cloud data, including: voxelizing the point cloud data to obtain voxelized point cloud data; extracting features from the voxelized point cloud data according to multiple preset downsampling scales to obtain a multi-scale feature map; and performing planar projection processing on the multi-scale feature map to generate a bird's-eye view.
[0057] It can be understood that the original point cloud data usually contains a large number of discrete points, each of which has a three-dimensional coordinate and possible additional information (such as color, intensity, etc.). In order to convert the discrete point cloud data into a regular structure that is easier to process, the point cloud data can be divided into a regular 3D voxel grid, where each voxel is a small 3D cube that can contain zero, one or more points. After voxelization, the point cloud data, each voxel may contain statistical information of all points in the voxel, such as the number of points, average coordinates, color histogram, etc. In order to capture features at different scales, multiple preset downsampling scales can be defined, which determine the resolution of the voxel grid, that is, the size of each voxel. Then, for each preset downsampling scale, the voxelized point cloud data is downsampled and features are extracted, wherein the extracted features may include the number, density, average height, color histogram, etc. of points in the voxel, which can reflect the local and global properties of the point cloud data. Further, the features at each downsampling scale are combined to form a multi-scale feature map. The multi-scale feature map can provide a comprehensive description of the point cloud data at different scales. Finally, the multi-scale feature map is projected from the 3D space to a 2D plane to generate a bird's-eye view. In one example, a specific viewing angle can be selected and a specific projection method can be used for projection according to actual conditions, such as vertical projection from above, and nearest neighbor projection or average projection.
[0058] In this way, the embodiment of the present application generates a bird's-eye view containing the target object from the point cloud data through steps such as voxelization processing, multi-scale feature extraction and plane projection processing, providing strong support for subsequent tasks such as target detection and scene understanding.
[0059] In an embodiment of the present application, a detection result of a target object is determined based on target key points in a key point set that match a target feature model, including: determining the number of target key points in a key point set that match a target feature model to obtain the number of target key points; when the number of target key points is greater than a preset number of matching points, performing point cloud registration processing on the target key points and the target feature model through a preset algorithm to obtain a position transformation matrix; determining the posture data of the target object based on the position transformation matrix and the target key points; the detection result of the target object includes the posture data of the target object.
[0060] Specifically, the preset matching point number is a preset threshold value used to determine whether there is a feature model that matches the target object in the model feature library, that is, an object of the same model as the target object. After determining the matching point number of each feature model and the key point set of the target object, the feature model corresponding to the maximum matching point number can be determined as the target feature model of the target object, and the number of target key points in the key point set that match the target feature model can be determined, that is, the target key point number. Further, the target key point number is compared with the preset matching point number. If the target key point number is greater than or equal to the preset matching point number, it is considered that the object corresponding to the target feature model and the target object may belong to the same type of object of the same model. At this time, the detection result of the target object can be determined based on the target feature model and the target key point.
[0061] Specifically, first, the target key points and the target feature model are processed by point cloud registration through a preset algorithm to obtain a position transformation matrix. Among them, the preset algorithm is an algorithm for point cloud registration, such as the iterative closest point (ICP) algorithm, the random sampling consistency (RANSAC) algorithm, etc. The preset algorithm can use the matching key points to calculate the position transformation relationship between the two point clouds, that is, the position transformation matrix. The position transformation matrix can be used to describe the position transformation relationship (such as rotation, translation, etc.) between the point cloud to be detected (target key points) and the target feature model. Further, the posture data of the target object can be determined according to the position transformation matrix and the target key points. In an example, the target feature model can be transformed into the coordinate system of the point cloud to be detected using the position transformation matrix, and then the posture data of the target object is determined according to the transformed model and the target key points. Finally, the detection result of the target object is determined according to the posture data of the target object. Based on the posture data of the target object, the position, direction, size and other information of the target object in the scene to be detected can be determined, thereby obtaining the final detection result.
[0062] In this way, the embodiment of the present application accurately determines the detection result of the target object through the steps of matching target key points, calculating the number of target key points, comparing the number of target key points with the preset matching points, performing point cloud registration processing, and determining the posture data of the target object, thereby providing strong support for subsequent target tracking, behavior analysis and other tasks.
[0063] In an embodiment of the present application, the method also includes: determining the number of target key points in the key point set that match the target feature model to obtain the number of target key points; when the number of target key points is less than or equal to the preset number of matching points, determining the candidate area containing the target object in the bird's-eye view according to a preset candidate area generation algorithm; optimizing the candidate area according to the target key points to obtain the target area, wherein the optimization processing includes pooling processing; and determining the detection result of the target object according to the target area.
[0064] Specifically, if the number of target key points is less than the preset number of matching points, it means that the object corresponding to the target feature model may be of a different model from the target object. At this time, the detection result of the target object can be determined by optimizing the candidate area in the bird's-eye view.
[0065] First, a candidate region is generated in a bird's-eye view according to a preset candidate region generation algorithm, wherein the preset candidate region generation algorithm is an algorithm for generating a candidate region that may contain a target object in a bird's-eye view, and this algorithm may generate a candidate region based on methods such as image segmentation, target detection, and clustering. Further, the target key points can be used to filter and optimize the candidate region to retain the region that is most likely to contain the target object and obtain the target region. The optimization process may include pooling, which can aggregate the features in the candidate region to reduce the amount of calculation and improve the robustness. The pooling process may be maximum pooling, average pooling, etc. It is understandable that in addition to the pooling process, filtering can also be performed based on the size, shape, position, and other features of the candidate region. Finally, the detection result of the target object is determined based on the target region, and the confidence can be calculated by a machine learning model, a statistical method, a rule matching, and the like. According to the confidence of the target region, if the confidence is higher than a certain threshold, it is considered that the target object is detected and the detection result is output.
[0066] In one example, the candidate area can be optimized according to the target key points. The entire point cloud area collected containing the target object is divided into multiple grid pool modules. Based on the candidate area, the center point of each candidate area grid pool is aggregated with the key point features within a certain radius using the PN++ method. If there is no key point feature information in this area, it is considered that there is no target object in this area. A coordinate information is spliced on each aggregated key point feature as the coordinate difference from the key point to the corresponding center point. A multi-dimensional multi-layer MLP network is used to vectorize and transform all grid pool features in the same candidate area, and the confidence of each candidate area is obtained to obtain the final detection result.
[0067] In this way, the method in the embodiment of the present application determines the candidate area, performs optimization processing, calculates the confidence, and determines the detection result. When the match between the key point set and the target feature model is not reliable enough, the detection result of the target object is accurately determined by optimizing the candidate area, thereby providing strong support for subsequent target tracking, behavior analysis and other tasks.
[0068] An embodiment of the present application also provides a processor configured to execute the method for target object detection in the above implementation.
[0069] Figure 2 This is a structural block diagram of a computing device provided in an embodiment of the present application. Figure 2 As shown, an embodiment of the present application also provides a computing device 200, including: a memory 210, configured to store instructions; and a processor 220 in the above embodiment, configured to call instructions from the memory 210 and to implement the method for target object detection in the above embodiment when executing the instructions.
[0070] An embodiment of the present application also provides a machine-readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, the method for detecting a target object in the above-mentioned embodiment is implemented.
[0071] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment or an embodiment in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0072] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems) and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0073] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0074] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0075] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0076] The memory may include non-permanent memory in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.
[0077] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.
[0078] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.
[0079] The above are only embodiments of the present application and are not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application should be included within the scope of the claims of the present application.
Claims
1. A method for detecting a target object, characterized in that: The method comprises: Acquire point cloud data, wherein the point cloud data includes a key point set of a target object; Determining a bird's-eye view including the target object according to the point cloud data; Determine, according to the key point set and the bird's-eye view, a target feature model corresponding to the target object in a pre-constructed model feature library, wherein the model feature library contains feature models of multiple different objects; The detection result of the target object is determined according to the target key points in the key point set that match the target feature model.
2. The method according to claim 1, characterized in that The method for determining the key point set includes: Acquire image data, wherein the image data includes an image of the target object; Performing target recognition on the image data to obtain a two-dimensional target frame of the target object in the image data; Based on a predetermined calibration relationship between the image data and the point cloud data, determining a point cloud corresponding to a pixel point in the two-dimensional target frame from the point cloud data; The key point set is determined based on the point cloud corresponding to the pixel points in the two-dimensional target frame.
3. The method according to claim 1, characterized in that Determining, according to the key point set and the bird's-eye view, a target feature model corresponding to the target object in a pre-constructed model feature library includes: Performing feature extraction on the key point set to obtain a first feature, wherein the first feature includes a multi-scale semantic feature and position information; Performing feature extraction on the bird's-eye view to obtain a second feature, wherein the second feature includes a bird's-eye view feature; Concatenating the first feature and the second feature to obtain a key point feature; Matching each feature model in the model feature library according to the key point features to determine the number of matching points between the key point set and each feature model; The feature model with the target maximum matching point number is determined as the target feature model corresponding to the target object.
4. The method according to claim 3, characterized in that: The feature model includes a plurality of feature points, each of which is associated with a feature point feature, and the matching of each feature model in the model feature library according to the key point feature to determine the number of matching points between the key point set and each of the feature models includes: Determining the Euclidean distance between each of the key points in the key point set and each of the feature points in each of the feature models according to the key point features and the feature point features; The key points and feature points whose Euclidean distances are the minimum are determined as matching points; The number of matching points between the key point set and each of the feature models is determined to obtain the number of matching points between the key point set and each of the feature models.
5. The method according to claim 1, characterized in that The method for constructing the model feature library includes: Acquire sample point cloud data of a plurality of sample models, where the plurality of sample models are models of a plurality of different objects, and the sample point cloud data includes feature point sets of the sample models; Determine regular voxels of multi-scale semantic features around each feature point in the feature point set through a 3D sparse convolutional network; Determine voxel features of each feature point at different levels according to regular voxels of multi-scale semantic features around each feature point; Aggregating and connecting the voxel features at different levels to obtain multi-scale semantic features and position information of each feature point; Determine a sample bird's-eye view of each of the sample models according to the sample point cloud data; Determine the bird's-eye view feature of each of the feature points according to the bird's-eye view; splicing the multi-scale semantic features, the position information and the bird's-eye view features into feature point features of each feature point; A corresponding feature model is generated according to the feature point features of each feature point in the feature point set of each sample model to form the model feature library.
6. The method according to claim 1, characterized in that The step of determining a bird's-eye view containing the target object according to the point cloud data comprises: Performing voxel processing on the point cloud data to obtain voxelized point cloud data; Performing feature extraction on the voxelized point cloud data according to a plurality of preset downsampling scales to obtain a multi-scale feature map; Performing plane projection processing on the multi-scale feature map to generate the bird's-eye view.
7. The method according to claim 1, characterized in that The step of determining the detection result of the target object according to the target key points in the key point set that match the target feature model comprises: Determine the number of target key points in the key point set that match the target feature model to obtain the number of target key points; When the number of target key points is greater than the preset number of matching points, performing point cloud registration processing on the target key points and the target feature model by a preset algorithm to obtain a position transformation matrix; The posture data of the target object is determined according to the position transformation matrix and the target key point, and the detection result of the target object includes the posture data of the target object.
8. The method according to claim 1, characterized in that The method further comprises: Determine the number of target key points in the key point set that match the target feature model to obtain the number of target key points; When the number of target key points is less than or equal to the preset number of matching points, determining a candidate area containing the target object in the bird's-eye view according to a preset candidate area generation algorithm; Optimizing the candidate region according to the target key point to obtain the target region, wherein the optimization process includes pooling process; A detection result of the target object is determined according to the target area.
9. A processor, characterized in that: The method is configured to execute the method for target object detection according to any one of claims 1 to 8.
10. A machine-readable storage medium storing a program or an instruction, characterized in that: When the program or the instruction is executed by a processor, the method for detecting a target object according to any one of claims 1 to 8 is implemented.