Multi-sensor parameter calibration method and target object detection method

By using a priori feature library to match RGB diagrams and depth diagrams in multi-sensor joint calibration, the target feature pair is determined, thereby improving the accuracy and efficiency of multi-sensor joint calibration, and solving the problem of low accuracy in the prior art.

CN120028804APending Publication Date: 2025-05-23ZOOMLION HEAVY INDUSTRY SCIENCE AND TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411875890.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-19
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

The method of multi-sensor joint calibration in the prior art has the problem of low accuracy, especially when parameter errors are large under external interference.

Method used

By acquiring the RGB map and point cloud data collected by the target area by the image acquisition device and the point cloud acquisition device, the depth map of the target area is determined based on the point cloud data, and the RGB map and depth map are matched with the prior feature templates in the pre-constructed prior feature library to determine the target feature pairs that match the RGB map and the depth map, thereby determining the target external parameters of the image acquisition device relative to the point cloud acquisition device.

Benefits of technology

The accuracy and efficiency of multi-sensor joint calibration is improved, the data processing volume is reduced, the accuracy of matching results is enhanced, and the accuracy of calibration results is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120028804A_ABST
    Figure CN120028804A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-sensor parameter calibration method and a target object detection method, and is used for parameter calibration between an image acquisition device and a point cloud acquisition device, and the method comprises the steps: obtaining an RGB image and point cloud data collected by the image acquisition device and the point cloud acquisition device for a target region; determining a depth map of the target area according to the point cloud data; respectively matching the RGB image and the depth image with a priori feature template in a pre-constructed priori feature library to obtain a matching result; under the condition that the matching result indicates that both the RGB image and the depth image are matched with the target prior feature template, determining a matched target feature pair of the RGB image and the depth image according to the target prior feature template; and determining a target external parameter of the image acquisition device relative to the point cloud acquisition device according to the target feature pair. Based on the prior feature model in the prior template feature library, the method of adding prior information is adopted to calibrate the joint external parameters of the multiple sensors, and the accuracy of the calibration result is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of multi-sensor joint calibration, and specifically to a multi-sensor parameter calibration method and a target object detection method. Background Art

[0002] In the prior art, when fusing the data of the laser radar and the monocular camera, joint calibration is only performed in the initial situation, without considering the parameter errors caused by external interference in the actual application process, and the strategy of fusing the sensors after separate detection is adopted, which easily leads to the loss of complementary information of the sensors, thereby affecting the accuracy of the calibration results. Therefore, the multi-sensor joint calibration method used in the prior art has the problem of low accuracy. Summary of the invention

[0003] The purpose of the embodiments of the present application is to provide a method and device for multi-sensor parameter calibration, a method and device for target object detection, and a machine-readable storage medium, so as to solve the problem of low accuracy of the multi-sensor joint calibration method used in the prior art.

[0004] In order to achieve the above-mentioned purpose, the first aspect of the embodiment of the present application provides a method for multi-sensor parameter calibration, which is used for parameter calibration between an image acquisition device and a point cloud acquisition device, and the method includes:

[0005] Obtaining RGB images and point cloud data collected by an image acquisition device and a point cloud acquisition device for a target area respectively;

[0006] Determine a depth map of the target area based on the point cloud data;

[0007] Match the RGB image and the depth image with the prior feature templates in the pre-built prior feature library to obtain the matching results;

[0008] When the matching result indicates that both the RGB image and the depth image match the target prior feature template, a target feature pair matching the RGB image and the depth image is determined according to the target prior feature template, and the target prior feature template is any prior feature template in the prior feature library;

[0009] The target extrinsic parameters of the image acquisition device relative to the point cloud acquisition device are determined according to the target feature pair.

[0010] In an embodiment of the present application, each prior feature template includes image features and depth map features of the corresponding prior object, and the RGB image and the depth map are respectively matched with the prior feature templates in a pre-constructed prior feature library to obtain a matching result, including: performing similarity calculations on the image features of the RGB image and the image features of each prior feature template to obtain corresponding image feature similarities; performing similarity calculations on the depth map features of the depth map and the depth map features of each prior feature template to obtain corresponding depth map feature similarities; determining a first prior feature template based on the image feature similarity, and determining a second prior feature template based on the depth map feature similarity; and determining that the RGB image and the depth map are both matched to the target prior feature template when the first prior feature template and the second prior feature template are the same prior feature templates.

[0011] In an embodiment of the present application, the image feature similarity between the image feature of the first prior feature template and the image feature of the RGB image is greater than a first preset similarity threshold, and the depth map feature similarity between the depth map feature of the second prior feature template and the depth map feature of the depth map is greater than a second preset similarity threshold.

[0012] In an embodiment of the present application, according to the target prior feature template, a target feature pair that matches the RGB image and the depth map is determined, including: extracting image features of the RGB image and depth map features of the depth map; determining a first feature vector in the image features of the RGB image that matches the target prior feature template; determining a second feature vector in the depth map features of the depth map that matches the target prior feature template; and matching the first feature vector and the second feature vector to obtain a target feature pair.

[0013] In an embodiment of the present application, the method also includes: in the process of matching the prior feature templates in the prior feature library, matching the prior feature templates in the prior feature library according to a preset initial feature image area size; if no prior feature template is matched, adjusting the size of the image area until the image area size reaches a size threshold, or matching any prior feature template, then stopping the matching.

[0014] In an embodiment of the present application, a method for constructing a prior feature library includes: determining a depth map and an RGB map of each prior object in a target area; pairing the depth map and the RGB map of each prior object with corresponding preset annotation information to form a training data set; training a preset neural network model according to the training data set to obtain a trained neural network model; performing feature extraction on the depth map and the RGB map of each prior object through the trained neural network model to obtain depth map features and image features of each prior object; associating and storing the depth map features and the image features of each prior object to obtain a prior feature template of each prior object to form a prior feature library.

[0015] A second aspect of an embodiment of the present application provides a method for detecting a target object, the method comprising:

[0016] Acquire point cloud data collected by a point cloud acquisition device and image data collected by an image acquisition device, wherein the image data includes an image of the target object, and the point cloud data includes a point cloud of the target object;

[0017] Performing target recognition on the image data to obtain a two-dimensional target frame of the target object in the image data;

[0018] Based on a predetermined calibration relationship between an image acquisition device and a point cloud acquisition device, determining a key point set corresponding to a pixel point in a two-dimensional target frame from the point cloud data, wherein the calibration relationship between the image acquisition device and the point cloud acquisition device is determined according to the above-mentioned method for multi-sensor parameter calibration;

[0019] The detection result of the target object is determined based on the key point set.

[0020] In an embodiment of the present application, the detection result of the target object is determined based on the key point set, including: determining a bird's-eye view containing the target object based on point cloud data; performing feature extraction on the key point set to obtain a first feature, the first feature including a multi-scale semantic feature and position information; performing feature extraction on the bird's-eye view to obtain a second feature, the second feature including a bird's-eye view feature; splicing the first feature and the second feature to obtain a key point feature; based on the key point feature, determining a target feature model corresponding to the target object in a pre-constructed model feature library, the model feature library including feature models of multiple different objects; and determining the detection result of the target object based on the target key points in the key point set that match the target feature model.

[0021] In an embodiment of the present application, a detection result of a target object is determined based on target key points in a key point set that match a target feature model, including: determining the number of target key points in the key point set that match a target feature model to obtain the number of target key points; when the number of target key points is less than or equal to a preset number of matching points, determining a candidate area in a bird's-eye view containing the target object according to a preset candidate area generation algorithm; optimizing the candidate area according to the target key points to obtain a target area, wherein the optimization processing includes pooling processing; and determining the detection result of the target object according to the target area.

[0022] A third aspect of an embodiment of the present application provides a processor configured to execute the above-mentioned method for multi-sensor parameter calibration, or the above-mentioned method for target object detection.

[0023] A fourth aspect of an embodiment of the present application provides a machine-readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, the above-mentioned method for multi-sensor parameter calibration or the above-mentioned method for target object detection is implemented.

[0024] The above technical scheme obtains the RGB image and point cloud data collected by the image acquisition device and the point cloud acquisition device for the target area respectively, and then determines the depth map of the target area based on the point cloud data, and then matches the RGB image and the depth map with the prior feature templates in the pre-constructed prior feature library respectively to obtain the matching results, and when the matching results indicate that both the RGB image and the depth map match the target prior feature template, the target feature pair that matches the RGB image and the depth map is determined based on the target prior feature template, wherein the target prior feature template is any prior feature template in the prior feature library, and finally the target external parameters of the image acquisition device relative to the point cloud acquisition device are determined based on the target feature pair. The present application can calibrate the joint external parameters of multiple sensors based on the prior feature model in the prior template feature library corresponding to the target area, and adopts the method of adding prior information, which is conducive to improving the accuracy and efficiency of the calibration results.

[0025] Other features and advantages of the embodiments of the present application will be described in detail in the subsequent specific implementation section. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] The accompanying drawings are used to provide a further understanding of the embodiments of the present application and constitute a part of the specification. Together with the following specific implementations, they are used to explain the embodiments of the present application, but do not constitute a limitation on the embodiments of the present application. In the accompanying drawings:

[0027] Figure 1 A flow chart of a method for multi-sensor parameter calibration provided in an embodiment of the present application;

[0028] Figure 2 A flowchart of a method for detecting a target object provided in an embodiment of the present application;

[0029] Figure 3 A structural block diagram of a computing device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0030] In order to make the purpose, technical scheme and advantages of the embodiments of the present application clearer, the technical scheme in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. It should be understood that the specific implementation methods described herein are only used to illustrate and explain the embodiments of the present application, and are not used to limit the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.

[0031] It should be noted that if the embodiments of the present application involve directional indications (such as up, down, left, right, front, back...), the directional indications are only used to explain the relative position relationship, movement status, etc. between the components under a certain specific posture (as shown in the accompanying drawings). If the specific posture changes, the directional indication will also change accordingly.

[0032] In addition, if there are descriptions involving "first", "second", etc. in the embodiments of the present application, the descriptions of "first", "second", etc. are only used for descriptive purposes and cannot be understood as indicating or suggesting their relative importance or implicitly indicating the number of technical features indicated. Therefore, the features defined as "first" and "second" may explicitly or implicitly include at least one of the features. In addition, the technical solutions between the various embodiments can be combined with each other, but they must be based on the ability of ordinary technicians in the field to implement them. When the combination of technical solutions is contradictory or cannot be implemented, it should be deemed that such combination of technical solutions does not exist and is not within the scope of protection required by this application.

[0033] Figure 1 A flow chart of a method for multi-sensor parameter calibration provided in an embodiment of the present application. Figure 1 As shown, an embodiment of the present application provides a method for multi-sensor parameter calibration, which is used for parameter calibration between an image acquisition device and a point cloud acquisition device. The method is described by applying the method to a processor as an example. The method may include the following steps.

[0034] Step S101, obtaining RGB images and point cloud data collected by an image acquisition device and a point cloud acquisition device for a target area respectively;

[0035] Specifically, the relative position between the image acquisition device and the point cloud acquisition device remains unchanged. In order to determine the calibration relationship between the image acquisition device and the point cloud acquisition device, it is first necessary to use the image acquisition device to obtain the RGB image of the target area, and use the point cloud acquisition device to obtain the point cloud data of the same target area, so as to provide a data basis for subsequent calibration.

[0036] Step S102, determining a depth map of the target area according to the point cloud data;

[0037] Specifically, based on the point cloud data, a depth map of the target area is generated through a reference coordinate system. In one example, each point in the point cloud data can be projected into the image coordinate system through the initial external parameters of the image acquisition device to obtain a depth map of the target area. It can be understood that the value of each pixel after the projection is the depth of the point cloud after the transformation from the radar coordinate system to the camera coordinate system, and is a null value when there is no mapping point.

[0038] Step S103, matching the RGB image and the depth image with the prior feature templates in the pre-built prior feature library respectively to obtain a matching result;

[0039] Specifically, the embodiment of the present application pre-constructs a priori feature library of the target area, and the prior feature library contains prior feature templates of various common objects in the target area, and each prior feature template includes a depth map and an RGB map of the corresponding prior object. In order to improve the accuracy of the calibration results of the image acquisition device and the point cloud acquisition device, the RGB map and the depth map can be matched with the prior feature templates in the pre-constructed prior feature library, respectively, and a matching result is obtained. In one example, the matching result may be that the RGB map and the depth map of the target area match the same prior feature template. In another example, the matching result may be that only one of the RGB map and the depth map of the target area matches the prior feature template, and the other does not match the prior feature template. In yet another example, the matching result may be that the RGB map and the depth map of the target area do not match the same prior feature template.

[0040] Step S104, when the matching result indicates that both the RGB image and the depth image match the target priori feature template, determine a target feature pair that matches the RGB image and the depth image according to the target priori feature template, where the target priori feature template is any priori feature template in the priori feature library;

[0041] Specifically, when the matching result indicates that the RGB image and the depth image match the same target prior feature template in the prior feature library, it means that the information of the target prior object corresponding to the target prior feature template is collected in both the RGB image and the depth image. Then, the area where the target prior object is located in the RGB image corresponds to the area where the target prior object is located in the depth image. In order to improve data processing efficiency, the target feature pairs that match the RGB image and the depth image can be determined based on the target prior feature template in the area corresponding to the target prior feature templates of the two. In this way, the size of the area to be matched can be reduced, thereby reducing the amount of data processing and improving matching efficiency.

[0042] Step S105, determining the target extrinsic parameters of the image acquisition device relative to the point cloud acquisition device according to the target feature pair.

[0043] Specifically, based on the matched target feature pairs, the target extrinsic parameters of the image acquisition device relative to the point cloud acquisition device can be calculated using geometric transformation or optimization algorithms. The target extrinsic parameters are used to describe the relative position and posture between the two acquisition devices, which is crucial for achieving data fusion and collaborative work between the two.

[0044] In this way, the prior feature library corresponding to the target area for data collection based on the image acquisition device and the point cloud acquisition device can narrow the feature matching range and reduce the amount of data processing. While improving the calibration efficiency, the matching feature pairs are determined according to the target prior feature template, which also improves the accuracy of the matching results, thereby improving the accuracy of the calibration results.

[0045] The above technical scheme obtains the RGB image and point cloud data collected by the image acquisition device and the point cloud acquisition device for the target area respectively, and then determines the depth map of the target area based on the point cloud data, and then matches the RGB image and the depth map with the prior feature templates in the pre-constructed prior feature library respectively to obtain the matching results, and when the matching results indicate that both the RGB image and the depth map match the target prior feature template, the target feature pair that matches the RGB image and the depth map is determined based on the target prior feature template, wherein the target prior feature template is any prior feature template in the prior feature library, and finally the target external parameters of the image acquisition device relative to the point cloud acquisition device are determined based on the target feature pair. The present application can calibrate the joint external parameters of multiple sensors based on the prior feature model in the prior template feature library corresponding to the target area, and adopts the method of adding prior information, which is conducive to improving the accuracy and efficiency of the calibration results.

[0046] In an embodiment of the present application, each prior feature template includes image features and depth map features of the corresponding prior object, and the RGB image and the depth map are respectively matched with the prior feature templates in a pre-constructed prior feature library to obtain a matching result, including: performing similarity calculations on the image features of the RGB image and the image features of each prior feature template to obtain corresponding image feature similarities; performing similarity calculations on the depth map features of the depth map and the depth map features of each prior feature template to obtain corresponding depth map feature similarities; determining a first prior feature template based on the image feature similarity, and determining a second prior feature template based on the depth map feature similarity; and determining that the RGB image and the depth map are both matched to the target prior feature template when the first prior feature template and the second prior feature template are the same prior feature templates.

[0047] In an embodiment of the present application, the image feature similarity between the image feature of the first prior feature template and the image feature of the RGB image is greater than a first preset similarity threshold, and the depth map feature similarity between the depth map feature of the second prior feature template and the depth map feature of the depth map is greater than a second preset similarity threshold.

[0048] It can be understood that in order to determine the matching result between the RGB image and the depth image of the target area and the prior feature model in the prior feature library, the similarity between the feature vectors can be used as a measure of the matching degree to determine the matching result.

[0049] Specifically, each prior feature template in the prior feature library includes the image features and depth map features of the corresponding prior object. Furthermore, the image features can be first extracted from the RGB image of the target area, and then the image features of the RGB image are similarly calculated with the image features of each prior feature template in the prior feature library to obtain a plurality of corresponding image feature similarities, and the first prior feature template matched to the RGB image is determined based on the plurality of image feature similarities. Among them, the similarity calculation method between image features may include Euclidean distance, cosine similarity, correlation coefficient, etc., to quantify the similarity between the image features in the RGB image and the image features in the prior feature template, and obtain the image feature similarity corresponding to the RGB image of the target area and each prior feature template. In one example, the image feature similarity corresponding to the first prior feature template is greater than the first preset similarity threshold, and is the maximum value among the plurality of image feature similarities.

[0050] At the same time, depth map features such as depth gradient, surface normal, depth change, etc. are extracted from the depth map, and then the depth map features of the depth map of the target area are respectively similarly calculated with the depth map features of each prior feature template in the prior feature library to obtain the depth map feature similarity corresponding to each prior feature template, and then determine the second prior feature template. In one example, the depth map feature similarity corresponding to the second prior feature template is greater than the second preset similarity threshold and is the maximum value among the multiple depth map feature similarities.

[0051] It can be understood that the first preset similarity threshold and the second preset similarity threshold are both set values ​​determined based on experiments or experience, and the first preset similarity threshold and the second preset similarity threshold can be the same. By setting the threshold and selecting the maximum value to determine the matched prior feature template, this dual matching judgment logic helps to improve the accuracy and reliability of the match and reduce the possibility of mismatching. Preferably, the setting of the preset similarity threshold needs to be adjusted according to the actual application scenario and data characteristics. When setting the threshold, you can consider using machine learning techniques such as cross-validation and grid search to find the optimal threshold combination. In addition, you can further adjust and optimize the similarity threshold based on the visual analysis of the matching results.

[0052] Furthermore, it is determined whether the first prior feature template and the second prior feature template are the same prior feature template. If the two are the same, it means that both the RGB image and the depth image have successfully matched the same target prior feature template, and the matching result is determined to be that both the RGB image and the depth image match the target prior feature template.

[0053] In this way, it is possible to accurately determine which prior feature template in the pre-built prior feature library best matches the RGB image and depth image, thereby providing a reliable foundation for subsequent tasks such as external parameter calculation and target recognition.

[0054] In an embodiment of the present application, according to the target prior feature template, a target feature pair that matches the RGB image and the depth map is determined, including: extracting image features of the RGB image and depth map features of the depth map; determining a first feature vector in the image features of the RGB image that matches the target prior feature template; determining a second feature vector in the depth map features of the depth map that matches the target prior feature template; and matching the first feature vector and the second feature vector to obtain a target feature pair.

[0055] Specifically, in order to quickly determine the feature pairs that match the RGB image and the depth image, the first feature vector that matches the image features of the RGB image with the image features of the target prior feature template can be determined, and at the same time, the second feature vector that matches the depth map features of the depth map of the target area with the depth map features of the target prior feature template can be determined. It can be understood that the first feature vector and the second feature vector each include multiple feature vectors, and then the first feature vector is matched with the second feature vector to obtain a target feature pair that matches each other. In one example, feature matching can be performed based on feature vector similarity, distance metrics, or other matching criteria.

[0056] In this way, the amount of data processing can be greatly reduced based on the target prior feature template, which is conducive to quickly finding matching target feature pairs in the RGB image and depth image, and also improves the accuracy of the matching results. These feature pairs not only provide key information for the subsequent image and point cloud data fusion, but also provide a basis for determining the relative position relationship between the image acquisition device and the point cloud acquisition device (i.e., external parameters).

[0057] In one example, when the RGB image and depth map of the target area do not match the same prior feature template, that is, the first prior feature template and the second prior feature template are not the same prior feature template, or only one of them matches the corresponding prior feature template, at this time, the feature pairs of the RGB image and depth map of the target area that match each other can be determined based on the first prior feature template and / or the second prior feature template respectively to perform external parameter calibration of the image acquisition device and the point cloud acquisition device.

[0058] In another example, when neither the RGB image nor the depth map of the target area matches the prior feature template, a matching feature pair is determined directly based on the image features of the RGB image and the depth map features of the depth map to perform external parameter calibration of the image acquisition device and the point cloud acquisition device.

[0059] In an embodiment of the present application, the method also includes: in the process of matching the prior feature templates in the prior feature library, matching the prior feature templates in the prior feature library according to a preset initial feature image area size; if no prior feature template is matched, adjusting the size of the image area until the image area size reaches a size threshold, or matching any prior feature template, then stopping the matching.

[0060] It can be understood that in order to improve the flexibility of matching and reduce the interference of environmental factors, the embodiment of the present application proposes a mechanism for dynamically adjusting the image area size in the process of matching the prior feature template in the prior feature library. Specifically, first, the preset initial feature image area size is used to extract features from the RGB image and the depth image, and match them with the prior feature template in the prior feature library. This initial size is set based on prior knowledge or experimental experience, and is intended to cover the area in the image that may contain the target feature. If no prior feature template is matched at the initial size, it indicates that the currently extracted feature area may not contain enough information to identify the target feature. At this time, the system adjusts the size of the image area according to the preset rules, such as increasing or decreasing the size, in an attempt to capture more or less feature information. After adjusting the size, the features are re-extracted and matched with the prior feature template. This process will be iterative until any of the following stop conditions are met: the image area size reaches the preset size threshold or matches any prior feature template. Among them, the image area size reaches the preset size threshold, indicating that all possible size ranges have been tried. Matching to any prior feature template indicates that the target feature has been successfully identified at the current size.

[0061] In one example, the specific strategy of resizing (such as the step size of increase or decrease, the direction of adjustment, etc.) needs to be set according to the actual application scenario and data characteristics. The size threshold is used to limit the range of resizing and prevent infinite loops, and can be set according to the actual situation of image size and feature distribution.

[0062] Thus, by introducing a mechanism to dynamically adjust the image region size, this method can more flexibly adapt to different image and data characteristics, and improve the accuracy and robustness of matching. At the same time, by setting reasonable stopping conditions, the amount of calculation can be effectively controlled, improving the practicality of the method.

[0063] In an embodiment of the present application, a method for constructing a prior feature library includes: determining a depth map and an RGB map of each prior object in a target area; pairing the depth map and the RGB map of each prior object with corresponding preset annotation information to form a training data set; training a preset neural network model according to the training data set to obtain a trained neural network model; performing feature extraction on the depth map and the RGB map of each prior object through the trained neural network model to obtain depth map features and image features of each prior object; associating and storing the depth map features and the image features of each prior object to obtain a prior feature template of each prior object to form a prior feature library.

[0064] Specifically, in order to construct the prior feature library corresponding to the target area, we first need to determine each prior object in the target area and obtain the depth map and RGB map of each prior object. The depth map provides the spatial position information of the object, while the RGB map provides the color and texture information of the object.

[0065] Furthermore, the depth map and RGB map of each prior object are paired with the corresponding preset annotation information. The annotation information may include the category, position and size of the prior object, etc., which is used to provide a supervisory signal in the subsequent neural network training process. In this way, the construction of a training data set containing depth maps, RGB maps and annotation information is completed through such pairing.

[0066] Next, a preset neural network model is trained using the training data set. The preset neural network model may be a convolutional neural network (CNN) or other neural network structures suitable for image processing. Through training, the neural network model can learn how to extract useful features from depth maps and RGB maps, and accurately classify or identify objects based on these features. Furthermore, the depth map and RGB map of each prior object are feature extracted using the trained neural network model to obtain the depth map features and image features of each prior object. Finally, the depth map features and image features of each prior object are associated and stored to form a prior feature template corresponding to the target area.

[0067] In this way, the embodiment of the present application builds a complete prior feature library by collecting feature templates of multiple prior objects, which can provide powerful feature support for subsequent object recognition, classification, tracking and other tasks. In addition, the embodiment of the present application combines the technology of deep learning and computer vision to improve the accuracy and efficiency of object recognition by building a prior feature library.

[0068] Figure 2 A flow chart of a method for detecting a target object provided in an embodiment of the present application is shown in FIG. Figure 2As shown, an embodiment of the present application provides a method for detecting a target object, which is described by taking the method applied to a processor as an example. The method may include the following steps.

[0069] Step S201, acquiring point cloud data acquired by a point cloud acquisition device and image data acquired by an image acquisition device, wherein the image data includes an image of a target object, and the point cloud data includes a point cloud of the target object.

[0070] Step S202: performing target recognition on the image data to obtain a two-dimensional target frame of the target object in the image data.

[0071] Step S203, based on a predetermined calibration relationship between the image acquisition device and the point cloud acquisition device, determine a key point set corresponding to the pixel points in the two-dimensional target frame from the point cloud data, wherein the calibration relationship between the image acquisition device and the point cloud acquisition device is determined according to the method for multi-sensor parameter calibration in the above-mentioned embodiment.

[0072] Step S204: determining the detection result of the target object based on the key point set.

[0073] Specifically, the processor can receive point cloud data collected by a point cloud acquisition device and image data collected by an image acquisition device, wherein the image data contains an image of the target object, and the point cloud data includes a point cloud of the target object. Furthermore, the image data is preprocessed, and the preprocessing includes denoising, graying, and edge detection. Then, target detection is performed on the image data according to the deep learning model to obtain a two-dimensional target frame of the target object. In one example, YOLOX can be used to perform target detection and recognition on image data, and for target objects within the field of view, the target detection of the target object is performed using the YOLOX target detection network. In addition, taking into account the complex environment of the detection scene of some target objects, the embodiment of the application itself inserts a channel attention module into the CSPDarknet backbone network of the YOLOX model to help the network better capture features at different levels, and finally obtain the two-dimensional target frame information, label information, and confidence information of the target object.

[0074] Next, considering that the two-dimensional target box does not belong to detailed pixel-level semantic segmentation, the box may contain other data points besides the target object. Therefore, it is necessary to determine the key point set based on the point cloud data and the two-dimensional target box based on the calibration relationship between the image acquisition device and the point cloud point set device determined in the above implementation. Among them, the calibration relationship is the calibration relationship between the image acquisition device and the point cloud acquisition device. In one example, the pixel points in the two-dimensional target box can be converted into three-dimensional space according to a predetermined calibration relationship to obtain the corresponding three-dimensional coordinate range, and then the corresponding point cloud can be screened out from the point cloud data. Based on the screened point cloud, the key points are extracted through the point cloud processing algorithm (such as the feature extraction method in the PCL library) to obtain the key point set corresponding to the target object. In another example, the pixel points in the target frame can be reverse mapped to the radar coordinate system and the point set can be screened using downsampling technology to obtain an appropriate number of point sets as key points. Here, voxel downsampling is considered to be used to screen the point set to ensure that the selected points are distributed more evenly, meet the quantity requirements and cover the key features of the original point cloud. At the same time, because the point set is reverse mapped from the two-dimensional target recognition frame, it is guaranteed that most of the point clouds are components of the target object, improving the accuracy and efficiency of subsequent processing. In this way, by combining image data and point cloud data, using deep learning, machine learning and other technologies for target recognition and point cloud processing, the key point set of the target object can be accurately determined.

[0075] Finally, the detection result of the target object is determined based on the key point set. In one example, the target object can be further analyzed and processed based on the key point set using a three-dimensional object detection algorithm to determine the detection result of the target object. The three-dimensional object detection algorithm can be a deep learning model based on point cloud.

[0076] In this way, the two-dimensional data of the target object is projected into the three-dimensional coordinate system through the external parameters jointly calibrated by the image acquisition device and the point cloud acquisition device, and the pixel position information is back-mapped into the point cloud map, and the key point set of the target object is further determined. Then, the detection result of the target object is determined based on the key point set. Compared with the mainstream pre-fusion or post-fusion recognition algorithm, this strategy can identify the target object more accurately and quickly, has higher robustness, and can be applied to complex working environments.

[0077] In some examples, the target object may be a hook, a sling, or other work accessories. In other examples, the target object may also be a pedestrian, a vehicle, or various reference objects or objects to be worked on in the work scene, etc., and examples are not given here one by one.

[0078] In an embodiment of the present application, the detection result of the target object is determined based on the key point set, including: determining a bird's-eye view containing the target object based on point cloud data; performing feature extraction on the key point set to obtain a first feature, the first feature including a multi-scale semantic feature and position information; performing feature extraction on the bird's-eye view to obtain a second feature, the second feature including a bird's-eye view feature; splicing the first feature and the second feature to obtain a key point feature; based on the key point feature, determining a target feature model corresponding to the target object in a pre-constructed model feature library, the model feature library including feature models of multiple different objects; and determining the detection result of the target object based on the target key points in the key point set that match the target feature model.

[0079] Specifically, a bird's-eye view is a view of a target object and its surroundings viewed vertically from above. This view helps capture the overall layout and shape of the target object, especially in three-dimensional space. In one example, a bird's-eye view can be generated by projecting or transforming point cloud data.

[0080] It can be understood that in order to determine the target feature model corresponding to the target object in the pre-built model feature library according to the key point set and the bird's-eye view, it is necessary to determine the key point features corresponding to the key point set, and determine the target feature model of the target object in the model feature library by the feature matching method. First, the key point set is feature extracted to obtain the first feature, and the first feature includes multi-scale semantic features and position information. Specifically, there are regular voxels of multi-scale semantic features around the key point. For each key point, the non-empty voxels of the minimum resolution level are first identified within the radius, and the voxels in the adjacent voxel set are transformed by PointNet to generate the key point features at this level. By performing the same operation at different levels and aggregating and connecting the features of different levels, the multi-scale semantic features of this key point are generated, and the accurate position information is retained. Further, the bird's-eye view is feature extracted to obtain the second feature, that is, the bird's-eye view feature of the key point. In one example, the bird's-eye view feature of the key point can be obtained by bilinear difference of the bird's-eye view mapping. In another example, a convolutional neural network or other image feature extraction method can be used to obtain the bird's-eye view feature. After obtaining the first and second features of the key point, the first feature and the second feature can be spliced ​​to form a joint feature vector. This feature vector contains both the local information of the key point and the global information of the scene to obtain the key point features of each key point. The three-dimensional structure of the key point features is richer, which is conducive to improving the accuracy of the feature matching results.

[0081] Furthermore, the number of matching points between the key point set and each feature model is determined according to each feature model in the feature library of the key point feature matching model. In one example, the Euclidean distance or cosine similarity between feature vectors can be used as a metric, and a threshold of similarity or distance can be set to determine the matching situation. Only when the similarity between feature points is higher than the threshold or the distance is lower than the threshold, they are considered to match. Then, for each feature model, the number of matching points between each feature model and the key point set is counted. Finally, the feature model with the target maximum number of matching points is determined as the target feature model, that is, among all feature models, the feature model with the most matching points is selected as the target feature model corresponding to the target object. It can be understood that if there are multiple feature models with the same maximum number of matching points, other factors can be further considered to make a decision, such as the quality of matching, the diversity of features, etc.

[0082] In this way, the embodiment of the present application can accurately determine the target feature model corresponding to the target object by combining the key point set and the bird's-eye view to perform feature extraction and matching.

[0083] Finally, the processor can determine the target key points in the key point set that match the target feature model through a feature matching algorithm. The feature matching algorithm may include neighbor search and Euclidean distance calculation. Then, based on the position, number and distribution of the target key points, the inspection result of the target object is determined. The inspection result may include the shape, size and position of the target object. In this way, the position information of the target object in the three-dimensional coordinate system and the size and shape of the target object are given to provide perception support for downstream tasks.

[0084] In an embodiment of the present application, a detection result of a target object is determined based on target key points in a key point set that match a target feature model, including: determining the number of target key points in the key point set that match a target feature model to obtain the number of target key points; when the number of target key points is less than or equal to a preset number of matching points, determining a candidate area in a bird's-eye view containing the target object according to a preset candidate area generation algorithm; optimizing the candidate area according to the target key points to obtain a target area, wherein the optimization processing includes pooling processing; and determining the detection result of the target object according to the target area.

[0085] Specifically, the preset matching point count is a preset threshold value used to determine whether there is a feature model matching the target object in the model feature library, that is, an object with the same model as the target object. After determining the matching point counts of each feature model and the key point set of the target object, the feature model corresponding to the maximum matching point count can be determined as the target feature model of the target object, and the number of target key points in the key point set that match the target feature model, that is, the target key point count, can be determined. Further, the target key point count is compared with the preset matching point count.

[0086] In one embodiment, if the number of target key points is less than the preset number of matching points, it means that the object corresponding to the target feature model may be of a different model from the target object. In this case, the detection result of the target object can be determined by optimizing the candidate area in the bird's-eye view.

[0087] First, a candidate region is generated in a bird's-eye view according to a preset candidate region generation algorithm, wherein the preset candidate region generation algorithm is an algorithm for generating a candidate region that may contain a target object in a bird's-eye view, and this algorithm may generate a candidate region based on methods such as image segmentation, target detection, and clustering. Further, the target key points can be used to filter and optimize the candidate region to retain the region that is most likely to contain the target object and obtain the target region. The optimization process may include pooling, which can aggregate the features in the candidate region to reduce the amount of calculation and improve the robustness. The pooling process may be maximum pooling, average pooling, etc. It is understandable that in addition to the pooling process, filtering can also be performed based on the size, shape, position, and other features of the candidate region. Finally, the detection result of the target object is determined based on the target region, and the confidence can be calculated by a machine learning model, a statistical method, a rule matching, and the like. According to the confidence of the target region, if the confidence is higher than a certain threshold, it is considered that the target object is detected and the detection result is output.

[0088] In one example, the candidate area can be optimized according to the target key points. The entire point cloud area collected containing the target object is divided into multiple grid pool modules. Based on the candidate area, the center point of each candidate area grid pool is aggregated with the key point features within a certain radius using the PN++ method. If there is no key point feature information in this area, it is considered that there is no target object in this area. A coordinate information is spliced ​​on each aggregated key point feature as the coordinate difference from the key point to the corresponding center point. A multi-dimensional multi-layer MLP network is used to vectorize and transform all grid pool features in the same candidate area, and the confidence of each candidate area is obtained to obtain the final detection result.

[0089] In another embodiment, if the number of target key points is greater than or equal to the preset number of matching points, it is considered that the object corresponding to the target feature model and the target object may belong to the same type of object of the same model. At this time, the detection result of the target object can be determined based on the target feature model and the target key points.

[0090] Specifically, first, the target key points and the target feature model are processed by point cloud registration through a preset algorithm to obtain a position transformation matrix. Among them, the preset algorithm is an algorithm for point cloud registration, such as the iterative closest point (ICP) algorithm, the random sampling consistency (RANSAC) algorithm, etc. The preset algorithm can use the matching key points to calculate the position transformation relationship between the two point clouds, that is, the position transformation matrix. The position transformation matrix can be used to describe the position transformation relationship (such as rotation, translation, etc.) between the point cloud to be detected (target key points) and the target feature model. Further, the posture data of the target object can be determined according to the position transformation matrix and the target key points. In an example, the target feature model can be transformed into the coordinate system of the point cloud to be detected using the position transformation matrix, and then the posture data of the target object is determined according to the transformed model and the target key points. Finally, the detection result of the target object is determined according to the posture data of the target object. Based on the posture data of the target object, the position, direction, size and other information of the target object in the scene to be detected can be determined, thereby obtaining the final detection result.

[0091] In this way, the embodiment of the present application accurately determines the detection result of the target object through the steps of matching target key points, calculating the number of target key points, comparing the number of target key points with the preset matching points, performing point cloud registration processing, and determining the posture data of the target object, thereby providing strong support for subsequent target tracking, behavior analysis and other tasks.

[0092] An embodiment of the present application also provides a processor configured to execute the method for multi-sensor parameter calibration in the above embodiment, or the method for target object detection in the above embodiment.

[0093] Figure 3 This is a structural block diagram of a computing device provided in an embodiment of the present application. Figure 3 As shown, an embodiment of the present application also provides a computing device 300, including: a memory 310, configured to store instructions; and a processor 320 in the above embodiment, configured to call instructions from the memory 310 and to implement the method for multi-sensor parameter calibration in the above embodiment, or the method for target object detection in the above embodiment when executing the instructions.

[0094] An embodiment of the present application also provides a machine-readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, the method for multi-sensor parameter calibration in the above-mentioned embodiment and / or the method for target object detection in the above-mentioned embodiment are implemented.

[0095] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment or an embodiment in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.

[0096] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems) and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0097] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0098] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0099] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.

[0100] The memory may include non-permanent memory in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.

[0101] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.

[0102] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.

[0103] The above are only embodiments of the present application and are not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application should be included within the scope of the claims of the present application.

Claims

1. A method for multi-sensor parameter calibration, characterized in that: Used for parameter calibration between an image acquisition device and a point cloud acquisition device, the method comprises: Acquire the RGB image and point cloud data respectively collected by the image acquisition device and the point cloud acquisition device for the target area; Determine a depth map of the target area according to the point cloud data; Matching the RGB image and the depth image with the prior feature templates in the pre-built prior feature library respectively to obtain a matching result; When the matching result indicates that both the RGB image and the depth image match a target priori feature template, determining a target feature pair matching the RGB image and the depth image according to the target priori feature template, wherein the target priori feature template is any priori feature template in the prior feature library; The target extrinsic parameters of the image acquisition device relative to the point cloud acquisition device are determined according to the target feature pair.

2. The method according to claim 1, characterized in that: Each of the prior feature templates includes image features and depth map features of the corresponding prior object, and the RGB image and the depth map are matched with the prior feature templates in the pre-constructed prior feature library to obtain a matching result, including: Calculate the similarity between the image features of the RGB image and the image features of each prior feature template to obtain the corresponding image feature similarity; Performing similarity calculations on the depth map features of the depth map and the depth map features of each of the prior feature templates to obtain corresponding depth map feature similarities; Determine a first priori feature template based on the image feature similarity, and determine a second priori feature template based on the depth map feature similarity; In a case where the first a priori feature template and the second a priori feature template are the same a priori feature template, it is determined that both the RGB image and the depth image match a target a priori feature template.

3. The method according to claim 2, characterized in that The image feature similarity between the image feature of the first prior feature template and the image feature of the RGB image is greater than a first preset similarity threshold, and the depth map feature similarity between the depth map feature of the second prior feature template and the depth map feature of the depth map is greater than a second preset similarity threshold.

4. The method according to claim 1, characterized in that The step of determining, according to the target prior feature template, a target feature pair that matches the RGB image and the depth image, comprises: Extracting image features of the RGB image and depth map features of the depth map; Determine a first feature vector in the image features of the RGB image that matches the target prior feature template; Determine a second feature vector in the depth map features of the depth map that matches the target prior feature template; The first feature vector and the second feature vector are matched to obtain the target feature pair.

5. The method according to claim 1, characterized in that The method further comprises: In the process of matching the prior feature templates in the prior feature library, matching the prior feature templates in the prior feature library according to a preset initial feature image area size; In the case that any priori feature template is not matched, the size of the image region is adjusted until the size of the image region reaches a size threshold, or any priori feature template is matched, then the matching is stopped.

6. The method according to claim 1, characterized in that The method for constructing the prior feature library includes: Determine the depth map and RGB map of each prior object in the target area; The depth map and the RGB map of each prior object are paired with the corresponding preset annotation information to form a training data set; Training a preset neural network model according to the training data set to obtain a trained neural network model; Performing feature extraction on the depth map and RGB map of each of the prior objects through the trained neural network model to obtain depth map features and image features of each of the prior objects; The depth map features and image features of each of the prior objects are associated and stored to obtain a priori feature template of each of the prior objects to form the prior feature library.

7. A method for detecting a target object, characterized in that: The method comprises: Acquire point cloud data acquired by a point cloud acquisition device and image data acquired by an image acquisition device, wherein the image data includes an image of the target object, and the point cloud data includes a point cloud of the target object; Performing target recognition on the image data to obtain a two-dimensional target frame of the target object in the image data; Based on a predetermined calibration relationship between an image acquisition device and a point cloud acquisition device, determining a key point set corresponding to a pixel point in the two-dimensional target frame from the point cloud data, wherein the calibration relationship between the image acquisition device and the point cloud acquisition device is determined according to the method for multi-sensor parameter calibration according to any one of claims 1 to 6; A detection result of the target object is determined based on the key point set.

8. The method according to claim 7, characterized in that The determining the detection result of the target object based on the key point set includes: Determining a bird's-eye view including the target object according to the point cloud data; Performing feature extraction on the key point set to obtain a first feature, wherein the first feature includes a multi-scale semantic feature and position information; Performing feature extraction on the bird's-eye view to obtain a second feature, wherein the second feature includes a bird's-eye view feature; Concatenating the first feature and the second feature to obtain a key point feature; Determine, according to the key point features, a target feature model corresponding to the target object in a pre-constructed model feature library, wherein the model feature library contains feature models of multiple different objects; The detection result of the target object is determined according to the target key points in the key point set that match the target feature model.

9. The method according to claim 8, characterized in that The step of determining the detection result of the target object according to the target key points in the key point set that match the target feature model comprises: Determine the number of target key points in the key point set that match the target feature model to obtain the number of target key points; When the number of target key points is less than or equal to the preset number of matching points, determining a candidate area containing the target object in the bird's-eye view according to a preset candidate area generation algorithm; Optimizing the candidate region according to the target key point to obtain the target region, wherein the optimization process includes pooling process; A detection result of the target object is determined according to the target area.

10. A processor, characterized in that: The method is configured to execute the method for multi-sensor parameter calibration according to any one of claims 1 to 6, or the method for target object detection according to any one of claims 7 to 9.

11. A machine-readable storage medium storing a program or an instruction, characterized in that: When the program or the instruction is executed by a processor, the method for multi-sensor parameter calibration according to any one of claims 1 to 6, or the method for target object detection according to any one of claims 7 to 9 is implemented.