Target object recognition method, device, system and medium integrated with laser radar
By combining image recognition and target constraint parameter optimization methods of point cloud data on a low-computing power platform, the problems of lack of three-dimensional information in image recognition and high resource consumption in point cloud processing are solved, and real-time recognition and extraction of multiple target objects are achieved.
Patent Information
- Application Number
- CN202411529566.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-30
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2044-10-30
AI Technical Summary
In the existing technology, image-based target object recognition methods lack three-dimensional information, resulting in recognition limitations. Although point cloud-based methods provide rich three-dimensional information, they are computationally intensive and consume large processing resources, making it difficult to achieve real-time multi-target recognition on low-computing power platforms.
By acquiring the initial point cloud data and the original image, image recognition is performed, target constraint parameters such as the grid scale are determined, and the point cloud data is allocated to the grid of the target image using a mapping matrix. The image recognition results and constraint parameters are combined to perform real-time target recognition and optimize the point cloud processing process.
While ensuring recognition accuracy, the calculation time is significantly reduced, and real-time spatial range recognition and extraction of multiple targets are achieved on a low-computing power platform, meeting application scenarios with high real-time requirements.
Smart Images

Figure CN119672381B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer vision technology, and in particular to a target object recognition method, device, system and medium integrated with laser radar. Background Art
[0002] Computer vision can realize automatic environmental perception to achieve tasks such as target object recognition, improve recognition accuracy, and greatly reduce personnel workload. While ensuring the recognition accuracy of target objects, it also improves the recognition speed and realizes real-time target object recognition, which is crucial for specific application scenarios such as power safety management and control, anti-extradition monitoring, intelligent monitoring, autonomous driving, human-computer interaction, and intelligent manufacturing.
[0003] The relevant technologies for target object recognition mainly include image-based semantic recognition and point cloud-based semantic recognition. The former can be applied to real-time recognition on low-computing power platforms, but cannot accurately represent the true three-dimensional distribution of the target. Point cloud data can provide rich environmental information. Point cloud clustering is usually used to extract point clouds related to the identified object and determine the three-dimensional spatial distribution of the target from them. However, this is time-consuming, does not meet real-time requirements, and cannot be applied to low-computing power platforms.
[0004] However, point cloud clustering usually requires traversing each point to achieve object recognition. As the amount of collected point cloud data increases dramatically, the time required for point cloud mapping extraction and clustering processing increases significantly, and the computing resources consumed are huge. Especially in multi-target object recognition scenarios, it is often difficult to achieve real-time recognition of multiple target objects on low-computing power platforms. Summary of the Invention
[0005] The present application provides a target object recognition method, device, system and medium with integrated laser radar, which are used for target object recognition methods based on images, which lack three-dimensional information and lead to recognition limitations. Although target object recognition methods based on point clouds provide rich three-dimensional information, they are difficult to achieve real-time recognition of multiple targets on low-computing power platforms due to computationally intensive processing and high resource consumption.
[0006] In a first aspect, the present application provides a target object recognition method, comprising: acquiring initial point cloud data and an original image; performing image recognition on the original image to obtain an image recognition result; determining target constraint parameters based on contour features of each object in the image recognition result, wherein the target constraint parameters include at least a grid scale of the target image; allocating the initial point cloud data to each target grid of the target image according to a preset mapping matrix, the image recognition result and the target constraint parameters to obtain target point cloud data; performing real-time target recognition on the target point cloud data based on the target constraint parameters to obtain a target object recognition result.
[0007] The second aspect of the present application provides a target object recognition device, comprising: an acquisition module for acquiring initial point cloud data and an original image. A first recognition module for performing image recognition on the original image to obtain an image recognition result. A determination module for determining target constraint parameters based on the contour features of each object in the image recognition result, wherein the target constraint parameters include at least the grid scale of the target image; an allocation module for allocating the initial point cloud data to each target grid of the target image according to a preset mapping matrix, the image recognition result and the target constraint parameters to obtain target point cloud data; and a second recognition module for performing real-time target recognition on the target point cloud data based on the target constraint parameters to obtain a target object recognition result.
[0008] The third aspect of the present application provides a system of integrated laser radar, comprising: a memory and at least one processor, wherein the memory stores instructions; the at least one processor calls the instructions in the memory so that the system of integrated laser radar executes the above-mentioned target object recognition method.
[0009] A fourth aspect of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores instructions, which, when executed on a computer, enable the computer to execute the above-mentioned target object recognition method.
[0010] In the technical solution provided by the present application, the target constraint parameters are determined by the contour features of each object in the image recognition results, and the constraint parameters are set for each object in a targeted manner, so that the subsequent point cloud data processing process is constrained by the key features of the target object; by introducing a new grid scale and other constraint parameters of the target image, the calculation time of the point cloud processing can be significantly reduced while ensuring the recognition accuracy; the grid scale is dynamically updated according to the actual observation conditions of the camera, so that the processing algorithm is insensitive to changes in the order of magnitude of the point cloud, thereby effectively reducing the computational complexity; the optimized method takes into account the advantages of small computational complexity and fast speed of image data processing, as well as the advantage of rich three-dimensional spatial information of point cloud data, and can realize real-time spatial range recognition and extraction of multiple targets, meeting application scenarios with high real-time requirements. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] Figure 1 This is a schematic diagram of an embodiment of a target object identification method in an embodiment of the present application;
[0012] Figure 2 This is a schematic diagram of another embodiment of the target object identification method in the embodiment of the present application;
[0013] Figure 3 This is a schematic diagram of an embodiment of a target object recognition device in an embodiment of the present application;
[0014] Figure 4 This is a schematic diagram of another embodiment of the target object recognition device in the embodiment of the present application;
[0015] Figure 5 This is a schematic diagram of an embodiment of a system integrating a laser radar in an embodiment of the present application. DETAILED DESCRIPTION
[0016] The present application provides a target object recognition method, device, system and medium integrated with laser radar, which are used to optimize the target object extraction process, shorten the operation time while ensuring the accuracy of the extraction results.
[0017] The terms "first," "second," "third," "fourth," and the like (if any) in the specification and claims of this application and in the accompanying drawings are used to distinguish similar objects and are not necessarily used to describe a particular order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate so that the embodiments described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "including" or "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product, or apparatus comprising a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0018] For ease of understanding, the specific process of the embodiment of the present application is described below. Figure 1 In one embodiment of the target object identification method of the present application, the method includes:
[0019] 101. Obtain initial point cloud data and original images.
[0020] It is understandable that the execution subject of this application can be a target object identification device, or a security monitoring system or a terminal such as a server, a vehicle or a robot, and the specific details are not limited here.
[0021] The above-mentioned execution entity is equipped with an image acquisition device (such as a camera, etc.) and a laser radar, or a communication connection is used to obtain the initial point cloud data and original image in the same scene to execute the target object recognition method provided by this application. This application can be applied to various target object recognition scenarios with high real-time requirements, especially real-time object recognition on low computing power platforms. The embodiment of this application is illustrated by taking the terminal as the execution entity as an example.
[0022] Specifically, for the same target scene, cameras and lidar are used to collect multi-frame data in real time to obtain initial point cloud data and original images.
[0023] It can be understood that due to the different data acquisition frequencies of the camera and the lidar, the acquisition devices are time synchronized. When the camera acquisition frequency is higher than the radar acquisition frequency, each frame of the lidar point cloud frame is matched with the camera image frame that is closest in matching time and whose time error is less than a certain preset threshold based on the time of the image frame and the radar point cloud frame.
[0024] 102. Perform image recognition on the original image to obtain an image recognition result.
[0025] This embodiment takes into account the real-time processing requirements and the computing power limitations of the computing platform. The original image is identified and segmented through image semantic recognition, and the possible locations of the target objects are preliminarily divided. The image recognition results of this embodiment include the object contours corresponding to each object.
[0026] The target scene includes at least one object to be identified, which may be a person, a car, an obstacle, etc. This embodiment does not limit the number and type of objects in the target scene, that is, the image recognition result includes the outline of each object and the label of each object.
[0027] It is understandable that the semantic recognition method of an image can be a feature-based target object recognition method: such as SIFT, SURF, etc., a target object recognition method based on image segmentation: such as threshold segmentation, region growing, edge detection, etc., and other applicable image recognition algorithms can also be used. This embodiment does not make specific restrictions.
[0028] 103. Determine target constraint parameters based on the contour features of each object in the image recognition result. The target constraint parameters at least include a grid scale of the target image.
[0029] In this embodiment, contour features are used to indicate contour characteristics corresponding to each object extracted based on image processing technology. These can be information such as the contour area, main direction, and shape of the contour on the original image. Target constraint parameters are used to indicate parameters for constraining and optimizing the target object recognition process of point cloud data. The target constraint parameters include at least the grid scale of the target image. Through the contour features of each object, this embodiment can set more precise constraint conditions for the target recognition process for each object, namely, target constraint parameters. This allows the subsequent point cloud data processing process to be constrained by the key features of the target object, thereby improving the speed, accuracy, and robustness of recognition and effectively reducing unnecessary data processing.
[0030] Optionally, target constraint parameters are determined based on the contour features of each object in the image recognition results, where the target constraint parameters include at least a grid scale for the target image. This includes determining the grid scale corresponding to each object in the target image based on the contour features corresponding to each object in the image recognition results. In this embodiment, a new grid scale is determined based on the contour features of each object in the image recognition results in the original image to generate the target image, thereby ensuring uniform distribution and data validity of the downsampled point cloud at the current camera observation angle.
[0031] In this embodiment, the grid scale of the target image refers to the grid size of the target image after rasterization processing under the camera angle. The grid scale of the target image is a two-dimensional constraint parameter, and the grid scales corresponding to each object in the target image may be the same or different. That is, the target image of this application can have multiple grids of the same scale or multi-scale grids. The grid scale of this application can effectively constrain the point cloud during both downsampling and point cloud clustering.
[0032] Specifically, the grid scale corresponding to each object in the target image is determined based on the contour features corresponding to each object in the image recognition result, including: determining the contour area corresponding to each object according to the contour corresponding to each object in the image recognition result; and determining the grid scale corresponding to each object in the target image by taking the ratio of the contour area corresponding to each object to the corresponding preset area threshold.
[0033] In this embodiment, the target constraint parameters may include two-dimensional constraint parameters such as grid scale, downsampling distance threshold, maximum point number threshold, and allowed point number range, that is, parameters for constraining downsampling and point cloud clustering of point cloud data from the perspective of two-dimensional images and cameras, and may also include three-dimensional constraint parameters such as direction constraint parameters and angle threshold, that is, parameters for constraining point cloud clustering mainly from the perspective of three-dimensional point clouds.
[0034] Optionally, after determining the grid scale corresponding to each object in the target image based on the contour features corresponding to each object in the image recognition results, the method further includes obtaining default constraint parameters for each object based on the object labels in the image recognition results. This embodiment further determines other constraint parameters based on the object labels, and together with the grid scale, further constrains the point cloud, significantly reducing the processing load of subsequent point cloud clustering and improving processing accuracy.
[0035] Among them, the default value of the constraint parameter is used to indicate the empirical value of the constraint parameter associated with the object label, that is, the parameter for constraining and optimizing each pre-labeled object in the target object recognition process of the point cloud data. For example, if the object label is a person, the default value of the downsampling distance corresponding to the pre-stored "person" label can be directly obtained, which is 0.1 meters, that is, the target constraint parameter also includes the downsampling distance threshold.
[0036] It should be further explained that, in addition to the grid scale, other constraint parameters can be fixed empirical values, that is, obtained by association through object labels, or can be further calculated and determined based on image recognition results and initial point cloud data to improve the adaptability of other constraint parameters to actual scenes. This embodiment does not impose specific restrictions.
[0037] The above-mentioned downsampling distance threshold is used to indicate the minimum distance between any two points in each grid. For example, the minimum distance between any two points in each two-dimensional grid in the target image is mapped to the target image and then two-dimensional angle downsampling is performed through the downsampling distance threshold to reduce the consumed computing resources. The downsampling distance threshold corresponding to each object may be the same or different. Downsampling through the downsampling distance threshold can maintain a certain distance between any two points in each grid, thereby reducing the number of points in each grid and significantly reducing the processing volume of subsequent clustering. In order to increase the downsampling rate, the downsampling distance threshold of this application is mainly through
[0038] The above-mentioned directional constraint parameters are used to indicate the main orientation or posture information of the object. The directional constraint parameters corresponding to each object may be the same or different. By restricting the point cloud clustering direction with the directional constraint parameters, the search space during clustering can be reduced, thereby improving clustering speed. For example, when identifying pedestrians, the main orientation of the point cloud is roughly perpendicular to the ground when the pedestrian is standing. When clustering the point cloud, the directional constraint parameters can focus on the point cloud data in this direction, while constraining the point cloud data in other directions, which can improve the point cloud clustering speed.
[0039] The maximum point count threshold indicates the maximum number of points allowed within each grid. For example, the maximum number of points allowed within each 2D grid in the target image may be the same or different for each object. Downsampling by setting a maximum point count threshold can limit the number of points within each grid, reducing the processing load for subsequent clustering.
[0040] The allowed point range is used to indicate the range of points allowed within each grid. For example, the allowed point range for each 2D grid in the target image may be the same or different for each object. Downsampling by setting the allowed point range provides an allowable range for the number of points within each grid, ensuring that the point cloud data within the grid is neither too dense, resulting in excessive computational complexity, nor too sparse, resulting in loss of important information.
[0041] The above-mentioned angle threshold is used to indicate the maximum angle difference between adjacent points in each grid, for example, the angle difference between two adjacent points compared to a reference direction (such as the depth direction of the acquisition device), or the maximum angle difference formed by the adjacent points on both sides of any point. By setting the angle threshold and only retaining those point pairs whose angle difference is lower than the set threshold, the number of points that need further analysis can be effectively reduced, thereby accelerating the subsequent clustering process.
[0042] It can be understood that the target constraint parameters can be constrained from the perspective of two-dimensional images or from the perspective of three-dimensional point clouds. For example, under appropriate circumstances, the above-mentioned angle threshold can also constrain the point cloud data from a two-dimensional perspective, and the maximum point count threshold and the allowed point count range can also constrain the point cloud data from a three-dimensional perspective. The present application can execute the target object recognition method of the present application through one or more of the above-mentioned target constraint parameters, which can achieve downsampling of point cloud data, reduce processing volume, and improve the speed and accuracy of point cloud clustering.
[0043] 104. Distribute the initial point cloud data to each target grid of the target image according to the preset mapping matrix, image recognition result and target constraint parameters to obtain target point cloud data.
[0044] In practical applications, the image recognition results of the original image may contain multiple object contours. The pixel values outside the contours are set to 0 to generate a target image of the same size as the original image. The point cloud coordinates corresponding to each point in the initial point cloud data are converted to the camera coordinate system based on the mapping matrix to obtain the two-dimensional pixel coordinates corresponding to each point. The initial point cloud data is traversed according to the contours corresponding to each object and the two-dimensional pixel coordinates corresponding to each point in the image recognition results to determine the point cloud data corresponding to each contour and obtain the target point cloud data. In this embodiment, the point cloud data corresponding to each contour area is obtained, and irrelevant point cloud data is eliminated to reduce the computational complexity of subsequent point cloud processing. Downsampling is performed based on the constraints of the grid scale to achieve a uniform distribution of point clouds based on the camera observation angle.
[0045] The above-mentioned target image can be an annotated image, that is, multiple contours are marked with serial numbers starting from 1, and the pixel value within each contour is its corresponding serial number. That is, an annotated image usually refers to an image with specific marks or annotations added to the original image, such as watermarks, text annotations, legends, highlighted areas, etc., in order to highlight certain features or information.
[0046] In this embodiment, the mapping matrix is used to indicate the conversion relationship between the coordinate systems of the camera and the lidar, that is, the coordinate conversion relationship between the camera coordinate system and the point cloud coordinate system. For the successfully matched point cloud frames and image frames, the coordinate conversion relationship between the image coordinate system and the point cloud coordinate system can be established through the corresponding parameter matrix based on the pre-performed camera lidar calibration.
[0047] Optionally, after traversing the initial point cloud data based on the contours corresponding to each object and the two-dimensional pixel coordinates corresponding to each point in the image recognition results to determine the point cloud data corresponding to each contour, the method further includes: downsampling the point cloud data corresponding to each contour according to the default values of each constraint parameter to obtain target point cloud data. Based on the new grid scale, this embodiment further constrains the point cloud according to the default values of the constraint parameters preset by the object label to reduce the processing load of subsequent point cloud clustering. By jointly optimizing the target object recognition results of the point cloud data through the grid scale and other constraint parameters, it is possible to reduce the processing load while ensuring recognition accuracy and achieve real-time recognition.
[0048] Optionally, a fixed empirical value can be set as the downsampling distance threshold, and the point cloud data corresponding to each contour is downsampled according to the default values of each constraint parameter to obtain the target point cloud data, including: downsampling the point cloud data within each grid according to the downsampling distance threshold to obtain the target point cloud data. After implementing point cloud downsampling in this way, the time consumption for subsequent target recognition is greatly reduced. At the same time, because the above-mentioned downsampling process relies on a fixed threshold, for situations where the point cloud density varies due to factors such as distance, the final point cloud density after downsampling is equivalent, and the time consumption is also equivalent.
[0049] 105. Real-time target recognition is performed on the target point cloud data based on the target constraint parameters to obtain the target object recognition result.
[0050] Specifically, the target point cloud data is clustered based on the grid constraints of the target image to obtain clustering results corresponding to each object; and the target object recognition result is output based on the clustering results corresponding to each object. In this embodiment, the point cloud clustering constraint based on the grid constraints can ensure that the point cloud clusters are traversed only within the range corresponding to the object, accelerating the clustering process and avoiding unnecessary traversal calculations.
[0051] In this embodiment, the target object recognition result includes at least one object, and various target object recognition algorithms can be used to perform target recognition on the clustering results. The clustering algorithm in this embodiment can be a hierarchical clustering algorithm, a Dbscan clustering algorithm, or other applicable clustering algorithms. By adjusting the appropriate clustering distance threshold, the target point cloud data can be grouped into different classes or clusters; the target recognition algorithm in this embodiment can be a target object recognition algorithm based on a deep learning algorithm, such as a three-dimensional convolutional neural network, VoteNet, or other applicable target object recognition algorithms, and this embodiment does not impose specific limitations.
[0052] In an embodiment of the present application, target constraint parameters are determined by the contour features of each object in the image recognition results, and constraint parameters can be set for each object in a targeted manner, so that the subsequent point cloud data processing process is constrained by the key features of the target object; by introducing a new grid scale and other constraint parameters of the target image, the calculation time of point cloud processing can be significantly reduced while ensuring recognition accuracy; the grid scale is dynamically updated according to the actual observation conditions of the camera, so that the processing algorithm is insensitive to changes in the order of magnitude of the point cloud, thereby effectively reducing the computational complexity; the optimized method takes into account the advantages of low computational complexity and high speed of image data processing, as well as the advantage of rich three-dimensional spatial information of point cloud data, and can realize real-time spatial range recognition and extraction of multiple targets, meet application scenarios with high real-time requirements, and realize real-time multi-target recognition on low computing power platforms.
[0053] See also Figure 2 Another embodiment of the target object identification method in the embodiment of the present application includes:
[0054] 201. Obtain initial point cloud data and original image.
[0055] 202. Perform image recognition on the original image to obtain an image recognition result.
[0056] Steps 201 and 202 may be performed with reference to steps 101 and 102 and will not be repeated here.
[0057] 203. Determine target constraint parameters based on the image recognition result and the initial point cloud data. The target constraint parameters at least include a grid scale and a downsampling distance threshold corresponding to each object.
[0058] Taking the target constraint parameters as the grid scale and the downsampling distance threshold as an example, specifically, the target area ratio corresponding to each object is determined based on the image recognition result, and the grid scale corresponding to each object in the target image is determined based on the target area ratio corresponding to each object; the point unit volume corresponding to each object is determined based on the image recognition result and the initial point cloud data, and the downsampling distance threshold corresponding to each object is determined according to the point unit volume.
[0059] In this embodiment, the contour feature can be represented by the ratio of the area of the rectangular region of the image where each object is located in the image recognition result to the area threshold corresponding to the object, that is, the grid scale corresponding to each object in the target image is determined according to the target area ratio corresponding to each object, wherein the target area ratio is the ratio of the contour area corresponding to each object to the corresponding preset area threshold.
[0060] The above-mentioned determination of the target area ratio corresponding to each object based on the image recognition result includes: determining the contour area corresponding to each object according to the contour corresponding to each object in the image recognition result; and determining the grid scale corresponding to each object in the target image by the ratio of the contour area corresponding to each object to the corresponding preset area threshold.
[0061] Optionally, determining the grid scale corresponding to each object in the target image based on the target area ratio corresponding to each object includes: determining whether the target area ratio corresponding to each object is greater than 1; if the target area ratio corresponding to any object is greater than 1, determining the target area ratio as the grid scale for any object; and if the target area ratio corresponding to any object is less than or equal to 1, setting the grid scale for any object to 1. This embodiment sets the lower limit of the grid scale to 1, avoids the generation of extremely fine grids due to overly small objects, reduces the consumption of computing resources, improves processing efficiency, and shortens processing time.
[0062] The above-mentioned determination of the point unit volume corresponding to each object based on the image recognition results and the initial point cloud data includes: determining the point cloud space volume corresponding to each object according to each contour and the initial point cloud data; determining the point unit volume corresponding to each object according to the point cloud space volume corresponding to each object and the corresponding preset point count threshold.
[0063] Optionally, determining the downsampling distance threshold corresponding to each object based on the point unit volume includes determining the side length of the point unit volume corresponding to each object as the downsampling distance threshold corresponding to each object. For example, assuming the object is a person, and the point cloud volume corresponding to the person is 1 cubic meter, and the preset upper limit of the number of points is 500, then the point unit volume of each point is 0.002 cubic meters. Calculated as a cube, the side length of the point unit volume corresponding to the person is 0.126 meters. In this case, 0.126 meters is used as the downsampling distance threshold within the grid when identifying people.
[0064] Optionally, the determining of the downsampling distance threshold corresponding to each object based on the point unit volume includes: determining the diagonal length of the point unit volume corresponding to each object as the downsampling distance threshold corresponding to each object.
[0065] 204. Determine candidate point cloud data corresponding to each grid in the target image based on the mapping matrix, the image recognition result, the initial point cloud data, and the grid scale corresponding to each object.
[0066] Based on the mapping matrix, the three-dimensional point cloud coordinates corresponding to each point in the initial point cloud data are converted to the camera coordinate system to obtain the two-dimensional pixel coordinates corresponding to each point; the multiple grids corresponding to each object are determined according to the contours corresponding to each object and the grid scale corresponding to each object in the image recognition results; based on the two-dimensional pixel coordinates corresponding to each point, each point is assigned to the multiple grids corresponding to each object to obtain the candidate point cloud data corresponding to each grid.
[0067] Specifically, the initial pixel coordinates corresponding to each point are traversed, the coordinates of the upper left corner of the rectangle are subtracted, and the result is divided by the corresponding grid scale to obtain the new coordinates under the new scale two-dimensional grid, that is, the two-dimensional pixel coordinates corresponding to each point in the target object, so as to facilitate the calculation of the Euclidean distance when downsampling the downsampling distance threshold of each object.
[0068] It is understandable that the target image may have a multi-scale grid, and the grid scale corresponding to each object in the target image is the grid scale determined in step 103. For example, the grid scale corresponding to object 1 is 502*502, and the grid scale corresponding to object 2 is 1000*1000. In this embodiment, the initial point cloud data is preliminarily processed through the target image to reduce the number of point clouds required for subsequent target recognition and improve accuracy. In addition, the initial point cloud data is processed from a two-dimensional perspective by the target image, which can save computing resources and improve processing speed compared to direct processing from a three-dimensional perspective. Furthermore, the processing of the initial point cloud data is constrained within the grid, greatly reducing the traversal calculations in the process. At the same time, due to the constraints of the two-dimensional grid, a uniform distribution of point clouds based on the camera observation perspective can be achieved. When there are objects to be identified with large differences in size in the target scene, multi-scale grids can improve the downsampling accuracy of the initial point clouds of different objects by downsampling the point clouds according to the grid scales of different objects. Smaller grids are used for small and delicate objects to retain more details, while larger grids are used for large and simple objects to reduce processing time. This can find a better balance between speed and accuracy.
[0069] 205. Downsample the candidate point cloud data corresponding to each grid according to the downsampling distance threshold corresponding to each object to obtain target point cloud data.
[0070] Specifically, one target point is randomly selected from the candidate point cloud data corresponding to each grid and assigned to the target grid; when the grid point does not exist in the target grid, the target point is retained as a grid point; when the grid point already exists in the target grid, it is determined whether the distance between the target point and the grid point is greater than the downsampling distance threshold; if so, the target point is deleted, otherwise the target point is retained as a grid point; the candidate point cloud data corresponding to each grid is traversed to obtain the target point cloud data.
[0071] This embodiment downsamples by using a downsampling distance threshold to maintain a certain distance between any two points within each grid. Repeating the above steps for all newly added points is equivalent to downsampling the point cloud within the grid based on the first Euclidean distance. Each point within the grid is the cluster center of each retained cluster, which can significantly reduce the processing load of subsequent clustering. Moreover, because the Euclidean distance traversal comparison between each point and other points is constrained within the grid, the traversal calculation process is greatly reduced. At the same time, due to the constraints of the two-dimensional grid, a uniform distribution of the point cloud based on the camera observation perspective is achieved.
[0072] Optionally, after the above-mentioned downsampling through the downsampling distance threshold, it also includes: determining whether the number of remaining points in each grid is greater than the maximum point threshold; if so, uniformly downsampling the remaining points in each grid; if not, retaining all the remaining points in each grid so that the remaining points in each grid are less than the maximum point threshold.
[0073] Optionally, after the above-mentioned downsampling through the downsampling distance threshold, it also includes: determining whether the remaining number of points in each grid is within the preset allowable point range; if so, retaining all the remaining points in each grid; if not, adjusting the remaining points in each grid so that the remaining number of points in each grid is within the preset allowable point range.
[0074] If the above is not the case, the remaining points of each grid are adjusted so that the number of remaining points in each grid is within the preset allowable point range: if the remaining points of each grid are too few, all points in the grid are deleted; if the remaining points of each grid are too many, the remaining points in each grid are uniformly downsampled.
[0075] Optionally, after the above-mentioned downsampling by downsampling distance threshold, it also includes: downsampling the remaining points in each grid based on the angle threshold, so that only point pairs with angle differences lower than the set threshold are retained in each grid, which can effectively reduce the number of points that need further analysis and accelerate the subsequent clustering process.
[0076] 206. Cluster the areas to be extracted corresponding to each object in the target point cloud data based on the target constraint parameters and a preset clustering distance threshold to obtain clustering results corresponding to each object. The areas to be extracted are all target grids corresponding to each object.
[0077] Specifically, based on the grid constraint of the target image and a preset clustering distance threshold, the to-be-extracted regions corresponding to the respective objects are clustered to obtain clustering results corresponding to the respective objects.
[0078] Optionally, the target image's grid constraints, directional constraint parameters, and a preset clustering distance threshold are used to cluster the areas to be extracted corresponding to each object, obtaining clustering results corresponding to each object. This embodiment reduces the search space during the clustering process by clustering point clouds based on two-dimensional grid constraints and directional constraint parameters, thereby increasing clustering speed.
[0079] The above cluster distance threshold is used to indicate the maximum distance between any point in each class and the cluster center. Similarly, the cluster distance threshold can also be expressed by Euclidean distance or Manhattan distance, or other methods, for example, by the second Euclidean distance.
[0080] It can be understood that both the first Euclidean distance and the second Euclidean distance indicate the distance between two points, and the first Euclidean distance is used to downsample the points in each grid, while the second Euclidean distance is used to cluster the point clouds in each area to be extracted.
[0081] In this embodiment, "class" is used to indicate a point cloud set in which the distance between each point and the cluster center is within the cluster distance threshold. The clustering results corresponding to each object include at least one class, each class includes multiple target points, and the distance between any target point in each class and the cluster center is less than the cluster distance threshold.
[0082] 207. Output the target object recognition result according to the clustering result of each area to be extracted.
[0083] Specifically, the class with the largest number of points in the clustering results of each area to be extracted is determined as the target class; if the number of points corresponding to the target class is greater than a preset point threshold, the target class is determined as the target object; the clustering results of each area to be extracted are traversed to obtain the target object recognition result.
[0084] To facilitate understanding, an example is provided. Under the downsampling of the first Euclidean distance (such as set to 10 cm), the number of point clouds in the area to be extracted corresponding to each target object is controlled within a smaller range. Under the constraint of the two-dimensional grid, the remaining point clouds are clustered into multiple categories through the second Euclidean distance threshold (such as set to 20 cm), and the category with the largest number of points is selected. If the number of points in this category is greater than a certain point threshold (such as set to 10 points), it can be determined as the final target object point cloud, and then the three-dimensional spatial range of the target object is calculated. The point threshold set in this embodiment can eliminate the situation of misidentification.
[0085] In an embodiment of the present application, target constraint parameters are determined by the contour features of each object in the image recognition results, and constraint parameters can be set for each object in a targeted manner, so that the subsequent point cloud data processing process is constrained by the key features of the target object; by introducing a new grid scale and downsampling distance threshold of the target image, the calculation time of point cloud processing can be significantly reduced while ensuring recognition accuracy; the grid scale and downsampling distance threshold are dynamically updated according to the actual observation conditions of the camera, so that the processing algorithm is insensitive to changes in the order of magnitude of the point cloud, thereby effectively reducing the computational complexity; the optimized method takes into account the advantages of low computational complexity and high speed of image data processing, as well as the advantage of rich three-dimensional spatial information of point cloud data, and can realize real-time spatial range recognition and extraction of multiple targets, meet application scenarios with high real-time requirements, and realize real-time multi-target recognition on low computing power platforms.
[0086] The above describes the target object recognition method in the embodiment of the present application. The following describes the target object recognition device in the embodiment of the present application. Figure 3 In one embodiment of the present application, a target object recognition device includes:
[0087] The acquisition module 301 is used to acquire initial point cloud data and original images.
[0088] The first recognition module 302 is configured to perform image recognition on the original image to obtain an image recognition result.
[0089] A determination module 303 is configured to determine target constraint parameters based on the contour features of each object in the image recognition result, where the target constraint parameters include at least a grid scale of the target image;
[0090] An allocation module 304 is configured to allocate the initial point cloud data to target grids of the target image according to a preset mapping matrix, image recognition results, and target constraint parameters to obtain target point cloud data.
[0091] The second recognition module 305 is used to perform real-time target recognition on the target point cloud data based on the target constraint parameters to obtain a target object recognition result.
[0092] In an embodiment of the present application, target constraint parameters are determined by the contour features of each object in the image recognition results, and constraint parameters can be set for each object in a targeted manner, so that the subsequent point cloud data processing process is constrained by the key features of the target object; by introducing a new grid scale and other constraint parameters of the target image, the calculation time of point cloud processing can be significantly reduced while ensuring recognition accuracy; the grid scale is dynamically updated according to the actual observation conditions of the camera, so that the processing algorithm is insensitive to changes in the order of magnitude of the point cloud, thereby effectively reducing the computational complexity; the optimized method takes into account the advantages of low computational complexity and high speed of image data processing, as well as the advantage of rich three-dimensional spatial information of point cloud data, and can realize real-time spatial range recognition and extraction of multiple targets, meeting application scenarios with high real-time requirements.
[0093] See also Figure 4 Another embodiment of the target object recognition device in the embodiment of the present application includes:
[0094] The acquisition module 301 is used to acquire initial point cloud data and original images.
[0095] The first recognition module 302 is configured to perform image recognition on the original image to obtain an image recognition result.
[0096] A determination module 303 is configured to determine target constraint parameters based on the contour features of each object in the image recognition result, where the target constraint parameters include at least a grid scale of the target image;
[0097] An allocation module 304 is configured to allocate the initial point cloud data to target grids of the target image according to a preset mapping matrix, image recognition results, and target constraint parameters to obtain target point cloud data.
[0098] The second recognition module 305 is used to perform real-time target recognition on the target point cloud data based on the target constraint parameters to obtain a target object recognition result.
[0099] Optionally, the determining module 303 includes:
[0100] A first determining unit 3031 is configured to determine a target area ratio corresponding to each object based on the image recognition result, and determine a grid scale corresponding to each object in the target image based on the target area ratio corresponding to each object;
[0101] The second determining unit 3032 is configured to determine a point unit volume corresponding to each object based on the image recognition result and the initial point cloud data, and determine a downsampling distance threshold corresponding to each object according to the point unit volume.
[0102] Optionally, the first determining unit 3031 is specifically configured to: determine, based on the contours corresponding to the objects in the image recognition result, a ratio of the contour area corresponding to the objects to a preset area threshold, to obtain a target area ratio corresponding to the objects;
[0103] If the target area ratio corresponding to any object is greater than 1, the target area ratio is determined as the grid scale of any object;
[0104] If the target area ratio corresponding to any object is less than or equal to 1, the grid scale of any object is set to 1.
[0105] Optionally, the second determining unit 3032 is specifically configured to: determine the point cloud space volume corresponding to each object according to each contour and the initial point cloud data;
[0106] Determine the point unit volume corresponding to each object based on the point cloud space volume corresponding to each object and the corresponding preset point number threshold;
[0107] The side length of the point unit volume corresponding to each object is determined as the downsampling distance threshold corresponding to each object.
[0108] Optionally, the allocation module 304 includes:
[0109] An allocating unit 3041 is configured to determine candidate point cloud data corresponding to each grid in the target image based on the mapping matrix, the image recognition result, the initial point cloud data, and the grid scale corresponding to each object;
[0110] The downsampling unit 3042 is configured to downsample the candidate point cloud data corresponding to each grid according to the downsampling distance threshold corresponding to each object to obtain target point cloud data.
[0111] Optionally, the downsampling unit 3042 is specifically configured to: select one target point from the candidate point cloud data corresponding to each grid and assign it to the target grid;
[0112] When a grid point does not exist in the target grid, the target point is retained as a grid point;
[0113] When a grid point already exists in the target grid, determine whether the distance between the target point and the grid point is greater than the downsampling distance threshold;
[0114] If yes, delete the target point, otherwise keep the target point as a grid point;
[0115] Traverse the candidate point cloud data corresponding to each grid to obtain the target point cloud data.
[0116] Optionally, the second recognition module 305 is specifically configured to: cluster the to-be-extracted regions corresponding to each object in the target point cloud data based on the target constraint parameter and a preset clustering distance threshold, and obtain clustering results corresponding to each object, where the to-be-extracted regions are all target grids corresponding to each object;
[0117] The target object recognition results are output according to the clustering results corresponding to each object.
[0118] In the embodiment of the present application, the target constraint parameters are determined by the contour features of each object in the image recognition results, and the constraint parameters can be set for each object in a targeted manner, so that the subsequent point cloud data processing process is constrained by the key features of the target object; by introducing a new grid scale and downsampling distance threshold of the target image, the calculation time of the point cloud processing can be significantly reduced while ensuring the recognition accuracy; the grid scale and downsampling distance threshold are dynamically updated according to the actual observation conditions of the camera, so that the processing algorithm is insensitive to changes in the order of magnitude of the point cloud, thereby effectively reducing the computational complexity; the optimized method takes into account the advantages of small computational complexity and fast speed of image data processing, as well as the advantage of rich three-dimensional spatial information of point cloud data, and can realize real-time spatial range recognition and extraction of multiple targets, meeting application scenarios with high real-time requirements.
[0119] See also Figure 5 As shown, the system of integrated laser radar includes an image acquisition device 500, a laser radar 501, a processor 502 and a memory 503. The image acquisition device 500 is used to acquire original images; the laser radar 501 is used to acquire initial point cloud data. The memory 503 stores machine executable instructions that can be executed by the processor 502. The processor 502 executes the machine executable instructions to implement the above-mentioned target object recognition method.
[0120] Further, Figure 5 The illustrated integrated lidar system further includes a bus 504 and a communication interface 505 , and the processor 502 , the communication interface 505 and the memory 503 are connected via the bus 504 .
[0121] Among them, the memory 503 may include a high-speed random access memory (RAM), and may also include a non-volatile memory (non-volatile memory), for example, at least one disk storage. The communication connection between the system network element and at least one other network element is realized through at least one communication interface 505 (which can be wired or wireless), and the Internet, wide area network, local area network, metropolitan area network, etc. can be used. The bus 504 can be an ISA bus, a PCI bus, or an EISA bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 5Only one bidirectional arrow is used in the diagram, but this does not mean that there is only one bus or one type of bus.
[0122] The processor 502 may be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method can be completed by hardware integrated logic circuits or software instructions in the processor 502. The above processor 502 may be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it may also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present disclosure. The general-purpose processor may be a microprocessor or any conventional processor. The steps of the method disclosed in conjunction with the embodiments of the present disclosure can be directly implemented and executed by a hardware decoding processor, or by a combination of hardware and software modules in the decoding processor. The software module can be located in a storage medium mature in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, etc. The storage medium is located in the memory 503 , and the processor 502 reads the information in the memory 503 and completes the method steps of the above embodiment in combination with its hardware.
[0123] The present application also provides a computer-readable storage medium, which may be a non-volatile computer-readable storage medium or a volatile computer-readable storage medium. The computer-readable storage medium stores instructions that, when executed on a computer, cause the computer to execute the steps of the target object recognition method.
[0124] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0125] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0126] As described above, the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A target object recognition method, characterized in that: The target object identification method includes: Obtain initial point cloud data and original images; Performing image recognition on the original image to obtain an image recognition result; Determining target constraint parameters based on the contour features of each object in the image recognition result, wherein the target constraint parameters at least include a grid scale of the target image; Allocating the initial point cloud data to each target grid of the target image according to a preset mapping matrix, the image recognition result, and the target constraint parameter to obtain target point cloud data; Performing real-time target recognition on the target point cloud data based on the target constraint parameters to obtain a target object recognition result; Determining target constraint parameters based on the contour features of each object in the image recognition result includes: Determining a target area ratio corresponding to each object based on the image recognition result, and determining a grid scale corresponding to each object in the target image based on the target area ratio corresponding to each object; Determining the target area ratio corresponding to each object based on the image recognition result, and determining the grid scale corresponding to each object in the target image based on the target area ratio corresponding to each object, includes: Determine the ratio of the contour area corresponding to each object to a preset area threshold value according to the contour corresponding to each object in the image recognition result, and obtain the target area ratio corresponding to each object; If the target area ratio corresponding to any one of the objects is greater than 1, the target area ratio is determined as the grid scale of the any one of the objects; If the target area ratio corresponding to any object is less than or equal to 1, the grid scale of the any object is set to 1.
2. The target object recognition method according to claim 1, characterized in that: The target constraint parameters also include a downsampling distance threshold corresponding to each object; The determining of target constraint parameters based on the contour features of each object in the image recognition result further includes: A point unit volume corresponding to each object is determined based on the image recognition result and the initial point cloud data, and a downsampling distance threshold corresponding to each object is determined according to the point unit volume.
3. The target object recognition method according to claim 2, characterized in that: The determining of the point unit volume corresponding to each object based on the image recognition result and the initial point cloud data, and determining the downsampling distance threshold corresponding to each object according to the point unit volume, includes: Determine the point cloud space volume corresponding to each object according to each of the outlines and the initial point cloud data; Determine the point unit volume corresponding to each object based on the point cloud space volume corresponding to each object and the corresponding preset point number threshold; The side length of the point unit volume corresponding to each object is determined as the downsampling distance threshold corresponding to each object.
4. The target object recognition method according to claim 1, characterized in that: The target constraint parameters also include: at least one of a maximum point threshold, an allowed point range, a direction constraint parameter, and an angle threshold.
5. The target object recognition method according to claim 1, characterized in that: The target constraint parameters include the grid scale and downsampling distance threshold corresponding to each object; The step of allocating the initial point cloud data to each target grid of the target image according to a preset mapping matrix, the image recognition result, and the target constraint parameter to obtain target point cloud data includes: Determine candidate point cloud data corresponding to each grid in the target image based on the mapping matrix, the image recognition result, the initial point cloud data, and the grid scale corresponding to each object; The candidate point cloud data corresponding to each grid is downsampled according to the downsampling distance threshold corresponding to each object to obtain the target point cloud data.
6. The target object recognition method according to claim 5, characterized in that: The step of downsampling the candidate point cloud data corresponding to each grid according to the downsampling distance threshold corresponding to each object to obtain target point cloud data includes: Select a target point from the candidate point cloud data corresponding to each grid and assign it to the target grid; When the grid point does not exist in the target grid, retaining the target point as a grid point; When a grid point already exists in the target grid, determining whether the distance between the target point and the grid point is greater than the downsampling distance threshold; If yes, delete the target point; otherwise, keep the target point as a grid point; Traverse the candidate point cloud data corresponding to each grid to obtain the target point cloud data.
7. The target object recognition method according to any one of claims 1 to 5, characterized in that: The performing real-time target recognition on the target point cloud data based on the target constraint parameters to obtain a target object recognition result includes: Clustering the to-be-extracted regions corresponding to each object in the target point cloud data based on the target constraint parameter and a preset clustering distance threshold to obtain clustering results corresponding to each object, wherein the to-be-extracted regions are all target grids corresponding to each object; The target object recognition results are output according to the clustering results corresponding to each object.
8. A target object recognition device, characterized in that: The target object recognition device includes: Acquisition module, used to obtain initial point cloud data and original images; A first recognition module is used to perform image recognition on the original image to obtain an image recognition result; a determination module, configured to determine target constraint parameters based on the contour features of each object in the image recognition result, wherein the target constraint parameters include at least a grid scale of the target image; an allocating module, configured to allocate the initial point cloud data to each target grid of the target image according to a preset mapping matrix, the image recognition result, and the target constraint parameter, to obtain target point cloud data; A second recognition module is used to perform real-time target recognition on the target point cloud data based on the target constraint parameters to obtain a target object recognition result; The determination module includes: a first determining unit, configured to determine a target area ratio corresponding to each object based on the image recognition result, and determine a grid scale corresponding to each object in the target image based on the target area ratio corresponding to each object; The first determining unit is specifically configured to: determine, based on the contours corresponding to the objects in the image recognition result, a ratio of the contour area corresponding to each object to a preset area threshold, to obtain a target area ratio corresponding to each object; If the target area ratio corresponding to any one of the objects is greater than 1, the target area ratio is determined as the grid scale of the any one of the objects; If the target area ratio corresponding to any object is less than or equal to 1, the grid scale of the any object is set to 1.
9. A system integrating laser radar, characterized in that: The system integrated with the laser radar includes: an image acquisition device, a laser radar, a memory and at least one processor, wherein the memory stores instructions; The image acquisition device is used to acquire original images; The laser radar is used to collect initial point cloud data; The at least one processor calls the instructions in the memory so that the system integrated with the lidar executes the target object recognition method according to any one of claims 1 to 7.
10. A computer-readable storage medium having instructions stored thereon, characterized in that: When the instruction is read and executed, the target object recognition method according to any one of claims 1 to 7 is executed.
Citation Information
Patent Citations
Non-uniform point cloud surface domain segmentation method based on improved region growing technology
CN113902688A
Down-sampling method, device and equipment suitable for point cloud positioning and storage medium
CN117994746A