Target screening method and system based on point cloud processing algorithm in port driving and medium

Through the target screening method based on point cloud data, the target box semantic graph and overlap threshold are used to screen the target box in port autonomous driving, which solves the problem of low recognition accuracy of two-dimensional NMS in complex scenarios, and achieves efficient and accurate target recognition.

CN120472198AActive Publication Date: 2025-08-12QINGDAO PORT INT CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510426133.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-07
Publication Date
2025-08-12
Estimated Expiration
2045-04-07

AI Technical Summary

Technical Problem

In port autonomous driving, the existing two-dimensional NMS method has a wide variety of targets, large size differences and significant confidence distribution differences, resulting in low recognition accuracy.

Method used

By determining the target box to be processed based on point cloud data and projecting it into the target box semantic map, determining whether the target box needs to be removed based on the semantic value in the pixel area, accurately filtering it with the target category and overlap threshold, forming an effective target point cloud collection and extracting feature information.

Benefits of technology

It improves the accuracy and efficiency of target box removal, reduces computing complexity, optimizes the performance of point cloud processing, is suitable for complex multi-category scenarios, and improves recognition accuracy and adaptability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120472198A_ABST
    Figure CN120472198A_ABST
Patent Text Reader

Abstract

The invention provides a target screening method and system based on a point cloud processing algorithm in port driving and a medium, and the method comprises the steps: determining a to-be-processed target frame based on point cloud data; projecting the to-be-processed target frame to the target frame semantic map, and determining a pixel region of the to-be-processed target frame in the target frame semantic map; based on the semantic values of the pixel points in the pixel region, judging whether the to-be-processed target frame needs to be removed or not; and processing the point cloud data corresponding to the to-be-processed target frame based on the judgment result. The method can be efficient, fast and suitable for complex multi-class scenes, comprehensively considers multiple dimensions such as calculation complexity, multi-class processing and target priority processing, is high in expansibility, and is simple and practical in engineering.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of data processing technology, and in particular relates to a target screening method, system and medium based on a point cloud processing algorithm in port driving. Background Art

[0002] Point cloud data can be used for obstacle identification and target detection. In fields such as autonomous driving and robotics, a core issue is how to perceive surrounding objects. In related technologies, the collected point cloud data can be projected onto a top-down view, and the top-down view frame can be obtained using two-dimensional (2D) detection technology. However, the original information of the point cloud is lost during quantization, and it is difficult to detect obscured objects when detecting from a 2D image.

[0003] Non-Maximum Suppression (NMS) is a crucial post-processing technique in object detection. It is primarily used to select the bounding box that best matches the target's true location from multiple overlapping candidate bounding boxes output by the model. Its core idea is to suppress redundant detection results and retain the candidate boxes with the highest confidence that best match the true target, thereby improving detection accuracy and efficiency.

[0004] In object detection tasks, the model typically generates a large number of candidate bounding boxes. These boxes may cover different regions of the same object or contain parts of multiple objects. Using these candidate boxes directly without filtering them can result in a large number of duplicate or incorrect bounding boxes in the detection results, impacting detection performance. NMS calculates the overlap between candidate boxes (typically using the Intersection over Union (IoU) metric) and confidence scores, gradually eliminating redundant boxes and ultimately retaining the optimal detection results.

[0005] Two-dimensional object detection in computer vision is currently the primary application of NMS. However, when it comes to three-dimensional object detection in scenarios such as autonomous port operations or from a BEV perspective, two-dimensional NMS methods may suffer from performance disadvantages. For example, in scenarios such as autonomous port operations, objects are often independent entities, and the objects to be detected are diverse, have large size variations, and have significantly different confidence distributions, resulting in low recognition accuracy. Summary of the Invention

[0006] In response to the problems in the prior art, the present invention provides a target screening method, system and medium based on a point cloud processing algorithm in port driving, which solves the problem of low recognition accuracy caused by the two-dimensional NMS method in the prior art when the targets to be detected are of various types, have large size differences, and have significant differences in confidence distribution.

[0007] The technical solution adopted in the present invention is as follows: In a first aspect, the present application provides a target screening method based on a point cloud processing algorithm in port driving, comprising: Acquire point cloud data in port environments; Based on the point cloud data, determine the target frame to be processed; Projecting the target frame to be processed onto a target frame semantic map, and determining a pixel area of the target frame to be processed in the target frame semantic map; Based on the semantic values of the pixels in the pixel area, it is determined whether the target frame to be processed needs to be eliminated; Based on the judgment result, processing the point cloud data corresponding to the target frame to be processed; Perform cluster analysis on the processed point cloud data to form a valid target point cloud set; Extract features from the point cloud in the valid target point cloud set to generate target feature information; Based on the target feature information, the category of the valid target is determined and the valid target after the category is determined is output.

[0008] Preferably, judging whether the target frame to be processed needs to be eliminated based on the semantic values of the pixels in the pixel area includes: Determine whether there is a non-zero pixel point in the pixel area; In response to the absence of non-zero pixels in the pixel area, determining that the target frame to be processed does not need to be eliminated; In response to the presence of non-zero pixels in the pixel area, determining an overlapping target frame that overlaps with the target frame to be processed based on the non-zero pixels; At least based on the target categories corresponding to the target frame to be processed and the overlapping target frame, it is determined whether the target frame to be processed needs to be eliminated.

[0009] Preferably, if it is determined that the target frame to be processed does not need to be eliminated, the semantic values of the pixels in the pixel area are assigned as identification values corresponding to the target frame to be processed, and the projection and judgment operations are performed on the next target frame to be processed.

[0010] Preferably, judging whether the target frame to be processed needs to be removed based at least on the target categories corresponding to the target frame to be processed and the overlapping target frame respectively includes: Determine the intersection area and the combined area of the target frame to be processed and the overlapping target frame; determining a degree of overlap based on the intersection area and the combined area; Determining an overlap threshold based on the target categories corresponding to the target frame to be processed and the overlapping target frame respectively; Based on the overlap degree and the overlap degree threshold, it is determined whether the target frame to be processed needs to be eliminated.

[0011] Preferably, determining the overlap threshold based on the target categories corresponding to the target frame to be processed and the overlapping target frame includes: If the target categories corresponding to the target frame to be processed and the overlapping target frame are the same, setting the overlap threshold to be less than the reference threshold; If the target categories corresponding to the target frame to be processed and the overlapping target frame are different, the overlap threshold is determined based on the proximity degree of the target categories corresponding to the target frame to be processed and the overlapping target frame.

[0012] Preferably, judging whether the target frame to be processed needs to be eliminated based on the overlap degree and the overlap degree threshold includes: In response to the overlap being less than the overlap threshold, determining that the target frame to be processed does not need to be eliminated; In response to the overlap being not less than the overlap threshold, it is determined that the target frame to be processed or the overlapping target frame needs to be removed.

[0013] Preferably, if it is determined that the target frame to be processed or the overlapping target frame needs to be removed, a first confidence level of the target frame to be processed and a second confidence level of the overlapping target frame are determined based on the target categories corresponding to the target frame to be processed and the overlapping target frame respectively; Based on the first confidence level and the second confidence level, a target frame to be eliminated is determined from the target frame to be processed and the overlapping target frame.

[0014] Preferably, the method further comprises: In response to the target frame to be removed being the overlapping target frame, obtaining a coverage range of the overlapping target frame on the target frame semantic graph; The semantic value of the pixel points within the coverage area is assigned to 0.

[0015] In a second aspect, the present application provides a point cloud processing system for implementing a target screening method based on a point cloud processing algorithm in port driving as described in the first aspect, comprising: A first determination module is configured to determine a target frame to be processed based on the point cloud data; A second determining module is configured to project the target frame to be processed onto a target frame semantic map, and determine a pixel area of the target frame to be processed in the target frame semantic map; A judgment module is configured to judge whether the target frame to be processed needs to be eliminated based on the semantic values of the pixels in the pixel area; The processing module is configured to process the point cloud data corresponding to the target box to be processed based on the judgment result.

[0016] In a third aspect, the present application provides a computer-readable storage medium storing computer instructions. When the computer reads the computer instructions in the storage medium, the computer executes a target screening method based on a point cloud processing algorithm in port driving as described in the first aspect.

[0017] It can be seen from the above technical solutions that this application has the following advantages: 1. By determining the target frame to be processed based on point cloud data and projecting it into the target frame semantic map, the pixel area of the target frame in the semantic map can be quickly and intuitively determined. Then, based on the semantic values of the pixels in this area, the target frame is intelligently judged whether it needs to be eliminated. This effectively improves the accuracy and efficiency of target frame elimination, avoids the subsequent processing of invalid data, thereby reducing the overall computational complexity and optimizing the performance of point cloud processing.

[0018] 2. By determining whether there are non-zero pixels in the pixel area, the initial target frame removal screening is quickly completed. When there are no non-zero pixels in the area, it can be quickly determined that there is no need to remove the target frame, reducing redundant calculations. When non-zero pixels exist, the overlapping target frame is further determined and accurately judged based on the target category, thereby improving the intelligence and accuracy of the target frame removal process.

[0019] 3. When it's clear that the target frame doesn't need to be removed, the semantic value of the pixel within the pixel area is assigned to the corresponding identification value of the target frame, avoiding repeated calculations and misjudgments of subsequent target frames. This approach not only ensures the accuracy of target frame recognition, but also greatly improves the continuity and overall efficiency of the processing process.

[0020] 4. By dynamically determining the degree of overlap based on the intersection and combined areas of the target frame and overlapping target frames, and setting different overlap thresholds based on target category, we can more accurately determine whether a target frame should be removed. This method effectively avoids misjudgments or missed detections due to overlapping target frames, significantly improving the accuracy and reliability of point cloud data processing results.

[0021] 5. By setting different overlap thresholds based on the target box category, the overlap threshold is set lower for target boxes of the same category, effectively avoiding redundant recognition of the same category; while for target boxes of different categories, the threshold is dynamically determined based on the proximity between the categories. This differentiated approach to target category processing significantly improves the accuracy and adaptability of the method in complex scenarios, enhancing its application effectiveness.

[0022] 9. Through the reasonable coordination between the first determination module, the second determination module, the judgment module and the processing module, the functions and advantages of the target screening method based on the point cloud processing algorithm in the above-mentioned port driving are realized. The modular design not only ensures the stability and efficiency of the system implementation, but also effectively improves the application value and scalability of the method, and significantly reduces the computational burden and manual intervention in the data processing process. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] In order to more clearly illustrate the technical solution of the present application, the following is a brief introduction to the drawings required for the description. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0024] Figure 1 is an exemplary module diagram of a point cloud processing system according to some embodiments; Figure 2 An exemplary flow chart of a target screening method based on a point cloud processing algorithm in port driving according to some embodiments; Figure 3 This is an exemplary flowchart of determining whether a target frame to be processed needs to be eliminated, as shown in some embodiments; Figure 4 This is an exemplary flowchart of some embodiments for determining whether a target frame to be processed needs to be eliminated by using overlap and overlap thresholds. DETAILED DESCRIPTION

[0025] In order to make the application objectives, features, and advantages of this application more obvious and easy to understand, the technical solutions protected by this application will be clearly and completely described below using specific embodiments and drawings. Obviously, the embodiments described below are only part of the embodiments of this application, not all of them. Based on the embodiments in this patent, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this patent.

[0026] Existing NMS methods perform well in 2D image tasks, but face challenges in 3D or BEV-view tasks. For example, in 3D and BEV-view tasks, due to changes in object occlusion relationships and missing semantic information, the increased class differences and reduced box overlap between objects make traditional NMS methods difficult to apply. For another example, in scenarios like autonomous port operations, large mechanical equipment such as gantry cranes, quay cranes, and forklifts are difficult to detect, with a wide distribution of confidence levels and significant class differences. A low-confidence category A may be more reliable than a high-confidence category B, and the two may have no overlapping target boxes, making traditional NMS methods ineffective. Furthermore, traditional NMS requires sorting and pairwise comparison of target boxes, resulting in high time complexity, large data processing volumes, and low data processing efficiency. Although some improved NMS methods can address multi-category and size differences, they also suffer from high computational complexity and unsuitability for real-time scenarios. Therefore, this application provides a target screening method and system based on point cloud processing algorithm in port driving, which can be efficient, fast and applicable to complex multi-category scenarios. It comprehensively considers multiple dimensions such as computational complexity, multi-category processing and target priority processing, has strong scalability, simple engineering implementation and is practical.

[0027] In some embodiments, as Figure 1 As shown, the point cloud processing system 100 may include a first determination module 110 , a second determination module 120 , a judgment module 130 and a processing module 140 .

[0028] The first determining module 110 is a module for acquiring point cloud data in a port environment and determining a target frame to be processed. In some embodiments, the first determining module 110 can be configured to determine a target frame to be processed based on the point cloud data.

[0029] The second determination module 120 is a module for determining the pixel area of the target frame to be processed. In some embodiments, the second determination module 120 can be configured to project the target frame to be processed onto the target frame semantic map and determine the pixel area of the target frame to be processed in the target frame semantic map.

[0030] The judgment module 130 is a module for judging whether the target frame to be processed needs to be eliminated. In some embodiments, the judgment module 130 can be configured to judge whether the target frame to be processed needs to be eliminated based on the semantic values of the pixels in the pixel area.

[0031] In some embodiments, the judgment module 130 can also be configured to: determine whether there are non-zero pixels in the pixel area; in response to the absence of non-zero pixels in the pixel area, determine that the target frame to be processed does not need to be eliminated; in response to the presence of non-zero pixels in the pixel area, determine an overlapping target frame that overlaps with the target frame to be processed based on the non-zero pixels; at least based on the target categories corresponding to the target frame to be processed and the overlapping target frame, determine whether the target frame to be processed needs to be eliminated.

[0032] In some embodiments, the judgment module 130 can also be configured to: if it is determined that the target frame to be processed does not need to be eliminated, assign the semantic value of the pixel point in the pixel area to the identification value corresponding to the target frame to be processed, and perform projection and judgment operations on the next target frame to be processed.

[0033] In some embodiments, the judgment module 130 can also be configured as follows: if the intersection area and the merged area of the target frame to be processed and the overlapping target frame are determined; based on the intersection area and the merged area, the degree of overlap is determined; based on the target categories corresponding to the target frame to be processed and the overlapping target frame respectively, the overlap threshold is determined; based on the overlap and the overlap threshold, whether the target frame to be processed needs to be eliminated.

[0034] In some embodiments, the judgment module 130 can also be configured as follows: if the target categories corresponding to the target frame to be processed and the overlapping target frame are the same, the overlap threshold is set to be less than the baseline threshold; if the target categories corresponding to the target frame to be processed and the overlapping target frame are different, the overlap threshold is determined based on the degree of proximity of the target categories corresponding to the target frame to be processed and the overlapping target frame.

[0035] In some embodiments, the judgment module 130 can also be configured to: in response to the overlap being less than the overlap threshold, determine that the target frame to be processed does not need to be eliminated; in response to the overlap being not less than the overlap threshold, determine that the target frame to be processed or the overlapping target frame needs to be eliminated.

[0036] In some embodiments, the judgment module 130 can also be configured as follows: if it is determined that the target frame to be processed or the overlapping target frame needs to be removed, based on the target categories corresponding to the target frame to be processed and the overlapping target frame respectively, determine the first confidence level of the target frame to be processed and the second confidence level of the overlapping target frame; based on the first confidence level and the second confidence level, determine the target frame to be removed from the target frame to be processed and the overlapping target frame.

[0037] In some embodiments, the judgment module 130 can also be configured to: in response to the target frame being removed as an overlapping target frame, obtain the coverage of the overlapping target frame on the target frame semantic map; and assign the semantic value of the pixel points within the coverage range to 0.

[0038] The processing module 140 is a module for processing point cloud data. In some embodiments, the processing module 140 can be configured to process the point cloud data corresponding to the target frame to be processed based on the judgment result.

[0039] For more information on the above, see Figures 2 to 4 Related description.

[0040] It should be understood that Figure 1 The system and its modules shown can be implemented in various ways. It should be noted that the above description of the point cloud processing system 100 and its modules is for convenience only and does not limit this specification to the scope of the embodiments. It is understandable that those skilled in the art, after understanding the principles of the system, may arbitrarily combine the modules or form subsystems connected with other modules without deviating from the principles. In some embodiments, Figure 1 The first determination module 110, second determination module 120, judgment module 130, and processing module 140 disclosed herein may be different modules within a system, or a single module may implement the functions of two or more of the aforementioned modules. For example, the modules may share a storage module, or each module may have its own storage module. Such variations are within the scope of protection of this specification.

[0041] Figure 2 This is an exemplary flow chart of a target screening method based on a point cloud processing algorithm in port driving according to some embodiments of this specification. Figure 2 As shown, the process 200 includes the following steps. In some embodiments, the process 200 may be executed by the point cloud processing system 100 .

[0042] Step 210 : Determine a target frame to be processed based on the point cloud data. In some embodiments, the first determination module 110 may execute step 210 .

[0043] For more information about the first determination module 110, please refer to Figure 1 Related description.

[0044] Point cloud data refers to a collection of multiple point data. Point data refers to data related to a point cloud point. In some embodiments, multiple point cloud points can be combined to form a point cloud image, and a point cloud point is the smallest unit that is combined to form a point cloud image.

[0045] In some embodiments, the point cloud data may be acquired by a detector, which may include at least one of a radar detector, a radar camera, a sensor, a laser scanner, and the like.

[0046] In some embodiments, the first determination module 110 may include a detector. In some embodiments, the first determination module 110 may be in communication with the detector. The detector may upload the point cloud data to the first determination module 110.

[0047] In some embodiments, the point cloud image corresponding to the point cloud data may include multiple frames of point cloud images, and each frame of the point cloud image may include at least part of the point data in the point cloud data.

[0048] In some embodiments, the point cloud data may include data collected by the detector in a unit time. For example, the unit time may include 100ms, 200ms, etc. In some embodiments, the detector may be mounted on a vehicle and move synchronously with the vehicle.

[0049] In some embodiments, the point data may include at least one of position information, color information, target category, intensity information, time information, etc. of the point cloud point.

[0050] Position information refers to information related to the position of a point cloud point in a coordinate system. For example, the coordinates of a point cloud point. In some embodiments, the coordinate system may include at least one of a two-dimensional coordinate system and a three-dimensional coordinate system. In some embodiments, the coordinate system may include a radar coordinate system, a vehicle coordinate system, a bird's-eye view (BEV) space coordinate system (referred to as the BEV image coordinate system), etc. As an example only, a vehicle coordinate system is a coordinate system with the center point of the vehicle as the origin, the vehicle's travel direction as the x-axis, the direction perpendicular to the vehicle's travel direction as the y-axis, and the direction perpendicular to the xy plane as the z-axis.

[0051] Color information refers to information related to the color of a point, such as red, blue, or the like. In some embodiments, the color information can be represented by RGB values, such as 0-255.

[0052] Target categories refer to information related to the categories of target objects when classifying multiple target objects. Target objects refer to objects that need to be detected and identified in point cloud data. For example, target objects may include at least one of a forklift, a container truck trailer, a pedestrian, or a vehicle in a port. In some embodiments, the target categories of target objects may include at least one of a stationary target and a non-stationary target. Stationary targets may include at least one of terrain, vegetation, and buildings. Non-stationary targets may include at least one of pedestrians and vehicles. For more information on target categories, please refer to the relevant description of step 332.

[0053] The intensity information is information related to the intensity of the point. For example, the intensity information may include at least one of a pixel value, a brightness value, etc. of the point.

[0054] Time information refers to information related to the time of point data, such as the time when the point data was acquired.

[0055] The pending target frame refers to the target frame waiting to be processed.

[0056] In some embodiments, the first determination module 110 can determine the target frame to be processed in various ways, such as by obtaining the target frame based on a first preset algorithm. In some embodiments, the first preset algorithm can include a three-dimensional target detection algorithm.

[0057] A 3D object detection algorithm refers to an algorithm that can determine the target box to be processed, for example, at least one of the PointPillars algorithm and the CenterPoint algorithm.

[0058] In some embodiments, the input of the three-dimensional object detection algorithm may include point cloud data, and the output may include a target frame to be processed and target frame information of the target frame to be processed.

[0059] The target frame information refers to information related to the target frame to be processed.

[0060] In some embodiments, the target frame information may include at least one of the center point, size, heading, target category, confidence level, etc. of the target frame to be processed.

[0061] The center point refers to the geometric center of the target frame to be processed. In some embodiments, the center point can be represented based on the position information of the center point, for example, the coordinates of the center point in the vehicle body coordinate system.

[0062] Size refers to information related to the size of the target frame to be processed. For example, when the target frame to be processed is a rectangular frame, the size may include the length and width of the target frame to be processed.

[0063] Heading refers to information related to the direction of travel of a vehicle. In some embodiments, heading can be represented by a heading angle. The heading angle is the angular difference between the vehicle's forward direction and the vehicle's longitudinal axis.

[0064] The target category of the target frame to be processed refers to the type or category of the target object in the target frame to be processed. For example, the target category can be at least one of a vehicle, a pedestrian, and cargo.

[0065] Confidence refers to the degree of certainty of the target box to be processed. Confidence can be expressed as a probability value.

[0066] In some embodiments, the first determination module 110 may classify the plurality of target frames to be processed based on the target categories of the target frames to be processed. The first determination module 110 may sort the plurality of target frames to be processed in each target category based on confidence, and subsequently process the target frames to be processed in sequence according to the sorting order.

[0067] In some implementations of this specification, target frame information is used to facilitate subsequent acquisition of relevant information of target frames such as the target frame to be processed, thereby facilitating subsequent determination of whether the target frame needs to be processed (such as removed, retained, etc.).

[0068] In some embodiments, the target frame to be processed determined above may include multiple redundant frames.

[0069] A redundant frame refers to a target frame that does not meet the requirements. For example, the target frame information of the redundant frame does not meet the requirements. In some embodiments, the first determination module 110 can perform a preliminary screening on the target frames to be processed determined above to determine the target frames to be processed after removing the redundant frames.

[0070] In some embodiments, the first determining module 110 may filter target frames to be processed that meet preset filtering conditions to achieve the effect of removing redundant frames.

[0071] The preset filtering condition refers to a preset condition for removing redundant frames from the target frame to be processed. In some embodiments, the preset filtering condition can be related to a confidence level. For example, the preset filtering condition can be to remove target frames to be processed whose confidence level is less than a confidence threshold.

[0072] In some embodiments, the first determination module 110 may obtain the confidence threshold in various ways, such as obtaining at least one of manual input and obtaining from historical data.

[0073] In some embodiments, the confidence threshold may be related to the target category. The confidence thresholds corresponding to different target categories may be the same or different. For example, the confidence threshold for a non-stationary target may be lower than the confidence threshold for a stationary target.

[0074] In some embodiments, the first determination module 110 may construct a first preset table based on historical target categories and historical confidence thresholds in the historical data. The first preset table may include historical target categories, historical confidence thresholds, and a correspondence between historical target categories and historical confidence thresholds. In some embodiments, the first determination module 110 may query the first preset table based on the current target category to determine the same or similar historical target categories, and use the historical confidence thresholds corresponding to the historical target categories as the confidence threshold corresponding to the current target category.

[0075] In some embodiments, the first determining module 110 may obtain the preset screening condition in various ways, for example, by obtaining at least one of manual input and historical data.

[0076] Processing each point cloud frame can generate tens of thousands, or even hundreds of thousands, of target frames, including a large number of redundant, low-confidence frames. Using NMS to process all target frames is computationally intensive and time-consuming. By identifying candidate target frames and then determining the target frames to be processed from them, we can perform a preliminary screening of the candidate frames, reducing the number of target frames to be processed and the subsequent data processing workload.

[0077] Using a confidence threshold as a screening criterion, candidate object boxes with confidence levels below the threshold can be removed. Depending on different situations or needs, such as different application scenarios, different network performance, and different threshold selection, the number of candidate object boxes can be reduced to tens, hundreds, or even thousands, thereby reducing the subsequent data processing workload and improving data processing efficiency and accuracy.

[0078] In step 220 , the target frame to be processed is projected onto the target frame semantic map, and a pixel area of the target frame to be processed in the target frame semantic map is determined. In some embodiments, the second determining module 120 may perform step 220 .

[0079] For more information about the second determination module 120, please refer to Figure 1 Related description.

[0080] The target frame semantic map refers to a map used to assign semantic labels to pixels.

[0081] In some embodiments, the initial target frame semantic image may be a blank image including a plurality of pixels with a pixel value of 0.

[0082] In some embodiments, the target frame semantic map can be an image of an image coordinate system from a vehicle vehicle perspective. In some embodiments, the coordinate origin of the image coordinate system from a vehicle perspective can be the same as the coordinate origin of the vehicle body coordinate system. For example, the coordinate origin can be the center point of the vehicle.

[0083] In some embodiments, the resolution of the target frame semantic map can be set to R. That is, the resolution of each pixel in the target frame semantic map in the image coordinate system under the BEV perspective is R meters × R meters. In some embodiments, the resolution is a preset value. The second determination module 120 can determine the resolution in various ways, such as by obtaining at least one of manual input and historical data.

[0084] In some embodiments, the second determination module 120 may set a range to be investigated from the perspective of the BEV. The range to be investigated can be determined in various ways, such as by the detection range of a detector that acquires point cloud data. For example, using a radar detector as an example, the forward range to be investigated FR is 100 meters, the rearward range to be investigated BR, the leftward range to be investigated LR, and the rightward range to be investigated RR are each 50 meters. FR, BR, LR, and RR refer to the distances forward, backward, left, and right, respectively, with the center of the vehicle as the coordinate origin and the vehicle's direction of travel or head orientation as the forward direction.

[0085] A semantic identifier is an identifier assigned to a target frame to be processed. The semantic identifier can be used to distinguish information about different target frames to be processed. For example, if the number of target frames to be processed is i, the second determination module 120 may use i as the semantic identifier of the i-th target frame to be processed.

[0086] In some embodiments, the semantic identifier may be related to the target category corresponding to the target frame to be processed. Different target categories may correspond to different semantic identifiers.

[0087] A pixel region refers to the area occupied by at least one pixel in the semantic map of the target frame in the image coordinate system under the BEV perspective. In some embodiments, the pixel region may include the area occupied by the corresponding pixel in the semantic map of the target frame in the image coordinate system under the BEV perspective after the target frame to be processed is projected.

[0088] In some embodiments, the second determining module 120 may project the target frame to be processed into the target frame semantic map, and use the area corresponding to the projection as the pixel area corresponding to the target frame to be processed.

[0089] Projection refers to converting the coordinates of the target frame to be processed in the vehicle coordinate system into the coordinates of the target frame semantic map in the image coordinate system under the BEV perspective. For example, the spatial coordinates of any point in the vehicle coordinate system are converted to , converted to the image coordinates of the target box semantic map in the image coordinate system under the BEV perspective .

[0090] In some embodiments, the second determining module 120 may be based on the spatial coordinates , calculate the image coordinates through the second preset algorithm The second preset algorithm may include the formula:

[0091] Where R is the resolution of the target frame semantic map. In some embodiments, the second determination module 120 can determine the resolution in a variety of ways. For example, at least one of obtaining manual input, obtaining from historical data, and obtaining based on experience can be used. For more information about LR and BR, please refer to the relevant description above.

[0092] In some embodiments, taking a rectangular target frame to be processed as an example, the second determining module 120 may obtain four target frame vertices of the rectangular target frame to be processed in a target frame semantic graph in an image coordinate system under a BEV perspective.

[0093] In some embodiments, the second determination module 120 may determine the vertices of the target frame in various ways. For example, the second determination module 120 may obtain a vertex matrix of the vertices of the target frame to be processed under non-rotational translation:

[0094] Wherein, L is the length of the target frame to be processed, and W is the width of the target frame to be processed.

[0095] The second determination module 120 can multiply the vertex matrix by the rotation matrix to rotate the target frame to be processed. Then, the rotated target frame to be processed is translated into the target frame semantic map in the image coordinate system under the BEV perspective. In some embodiments, the rotation matrix may include:

[0096] in, Refers to the heading angle. For more information about the heading angle, please refer to the relevant description of step 210. For more information about the image coordinate system, projection, and pixel area under the BEV perspective, please refer to the relevant description above.

[0097] In some embodiments, the second determining module 120 may determine the target frame coverage obtained by the above four target frame vertices as the pixel area of the target frame to be processed in the target frame semantic map.

[0098] In step 230 , based on the semantic values of the pixels in the pixel area, it is determined whether the target frame to be processed needs to be eliminated. In some embodiments, the determination module 130 may execute step 230 .

[0099] For more information about the judgment module 130, please refer to Figure 1 Related description.

[0100] A semantic value refers to the meaning assigned to a pixel. In some embodiments, a semantic value may correspond to a semantic identifier. For example, after the second determination module 120 projects the i-th target frame to be processed onto the target frame semantic map, it determines the pixel region corresponding to the i-th target frame to be processed. The second determination module 120 may assign a semantic value of i to the pixel points in the pixel region, indicating that the pixel points in the pixel region are occupied by the i-th target frame to be processed, and the pixel region can be used as the i-th pixel region. This forms a correspondence between the frame semantic map and the target frame to be processed.

[0101] In some embodiments, the judgment module 130 determines whether the target frame to be processed needs to be eliminated based on the semantic values of the pixels in the pixel area in a variety of ways. For example, the judgment module 130 can determine whether there are non-zero pixels in the semantic values of the pixels in the pixel area. In response to the absence of non-zero pixels in the pixel area, the judgment module 130 can determine that the target frame to be processed corresponding to the pixel area does not need to be eliminated. In response to the presence of non-zero pixels in the pixel area, the judgment module 130 can determine the overlapping target frame that overlaps with the target frame to be processed based on the non-zero pixels; at least based on the target categories corresponding to the target frame to be processed and the overlapping target frame, respectively, determine whether the target frame to be processed needs to be eliminated. For more information, please refer to Figure 3 and / or Figure 4 Related description.

[0102] In step 240 , based on the judgment result, the point cloud data corresponding to the target frame to be processed is processed. In some embodiments, the processing module 140 may execute step 240 .

[0103] For more information about the processing module 140, please refer to Figure 1 Related description.

[0104] The judgment result refers to the result determined by the judgment module 130 as to whether the target frame to be processed needs to be eliminated.

[0105] In some embodiments, the determination module 130 may transmit the determination result to the processing module 140 .

[0106] In some embodiments, the processing module 140 may process target frames to be processed that do not need to be eliminated and / or target frames to be processed that need to be eliminated based on the judgment result. For example, the first target frame to be processed does not need to be eliminated, and the processing module 140 may draw the first target frame to be processed in the corresponding pixel area in the target frame semantic map. That is, the processing module 140 may retain the first target frame to be processed in the corresponding pixel area in the target frame semantic map. Retention may include retaining the semantic value corresponding to each pixel point in the pixel area. For another example, the second target frame to be processed needs to be eliminated, and the processing module 140 may determine and record the semantic identifier 2 of the second target frame to be processed, and continue to perform the operation of determining whether the next target frame to be processed needs to be eliminated.

[0107] In some embodiments, the point cloud processing system may traverse multiple target frames to be processed once until all target frames to be processed are determined by the judgment module 130 to determine whether they need to be eliminated.

[0108] Some embodiments of this specification provide a target screening method based on a point cloud processing algorithm in port driving, which determines whether a target frame to be processed needs to be eliminated based on semantic values. When making the judgment, it is only necessary to traverse all target frames to be processed once, avoiding pairwise comparison of target frames to be processed, improving data processing efficiency, and effectively shortening data processing time, thereby improving detection work efficiency and quality.

[0109] Common sorting algorithms have a time complexity of O(NlogN), requiring pairwise comparison and deduplication after sorting. Some embodiments of this specification provide target screening methods for port driving based on point cloud processing algorithms. These methods use NMS to determine whether there are target boxes to be eliminated, eliminating the need to sort target boxes by confidence level and requiring only a single pass. This method has a low time complexity of O(N).

[0110] Figure 3 This is an exemplary flow chart for determining whether a target frame to be processed needs to be eliminated according to some embodiments of this specification. Figure 3 As shown, the process 300 includes the following steps: In some embodiments, the process 300 may be executed by the determination module 130 .

[0111] Step 310: Determine whether there is a non-zero pixel in the pixel area. Figure 2 Related description.

[0112] A non-zero pixel point refers to a pixel point in a pixel region whose pixel value is not 0. In some embodiments, a non-zero pixel point may be a pixel point in the pixel region where a target object or any other meaningful information is located.

[0113] In some embodiments, the determination module 130 can determine whether there are non-zero pixels in the pixel region in a variety of ways. In some embodiments, the determination module 130 can obtain a set of pixels on the target frame semantic map within the coverage range of the target frame to be processed using the four target frame vertices of the target frame to be processed in the image coordinate system under the BEV perspective. The determination module 130 can examine each pixel in the set of pixels to determine whether there are non-zero pixels in the pixel region.

[0114] Step 321 : In response to the absence of non-zero pixels in the pixel area, determining that the target frame to be processed does not need to be eliminated.

[0115] For more information about the target frame to be processed, see Figure 2 Related instructions.

[0116] In some embodiments, the target screening method based on the point cloud processing algorithm in port driving also includes: if it is determined that the target frame to be processed does not need to be eliminated, the judgment module 130 can assign the semantic value of the pixel point in the pixel area as the identification value corresponding to the target frame to be processed, and perform projection and judgment operations on the next target frame to be processed.

[0117] For more information about pixels, semantic values, and projections, see Figure 2 Related instructions.

[0118] An identifier is a label used to distinguish and mark target frame information. An identifier value is a data value associated with an identifier. In some embodiments, the identifier value can be a number, a letter, or the like.

[0119] Assigning a value refers to assigning an identification value to pixels within a pixel region. In some embodiments, the value (i) is used as the semantic value of the target frame to be processed. If it is determined that the target frame to be processed does not need to be removed, the determination module 130 may assign the semantic value of the pixels within the pixel region to the identification value (i), indicating that the pixel region is occupied by the target frame to be processed i.

[0120] The judgment operation refers to judging whether the target box to be processed needs to be eliminated based on the semantic value of the pixel points in the pixel area. The specific process of the judgment operation can be found in Figure 2 Related instructions for step 230 in .

[0121] In some embodiments of this specification, by marking the processed pixel area, repeated judgment and calculation on the same area in subsequent processing can be avoided, thereby saving computing resources and improving processing efficiency.

[0122] Step 322 : In response to the presence of non-zero pixels in the pixel area, an overlapping target frame that overlaps with the target frame to be processed is determined based on the non-zero pixels.

[0123] The overlapping target frame refers to a target frame in the pixel area that has pixel overlap with the target frame to be processed. In some embodiments, the number of overlapping target frames can be one or more.

[0124] In some embodiments, the determination module 130 may determine the overlapping target frame based on the non-zero pixels in a variety of ways.

[0125] For example, in some embodiments, the judgment module 130 may traverse the pixel area covered by the target frame to be processed, count the semantic values of all non-zero pixels and their corresponding pixel counts, and form a statistical result.

[0126] The pixel count refers to the number of pixels with the same semantic value. The statistical result can be a tuple {(j1,k1),(j2,k2),...,(jM,kM)}, where j1, j2,...jM∈{1, 2,..., N}. j1, j2,...jM represent the semantic values of different non-zero pixels, and kM is the pixel count, which represents the number of non-zero pixels with the semantic value jM among all non-zero pixels in the pixel region.

[0127] In some embodiments, when the number of elements in the tuple corresponding to the statistical result is 0, that is, there is no overlapping target frame that overlaps with the target frame to be processed in the pixel area.

[0128] In some embodiments, the value of M can be 1, indicating that in the pixel area covered by the target frame to be processed, there is only one non-zero pixel point with the corresponding semantic value j1, and the corresponding pixel count is k1, that is, there is only one overlapping target frame in the target frame to be processed.

[0129] In some embodiments, the value of M can also be other numerical values N, indicating that in the pixel area covered by the target frame to be processed, there are N non-zero pixel points with corresponding semantic values j1, j2...jN, and the corresponding pixel counts are k1, k2...kN, that is, there are N overlapping target frames in the target frame to be processed.

[0130] In some embodiments, the determination module 130 may sequentially determine the number of overlapping target frames based on the semantic values j1, j2, ..., jM in the statistical results. For example, assuming the statistical results are {(1, 5), (2, 6), (3, 7), (4, 8)}, it can be determined that there are four types of non-zero pixels in the target processing frame, corresponding to four overlapping target frames; based on the four types of non-zero pixels with semantic values of 1, 2, 3, and 4, four overlapping target frames can be determined in the target processing frame.

[0131] Step 332 : Determine whether the target frame to be processed needs to be removed based on at least the target categories corresponding to the target frame to be processed and the overlapping target frames.

[0132] The target category refers to the type or category of the target object in the pixel area. For example, the target category can be vehicles, pedestrians, cargo, etc.

[0133] In some embodiments, the determination module 130 can obtain the target category based on various methods. For example, the determination module 130 can query the corresponding target frame information in the target frame sequence based on the semantic values corresponding to the target frame to be processed and the overlapping target frame, and obtain the target category corresponding to the target frame to be processed and the overlapping target frame based on the target frame information.

[0134] A target frame sequence refers to a preset sequence of multiple target frames. Target frame information can include, for example, the target frame's target category. The target frame's target category refers to the category of the target object represented by each target frame. For example, if target frame 1 has a target category of vehicle, all target objects represented by target frames 1 corresponding to this target category are vehicles. Target frame 1 under this target category can represent the 11th, 13th, or 15th vehicle, and so on.

[0135] For more information about the target frame and target frame information, see Figure 2 Related instructions.

[0136] In some embodiments, the determination module 130 may determine whether the target frame to be processed needs to be removed in a variety of ways based at least on the target categories corresponding to the target frame to be processed and the overlapping target frames.

[0137] In some embodiments, the judgment module 130 can determine the intersection area and merged area of the target frame to be processed and the overlapping target frame; determine the degree of overlap based on the intersection area and merged area; determine the overlap threshold based on the target categories corresponding to the target frame to be processed and the overlapping target frame respectively; and determine whether the target frame to be processed needs to be eliminated based on the overlap and the overlap threshold.

[0138] For detailed information on how to determine whether the target frame to be processed needs to be eliminated, see Figure 4 Related instructions.

[0139] In some embodiments of this specification, the presence of non-zero pixels within a pixel region is used to determine whether to exclude a target frame. This can further improve the accuracy of the result of determining whether to exclude a target frame, thereby optimizing the system's detection capabilities. Furthermore, by determining whether to exclude a target frame according to a set procedure, the execution logic of each step is clarified, simplifying the processing flow of complex tasks and improving the overall efficiency and maintainability of the system.

[0140] Figure 4 This is an exemplary flow chart of determining whether a target frame to be processed needs to be eliminated by using overlap and overlap threshold values according to some embodiments of this specification. Figure 4 As shown, the process 400 includes the following steps: In some embodiments, the process 400 may be executed by the determination module 130 .

[0141] Step 410: Determine the intersection area and the combined area of the target frame to be processed and the overlapping target frame.

[0142] For more information about the target frame to be processed, please refer to the description of step 210. For more information about the overlapping target frames, please refer to the description of step 322.

[0143] The intersection area is the area of the overlap between two plane shapes. For example, the area of the overlap between the target frame to be processed and the overlapping target frame.

[0144] In some embodiments, the determination module 130 may determine a first pixel count of pixels corresponding to the portion where the target frame to be processed overlaps the overlapped target frame. The determination module 130 may calculate a first product of the first pixel count and the resolution corresponding to the pixel, and use the first product as the intersection area.

[0145] The merged area refers to the sum of the areas of the overlapping and non-overlapping parts of the target frame to be processed and the overlapping target frame.

[0146] In some embodiments, the determination module 130 may statistically determine a second pixel count of pixels corresponding to the non-overlapping portion of the target frame to be processed and the overlapped target frame. The determination module 130 may calculate a second product of the second pixel count and the resolution corresponding to the pixel, and use the sum of the second product and the intersection area as the merged area.

[0147] For more information about pixel count, please refer to the description of step 322. For more information about resolution, please refer to the description of step 220.

[0148] In step 420 , the degree of overlap is determined based on the intersection area and the combined area.

[0149] The overlap degree refers to the degree of overlap between the target frame to be processed and the overlapping target frame. In some embodiments, the overlap degree can be represented by the intersection over union (IoU).

[0150] In some embodiments, the judgment module 130 may calculate the ratio of the intersection area to the combined area and use the ratio as the degree of overlap. In some embodiments, the same target frame to be processed may correspond to multiple different overlapping target frames. The judgment module 130 may separately calculate the intersection area of the multiple overlapping target frames and the target frame to be processed, and separately calculate the combined area of the multiple overlapping target frames and the target frame to be processed. The judgment module 130 may separately calculate the ratio of the intersection area to the corresponding combined area to determine the degree of overlap corresponding to the multiple overlapping target frames.

[0151] Step 430 : Determine an overlap threshold based on the target categories corresponding to the target frame to be processed and the overlapping target frame.

[0152] For more information about the target category, please refer to the description of step 332.

[0153] The overlap threshold refers to a preset value used for comparison with the overlap.

[0154] In some embodiments, the judgment module 130 can determine the overlap threshold in a variety of ways. For example, the judgment module 130 can construct a second preset table based on the historical target categories corresponding to the target frame to be processed and the overlapping target frame in the historical data, as well as the corresponding historical overlap thresholds. The second preset table can include historical target categories, historical overlap thresholds, and their corresponding relationships. Based on the current target category, the judgment module 130 can query the second preset table to determine the same or similar historical target categories, and determine the historical overlap threshold corresponding to the historical target category as the current overlap threshold.

[0155] In some embodiments, the determination module 130 may also obtain the overlap threshold value in other ways, such as obtaining it by at least one of manual input and obtaining it from historical data.

[0156] In some embodiments, the target screening method based on the point cloud processing algorithm in port driving also includes: if the target categories corresponding to the target frame to be processed and the overlapping target frame are the same, the judgment module 130 can set the overlap threshold to be less than the benchmark threshold; if the target categories corresponding to the target frame to be processed and the overlapping target frame are different, the judgment module 130 can determine the overlap threshold based on the degree of proximity of the target categories corresponding to the target frame to be processed and the overlapping target frame.

[0157] The baseline threshold is a fixed parameter value that determines the overlap threshold between the target frame to be processed and the overlapping target frame. For example, the baseline threshold is set to 0.5.

[0158] When the target categories corresponding to the target frame to be processed A and the overlapping target frame B are the same, the overlap threshold corresponding to the target frame to be processed A and the overlapping target frame B is less than the reference threshold of 0.5. For example, the overlap threshold is 0.4, 0.3, etc.

[0159] In some embodiments, the baseline threshold may be determined based on historical experience.

[0160] The degree of proximity refers to the similarity or correlation between the two target categories corresponding to the target frame to be processed and the overlapping target frame.

[0161] In some embodiments, the degree of proximity can be quantified by the semantic or hierarchical relationship between the object categories. For example, if two object categories belong to the same semantic group or category set, the degree of proximity is high, and vice versa.

[0162] In some embodiments, the proximity degree can be determined in a variety of ways. For example, it can be determined by comparing the semantic similarity between target categories.

[0163] In some embodiments, the judgment module 130 may determine the overlap threshold by querying a third preset table based on the proximity of the target categories corresponding to the target frame to be processed and the overlapping target frame.

[0164] The third preset table refers to a correspondence table between proximity and overlap thresholds, and can be determined based on manual presetting or the like.

[0165] In some embodiments, the overlap threshold can be positively correlated with the degree of proximity. For example, if the lock station and lock frame are typically placed in close proximity, indicating a high degree of proximity, the overlap threshold can be set to a higher value. For another example, if the truck head and trailer have a large structural overlap, indicating a high degree of proximity, the overlap threshold can be set to the highest value.

[0166] In some embodiments of the present specification, the determination of the overlap threshold depends on the category relationship and proximity between the target frames. In this way, appropriate overlap thresholds can be set for different target frames to be processed and overlapping target frames, which is conducive to further improving the accuracy of determining whether the target frames to be processed need to be eliminated, and further improving the efficiency and quality of the detection work.

[0167] Step 440 : Based on the overlap degree and the overlap degree threshold, determine whether the target frame to be processed needs to be eliminated.

[0168] In some embodiments, the determination module 130 may compare the degree of overlap with an overlap threshold to determine whether the overlapping target frame needs to be removed. For example, when the degree of overlap is less than the overlap threshold, the determination module 130 may determine that the corresponding overlapping target frame does not need to be removed.

[0169] In some embodiments, the judgment module 130 may determine that the target frame to be processed does not need to be eliminated in response to the overlap being less than the overlap threshold; the judgment module 130 may determine that the target frame to be processed or the overlapping target frame needs to be eliminated in response to the overlap being not less than the overlap threshold.

[0170] In some embodiments, in response to the overlap being not less than the overlap threshold, the determination module 130 may determine in various ways that the target frame to be processed or the overlapping target frame needs to be removed.

[0171] For example, when it is determined that the target frame to be processed or the overlapping target frame needs to be removed, the determination module 130 may remove the target frame to be processed by default.

[0172] In some embodiments, the target screening method based on the point cloud processing algorithm in port driving also includes: if the judgment module 130 determines that the target frame to be processed or the overlapping target frame needs to be eliminated, the judgment module 130 can determine the first confidence level of the target frame to be processed and the second confidence level of the overlapping target frame based on the target categories corresponding to the target frame to be processed and the overlapping target frame respectively.

[0173] In some embodiments, the determination module 130 may determine the target frame to be removed from the target frame to be processed and the overlapping target frames based on the first confidence level and the second confidence level.

[0174] The first confidence level refers to the confidence level corresponding to the target frame to be processed.

[0175] The second confidence level refers to the confidence level corresponding to the overlapping target box.

[0176] For more information about the confidence level, please refer to the description of step 210 .

[0177] The target frame to be removed refers to the target frame to be processed and the target frame that needs to be removed from the overlapping target frames.

[0178] In some embodiments, the judgment module 130 may determine, based on the first confidence level and the second confidence level, that at least one of the target frame to be processed or the overlapping target frame is a target frame to be removed. For example, when the first confidence level is greater than the second confidence level, the judgment module 130 may determine to remove the overlapping target frame. When the first confidence level is less than the second confidence level, the judgment module 130 may determine to remove the target frame to be processed.

[0179] In some embodiments, the judgment module 130 may determine a first priority score for the target frame to be processed and a second priority score for the overlapping target frame based on the first confidence level and the second confidence level, as well as the target categories corresponding to the target frame to be processed and the overlapping target frame, respectively. The judgment module 130 may determine the target frame with the lower priority score as the target frame to be removed.

[0180] The first priority score refers to a score for scoring the priority of the target box to be processed.

[0181] The second priority score is a score for scoring the priority of overlapping object boxes.

[0182] In some embodiments, different target categories and different confidence levels represent different degrees of reliability for the target frame to be processed or the overlapping target frame. The judgment module 130 may preset a fourth preset table in advance. The fourth preset table may include preset reliability levels for different target categories at different confidence levels.

[0183] In some embodiments, the judgment module 130 can obtain the reliability corresponding to the target frame to be processed and the overlapping target frame respectively by querying the fourth preset table based on the first confidence level and the second confidence level, as well as the target categories corresponding to the target frame to be processed and the overlapping target frame respectively.

[0184] In some embodiments, the judgment module 130 may determine a priority score based on a weighted sum of confidence and reliability. For example, for a target frame to be processed, the judgment module 130 may determine a first priority score based on the first confidence of the target frame to be processed and the weighted value of the reliability of the target frame to be processed under the first confidence. For another example, for an overlapping target frame, the judgment module 130 may determine a second priority score based on the second confidence of the overlapping target frame and the weighted value of the reliability of the overlapping target frame under the second confidence. In some embodiments, the weights corresponding to the first confidence, the second confidence, and the reliability corresponding to the target frame to be processed and the overlapping target frame, respectively, may be preset values. The judgment module 130 may determine the weights in a variety of ways. For example, obtaining at least one of manual input, obtaining from historical data, and the like.

[0185] In some embodiments, the judgment module 130 may determine whether to remove pending target frames or overlapping target frames based on the magnitude of the first priority score and the second priority score. For example, when the first confidence level differs from the second confidence level, the judgment module 130 may determine the target frame with the lower confidence level between the pending target frame and the overlapping target frame as the target frame to be removed. For another example, when the first confidence level and the second confidence level are the same, but the first priority score differs from the second priority score, the judgment module 130 may determine the target frame with the lower priority score between the pending target frame and the overlapping target frame as the target frame to be removed. For example, pending target frames and overlapping target frames of different target categories may have the same confidence level, but due to the different target categories, the corresponding reliability levels of the pending target frame and the overlapping target frame may differ, resulting in different first and second priority scores. For example, the confidence level of a common obstacle, a container truck trailer, and a rare and highly variable forklift truck both have a confidence level of 0.4. Since the first priority score corresponding to the forklift truck trailer is greater than the second priority score corresponding to the container truck trailer, when determining to remove pending target frames, the judgment module 130 may prioritize the pending target frames corresponding to the container truck trailer. For another example, regardless of whether the first confidence level and the second confidence level are the same, the judgment module may directly select the target frame to be processed and the overlapping target frame with the lower priority score as the target frame to be removed.

[0186] In some embodiments of this specification, a first priority score and a second priority score can be determined based on the confidence level and target category of the target frame to be processed and the overlapping target frame. By comparing the first priority score and the second priority score, the target frame to be removed is determined. When determining the target frame to be removed, both the confidence level and the target category can be taken into account, which can improve the accuracy of determining the target frame to be removed.

[0187] In some embodiments of the present specification, target frames to be processed or overlapping target frames of different target categories may have the same or different confidence levels. When determining target frames to be eliminated, the accuracy of determining target frames to be eliminated can be improved by considering the confidence levels.

[0188] In some embodiments of the present specification, by comparing the overlap between target frames with a preset overlap threshold, it is possible to effectively determine whether a target frame should be eliminated, thereby reducing the existence of overlapping and redundant target frames and optimizing the accuracy and precision of target detection.

[0189] In some embodiments, in response to the target frame being removed as an overlapping target frame, the determination module 130 may obtain the coverage of the overlapping target frame on the target frame semantic map and assign the semantic value of the pixel points within the coverage to 0.

[0190] For more information about the target frame semantic graph, please refer to the description of step 220.

[0191] The coverage range refers to the range on the target frame semantic map corresponding to the overlapping target frame. For more information about the overlapping target frame, please refer to the relevant description of step 322.

[0192] In some embodiments, the determination module 130 may use the area formed by the pixels corresponding to the overlapping target frames as the coverage range.

[0193] In some embodiments, the determination module 130 can assign the semantic values of pixels within the coverage range to 0 in various ways. For example, when the target frame to be removed is the j1th overlapping target frame, the determination module 130 can determine the semantic identifier j1 of the j1th overlapping target frame. The determination module 130 can determine the coverage range based on the vertices of the j1th overlapping target frame and assign the semantic values of pixels within the coverage range to 0 on the target frame semantic map.

[0194] In some embodiments, the judgment module 130 may perform an operation to determine whether the j1-jMth overlapping target frames need to be removed. That is, the i-th pixel region is sequentially compared with the j1-jMth pixel regions until it is determined that neither the i-th target frame to be processed nor the jMth overlapping target frame needs to be removed or the i-th target frame to be processed is removed. Then, the operation to determine whether the i+1th target frame to be processed needs to be removed is performed.

[0195] For more information on determining whether the target frame to be processed needs to be removed, see steps 230, Figure 3 、 Figure 4 For more information about the pixel area, please refer to the description of step 220.

[0196] By assigning the semantic value of the pixel points within the coverage range of the overlapping target frame on the target frame semantic map to 0, the non-zero pixel points within the coverage range on the target frame semantic map can be erased, which is beneficial to reducing the subsequent data processing amount.

[0197] Some embodiments of this specification provide a method for determining whether a target frame to be processed needs to be eliminated. By determining the intersection area and the combined area of the target frame to be processed and the overlapping target frame, the degree of overlap can be calculated, and whether the target frame to be processed or the overlapping target frame needs to be eliminated is determined based on the degree of overlap. There is no need to rely on confidence for judgment, and different overlapping situations of different target objects due to different shapes, the priority and detection level of the target objects can be considered, making the judgment more flexible and accurate, and more suitable for deduplication tasks of target frames to be processed of multiple target categories, and can also provide more accurate post-processing results.

[0198] In some embodiments, the target screening method based on a point cloud processing algorithm for port driving further includes: in response to the pending target frame being the last pending target frame, the judgment module 130 may retrieve a retained target frame based on the pending deletion list. The judgment module 130 may convert the point cloud data into pixels on the target frame semantic map, traverse the semantic value of each pixel, determine the retained target frame, and obtain a correspondence between each pixel and the retained target frame.

[0199] The last pending target frame is the last pending target frame in the order of the multiple pending target frames. For example, if there are N pending target frames, the pending target frames are ordered from 1 to N and processed sequentially, with the Nth pending target frame being the last pending target frame.

[0200] The to-be-deleted list refers to a list consisting of at least one target frame to be processed that needs to be removed.

[0201] In some embodiments, the list of pending targets may include a sequence of at least one target frame to be removed. In some embodiments, the judgment module 130 may determine the target frame to be removed based on the judgment result and then generate the list of pending targets. In some embodiments, the list of pending targets may include sequence data consisting of semantic identifiers of the target frames to be removed.

[0202] For more information about the target frame to be removed, please refer to the description of step 440.

[0203] The retained target frame refers to the target frame to be processed that is determined by the judgment module 130 to not need to be eliminated.

[0204] In some embodiments, the judgment module 130 may determine the target frames to be retained based on the target frames to be removed. For example, the target frames to be retained may include the target frames remaining after removing the target frames to be removed from the plurality of target frames to be processed. In some embodiments, the judgment module 130 may generate a retained list based on the semantic identifiers of the retained target frames. The retained list may include sequence data consisting of the semantic identifiers of the retained target frames.

[0205] In some embodiments, the judgment module 130 may traverse all point cloud points and determine the corresponding pixel point of each point cloud point projected on the target frame semantic map. The judgment module 130 may determine the semantic value corresponding to the pixel point. If the semantic value corresponding to the pixel point is non-zero, it may indicate that the pixel point is a pixel point of the retained target frame corresponding to the semantic value, and the judgment module 130 may assign the point cloud point to the retained target frame.

[0206] For example, the retained target frame is the i-th target frame to be processed, and the judgment module 130 can traverse all point cloud points, determine the corresponding pixel points of each point cloud point projected on the target frame semantic map, and determine all pixel points with a semantic value of i, and assign the point cloud points corresponding to the pixel points with a semantic value of i to the i-th target frame to be processed.

[0207] The correspondence relationship refers to the relationship between the pixel points on the target frame semantic map and the retained target frame. In some embodiments, the correspondence relationship may include whether the pixel points on the target frame semantic map and the retained target frame are of the same target category.

[0208] In some embodiments, the determination module 130 may determine whether the pixel on the target frame semantic map and the retained target frame are of the same target category based on the semantic value of the pixel on the target frame semantic map and the semantic value of the retained target frame. For example, if the semantic values are the same, the pixel on the target frame semantic map and the retained target frame are of the same target category. If the semantic values are different, the pixel on the target frame semantic map and the retained target frame are not of the same target category.

[0209] For more information about the point cloud, please refer to the description of step 210.

[0210] In some embodiments, the determination module 130 further obtains the point cloud points corresponding to the retained target frame based on the pixel points corresponding to the retained target frame, and further determines the convex hull feature corresponding to the retained target frame. The convex hull feature refers to the convex hull formed by the point cloud points corresponding to the retained target frame.

[0211] A convex hull is a convex polygon or polyhedron with the smallest area and volume that can contain all points in a point set. For example, it contains all point cloud points corresponding to the retained target box. At least some of the point cloud points may be located on the edge of the convex polygon or on the face of the convex polyhedron. Another portion of the point cloud points may be located within the convex polygon or the convex polyhedron. In some embodiments, an edge or face of the convex hull may include at least one edge or face of the retained target box.

[0212] In some embodiments, the determination module 130 may determine a convex hull feature of the retained target frame based on the retained target frame and its corresponding point cloud points. For example, at least one edge or at least one face of the retained target frame may be modified to form a convex hull that may include point cloud points outside the retained target frame.

[0213] In some embodiments, the determination module 130 may determine the convex hull feature corresponding to the retained target frame based on a third preset algorithm, which may include at least one of a Graham scan algorithm, a fast convex hull algorithm, and the like.

[0214] In some embodiments, when the retained target frame is a planar graphic, a Graham scan algorithm or the like may be used to determine the convex hull feature. In some embodiments, when the retained target frame is a three-dimensional graphic, a fast convex hull algorithm or the like may be used to determine the convex hull feature.

[0215] By determining the convex hull features corresponding to the retained target frame, the retained target frame can be updated so that the retained target frame contains all point cloud points of the same target category, thereby improving the segmentation accuracy of the retained target frame.

[0216] One or more embodiments of the present specification provide a computer-readable storage medium that stores computer instructions. When a computer reads the computer instructions in the storage medium, the computer can execute a target screening method based on a point cloud processing algorithm in port driving.

[0217] While the basic concepts have been described above, it will be apparent to those skilled in the art that the detailed disclosure is merely illustrative and does not limit this specification. Although not explicitly stated herein, various modifications, improvements, and revisions to this specification may be made by those skilled in the art. Such modifications, improvements, and revisions are suggested in this specification and remain within the spirit and scope of the exemplary embodiments of this specification.

[0218] In addition, unless expressly stated in the claims, the order of the processing elements and sequences, the use of alphanumeric characters, or the use of other names described in this specification are not intended to limit the order of the processes and methods of this specification. Although the above disclosure discusses some of the invention embodiments currently considered useful through various examples, it should be understood that such details are for illustrative purposes only, and the appended claims are not limited to the disclosed embodiments. On the contrary, the claims are intended to cover all modifications and equivalent combinations that are consistent with the spirit and scope of the embodiments of this specification. For example, although the system components described above can be implemented by hardware devices, they can also be implemented only by software solutions, such as installing the described system on an existing server or mobile device.

[0219] Similarly, it should be noted that, in order to simplify the presentation of this specification and thus facilitate understanding of one or more embodiments of the invention, the foregoing descriptions of the embodiments of this specification sometimes combine multiple features into a single embodiment, figure, or description thereof. However, this disclosure method does not imply that the subject matter of this specification requires more features than those recited in the claims. In fact, an embodiment may have fewer features than all of the features of a single disclosed embodiment.

[0220] In some embodiments, numbers describing the quantity of components and attributes are used. It should be understood that such numbers used in the description of the embodiments are modified by the modifiers "about", "approximately" or "substantially" in some examples. Unless otherwise stated, "about", "approximately" or "substantially" indicate that the numbers are allowed to vary by ±20%. Accordingly, in some embodiments, the numerical parameters used in the description and claims are approximate values, which may vary according to the required features of the individual embodiments. In some embodiments, the numerical parameters should take into account the specified significant digits and adopt the general method of retaining digits. Although the numerical domains and parameters used to confirm the breadth of their range in some embodiments of this specification are approximate values, in specific embodiments, the settings of such numerical values are as accurate as possible within the feasible range.

[0221] Finally, it should be understood that the embodiments described in this specification are intended only to illustrate the principles of the embodiments of this specification. Other variations may also fall within the scope of this specification. Therefore, by way of example and not limitation, alternative configurations of the embodiments of this specification may be considered consistent with the teachings of this specification. Accordingly, the embodiments of this specification are not limited to the embodiments explicitly described and illustrated in this specification.

Claims

1. A target screening method based on point cloud processing algorithm in port driving, characterized in that: include: Acquire point cloud data in port environments; Based on the point cloud data, determine the target frame to be processed; Projecting the target frame to be processed onto a target frame semantic map, and determining a pixel area of the target frame to be processed in the target frame semantic map; Based on the semantic values of the pixels in the pixel area, it is determined whether the target frame to be processed needs to be eliminated; Based on the judgment result, processing the point cloud data corresponding to the target frame to be processed; Perform cluster analysis on the processed point cloud data to form a valid target point cloud set; Extract features from the point cloud in the valid target point cloud set to generate target feature information; Based on the target feature information, the category of the valid target is determined and the valid target after the category is determined is output.

2. The target screening method based on point cloud processing algorithm in port driving according to claim 1 is characterized in that: The determining whether the target frame to be processed needs to be eliminated based on the semantic values of the pixels in the pixel area includes: Determine whether there is a non-zero pixel point in the pixel area; In response to the absence of non-zero pixels in the pixel area, determining that the target frame to be processed does not need to be eliminated; In response to the presence of non-zero pixels in the pixel area, determining an overlapping target frame that overlaps with the target frame to be processed based on the non-zero pixels; At least based on the target categories corresponding to the target frame to be processed and the overlapping target frame, it is determined whether the target frame to be processed needs to be eliminated.

3. The target screening method based on point cloud processing algorithm in port driving according to claim 2 is characterized in that: If it is determined that the target frame to be processed does not need to be eliminated, the semantic values of the pixels in the pixel area are assigned as identification values corresponding to the target frame to be processed, and the projection and judgment operations are performed on the next target frame to be processed.

4. The target screening method based on point cloud processing algorithm in port driving according to claim 2 is characterized in that: The determining whether the target frame to be processed needs to be eliminated based on at least the target categories corresponding to the target frame to be processed and the overlapping target frame respectively includes: Determine the intersection area and the combined area of the target frame to be processed and the overlapping target frame; determining a degree of overlap based on the intersection area and the combined area; Determining an overlap threshold based on the target categories corresponding to the target frame to be processed and the overlapping target frame respectively; Based on the overlap degree and the overlap degree threshold, it is determined whether the target frame to be processed needs to be eliminated.

5. The target screening method based on point cloud processing algorithm in port driving according to claim 4 is characterized in that: The determining of the overlap threshold based on the target categories respectively corresponding to the target frame to be processed and the overlapping target frame includes: If the target categories corresponding to the target frame to be processed and the overlapping target frame are the same, setting the overlap threshold to be less than the reference threshold; If the target categories corresponding to the target frame to be processed and the overlapping target frame are different, the overlap threshold is determined based on the proximity degree of the target categories corresponding to the target frame to be processed and the overlapping target frame.

6. The target screening method based on point cloud processing algorithm in port driving according to claim 4 is characterized in that: The determining whether the target frame to be processed needs to be eliminated based on the overlap and the overlap threshold includes: In response to the overlap being less than the overlap threshold, determining that the target frame to be processed does not need to be eliminated; In response to the overlap being not less than the overlap threshold, it is determined that the target frame to be processed or the overlapping target frame needs to be removed.

7. The target screening method based on point cloud processing algorithm in port driving according to claim 6 is characterized in that: If it is determined that the target frame to be processed or the overlapping target frame needs to be removed, determining a first confidence level of the target frame to be processed and a second confidence level of the overlapping target frame based on the target categories corresponding to the target frame to be processed and the overlapping target frame respectively; Based on the first confidence level and the second confidence level, a target frame to be eliminated is determined from the target frame to be processed and the overlapping target frame.

8. The target screening method based on point cloud processing algorithm in port driving according to claim 7 is characterized in that: The method further comprises: In response to the target frame to be removed being the overlapping target frame, obtaining a coverage range of the overlapping target frame on the target frame semantic graph; The semantic value of the pixel points within the coverage area is assigned to 0.

9. A point cloud processing system, characterized in that: A method for implementing a target screening method based on a point cloud processing algorithm in port driving according to any one of claims 1 to 8, comprising: A first determination module is configured to determine a target frame to be processed based on the point cloud data; A second determining module is configured to project the target frame to be processed onto a target frame semantic map, and determine a pixel area of the target frame to be processed in the target frame semantic map; A judgment module is configured to judge whether the target frame to be processed needs to be eliminated based on the semantic values of the pixels in the pixel area; The processing module is configured to process the point cloud data corresponding to the target box to be processed based on the judgment result.

10. A computer-readable storage medium, characterized in that The storage medium stores computer instructions. When the computer reads the computer instructions in the storage medium, the computer executes the target screening method based on point cloud processing algorithm in port driving as described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • Semantic object-based semantic dimension chain positioning method and system

    CN116740171A

  • Semantic map construction method and system for scene with dynamic target

    CN118411507A

  • Method For Generating Point Cloud Data And Data Generating Apparatus

    US20230237735A1

  • Data processing method and apparatus

    WO2023005797A1

  • Power transmission line detection method and apparatus, computer device, and storage medium

    WO2023174020A1