A visual target detection method, apparatus and computer program product

CN121190757BActive Publication Date: 2026-09-29GUANGZHOU AUTOMOBILE GROUP CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410772483.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-06-14
Publication Date
2026-09-29
Estimated Expiration
2044-06-14

AI Technical Summary

Technical Problem

但是,感知算法对设备要求较高,需要GPU渲染;栅格化方法实际上避开了GPU硬件,计算量非常高

Benefits of technology

[0045]实施本发明具有如下有益效果:本发明通过将图像分割为小的网格单元并为每个单元编号,可以迅速判断目标多边形覆盖哪些网格,判断过程相对简单且计算成本低,避免了复杂的多边形布尔运算,极大地加速了处理流程。本发明通过网格化策略,有效缩短了目标可见性分析的时间,使得即使在复杂多变的交通环境下,也能迅速完成对目标遮挡关系的判断和处理;不仅提高了仿真系统的响应速度,还使得模拟结果更加贴近实际情况,有助于提升自动驾驶算法的训练质量和安全性。本发明利用网格编号进行遮挡分析,简化了判断逻辑。当需要确定一个目标是否遮挡另一个目标时,只需比较覆盖的网格编号是否有交集,而非直接计算多边形之间的交集,这样既减少了计算负担,又避免了布尔运算可能带来的不稳定耗时问题。此外,由于本发明减轻了对计算资源的依赖,尤其是减少了对GPU的直接需求,因此,可以采用成本更低、能耗更低的解决方案,同时不牺牲系统的实时性能和可靠性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121190757B_ABST
    Figure CN121190757B_ABST
Patent Text Reader

Abstract

The application discloses a visual target detection method, device and computer program product, wherein the method comprises: calculating the three-dimensional coordinates of each sparse point based on the sparse point model of the visual target to be detected; projecting and converting the three-dimensional coordinates into pixel coordinates on a two-dimensional image, and calculating the convex hull containing all pixel coordinates of each visual target respectively; performing grid segmentation on the two-dimensional image to determine the grid covered by the convex hull of each visual target; detecting the occlusion relationship between visual targets according to the distance between each visual target and the ego vehicle, and updating the grid of the convex hull of the occluded visual target; and determining the visibility of each visual target according to a preset visibility filtering mode. Through grid segmentation and grid numbering, the application effectively avoids complex Boolean operations on long-distance targets, significantly reduces the number of Boolean operations, and further ensures the real-time performance of target simulation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of autonomous driving technology, and specifically to a visual target detection method, device, and computer program product. Background Technology

[0002] With the advancement of automotive intelligence and electrification, autonomous driving has become one of the mainstream trends in future automotive development. Key technologies for autonomous vehicles mainly include environmental perception, precise positioning, decision-making and planning, and control and execution. Environmental perception plays a crucial role in autonomous driving, much like human eyes.

[0003] Currently, the most commonly used methods in this field are perceptual algorithms and rasterization methods. However, perceptual algorithms have high equipment requirements, requiring GPU rendering; rasterization methods, while bypassing GPU hardware, have very high computational demands. Compared to the former two, the polygon Boolean method is faster, but its processing time is unstable, increasing with the amount of target data, making it difficult to meet real-time requirements. Summary of the Invention

[0004] The technical problem to be solved by the embodiments of the present invention is to provide a visual target detection method, device and computer program product to improve the real-time performance of target simulation.

[0005] To address the aforementioned technical problems, this invention provides a simulation method for visual target detection, comprising the following steps:

[0006] Based on the sparse point model of the visual target to be detected, the three-dimensional coordinates of each sparse point are calculated.

[0007] The three-dimensional coordinates are projected into pixel coordinates on a two-dimensional image, and the convex hull containing all the pixel coordinates of each visual target is calculated separately.

[0008] The two-dimensional image is segmented into a grid to determine the grid covered by the convex hull of each visual target;

[0009] Based on the distance between each visual target and the vehicle, the occlusion relationship between each visual target is detected, and the mesh of the convex hull of the occluded visual target is updated.

[0010] The visibility of each visual target is determined based on a preset visibility filtering method.

[0011] Preferably, the step of performing grid-based segmentation on the two-dimensional image specifically includes any of the following methods:

[0012] Different grids are divided by assigning a unique number to each grid.

[0013] Different grids are divided based on the visual features contained within them;

[0014] Different grids are divided based on the grid's position in the two-dimensional image.

[0015] Preferably, the step of detecting the occlusion relationship between visual targets based on the distance between each visual target and the vehicle specifically includes:

[0016] Based on the distance between each visual target and the vehicle, the visual targets are sorted in order from farthest to closest.

[0017] Starting with the next most distant visual target, the system detects whether each visual target occludes a more distant visual target in order of increasing distance, until the occlusion detection of the nearest visual target is completed.

[0018] Preferably, detecting whether each visual target occludes a visual target that is farther away specifically includes:

[0019] Determine if the grid numbers of the coverage areas of two visual targets overlap;

[0020] If the grid numbers of the two visual targets overlap, the polygon difference operation is used to determine whether the convex hulls of the two visual targets overlap; if the difference set obtained after the operation is not empty, it is determined that the relatively closer visual target occludes the relatively distant visual target.

[0021] Preferably, updating the mesh number of the convex hull of the occluded visual target specifically includes:

[0022] When a visual target is occluded, remove the coordinates of the occluded pixel from the current convex hull;

[0023] The convex hull of the visual target is recalculated based on the coordinates of the remaining unoccluded pixels;

[0024] Determine and update the mesh number for the recalculated convex hull cover.

[0025] Preferably, the step of dividing the image into different grids by assigning a unique number to each grid specifically involves: uniformly dividing the two-dimensional image into multiple grids and numbering each grid.

[0026] The determination of the grid number covered by the convex hull of each visual target specifically includes:

[0027] The grid number covered by the convex hull of each visual target is determined by judging whether the vertex or edge of the convex hull intersects with a certain grid, or whether the interior of the convex hull contains the center point of a certain grid.

[0028] Preferably, the calculation of the three-dimensional coordinates of each sparse point specifically involves: calculating the three-dimensional coordinates of each sparse point based on the center position and orientation of the visual target and the relative position of each sparse point with respect to the center of the visual target.

[0029] Preferably, determining the visibility of each visual target according to a preset visibility filtering method specifically involves using any of the following visibility filtering methods to determine the visibility of each visual target:

[0030] Calculate the ratio of the visible area of ​​the occluded visual target to the total area when it is completely unoccluded. If the calculated ratio is lower than a preset threshold, the visual target is determined to be invisible.

[0031] Calculate the visible area of ​​the occluded visual target. If the visible area is less than a preset threshold, the visual target is determined to be invisible.

[0032] If the visual target formed by the occlusion is composed of multiple discontinuous small areas, then the visual target is determined to be invisible.

[0033] If the obscured visual target is a vehicle, and the front of the vehicle is obscured, then the visual target is determined to be invisible.

[0034] The present invention also provides a simulation device for visual target detection, comprising:

[0035] The sparsity processing module is used to calculate the three-dimensional coordinates of each sparse point based on the sparse point model of the visual target to be detected.

[0036] The convex hull calculation module is used to convert the three-dimensional coordinate projection into pixel coordinates on the two-dimensional image, and to calculate the convex hull containing all its pixel coordinates for each visual target.

[0037] The meshing processing module is used to perform meshing segmentation on the two-dimensional image and determine the mesh covered by the convex hull of each visual target;

[0038] The occlusion detection module is used to detect the occlusion relationship between each visual target based on the distance between each visual target and the vehicle, and update the mesh of the convex hull of the occluded visual target.

[0039] The visibility determination module is used to determine the visibility of each visual target according to a preset visibility filtering method.

[0040] The present invention also provides a visual target detection device, comprising:

[0041] One or more processors;

[0042] Memory;

[0043] One or more applications, wherein the one or more applications are stored in the memory and configured to be executed by the one or more processors, and the one or more applications are configured to perform the visual target detection method.

[0044] The present invention also provides a computer program product, including computer instructions that instruct a computer device to perform an operation corresponding to the method.

[0045] The present invention offers the following advantages: By segmenting an image into small grid units and numbering each unit, it can quickly determine which grids a target polygon covers. This determination process is relatively simple and computationally inexpensive, avoiding complex polygon Boolean operations and significantly accelerating the processing flow. Through its gridding strategy, the invention effectively shortens the time required for target visibility analysis, enabling rapid judgment and processing of target occlusion relationships even in complex and changing traffic environments. This not only improves the response speed of the simulation system but also makes the simulation results more realistic, contributing to improved training quality and safety of autonomous driving algorithms. The invention utilizes grid numbering for occlusion analysis, simplifying the judgment logic. When determining whether one target occludes another, it only requires comparing the intersection of the covering grid numbers, rather than directly calculating the intersection between polygons. This reduces computational burden and avoids the instability and time-consuming issues that Boolean operations may cause. Furthermore, because the invention reduces reliance on computing resources, especially the direct demand for GPUs, a lower-cost, lower-energy-consumption solution can be adopted without sacrificing the system's real-time performance and reliability. Attached Figure Description

[0046] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0047] Figure 1 This is a flowchart illustrating a visual target detection method according to an embodiment of the present invention.

[0048] Figure 2a This is a schematic diagram of the sparse point model of the target vehicle in an embodiment of the present invention.

[0049] Figure 2b This is a schematic diagram of the sparse point model of the target pedestrian in an embodiment of the present invention.

[0050] Figure 3a This is a schematic diagram of the minimum convex hull of the target vehicle in an embodiment of the present invention.

[0051] Figure 3b This is a schematic diagram of the minimum convex hull of the target pedestrian in an embodiment of the present invention.

[0052] Figure 4This is a schematic diagram of visual target occlusion relationship detection in an embodiment of the present invention. Detailed Implementation

[0053] The following description of the embodiments is taken with reference to the accompanying drawings, which illustrate specific embodiments in which the invention can be implemented.

[0054] Please refer to Figure 1 As shown, Embodiment 1 of the present invention provides a visual target detection method, including the following steps:

[0055] Based on the sparse point model of the visual target to be detected, the three-dimensional coordinates of each sparse point are calculated.

[0056] The three-dimensional coordinates are projected into pixel coordinates on a two-dimensional image, and the convex hull containing all the pixel coordinates of each visual target is calculated separately.

[0057] The two-dimensional image is segmented into a grid to determine the grid covered by the convex hull of each visual target;

[0058] Based on the distance between each visual target and the vehicle, the occlusion relationship between each visual target is detected, and the mesh of the convex hull of the occluded visual target is updated.

[0059] The visibility of each visual target is determined based on a preset visibility filtering method.

[0060] As can be seen from the above steps, by introducing grid-based segmentation, the present invention effectively avoids complex Boolean operations on distant targets, significantly reduces the number of Boolean operations, and thus greatly improves the calculation speed; it not only optimizes the calculation efficiency of target detection, but also ensures the real-time performance of target simulation.

[0061] Specifically, the application scenario of this invention is autonomous driving PNC simulation. The PNC algorithm module is the test object of the simulation, responsible for outputting the position and power data of the main vehicle to drive its movement; while the traffic flow module is responsible for driving the movement of traffic vehicles, providing the simulation system with traffic behavior, such as overtaking, cutting in line, and emergency stopping. This invention is responsible for simulating targets based on image vision and outputting the data to the PNC algorithm module. It is understood that since the vehicle's sensors include not only cameras but also LiDAR, millimeter-wave radar, and other sensors, all of which can perceive target data and output it to the PNC algorithm module, the target data for the PNC algorithm module is multi-source. This invention mainly involves the simulation of visual targets.

[0062] Each as Figure 2a , Figure 2bAs shown, a sparse point model is a 3D representation of visual targets such as vehicles or pedestrians. They are represented as a set of sparse points, rather than a dense point cloud or a complete 3D model. These sparse points typically represent the key feature points of the visual target.

[0063] When calculating the 3D coordinates of each sparse point, both position and orientation must be considered. The position of each sparse point depends not only on its position relative to the center of the visual target but also on the target's orientation. For example, the 3D coordinates of the front end of a vehicle model will change accordingly as the vehicle's orientation changes. Therefore, to obtain the accurate 3D coordinates of these sparse points, the coordinates must be calculated based on the center position and orientation of the visual target, as well as the relative position of each sparse point to the center.

[0064] After obtaining the 3D coordinates of each sparse point, these coordinates need to be projected using the camera's intrinsic and extrinsic parameters to convert them into 2D image pixel coordinates. It's understandable that the camera's intrinsic parameters are mainly related to its internal structure and performance, such as focal length, principal point coordinates, and distortion coefficients; extrinsic parameters mainly describe the camera's position and orientation in the world coordinate system, typically including rotation and translation matrices.

[0065] The main steps of projection transformation are:

[0066] Coordinate transformation: First, the 3D coordinates of sparse points are transformed into coordinates in the camera coordinate system using the camera's extrinsic parameters (rotation matrix and translation matrix).

[0067] Projecting onto the image plane: Then, the three-dimensional coordinates are projected onto the two-dimensional image plane using the camera's intrinsic parameters (such as focal length).

[0068] Converting to pixel coordinates: Finally, by taking into account the image resolution and pixel size, the projected 2D coordinates are converted to the image's pixel coordinates.

[0069] The above steps can project the sparse point model of the target vehicle or pedestrian from three-dimensional space onto a two-dimensional image plane and calculate the corresponding image pixel coordinates, thereby realizing the conversion from the real world to a digital image.

[0070] After converting the three-dimensional coordinates into pixel coordinates on a two-dimensional image through camera projection, the next step in this embodiment of the invention is to calculate the convex hull of all pixel coordinates containing the visual target, as shown below. Figure 3a , Figure 3b As shown. It can be understood that for a set of points on a two-dimensional plane, the convex hull is the smallest convex polygon formed by connecting the outermost points.

[0071] Common algorithms for calculating the convex hull include Graham's scan method and Jarvis's step method. Taking Graham's scan method as an example, the main steps are as follows:

[0072] a. Find the point with the smallest y-coordinate. If there are multiple points, take the leftmost point as the starting point.

[0073] b. Sort the remaining points according to the angle between the line connecting them to the starting point and the x-axis. If the angles are the same, sort them by distance.

[0074] c. Initialize an empty stack, and push the starting point and the first sorted point onto the stack.

[0075] d. Traverse the remaining points after sorting. For each point, if the angle formed by it and the two points at the top of the stack is a concave angle, pop the point at the top of the stack until a convex angle is formed or only one point remains in the stack. Then push the current point onto the stack.

[0076] e. After the traversal is complete, the remaining points in the stack are the vertices of the convex hull.

[0077] Finally, connecting the points in the stack forms the smallest convex polygon containing all pixel coordinates. If needed, the edges of this smallest convex polygon can be smoothed or optimized.

[0078] As mentioned earlier, when the three-dimensional coordinates of each sparse point are converted into pixel coordinates on a two-dimensional image through camera projection, a two-dimensional image is obtained that reflects the projection effect of the target in the three-dimensional scene onto the two-dimensional plane.

[0079] This invention will introduce a grid partitioning mechanism, specifically including one of the following methods:

[0080] Different grids are divided by assigning a unique number to each grid.

[0081] Different grids are divided based on the visual features contained within the grid (such as color, texture, brightness, etc.);

[0082] Different grids are divided based on the grid's position in the 2D image (such as its distance from the center point, its position relative to other grids, etc.);

[0083] Considering that grid numbering requires relatively little computation in subsequent calculations, this embodiment of the invention uses grid numbering as an example for illustration.

[0084] Specifically, the two-dimensional image is uniformly divided into multiple grids, for example, the size of each grid is set to 100×100 pixels.

[0085] After dividing the grid, each grid needs to be numbered. Numbering usually follows certain rules, such as from left to right or from top to bottom, to ensure that each grid has a unique identifier, facilitating subsequent management and retrieval.

[0086] Next, it is necessary to determine which meshes are covered by the previously calculated convex hull. This is usually achieved by determining whether a vertex or edge of the convex hull intersects with a certain mesh, or whether the interior of the convex hull contains the center point of a certain mesh.

[0087] After the above steps, a convex hull and the grid number covered by the convex hull will be generated for each target (vehicle or pedestrian), thereby discretizing the complex image space and simplifying subsequent calculation and analysis tasks.

[0088] It should be noted that in the simulation of autonomous driving environments, due to occlusion by other objects, some targets (such as vehicles or pedestrians) may be partially or completely invisible in the camera's field of view. Therefore, the simulation needs to accurately identify which targets are visible and which are invisible, that is, to detect the occlusion relationships of the targets.

[0089] Please combine Figure 4 As shown, firstly, based on the distance between each visual target and the vehicle's observation point (e.g., the vehicle's camera), each visual target ( Figure 4 The vehicles A, B, and C shown are arranged in order from farthest to closest: vehicle A is the farthest target, and vehicle C is the closest target. The order from farthest to closest is: vehicle A (farthest) - vehicle B - vehicle C (closest).

[0090] In this embodiment of the invention, occlusion detection begins with the second most distant target, i.e., vehicle B. This is because the most distant target (vehicle A) is usually not occluded by other targets, so we start by considering vehicles that might occlude other targets. For vehicle B, we first check and attempt to cover targets that are even more distant than it (i.e., vehicle A). If vehicle B does not occlude vehicle A, then the visibility of vehicle A is unaffected.

[0091] Next, move to a target closer to car B (i.e., car C). For all targets farther than car C (car A and car B), car C may cause occlusion. Similarly, check and try to cover targets farther away in turn, first car A, then car B. Since car A has been confirmed not to be occluded by car B, car C will not occlude car A either (car A's convex hull remains unchanged, meaning its visibility remains unchanged). Then check the effect of car C on car B, and find that part of car B is occluded by car C, causing car B's convex hull area to decrease. This means that car B's visibility is affected, and visually, part of car B is no longer visible.

[0092] It should be noted that occlusion detection includes two main processes:

[0093] The first step is mesh overlap detection, which determines whether the visible portions (represented by mesh numbers) of two targets overlap. This is achieved by comparing the mesh numbers covered by the convex hull polygons of the two targets. If the two sets of mesh numbers intersect, meaning they cover the same mesh area, then the visible portions of the two targets overlap. If the two sets of mesh numbers do not intersect, then the visible portions of the two targets do not overlap.

[0094] Once the mesh numbering determines that overlap exists, the next step is to perform a polygon difference operation (used to find the different portions of the two convex hulls) to determine whether the convex hulls of the two targets actually overlap; that is, removing the portion covered by one convex polygon from another. If this operation executes successfully and produces a non-empty result (i.e., there is still a remaining portion after removal), this indicates that the two polygons have at least a portion of their spatial regions that are different; in other words, they have an overlapping region. Figure 4 As shown, target car B is relatively close to target car A. We can check whether they overlap by calculating the difference between the convex hull polygons of car A and car B. If the difference set obtained after the calculation is not empty, it means that car B does indeed occlude part of car A, that is, an occlusion relationship between the two targets is detected: the presence of car B reduces the visibility of car A (i.e., the convex hull), which is manifested as a reduction in the size of the polygon.

[0095] During the occlusion detection process described above, whenever the convex hull of a visual target changes, the grid number covering that convex hull is recalculated. For example, when it is determined that car B is occluded, the shape of its corresponding convex hull will change (shrink). At this time, the grid number covering this shrunken convex hull needs to be recalculated. Updating the grid number is to ensure accurate tracking of the visible portion of each target.

[0096] Specifically, firstly, the occluded pixel coordinates are removed from the current convex hull of the occluded visual target, retaining only the still visible pixel coordinates. Then, the convex hull is recalculated based on the pixel coordinates of the remaining unoccluded portion. Next, the recalculated convex hull determines which grids it covers and updates the grid numbers to reflect the target's current visibility.

[0097] Occlusion detection continues until the occlusion relationships and visibility of all targets are determined. Figure 4 Once the occlusion check of vehicle C is completed, the entire visibility calculation process ends.

[0098] In this way, the visibility of each target vehicle from the current perspective can be accurately determined, thus providing crucial information for subsequent autonomous driving decisions and path planning. This calculation order, starting with the next most distant target, helps optimize the processing flow, as closer targets are more likely to occlude more distant targets.

[0099] After determining the occlusion relationship between targets and the visible part of each target through the above process, the embodiments of the present invention further make a final judgment on the visibility of the targets based on a preset visibility filter, so as to determine the visibility state of the targets in the final visual simulation.

[0100] Specifically, visibility filtering can use the following logic:

[0101] Area ratio: Calculate the ratio of the visible area of ​​the target after it is occluded to the total area of ​​the target when it is completely unoccluded. If this ratio is lower than a certain preset threshold, the target is considered to be invisible.

[0102] Minimum area: A minimum visible area standard is set. If the remaining visible area after a target is occluded is less than this minimum value, then the target is considered invisible. This is because the exposed portion is too small to provide meaningful information.

[0103] Fragmentation: If the visible portion of a target after being occluded is too scattered, forming multiple discontinuous small areas (i.e., fragmentation), even if the total visible area is large, it will be considered invisible because it does not meet the requirement of coherence. This is because the visible portion in this case cannot provide overall information about the target.

[0104] Local priority: In certain specific situations, such as vehicle recognition in autonomous driving, if the front of the vehicle (which usually contains important identification information such as license plate, information about the vehicle's direction of travel and intention) is obscured, the entire vehicle will be judged as invisible even if the rear or other parts of the vehicle are still visible, because key identification information is missing.

[0105] Implementing visibility filtering after occlusion analysis and mesh coverage update can more closely approximate the visual recognition limitations in the real world, improving the realism and practicality of simulation results; it can more accurately simulate the visual perception capabilities of autonomous driving systems in complex traffic environments, helping to optimize decision-making algorithms and improve safety performance.

[0106] Corresponding to the visual target detection method described in Embodiment 1 of the present invention, Embodiment 2 of the present invention further provides a visual target detection device, comprising:

[0107] The sparsity processing module is used to calculate the three-dimensional coordinates of each sparse point based on the sparse point model of the visual target to be detected.

[0108] The convex hull calculation module is used to convert the three-dimensional coordinate projection into pixel coordinates on the two-dimensional image, and to calculate the convex hull containing all its pixel coordinates for each visual target.

[0109] The meshing processing module is used to perform meshing segmentation on the two-dimensional image and determine the mesh covered by the convex hull of each visual target;

[0110] The occlusion detection module is used to detect the occlusion relationship between each visual target based on the distance between each visual target and the vehicle, and update the mesh of the convex hull of the occluded visual target.

[0111] The visibility determination module is used to determine the visibility of each visual target according to a preset visibility filtering method.

[0112] Corresponding to the visual target detection method described in Embodiment 1 of the present invention, Embodiment 3 of the present invention also provides a visual target detection device, comprising:

[0113] One or more processors;

[0114] Memory;

[0115] One or more applications, wherein the one or more applications are stored in the memory and configured to be executed by the one or more processors, and the one or more applications are configured to perform the visual target detection method described in Embodiment 1 of the present invention.

[0116] Corresponding to the visual target detection method described in Embodiment 1 of the present invention, Embodiment 4 of the present invention also provides a computer program product, including computer instructions, which instruct a computer device to perform the operation corresponding to the visual target detection method described in Embodiment 1 of the present invention.

[0117] Preferably, the processor can be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor, or the processor can be any conventional processor. The processor is the control center of the device, connecting various parts of the device through various interfaces and lines.

[0118] The memory mainly includes a program storage area and a data storage area. The program storage area can store the operating system, applications required for at least one function, etc., while the data storage area can store related data, etc. Furthermore, the memory can be a high-speed random access memory, or a non-volatile memory, such as a plug-in hard drive, a SmartMedia Card (SMC), a Secure Digital (SD) card, and a Flash Card, or other volatile solid-state storage devices.

[0119] It should be noted that the above-mentioned devices may include, but are not limited to, processors and memory, as will be understood by those skilled in the art.

[0120] For the working principle and process of the above embodiments, please refer to the description of Embodiment 1 of the present invention, which will not be repeated here.

[0121] As can be seen from the above description, compared with the prior art, the beneficial effects of the present invention are as follows:

[0122] (1) Improved computational efficiency: Traditional polygon Boolean operations suffer from rapidly increasing computational costs when dealing with occlusion problems in complex scenes, especially when dealing with distant targets. These targets often occupy a large visual space, leading to higher Boolean operation complexity and computational delays. The grid numbering method used in this invention avoids this problem. By dividing the image into small grid units and numbering each unit, it is possible to quickly determine which grids the target polygons cover. This determination process is relatively simple and computationally inexpensive, avoiding complex polygon Boolean operations and greatly accelerating the processing flow.

[0123] (2) Real-time performance guarantee: The present invention effectively shortens the target visibility analysis time through the gridding strategy, so that even in complex and ever-changing traffic environments, the judgment and processing of target occlusion relationships can be completed quickly; it not only improves the response speed of the simulation system, but also makes the simulation results closer to the actual situation, which helps to improve the training quality and safety of autonomous driving algorithms.

[0124] (3) Simplified occlusion logic: This invention uses grid numbers for occlusion analysis, which simplifies the judgment logic. When it is necessary to determine whether one target occludes another target, it is only necessary to compare whether the covered grid numbers have an intersection, rather than directly calculating the intersection between polygons. This reduces the computational burden and avoids the instability and time-consuming problems that may be caused by Boolean operations.

[0125] Furthermore, since this invention reduces the reliance on computing resources, especially the direct demand on GPUs, a lower-cost and lower-energy-consumption solution can be adopted without sacrificing the real-time performance and reliability of the system.

[0126] The above description is merely a preferred embodiment of the present invention and should not be construed as limiting the scope of the invention. Therefore, any equivalent variations made in accordance with the claims of the present invention are still within the scope of the present invention.

Claims

1. A visual target detection method, characterized in that, Includes the following steps: Based on the sparse point model of the visual target to be detected, the three-dimensional coordinates of each sparse point are calculated. The three-dimensional coordinates are projected into pixel coordinates on a two-dimensional image, and the convex hull containing all the pixel coordinates of each visual target is calculated separately. The two-dimensional image is segmented into a grid to determine the grid covered by the convex hull of each visual target; Based on the distance between each visual target and the vehicle, the occlusion relationship between each visual target is detected, and the mesh of the convex hull of the occluded visual target is updated. The visibility of each visual target is determined based on a preset visibility filtering method.

2. The method according to claim 1, characterized in that, The grid segmentation of the two-dimensional image specifically includes any of the following methods: Different grids are divided by assigning a unique number to each grid. Different grids are divided based on the visual features contained within them; Different grids are divided based on the grid's position in the two-dimensional image.

3. The method according to claim 2, characterized in that, The step of detecting the occlusion relationship between visual targets based on the distance between each visual target and the vehicle specifically includes: Based on the distance between each visual target and the vehicle, the visual targets are sorted in order from farthest to closest. Starting with the next most distant visual target, the system detects whether each visual target occludes a more distant visual target in order of increasing distance, until the occlusion detection of the nearest visual target is completed.

4. The method according to claim 3, characterized in that, The detection of whether each visual target occludes a visual target that is farther away specifically includes: Determine if the grid numbers of the coverage areas of two visual targets overlap; If the grid numbers of the two visual targets overlap, the polygon difference operation is used to determine whether the convex hulls of the two visual targets overlap; if the difference set obtained after the operation is not empty, it is determined that the relatively closer visual target occludes the relatively distant visual target.

5. The method according to claim 3, characterized in that, The step of updating the mesh of the convex hull of the occluded visual target specifically involves updating the mesh number of the convex hull of the occluded visual target, which specifically includes: When a visual target is occluded, remove the coordinates of the occluded pixel from the current convex hull; The convex hull of the visual target is recalculated based on the coordinates of the remaining unoccluded pixels; Determine and update the mesh number for the recalculated convex hull cover.

6. The method according to claim 2, characterized in that, The specific method of dividing the image into different grids by assigning a unique number to each grid is as follows: the two-dimensional image is evenly divided into multiple grids, and each grid is numbered. The determination of the grid number covered by the convex hull of each visual target specifically includes: The grid number covered by the convex hull of each visual target is determined by judging whether the vertex or edge of the convex hull intersects with a certain grid, or whether the interior of the convex hull contains the center point of a certain grid.

7. The method according to claim 1, characterized in that, The calculation of the three-dimensional coordinates of each sparse point specifically involves: calculating the three-dimensional coordinates of each sparse point based on the center position and orientation of the visual target and the relative position of each sparse point with respect to the center of the visual target.

8. The method according to claim 1, characterized in that, The process of determining the visibility of each visual target according to a preset visibility filtering method specifically involves using any of the following visibility filtering methods to determine the visibility of each visual target: Calculate the ratio of the visible area of ​​the occluded visual target to the total area when it is completely unoccluded. If the calculated ratio is lower than a preset threshold, the visual target is determined to be invisible. Calculate the visible area of ​​the occluded visual target. If the visible area is less than a preset threshold, the visual target is determined to be invisible. If the visual target formed by the occlusion is composed of multiple discontinuous small areas, then the visual target is determined to be invisible. If the obscured visual target is a vehicle, and the front of the vehicle is obscured, then the visual target is determined to be invisible.

9. A visual target detection device, characterized in that, include: The sparsity processing module is used to calculate the three-dimensional coordinates of each sparse point based on the sparse point model of the visual target to be detected. The convex hull calculation module is used to convert the three-dimensional coordinate projection into pixel coordinates on the two-dimensional image, and to calculate the convex hull containing all its pixel coordinates for each visual target. The meshing processing module is used to perform meshing segmentation on the two-dimensional image and determine the mesh covered by the convex hull of each visual target; The occlusion detection module is used to detect the occlusion relationship between each visual target based on the distance between each visual target and the vehicle, and update the mesh of the convex hull of the occluded visual target. The visibility determination module is used to determine the visibility of each visual target according to a preset visibility filtering method.

10. A simulation device for visual target detection, characterized in that, include: One or more processors; Memory; One or more applications, wherein the one or more applications are stored in the memory and configured to be executed by the one or more processors, the one or more applications being configured to perform the visual target detection method as described in any one of claims 1 to 8.

11. A computer program product, characterized in that, Includes computer instructions that instruct a computer device to perform an operation corresponding to the method as described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • Target occlusion assessment method and system, medium and device

    CN112634313A

  • Volume calculation method, system and equipment based on point cloud data in sinkhole and medium

    CN117876465A