Method and device for generating visible area mask for vehicle perception

By generating visible area masks, the problem of sparse voxel data in scene semantic completion model training is solved, and effective supervision of false voxels and efficient training of models is achieved.

CN120220084APending Publication Date: 2025-06-27EACON TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510180343.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-18
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

In the training process of the scene semantic completion model, it is difficult to process sparse voxel data, resulting in unreasonable supervision of false voxels, which in turn affects the convergence and semantic completion capabilities of the model.

Method used

By acquiring the voxel data collected by the current vehicle, the initial mask, foreground top mask and background top mask are determined based on the voxel data and vehicle location, and a visible area mask is generated to adapt to sparse voxel data and effectively supervise the scene.

Benefits of technology

Effective supervision of the scene semantic completion model in sparse voxel data scenarios is realized, unreasonable supervision of false voxels is avoided, and the convergence and semantic completion ability of the model are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120220084A_ABST
    Figure CN120220084A_ABST
Patent Text Reader

Abstract

The invention provides a method and a device for generating a visible area mask for vehicle perception, and relates to the technical field of automatic driving and unmanned vehicles. The method comprises the following steps: acquiring voxel data corresponding to a target area acquired by a current vehicle, wherein the voxel data comprises effective voxels; determining an initial mask corresponding to the voxel data based on the voxel data and the position of the current vehicle; determining a foreground top mask corresponding to the foreground target based on a labeling box corresponding to the foreground target in the voxel data; determining a background top mask corresponding to the background target based on a background voxel corresponding to the background target in the voxel data; and generating a visible area mask corresponding to the voxel data based on the initial mask, the foreground top mask and the background top mask. The method for generating the visible region mask for vehicle perception can adapt to sparse voxel data, so that the voxel data can be fully supervised based on the visible region mask in the training process of the scene semantic completion model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the technical fields of autonomous driving and driverless vehicle technologies, and particularly to a method and apparatus for generating a visible area mask for vehicle perception. Background Art

[0002] The visible area mask plays a key role in the scene semantic completion process of the occupancy network. Although existing occupancy network ground truth generation methods can generate dense ground truth data, in some special cases, such as when the scene is severely occluded or in bad weather, the generated data usually has holes and a certain degree of sparsity, resulting in unreasonable empty voxels. During the training process of the scene semantic completion model, these unreasonable empty voxels will make it difficult for the scene semantic completion model to converge, and further lead to a serious degradation of the scene completion ability. Summary of the Invention

[0003] In view of this, the present disclosure provides a method and apparatus for generating a visible area mask for vehicle perception.

[0004] In a first aspect, a method for generating a visible area mask for vehicle perception is provided, including: obtaining voxel data corresponding to a target area collected by a current vehicle, where the voxel data includes valid voxels; determining an initial mask corresponding to the voxel data based on the voxel data and the position of the current vehicle; determining a foreground top mask corresponding to a foreground target based on a bounding box corresponding to the foreground target in the voxel data; determining a background top mask corresponding to a background target based on background voxels corresponding to the background target in the voxel data; and generating a visible area mask corresponding to the voxel data based on the initial mask, the foreground top mask, and the background top mask.

[0005] In combination with the first aspect, in some implementation manners of the first aspect, determining a foreground top mask corresponding to a foreground target based on a bounding box corresponding to the foreground target in the voxel data includes: determining a projection of the bounding box corresponding to the foreground target in the horizontal direction in the spatial coordinate system corresponding to the voxel data, and determining a projected horizontal coordinate of the projection coverage range; determining a spatial position in the spatial coordinate system where the horizontal coordinate is the same as the projected horizontal coordinate and the height coordinate is greater than the height coordinate of the top of the bounding box as a foreground supervision area; and generating a foreground top mask based on the foreground supervision area.

[0006] In combination with the first aspect, in some implementations of the first aspect, determining a background top mask corresponding to a background target based on background voxels corresponding to the background target in voxel data includes: in the spatial coordinate system corresponding to the voxel data, respectively determining the highest valid voxel for each of a plurality of horizontal coordinates; performing class filtering on the highest valid voxels for each of the plurality of horizontal coordinates to obtain a plurality of first background voxels; generating a first height supervision mask at the spatial positions where the plurality of first background voxels are located, and deleting the partial mask that overlaps with the foreground top mask in the first height supervision mask to obtain a second height supervision mask; determining the spatial positions of the empty voxels above the second height supervision mask as the background supervision region; and generating a background top mask based on the background supervision region.

[0007] In combination with the first aspect, in some implementations of the first aspect, determining an initial mask corresponding to the voxel data based on the voxel data and the position of the current vehicle includes: for each valid voxel, determining the connection line between the position of the current vehicle and the valid voxel; based on the voxel data, determining a plurality of target voxels through which the connection line passes, and the class of each of the plurality of target voxels; when the classes of the plurality of target voxels meet a preset class condition, determining the spatial positions where the plurality of target voxels are located and the spatial position where the valid voxel is located as the initial mask supervision region corresponding to the valid voxel; when the classes of the plurality of target voxels do not meet the preset class condition, determining the spatial position where the valid voxel is located as the initial mask supervision region corresponding to the valid voxel; and generating an initial mask based on the initial mask supervision region corresponding to each valid voxel.

[0008] In combination with the first aspect, in some implementations of the first aspect, the preset class condition includes: the classes of the plurality of voxels are empty voxels or noise voxels, and the noise voxels include dust voxels and / or rain and snow voxels.

[0009] In combination with the first aspect, in some implementations of the first aspect, the method for generating a visible region mask for vehicle perception further includes: in the spatial coordinate system corresponding to the voxel data, partitioning the spatial positions corresponding to the voxel data to obtain coarse-resolution blocks and high-resolution blocks, where a coarse-resolution block consists of multiple high-resolution blocks; respectively determining the most significant voxel in each coarse-resolution block, and respectively determining the top height value corresponding to each coarse-resolution block based on the height coordinate of the most significant voxel in each coarse-resolution block; traversing the high-resolution blocks, and when the target high-resolution block does not include valid voxels, determining the target coarse-resolution block corresponding to the target high-resolution block and the target top height value corresponding to the target coarse-resolution block; determining, in the spatial coordinate system, the spatial positions whose horizontal coordinates are the same as those of the target high-resolution block and whose height coordinates are greater than the target top height value as the occlusion supervision region; generating an occlusion region top mask based on the occlusion supervision region. Determining the visible region mask corresponding to the voxel data based on the initial mask, the foreground top mask, and the background top mask includes: taking the union of the initial mask, the foreground top mask, the background top mask, and the occlusion region top mask to obtain the visible region mask.

[0010] In combination with the first aspect, in some implementations of the first aspect, the method for generating a visible region mask for vehicle perception further includes: in the spatial coordinate system corresponding to the voxel data, traversing the horizontal coordinates of the spatial coordinate system, and when there are no valid voxels in the voxels corresponding to the target horizontal coordinate, deleting the visible region mask with the lowest height in the visible region mask corresponding to the target horizontal coordinate.

[0011] In a second aspect, there is provided a device for generating a visible region mask for vehicle perception, including: an acquisition module configured to acquire voxel data corresponding to a target region collected by a current vehicle, where the voxel data includes valid voxels; a first determination module configured to determine an initial mask corresponding to the voxel data based on the voxel data and the position of the current vehicle; a second determination module configured to determine a foreground top mask corresponding to a foreground target based on the annotation box corresponding to the foreground target in the voxel data; a third determination module configured to determine a background top mask corresponding to a background target based on the background voxels corresponding to the background target in the voxel data; and a generation module configured to generate a visible region mask corresponding to the voxel data based on the initial mask, the foreground top mask, and the background top mask.

[0012] In a third aspect, there is provided an electronic device, including: a processor; and a memory for storing executable instructions of the processor; wherein, the processor is configured to execute the method for generating a visible region mask for vehicle perception provided in the first aspect above by executing the executable instructions.

[0013] Fourthly, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the method for generating a visible area mask for vehicle perception provided in the first aspect is implemented.

[0014] In the embodiments of the present disclosure, prior information of foreground objects and background objects in voxel data is used to generate a foreground top mask and a background top mask. The visible area mask obtained by combining the foreground top mask, the background top mask and the initial mask can adapt to sparse voxel data. Training a scene semantic completion model based on the visible area mask can enable the model to learn semantic completion ability while avoiding unreasonable supervision of the model on false empty voxels. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 The figure shows a schematic diagram of a scene applicable to an embodiment of the present disclosure.

[0016] Figure 2 The figure shows a schematic flowchart of a method for generating a visible area mask for vehicle perception provided in an embodiment of the present disclosure.

[0017] Figure 3 The figure shows a schematic flowchart of steps for determining a foreground top mask corresponding to a foreground object based on a bounding box corresponding to the foreground object in voxel data provided in an embodiment of the present disclosure.

[0018] Figure 4 The figure shows a schematic flowchart of steps for determining a background top mask corresponding to a background object based on background voxels corresponding to the background object in voxel data provided in an embodiment of the present disclosure.

[0019] Figure 5 The figure shows a schematic flowchart of steps for determining an initial mask corresponding to voxel data based on voxel data and the position of the current vehicle provided in an embodiment of the present disclosure.

[0020] Figure 6 The figure shows a schematic flowchart of a method for generating a visible area mask for vehicle perception provided in another embodiment of the present disclosure.

[0021] Figure 7 The figure shows a schematic structural diagram of a device for generating a visible area mask for vehicle perception provided in an embodiment of the present disclosure.

[0022] Figure 8 The figure shows a schematic structural diagram of an electronic device provided in an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0023] Next, the technical solutions in the embodiments of the present disclosure will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present disclosure. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present disclosure without creative efforts belong to the scope of protection of the present disclosure.

[0024] Occupancy Networks is an advanced environmental perception technology. During the process of an autonomous vehicle perceiving the surrounding environment, through Occupancy Networks, the vehicle's autonomous driving system can generate a high-resolution and continuous three-dimensional spatial representation, which is crucial for achieving precise navigation and obstacle avoidance.

[0025] The visible region mask plays an important role in the scene semantic completion process of Occupancy Networks. During the training process of the scene semantic completion model, only the voxels covered by the visible region mask are supervised; when calculating the loss value, the loss function only considers the voxels within the visible region mask and ignores the voxels outside the visible region mask to improve the prediction effect of the model.

[0026] Facing problems such as scene occlusion and bad weather, the quality of the point cloud data collected by the radar deteriorates, often having holes and a certain degree of sparsity, resulting in false empty voxels in the generated Occupancy Networks. Among them, false empty voxels refer to the situation where there should be valid voxels at this position, but due to reasons such as occlusion, the point cloud data is not collected, resulting in this position being identified as an empty voxel. If this part of the false empty voxels is directly learned by the scene semantic completion model, it will cause the model to be difficult to converge, and then lead to a serious degradation of the scene semantic completion ability.

[0027] The visible region mask generation methods in the related art usually default that the Occupancy Networks scene is a dense mapping result and do not consider the processing of sparse voxel data. During the training process of the scene semantic completion model, using the visible region mask generation methods in the related art to process sparse Occupancy Networks scenes will lead to unreasonable supervision of false empty voxels, and then cause the model to generate unreasonable predictions.

[0028] In view of the above technical problems, the present disclosure provides a method for generating a visible region mask for vehicle perception, including: obtaining voxel data corresponding to a target region collected by a current vehicle, where the voxel data includes valid voxels; determining an initial mask corresponding to the voxel data based on the voxel data and the position of the current vehicle; determining a foreground top mask corresponding to a foreground target based on a bounding box corresponding to the foreground target in the voxel data; determining a background top mask corresponding to a background target based on background voxels corresponding to the background target in the voxel data; and generating a visible region mask corresponding to the voxel data based on the initial mask, the foreground top mask, and the background top mask. The method for generating a visible region mask for vehicle perception in the present disclosure can adapt to sparse voxel data, so that during the training process of a scene semantic completion model, the voxel data can be fully supervised based on the visible region mask.

[0029] Figure 1 The following is a schematic diagram of a scenario applicable to an embodiment of the present disclosure. As Figure 1 shown, this scenario includes a vehicle 10 and a server 20. The vehicle 10 is communicatively connected to the server 20, for example, through a wired or wireless network connection, etc.

[0030] The vehicle 10 is equipped with a radar device for collecting point cloud data of a target region. Then, the vehicle 10 sends the point cloud data to the server 20.

[0031] Based on the point cloud data sent by the vehicle 10, the server 20 generates voxel data corresponding to the target region and generates a visible region mask corresponding to the voxel data based on the method for generating a visible region mask for vehicle perception provided in the embodiment of the present disclosure. Among them, the server 20 can be an interoperability server or a background server between multiple heterogeneous systems, or an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as a big data and artificial intelligence platform, etc.

[0032] The following Figures 2 to 6 illustrates the method for generating a visible region mask for vehicle perception provided in the embodiment of the present disclosure by way of example.

[0033] Figure 2 The following is a schematic flowchart of the method for generating a visible region mask for vehicle perception provided in an embodiment of the present disclosure. As Figure 2 shown, the method for generating a visible region mask for vehicle perception provided in the embodiment of the present disclosure includes the following steps.

[0034] S210, obtain voxel data corresponding to a target region collected by a current vehicle.

[0035] The current vehicle is used to collect point cloud data of the target area. The three-dimensional space of the target area is divided into multiple regular voxel units, and based on the point cloud data, the attribute information of each voxel unit regarding the spatial position is determined to obtain the voxel data corresponding to the target area. The voxel data includes the position information and attribute information of multiple voxel units respectively. The position information of the voxel unit may include the position coordinates of the voxel unit.

[0036] Among them, the voxel data includes valid voxels and empty voxels. Specifically, according to the different attribute information of the voxel units, the multiple voxel units in the voxel data are distinguished into valid voxels and empty voxels; valid voxels represent the voxel units occupied by actual objects in the three-dimensional space, and empty voxels represent the voxel units not occupied by any actual objects in the three-dimensional space.

[0037] Exemplarily, a radar device is installed in the current vehicle, and the point cloud data of the target area can be obtained through the radar device.

[0038] S220. Based on the voxel data and the position of the current vehicle, determine the initial mask corresponding to the voxel data.

[0039] Using the ray sampling method, determine the initial mask corresponding to each valid voxel in the voxel data respectively.

[0040] Specifically, traverse each valid voxel. For each valid voxel, connect the position of the current vehicle with the position where the valid voxel is located to obtain a connection line segment.

[0041] Perform segmented sampling on the connection line segment to determine all the voxels intersecting with the connection line segment. Among all the voxels intersecting with the connection line segment, if there are other valid voxels in addition to the valid voxel, determine the position where the valid voxel is located as the supervision area. If all the voxels intersecting with the connection line segment are empty voxels, determine the positions of all the voxels intersecting with the connection line segment as the supervision area.

[0042] After traversing the valid voxels, determine the supervision areas corresponding to all the valid voxels respectively as the initial mask corresponding to the voxel data.

[0043] The initial mask obtained in the above way conforms to the perspective logic, but the generation process is too dependent on valid voxels. For the areas not scanned by the radar device of the current vehicle, or the areas where the distance from the radar device is too far to return valid points, effective supervision cannot be obtained.

[0044] S230. Based on the bounding box corresponding to the foreground target in the voxel data, determine the foreground top mask corresponding to the foreground target.

[0045] For foreground objects with a relatively high height in the target area, due to the viewing angle limitation of the radar device of the current vehicle, it may not be possible to scan the top and above positions of the foreground object. As a result, when predicting above the foreground object, there may be a problem of chaotic prediction of the voxel category above the foreground object.

[0046] To avoid the situation where the area above the foreground object cannot be effectively supervised, in the embodiments of the present disclosure, first, foreground and background separation is performed on the voxel data to determine the foreground objects and background objects included in the voxel data. Among them, the foreground object refers to the object of interest in the target area. In the field of autonomous driving, foreground objects usually include moving objects such as vehicles and pedestrians. Background objects usually include static environmental features, such as roads, buildings, retaining walls, etc., and also include dynamic targets such as rain, snow, and dust.

[0047] Based on the annotation box corresponding to the foreground object, the area above the annotation box is used as the supervision area, and a foreground top mask corresponding to the foreground object is generated at the position of the supervision area.

[0048] S240. Based on the background voxels corresponding to the background objects in the voxel data, determine the background top mask corresponding to the background objects.

[0049] For scenarios such as mining areas where the background height is relatively low, there are usually no valid voxels in the upper part of the scenario, resulting in the initial mask being unable to cover the upper layer area of the scenario. This makes the upper layer area always in an unsupervised state during the training process of the scenario semantic completion model, resulting in chaotic voxel categories in the prediction results of the upper layer area.

[0050] To avoid the situation where the upper layer area of the scenario cannot be effectively supervised, in the embodiments of the present disclosure, the empty voxel positions above the background voxels are determined as the supervision areas, and a background top mask corresponding to the background objects is generated at the positions of the supervision areas.

[0051] S250. Based on the initial mask, the foreground top mask, and the background top mask, generate a visible area mask corresponding to the voxel data.

[0052] Since the foreground top mask covers the area above the foreground object and the background top mask covers the area above the background object. Therefore, the foreground top mask and the background top mask can avoid the problem of chaotic voxel category prediction due to lack of supervision above the foreground object and background object in the scenario.

[0053] Combine the foreground top mask, the background top mask with the initial mask to obtain the visible area mask. Based on the visible area mask, the voxel data in the scenario can be fully supervised, avoiding the semantic ambiguity problem in the voxel data, and preventing the scenario semantic completion model from learning too many unreasonable empty voxels during training.

[0054] In the disclosed embodiment, a foreground top mask and a background top mask are generated using prior information of foreground targets and background targets in voxel data. Among them, based on the foreground top mask, the area above the foreground target can be supervised, effectively suppressing the unreasonable prediction phenomenon caused by the scene semantic completion model failing to supervise the area above the foreground target. Based on the background top mask, the upper area in the scene with a low background height can be supervised to avoid the problem that the upper area of ​​the scene is not effectively supervised. The visible area mask obtained by combining the foreground top mask, the background top mask and the initial mask can adapt to sparse voxel data, and training the scene semantic completion model based on the visible area mask can enable the model to learn semantic completion capabilities while avoiding the model's unreasonable supervision of false empty voxels.

[0055] The following will continue to introduce how to generate the foreground top mask, background top mask, and initial mask.

[0056] Figure 3 FIG. 1 is a flow chart of a step of determining a foreground top mask corresponding to a foreground target based on a label box corresponding to a foreground target in voxel data provided by an embodiment of the present disclosure. Figure 3 As shown, the step of determining the foreground top mask corresponding to the foreground target based on the labeling box corresponding to the foreground target in the voxel data provided in the embodiment of the present disclosure includes the following steps.

[0057] S231, determining the projection of the annotation box corresponding to the foreground object in the horizontal direction in the spatial coordinate system corresponding to the voxel data, and determining the projection horizontal coordinates of the projection coverage range.

[0058] The spatial coordinate system corresponding to the voxel data is the coordinate system of the three-dimensional space where the voxels (including valid voxels and empty voxels) in the voxel data are located. In three-dimensional space, voxels appear as small cubic units with fixed size and position, and each voxel has different corresponding position coordinates. The horizontal direction of the spatial coordinate system is parallel to the ground plane, and the height direction is perpendicular to the ground plane; illustratively, the ground plane can be the XY plane of the spatial coordinate system. In the spatial coordinate system, the position of the voxel in the horizontal direction is represented by the horizontal coordinate in the position coordinate, and the position in the height direction is represented by the height coordinate in the position coordinate.

[0059] The annotation box corresponding to the foreground object is used to describe the position of the foreground object in the spatial coordinate system. It is understandable that the annotation box can be obtained by manual annotation, or can be automatically generated by a annotation algorithm, or can be determined by manual annotation and an annotation algorithm, and the present disclosure does not make specific restrictions on this.

[0060] Determine the projection of the annotation box in the horizontal direction, and determine the projected horizontal coordinates of the projection coverage range. The position shown by the projected horizontal coordinates is within the projection coverage range. Multiple projected horizontal coordinates together form a set of projected horizontal coordinates.

[0061] Specifically, the projection of the annotation box in the horizontal direction can be the orthographic projection of the annotation box on the XY plane. Based on the projection, determine the projection coverage range of the annotation box on the XY plane, and obtain the projected horizontal coordinates of the projection coverage range.

[0062] S232. In the spatial coordinate system, determine the spatial positions where the horizontal coordinates are the same as the projected horizontal coordinates and the height coordinates are greater than the height coordinate of the top of the annotation box as the foreground supervision regions.

[0063] If there are multiple foreground objects in the scene, the annotation boxes corresponding to the multiple foreground objects are different, and the top of each annotation box corresponds to its respective height coordinate. In this step, traverse all the projected horizontal coordinates in the set of projected horizontal coordinates. For each projected horizontal coordinate, determine the annotation box to which the projected horizontal coordinate belongs; then, determine the spatial positions where the horizontal coordinates are the same as the projected horizontal coordinates and the height coordinates are greater than the height coordinate of the top of the annotation box corresponding to the projected horizontal coordinate in the spatial coordinate system as the foreground supervision region corresponding to the projected horizontal coordinate. In this way, obtain the foreground supervision regions corresponding to each of the multiple projected horizontal coordinates respectively, and the foreground supervision regions corresponding to each of the multiple projected horizontal coordinates together form the foreground supervision region.

[0064] Among them, the "spatial positions where the horizontal coordinates are the same as the projected horizontal coordinates and the height coordinates are greater than the height coordinate of the top of the annotation box corresponding to the projected horizontal coordinate" can be understood as the three-dimensional spatial region composed of three-dimensional spatial points whose horizontal coordinates and height coordinates both meet the above conditions. The height coordinate of the top of the annotation box can be the height coordinate of the top vertex of the annotation box. Exemplarily, in the set of projected horizontal coordinates, a projected horizontal coordinate is (x0, y0), and the height coordinate of the top of the annotation box corresponding to the projected horizontal coordinate is z0. In the process of determining the foreground supervision region, take the spatial positions where the horizontal coordinates are (x0, y0) and the height coordinates are greater than z0 in the spatial coordinate system as the foreground supervision region corresponding to the projected horizontal coordinate (x0, y0). For example, determine the spatial position with coordinates (x0, y0, z1) as the foreground supervision region corresponding to the projected horizontal coordinate (x0, y0), where z1 > z0.

[0065] S233. Generate a foreground top mask based on the foreground supervision region.

[0066] In this step, a foreground top mask is generated within the foreground supervision region, that is, the foreground top mask covers all spatial positions of the foreground supervision region. Exemplarily, in the spatial coordinate system, the coordinates corresponding to the foreground supervision region are assigned a value of 1, and the coordinates of the remaining positions are assigned a value of 0 to obtain the foreground top mask.

[0067] In the embodiments of the present disclosure, based on the prior information of the foreground object in the scene, a foreground top mask covering the region above the foreground object is generated. Based on the foreground top mask, effective supervision of the voxels above the foreground object can be formed. Moreover, based on the annotation box of the foreground object, the shape of the foreground object can be effectively constrained, promoting the processing accuracy of the foreground object in complex scenes.

[0068] Figure 4 The following is a schematic flowchart of the steps for determining the background top mask corresponding to the background object based on the background voxels corresponding to the background object in the voxel data provided by an embodiment of the present disclosure. As Figure 4 shown, the steps for determining the background top mask corresponding to the background object based on the background voxels corresponding to the background object in the voxel data provided by the embodiments of the present disclosure include the following steps.

[0069] S241, in the spatial coordinate system corresponding to the voxel data, respectively determine the highest valid voxels of each of the multiple horizontal coordinates.

[0070] Traverse all the horizontal coordinates in the spatial coordinate system and respectively determine the highest valid voxels of each of the multiple horizontal coordinates.

[0071] Taking the horizontal coordinate (x1, y1) as an example, first, find all the valid voxels with the horizontal coordinate (x1, y1) as candidate valid voxels. Exemplarily, the coordinates of the obtained candidate valid voxels are respectively (x1, y1, z2), (x1, y1, z3), (x1, y1, z4), where z2 < z3 < z4.

[0072] Next, among the candidate valid voxels, determine the candidate valid voxel with the largest value of the height coordinate as the highest valid voxel of the horizontal coordinate (x1, y1); that is, take the candidate valid voxel with the coordinate (x1, y1, z4) as the highest valid voxel of the horizontal coordinate (x1, y1).

[0073] Based on the above method, the highest valid voxels of each of the multiple horizontal coordinates are obtained. It can be understood that the candidate valid voxels corresponding to the above-listed horizontal coordinates are only illustrative.

[0074] S242, perform category filtering on the highest valid voxels of each of the multiple horizontal coordinates to obtain multiple first background voxels.

[0075] Perform class filtering on the most significant voxels for each of multiple horizontal coordinates. The classes of valid voxels can include a "foreground voxel" class, a "background voxel" class, etc. Among them, the class of valid voxels corresponding to the foreground target is the "foreground voxel" class, and the class of valid voxels corresponding to the background target is the "background voxel" class. Filter out the valid voxels with the background class from the multiple most significant voxels as multiple first background voxels.

[0076] Alternatively, the classes of valid voxels can be further subdivided, and the subdivided classes can be divided into foreground and background. Exemplarily, the classes of valid voxels can include a "vehicle voxel" class, a "pedestrian voxel" class, a "ground voxel" class, a "retaining wall voxel" class, etc. Among them, the "vehicle voxel" class and the "pedestrian voxel" class are foreground, and the "ground voxel" class and the "retaining wall voxel" class are background. During the process of performing class filtering on the multiple most significant voxels, the valid voxels with the "ground voxel" class and the "retaining wall voxel" class are used as multiple first background voxels.

[0077] S243, Generate a first height supervision mask at the spatial positions of the multiple first background voxels, and delete the partial mask that coincides with the foreground top mask in the first height supervision mask to obtain a second height supervision mask.

[0078] The multiple first background voxels obtained based on the above steps cover the background targets in the scene. Since the foreground target is above the ground voxels, if a background top mask is directly generated in the area above the first background voxels, then the background top mask will partially coincide with the foreground top mask.

[0079] Therefore, after obtaining the multiple first background voxels, generate a first height supervision mask at the spatial positions of the multiple first background voxels, and delete the partial mask that coincides with the foreground top mask in the first height supervision mask to obtain a second height supervision mask. Among them, the first height supervision mask coincides with the foreground top mask means that the coordinates of the first height supervision mask are the same as those of the foreground top mask.

[0080] S244, Determine the background supervision area as the spatial positions of the empty voxels located above the second height supervision mask.

[0081] After determining the second height supervision mask, determine the background supervision area as the spatial positions of the empty voxels located above the second height supervision mask.

[0082] Specifically, traverse the first background voxels inside the second height supervision mask. For each first background voxel, if there is an empty voxel with the same horizontal coordinate as this first background voxel and a height coordinate greater than the height coordinate of this first background voxel, then determine the spatial position of this empty voxel as the background supervision area.

[0083] Exemplarily, the coordinates of the first background voxel inside the second height supervision mask are (x2, y2, z5). The coordinates of the empty voxels above the first background voxel are (x2, y2, z6) and (x2, y2, z7) respectively, where z5 < z6 < z7. At this time, the positions (x2, y2, z6) and (x2, y2, z7) are determined as the background supervision regions.

[0084] It can be understood that each first background voxel inside the second height supervision mask can correspond to multiple empty voxels, and the empty voxels corresponding to each first background voxel together form all the background supervision regions.

[0085] S245, generate a background top mask based on the background supervision regions.

[0086] In this step, a background top mask is generated within the background supervision regions. Exemplarily, in the spatial coordinate system, the coordinates corresponding to the background supervision regions are assigned a value of 1, and the coordinates of the remaining positions are assigned a value of 0 to obtain the background top mask.

[0087] In the embodiments of the present disclosure, based on the prior information of the background objects in the scene and the pre-generated foreground top mask, a background top mask covering the region above the background objects is generated. Moreover, it is possible to avoid the background top mask and the foreground top mask overlapping with each other. Based on the background top mask, effective supervision of the voxels above the background objects can be formed, avoiding the problem that the upper region cannot be effectively supervised in a scene where the background height is low.

[0088] Figure 5 The figure shows a schematic flowchart of the steps for determining the initial mask corresponding to the voxel data based on the voxel data and the position of the current vehicle provided by an embodiment of the present disclosure. As Figure 5 shown, the steps for determining the initial mask corresponding to the voxel data based on the voxel data and the position of the current vehicle provided by the embodiments of the present disclosure include the following steps.

[0089] S221, for each valid voxel, determine the connection line between the position of the current vehicle and the valid voxel.

[0090] For each valid voxel in the voxel data, a line segment is generated with the position of the valid voxel and the position of the current vehicle as endpoints, serving as the connection line corresponding to the valid voxel.

[0091] S222, based on the voxel data, determine the multiple target voxels passed through by the connection line, and the categories of the multiple target voxels respectively.

[0092] In the spatial coordinate system corresponding to the voxel data, each voxel represents a small area in space. If a connection line passes through the small area represented by a certain voxel, then that voxel is determined as the target voxel. The connection line may pass through multiple voxels, and the multiple voxels may include valid voxels or empty voxels.

[0093] It can be understood that in this embodiment, the voxels at the connection line endpoint positions are not included in the multiple target voxels, that is, the valid voxels serving as endpoints and the voxels corresponding to the position of the current vehicle.

[0094] After obtaining the multiple target voxels, based on the voxel data, the categories of the multiple target voxels are respectively determined. Exemplarily, if the target voxel is a valid voxel, the category can be a "vehicle voxel" category, a "retaining wall voxel" category, a "dust voxel" category, etc.; if the target voxel is an empty voxel, the category of the empty voxel is regarded as the "empty voxel" category.

[0095] S223, when the categories of the multiple target voxels respectively meet the preset category conditions, the spatial positions where the multiple target voxels are located, and the spatial position where the valid voxel is located, are determined as the initial mask supervision region corresponding to the valid voxel.

[0096] The preset category conditions include: the categories of the multiple target voxels are all specified categories. If the categories of the multiple target voxels respectively meet the preset category conditions, it means that the multiple target voxels are all voxels of the specified categories. At this time, the spatial positions where the multiple target voxels are located, and the spatial position of the valid voxel serving as the connection line endpoint, are used as the initial mask supervision region corresponding to the valid voxel.

[0097] There can be multiple specified categories. If the category of the target voxel is one of the multiple specified categories, then the category of the target voxel meets the preset category conditions.

[0098] Exemplarily, the specified category can be the "empty voxel" category. If the category of the target voxel meets the preset category conditions, it means that the multiple target voxels are all empty voxels; that is, there is no occlusion between the target voxels and the current vehicle. At this time, the area between the target voxels and the current vehicle can be determined as the initial mask supervision region. Of course, the spatial position where the valid voxel is located also needs to be determined as the initial mask supervision region.

[0099] S224, when the categories of the multiple target voxels do not meet the preset category conditions, the spatial position where the valid voxel is located is determined as the initial mask supervision region corresponding to the valid voxel.

[0100] If the categories of multiple target voxels do not meet the preset category conditions, it indicates that the multiple target voxels include voxels other than the voxels of the specified category. At this time, the spatial position where the valid voxel serving as the connection endpoint is located is used as the initial mask supervision region corresponding to the valid voxel.

[0101] S225, generate an initial mask based on the initial mask supervision region corresponding to each valid voxel.

[0102] Based on the above steps, generate the initial mask supervision region corresponding to each valid voxel. The initial mask supervision regions corresponding to each valid voxel together form all the initial mask supervision regions.

[0103] Next, generate an initial mask within the initial mask supervision region. Exemplarily, in the spatial coordinate system, assign the coordinates corresponding to the initial mask supervision region as 1, and assign the coordinates of the remaining positions as 0 to obtain the initial mask.

[0104] In the embodiments of the present disclosure, the generated initial mask conforms to the perspective logic and is the basis of the visible region mask. Based on the initial mask, combined with the foreground top mask and the background top mask, it can make up for the problem that the initial mask is difficult to cover the area not scanned by the radar device or the area where there is no effective point return due to being too far away from the radar device, so as to realize reasonable and effective supervision of the voxels in the scene.

[0105] In some embodiments, the preset category conditions include: the categories of multiple voxels are "empty voxels" or "noise voxels"; among them, the category of "noise voxels" includes dust voxels and / or rain and snow voxels.

[0106] In this embodiment, if the categories of multiple target voxels are all "empty voxels" or "noise voxels", then the categories of multiple target voxels meet the preset category conditions.

[0107] In scenes with relatively harsh natural environments such as mining areas, the dust, dust, or extreme rain and snow weather in the mining area will affect the performance of the radar device, resulting in noise in the return signal of the radar device. In the subsequent process of generating voxel data, noise voxels such as dust voxels and rain and snow voxels may be generated. During the process of generating the initial mask, these noise voxels will be misidentified as the occlusion between the target voxel and the current vehicle, resulting in the target voxel not being supervised. To solve this problem, in this embodiment, if there are only noise voxels such as dust voxels and rain and snow voxels between the target voxel and the current vehicle and no other valid voxels, it is considered that there is no occlusion between the target voxel and the current vehicle, so as to improve the reliability of the initial mask and make the initial mask more adaptable to scenes with harsh environments and extreme weather.

[0108] The above has introduced a method for generating a visible region mask for vehicle perception. However, in some scenarios, there may be regions with frequent occlusions, and the radar device cannot scan the occluded regions, resulting in a large number of empty voxels in the occluded regions. Since these empty voxels do not exist above any valid voxels or on the line connecting any valid voxels to the current vehicle, the visible region mask generated based on the method in the above embodiments cannot effectively supervise the occluded regions. In addition, there is also semantic ambiguity in the occluded regions themselves; the semantics of the occluded regions are unknown, but for application considerations, the priority of predicting the occluded regions as empty voxels is higher than the priority of predicting them as confused voxels.

[0109] To effectively supervise the voxels in the un-scanned regions, another embodiment of the present disclosure provides a method for generating a visible region mask for vehicle perception. As Figure 6 shown, in addition to the steps in the above embodiments, the method for generating a visible region mask for vehicle perception provided in the embodiments of the present disclosure further includes the following steps.

[0110] S610, in the spatial coordinate system corresponding to the voxel data, partition the spatial positions corresponding to the voxel data to obtain coarse-resolution blocks and high-resolution blocks.

[0111] Based on the distribution of valid voxels in the voxel data, partition the spatial positions corresponding to the voxel data. Among them, the spatial positions corresponding to the voxel data are the regions where valid voxels are distributed, and this region has a corresponding relationship with the target region in the real scene.

[0112] The size of the coarse-resolution blocks is larger than that of the high-resolution blocks. If there are multiple coarse-resolution blocks in the spatial coordinate system, the sizes of the multiple coarse-resolution blocks are the same. The coarse-resolution blocks are composed of multiple high-resolution blocks, and the sizes of the multiple high-resolution blocks are the same. And, the size of the high-resolution blocks should be larger than the size of the voxels, that is, one high-resolution block includes multiple voxels.

[0113] The sizes of the coarse-resolution blocks and high-resolution blocks are determined by the distribution of valid voxels in the voxel data. During the process of partitioning the coarse-resolution blocks, ensure that there is at least one valid voxel in each coarse-resolution block.

[0114] S620, respectively determine the highest valid voxel in each coarse-resolution block, and respectively determine the top height value corresponding to each coarse-resolution block based on the height coordinate of the highest valid voxel in each coarse-resolution block.

[0115] Traverse all the coarse-resolution blocks. For each coarse-resolution block, determine the highest valid voxel in this coarse-resolution block. Among them, the highest valid voxel is the valid voxel included in this coarse-resolution block and having the largest height coordinate.

[0116] After determining the most significant voxel in the coarse-resolution block, the height coordinate of the most significant voxel is used as the top height value corresponding to the coarse-resolution block.

[0117] In this way, the most significant voxel in each coarse-resolution block is obtained, and the top height value corresponding to each coarse-resolution block is determined.

[0118] S630. Traverse the high-resolution blocks. When the target high-resolution block does not include valid voxels, determine the target coarse-resolution block corresponding to the target high-resolution block and the target top height value corresponding to the target coarse-resolution block.

[0119] Traverse all the high-resolution blocks. For the convenience of description, the traversed high-resolution block is called the target high-resolution block, and the target high-resolution block is any one of the multiple high-resolution blocks.

[0120] For the target high-resolution block, determine the voxels included in the target high-resolution block. Among them, the voxels included in the target high-resolution block can be valid voxels or empty voxels.

[0121] If the target high-resolution block includes valid voxels, continue traversing.

[0122] If the target high-resolution block does not include valid voxels, determine the coarse-resolution block to which the target high-resolution block belongs as the target coarse-resolution block corresponding to the target high-resolution block, that is, the target coarse-resolution block includes the target high-resolution block. Determine the top height value corresponding to the target coarse-resolution block as the target top height value.

[0123] S640. In the space coordinate system, determine the space position where the horizontal coordinate is the same as the horizontal coordinate of the target high-resolution block and the height coordinate is greater than the target top height value as the occlusion supervision area.

[0124] After determining the target top height value, determine the space position where the horizontal coordinate is the same as the horizontal coordinate of the target high-resolution block and the height coordinate is greater than the target top height value as the occlusion supervision area.

[0125] Exemplarily, the horizontal coordinates of the empty voxels included in the target high-resolution block are (x3, y3) and (x4, y4) respectively, and the target top height value is z8. At this time, determine the space position above the coordinate (x3, y3, z8) and the space position above the coordinate (x4, y4, z8) as the occlusion supervision area corresponding to the target high-resolution block.

[0126] S650. Generate an occlusion area top mask based on the occlusion supervision area.

[0127] In this step, an occlusion region top mask is generated within the occlusion supervision region, that is, the occlusion region top mask covers all spatial positions of the occlusion supervision region. Exemplarily, in a spatial coordinate system, the coordinates corresponding to the occlusion supervision region are assigned a value of 1, and the coordinates of the remaining positions are assigned a value of 0 to obtain the occlusion region top mask.

[0128] After determining the occlusion region top mask, it is combined with the initial mask, foreground top mask, and background top mask to obtain a visible region mask.

[0129] Specifically, in the embodiments of the present disclosure, the step of generating a visible region mask corresponding to the voxel data based on the initial mask, foreground top mask, and background top mask includes: taking the union of the initial mask, foreground top mask, background top mask, and occlusion region top mask to obtain the visible region mask.

[0130] In this step, by taking the union of the initial mask, foreground top mask, background top mask, and occlusion region top mask, the regions represented by the above-mentioned multiple masks are merged to obtain the visible region mask.

[0131] Exemplarily, in the above-mentioned multiple masks, the value of the supervision region corresponding to each mask is 1, and the remaining positions are 0. During the process of taking the union, the regions with a value of 1 in the multiple masks are merged to generate the visible region mask. Specifically, it can be implemented by performing an "OR" operation on the initial mask, foreground top mask, background top mask, and occlusion region top mask.

[0132] In the embodiments of the present disclosure, the spatial coordinate system is divided into coarse-resolution blocks and high-resolution blocks. If a high-resolution block does not include valid voxels, the high-resolution block is regarded as being located in the occluded region, and an occlusion region top mask corresponding to the high-resolution block is generated according to the coarse-resolution block to which the high-resolution block belongs, so as to effectively supervise the upper space of the occluded region.

[0133] For regions far from the current vehicle, the number of valid points collected by the radar device of the current vehicle is too small, resulting in sparse valid voxels in the far region and the existence of false empty voxels. Or, after generating the visible region mask, it may be necessary to scale the visible region mask by means such as perspective sampling and downsampling; this may cause the space between ground voxels to be recognized as empty, and the false empty voxels generated in this way may be supervised; however, the actual position of the false empty voxels should be the ground, and theoretically, this position should not be supervised. To avoid the problem of unreasonable supervision caused by false empty voxels, the visible region mask can be further processed.

[0134] In some embodiments, in addition to the steps in the above embodiments, the method for generating a visible region mask for vehicle perception provided in the embodiments of the present disclosure further includes: in the spatial coordinate system corresponding to the voxel data, traversing the horizontal coordinates of the spatial coordinate system, and when there are no valid voxels in the voxels corresponding to the target horizontal coordinate, deleting the visible region mask with the lowest height in the visible region mask corresponding to the target horizontal coordinate.

[0135] In the spatial coordinate system, traverse all the horizontal coordinates. For ease of description, the traversed horizontal coordinate is referred to as the target horizontal coordinate, and the target horizontal coordinate is any one of the multiple horizontal coordinates in the spatial coordinate system.

[0136] For the voxels corresponding to the target horizontal coordinate, their horizontal coordinates are the same as the target horizontal coordinate. The voxels corresponding to the target horizontal coordinate may include empty voxels or valid voxels, and the visible target horizontal coordinate may correspond to multiple voxels.

[0137] If there are no valid voxels in the voxels corresponding to the target horizontal coordinate, it means that in the columnar region composed of the voxels corresponding to the target horizontal coordinate, the columnar region may be empty because the radar device fails to collect valid points, but theoretically, this columnar region should not be supervised. At this time, it is necessary to process the visible region mask at the position of this columnar region.

[0138] Specifically, first search in the visible region mask according to the target horizontal coordinate to obtain the visible region mask with the same horizontal coordinate as the target horizontal coordinate, that is, the visible region mask corresponding to the target horizontal coordinate. In the visible region mask corresponding to the target horizontal coordinate, determine the visible region mask with the lowest height according to the height coordinates of the visible region masks respectively, and delete the visible region mask with the lowest height.

[0139] The ground is usually located at the lowest position in the scene. Therefore, deleting the visible region mask with the lowest height can prevent the ground at this horizontal coordinate position from being supervised, thereby avoiding the unreasonable supervision problem of the ground caused by data sparsity or the scaling process of the visible region mask.

[0140] It can be understood that for the unreasonable supervision problem caused by scaling the visible region mask, the steps in this embodiment need to be executed after operations such as perspective sampling and downsampling are completed.

[0141] The above text combines Figures 1 to 6 and describes the method embodiments of the present disclosure in detail. Next, the device embodiments of the present disclosure will be described in combination with Figure 7 in detail. It should be understood that the descriptions of the method embodiments and the device embodiments correspond to each other. Therefore, the parts not described in detail can refer to the previous method embodiments.

[0142] Figure 7 The following is a schematic structural diagram of a visible area mask generation device for vehicle perception provided by an embodiment of the present disclosure. As Figure 7 shown, the visible area mask generation device 700 for vehicle perception according to an embodiment of the present disclosure includes: an acquisition module 710, a first determination module 720, a second determination module 730, a third determination module 740, and a generation module 750.

[0143] Specifically, the acquisition module 710 is configured to acquire voxel data corresponding to a target area collected by a current vehicle, and the voxel data includes valid voxels; the first determination module 720 is configured to determine an initial mask corresponding to the voxel data based on the voxel data and the position of the current vehicle; the second determination module 730 is configured to determine a foreground top mask corresponding to a foreground target based on a bounding box corresponding to the foreground target in the voxel data; the third determination module 740 is configured to determine a background top mask corresponding to a background target based on background voxels corresponding to the background target in the voxel data; the generation module 750 is configured to generate a visible area mask corresponding to the voxel data based on the initial mask, the foreground top mask, and the background top mask.

[0144] In some embodiments, the second determination module 730 is further configured to determine a projection of a bounding box corresponding to a foreground target in a horizontal direction in a spatial coordinate system corresponding to the voxel data, and determine a projected horizontal coordinate of a projection coverage range; determine a spatial position in the spatial coordinate system where the horizontal coordinate is the same as the projected horizontal coordinate and the height coordinate is greater than the height coordinate of the top of the bounding box as a foreground supervision area; and generate a foreground top mask based on the foreground supervision area.

[0145] In some embodiments, the third determination module 740 is further configured to respectively determine a highest valid voxel of each of a plurality of horizontal coordinates in a spatial coordinate system corresponding to the voxel data; perform category filtering on the highest valid voxels of each of the plurality of horizontal coordinates to obtain a plurality of first background voxels; generate a first height supervision mask at spatial positions where the plurality of first background voxels are located, and delete a partial mask that coincides with the foreground top mask in the first height supervision mask to obtain a second height supervision mask; determine a spatial position where empty voxels are located above the second height supervision mask as a background supervision area; and generate a background top mask based on the background supervision area.

[0146] In some embodiments, the first determination module 720 is further configured to, for each valid voxel, determine a connection line between the position of the current vehicle and the valid voxel; based on the voxel data, determine a plurality of target voxels through which the connection line passes, and the respective categories of the plurality of target voxels; when the respective categories of the plurality of target voxels meet the preset category conditions, determine the spatial positions where the plurality of target voxels are located and the spatial position where the valid voxel is located as the initial mask supervision region corresponding to the valid voxel; when the respective categories of the plurality of target voxels do not meet the preset category conditions, determine the spatial position where the valid voxel is located as the initial mask supervision region corresponding to the valid voxel; and generate an initial mask based on the initial mask supervision region corresponding to each valid voxel.

[0147] In some embodiments, the preset category conditions include: the respective categories of the plurality of voxels are empty voxels or noise voxels, and the noise voxels include dust voxels and / or rain and snow voxels.

[0148] In some embodiments, the generation device 700 for the visible region mask for vehicle perception further includes an occlusion region top mask generation module. The occlusion region top mask generation module is configured to, in the spatial coordinate system corresponding to the voxel data, partition the spatial positions corresponding to the voxel data to obtain a coarse resolution block and high resolution blocks, where the coarse resolution block is composed of a plurality of high resolution blocks; respectively determine the highest valid voxel in each coarse resolution block, and respectively determine the top height value corresponding to each coarse resolution block based on the height coordinate of the highest valid voxel in each coarse resolution block; traverse the high resolution blocks, and when the target high resolution block does not include valid voxels, determine the target coarse resolution block corresponding to the target high resolution block and the target top height value corresponding to the target coarse resolution block; determine the spatial positions in the spatial coordinate system where the horizontal coordinate is the same as the horizontal coordinate of the target high resolution block and the height coordinate is greater than the target top height value as the occlusion supervision region; and generate an occlusion region top mask based on the occlusion supervision region. The generation module 750 is further configured to take the union of the initial mask, the foreground top mask, the background top mask, and the occlusion region top mask to obtain the visible region mask.

[0149] In some embodiments, the generation device 700 for the visible region mask for vehicle perception further includes a sparse cropping module. The sparse cropping module is configured to, in the spatial coordinate system corresponding to the voxel data, traverse the horizontal coordinates of the spatial coordinate system, and when there are no valid voxels in the voxels corresponding to the target horizontal coordinate, delete the visible region mask with the lowest height in the visible region mask corresponding to the target horizontal coordinate.

[0150] Next, Figure 8 an electronic device according to an embodiment of the present disclosure will be described. Figure 8The following is a schematic structural diagram of an electronic device provided by an embodiment of the present disclosure. As Figure 8 shown, the electronic device 800 includes one or more processors 810 and a memory 820.

[0151] The processor 810 may be a central processing unit (CPU) or other forms of processing units with data processing capabilities and / or instruction execution capabilities, and may control other components in the electronic device 800 to perform desired functions.

[0152] The memory 820 may include one or more computer program products, and the computer program products may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. Volatile memory may include, for example, random access memory (RAM) and / or cache memory, etc. Non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc. One or more computer program instructions may be stored on the computer-readable storage medium, and the processor 810 may run the program instructions to implement the point cloud processing methods of various embodiments of the present disclosure described above and / or other desired functions.

[0153] In some embodiments, the electronic device 800 may further include: an input device 830 and an output device 840, and these components are interconnected through a bus system and / or other forms of connection mechanisms (not shown).

[0154] The input device 830 may include, for example, a touch screen, a microphone, a keyboard, a mouse, etc. The output device 840 may include, for example, a display, a speaker, and a communication network and its connected remote output devices, etc.

[0155] Of course, for simplicity, Figure 8 only some of the components related to the present disclosure in the electronic device 800 are shown, and components such as buses, input / output interfaces, etc. are omitted. In addition, according to specific application scenarios, the electronic device 800 may further include any other appropriate components.

[0156] In addition to the above methods and devices, the embodiments of the present disclosure may also be computer program products, which include computer program instructions that, when run by a processor, cause the processor to execute the steps in the method for generating a visible area mask for vehicle perception according to various embodiments of the present disclosure described above in this specification.

[0157] A computer program product may write program code for performing the operations of the embodiments of the present disclosure in any combination of one or more programming languages. The programming languages include object-oriented programming languages such as Java, C++, etc., and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computing device, partially on the user's device, executed as a stand-alone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.

[0158] In addition, an embodiment of the present disclosure may also be a computer-readable storage medium having computer program instructions stored thereon, and when the computer program instructions are run by a processor, the processor is caused to execute the steps in the method for generating a visible region mask for vehicle perception according to various embodiments of the present disclosure described above in this specification.

[0159] The computer-readable storage medium may employ any combination of one or more readable media. The readable media may be a readable signal medium or a readable storage medium. The readable storage medium may, for example, include but is not limited to an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (a non-exhaustive list) of the readable storage medium include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0160] The basic principles of the present disclosure have been described above in conjunction with specific embodiments. However, it should be noted that the advantages, benefits, effects, etc. mentioned in the present disclosure are only examples and not limitations, and it cannot be considered that these advantages, benefits, effects, etc. are essential for each embodiment of the present disclosure. In addition, the above-described specific details are only for the purposes of illustration and facilitating understanding, and are not limitations. The above details do not limit the present disclosure to necessarily adopt the above specific details for implementation.

[0161] The block diagrams of the devices, apparatuses, equipment, and systems involved in this disclosure are merely illustrative examples and are not intended to require or imply that they must be connected, arranged, and configured in the manner shown in the block diagrams. As those skilled in the art will recognize, these devices, apparatuses, equipment, and systems can be connected, arranged, and configured in any manner. Words such as "comprising," "including," "having," etc. are open-ended terms meaning "including but not limited to" and can be used interchangeably with each other. The word "or" and "and" used herein refer to the phrase "and / or" and can be used interchangeably with it, unless the context clearly indicates otherwise. The phrase "such as" used herein refers to the phrase "such as but not limited to" and can be used interchangeably with it.

[0162] It should also be noted that in the systems, equipment, and methods of this disclosure, each component or each step can be decomposed and / or recombined. These decompositions and / or recombinations should be regarded as equivalent solutions of this disclosure.

[0163] The above description of the disclosed aspects is provided to enable any person skilled in the art to make or use this disclosure. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein can be applied to other aspects without departing from the scope of this disclosure. Therefore, this disclosure is not intended to be limited to the aspects shown herein, but rather to the broadest scope consistent with the principles and novel features disclosed herein.

[0164] The above description has been given for purposes of illustration and description. In addition, this description is not intended to limit the embodiments of this disclosure to the forms disclosed herein. Although multiple example aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, changes, additions, and sub-combinations thereof.

Claims

1. A method for generating a visible area mask for vehicle perception, characterized in that: include: Acquire voxel data corresponding to the target area collected by the current vehicle, wherein the voxel data includes valid voxels; Based on the voxel data and the current position of the vehicle, determining an initial mask corresponding to the voxel data; Determining a foreground top mask corresponding to the foreground object based on a label box corresponding to the foreground object in the voxel data; Determining a background top mask corresponding to the background target based on background voxels corresponding to the background target in the voxel data; A visible area mask corresponding to the voxel data is generated based on the initial mask, the foreground top mask, and the background top mask.

2. The method according to claim 1, characterized in that The determining, based on the labeling box corresponding to the foreground object in the voxel data, a foreground top mask corresponding to the foreground object comprises: In the spatial coordinate system corresponding to the voxel data, determining the projection of the annotation box corresponding to the foreground object in the horizontal direction, and determining the projection horizontal coordinates of the projection coverage range; Determine a spatial position in the spatial coordinate system whose horizontal coordinate is consistent with the projection horizontal coordinate and whose height coordinate is greater than the height coordinate of the top of the annotation box as a foreground supervision area; Based on the foreground supervision region, the foreground top mask is generated.

3. The method according to claim 1, characterized in that: The determining, based on the background voxels corresponding to the background target in the voxel data, a background top mask corresponding to the background target, comprises: In the spatial coordinate system corresponding to the voxel data, respectively determine the highest valid voxel of each of the plurality of horizontal coordinates; Performing category filtering on the most significant voxels of each of the plurality of horizontal coordinates to obtain a plurality of first background voxels; Generate a first high-level supervision mask at the spatial position where the plurality of first background voxels are located, and delete a portion of the first high-level supervision mask that overlaps with the foreground top mask to obtain a second high-level supervision mask; determine the spatial position where the empty voxels located above the second high-level supervision mask are located as a background supervision area; Based on the background supervision region, a background top mask is generated.

4. The method according to claim 1, characterized in that: The determining, based on the voxel data and the current vehicle position, an initial mask corresponding to the voxel data comprises: For each of the valid voxels, determining a line connecting the position of the current vehicle and the valid voxel; Based on the voxel data, determining a plurality of target voxels through which the connecting line passes, and categories of each of the plurality of target voxels; In the case where the categories of the plurality of target voxels respectively meet the preset category conditions, determining the spatial positions of the plurality of target voxels and the spatial position of the valid voxel as the initial mask supervision region corresponding to the valid voxel; When the categories of the plurality of target voxels do not satisfy the preset category condition, determining the spatial position of the valid voxel as the initial mask supervision region corresponding to the valid voxel; The initial mask is generated based on the initial mask supervision region corresponding to each of the valid voxels.

5. The method according to claim 4, characterized in that The preset category condition includes: each category of the plurality of voxels is an empty voxel or a noise voxel, and the noise voxel includes a dust voxel and / or a rain and snow voxel.

6. The method according to claim 1, characterized in that Also includes: In the spatial coordinate system corresponding to the voxel data, the spatial position corresponding to the voxel data is divided into blocks to obtain a coarse resolution block and a high resolution block, wherein the coarse resolution block is composed of a plurality of the high resolution blocks; Determine the highest valid voxel in each coarse resolution block respectively, and determine the top height value corresponding to each coarse resolution block respectively based on the height coordinate of the highest valid voxel in each coarse resolution block; Traversing the high-resolution block, and determining a target coarse-resolution block corresponding to the target high-resolution block and a target top height value corresponding to the target coarse-resolution block when the target high-resolution block does not include valid voxels; Determine a spatial position in the spatial coordinate system whose horizontal coordinate is consistent with the horizontal coordinate of the target high-resolution block and whose height coordinate is greater than the height value of the top of the target as an occlusion supervision area; Based on the occluded supervision area, generating a top mask of the occluded area; The determining, based on the initial mask, the foreground top mask and the background top mask, a visible area mask corresponding to the voxel data comprises: The visible area mask is obtained by taking the union of the initial mask, the foreground top mask, the background top mask, and the occluded area top mask.

7. The method according to claim 1, characterized in that Also includes: In the spatial coordinate system corresponding to the voxel data, the horizontal coordinates of the spatial coordinate system are traversed, and when there are no valid voxels in the voxels corresponding to the target horizontal coordinates, the visible area mask with the lowest height in the visible area mask corresponding to the target horizontal coordinates is deleted.

8. A device for generating a visible area mask for vehicle perception, characterized in that: include: An acquisition module is configured to acquire voxel data corresponding to the target area currently collected by the vehicle, wherein the voxel data includes valid voxels; A first determination module is configured to determine an initial mask corresponding to the voxel data based on the voxel data and the current vehicle position; A second determination module is configured to determine a foreground top mask corresponding to the foreground object based on a label box corresponding to the foreground object in the voxel data; A third determination module is configured to determine a background top mask corresponding to the background target based on background voxels corresponding to the background target in the voxel data; A generating module is configured to generate a visible area mask corresponding to the voxel data based on the initial mask, the foreground top mask and the background top mask.

9. An electronic device, characterized in that: include: processor; as well as A memory, configured to store executable instructions of the processor; Wherein, the processor is configured to execute the method for generating a visible area mask for vehicle perception as described in any one of claims 1 to 7 by executing the executable instructions.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for generating a visible area mask for vehicle perception according to any one of claims 1 to 7 is implemented.