Task alignment based bev feature dynamic sample assignment method and system

CN122841935APending Publication Date: 2026-09-29TIANJIN QINGZHI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611102541.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-23
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

[0009]本发明的一个目的在于,解决现有静态高斯分配策略中分配规则僵化导致低质量样本过多、分类与回归任务目标不一致、正样本数量过大且冗余及多目标冲突处理粗糙的问题

Benefits of technology

[0036]基于前文的描述,本领域技术人员能够理解的是,本发明通过获取BEV特征图和真实目标框,通过检测头预测各网格点的分类结果和3D预测框;再根据真实目标框在BEV空间的投影范围确定候选区域,筛选候选点集;对每个候选点,融合其分类预测得分与定位匹配度,计算任务对齐得分;然后,按任务对齐得分从候选点集中选取Top-K个正样本;当同一网格点被多个真实目标框选为正样本时,将其分配给定位匹配度最高的真实目标框,其余网格点为负样本;最后,根据分配结果,构建与任务对齐得分正相关的分类训练目标和回归训练目标,基于分类损失和回归损失迭代更新检测头参数。本发明通过任务对齐得分实现分类与定位的联合优化,提升检测精度;通过动态Top-K选取实现正样本的自适应控制,减少冗余并改善样本平衡;通过定位匹配度驱动的冲突解决机制,提升密集场景下的分配合理性;整体方法与主流BEV检测框架兼容,支持端到端训练。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122841935A_ABST
    Figure CN122841935A_ABST
Patent Text Reader

Abstract

The application belongs to the technical field of intelligent vehicles, and relates to a BEV feature dynamic sample distribution method and system based on task alignment. The BEV feature map and the real target frame are obtained to predict the classification result and the 3D prediction frame of each grid point. The candidate region is determined according to the projection range of the real target frame in the BEV space, and the candidate point set is screened. The task alignment score is calculated by fusing each candidate point. The positive and negative samples are divided according to the task alignment score. The classification training target and the regression training target positively correlated with the task alignment score are constructed, and the detection head parameters are updated based on the classification training target and the regression training target. The application aims to solve the problems of excessive low-quality samples caused by rigid allocation rules in the existing static Gaussian allocation strategy, inconsistent classification and regression task targets, excessive and redundant positive samples, and rough multi-target conflict processing. The application can realize joint optimization of classification and positioning, adaptive control of positive samples, improve the rationality of dense scene distribution, and support end-to-end training.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of intelligent vehicle technology, specifically relating to a method and system for dynamic sample allocation of BEV features based on task alignment. Background Technology

[0002] Autonomous driving environmental perception systems require real-time and accurate detection of 3D targets around the vehicle. Multimodal fusion methods based on bird's-eye view (BEV) have become the mainstream approach. This involves projecting camera images and LiDAR point clouds onto the BEV space to extract features and output 3D detection results. Sample allocation is a crucial step in the training process, directly impacting detection accuracy and convergence speed.

[0003] Existing methods typically employ a static assignment strategy based on Gaussian heatmaps. This involves calculating a Gaussian radius for each ground truth target, centered at its BEV center point and calculated according to the target's size. Within this radius, all grid points are considered positive samples. Classification targets are assigned values ​​using a two-dimensional Gaussian function, while regression targets receive the same ground truth (GT) encoding values. However, this method still has the following drawbacks:

[0004] First, the assignment rules are rigid, resulting in a large number of low-quality samples. The IoU between the predicted bounding boxes and the ground truth bounding boxes of Gaussian radius edge points is often very low, and forcing these points to learn regression tasks will introduce noise. At the same time, the radius size is determined only by the target size and is unrelated to the model's current predictive ability, so a large number of points that are not suitable as positive samples are still forcibly assigned in the early stages of training;

[0005] Second, the objectives of classification and regression tasks are inconsistent. Classification aims to fit a Gaussian distribution; as long as a point falls within the radius, even poor regression predictions result in minimal classification loss. Regression tasks, however, require high localization quality from the samples. This inconsistency leads to the model generating bounding boxes with high classification confidence but low localization accuracy, impacting the final detection accuracy.

[0006] Third, an excessive number of positive samples exacerbates the imbalance between positive and negative samples. The Gaussian radius of a large target typically contains hundreds of positive sample points, most of which are redundant, increasing computational overhead and easily leading to overfitting.

[0007] Fourth, there is a lack of effective handling of multi-target conflicts. The radii of nearby targets may overlap, and the same grid point may be assigned as a positive sample to multiple ground truths (GTs). Existing methods simply assign it to the GT with the largest IoU without considering the potential value of that point to other GTs.

[0008] In view of this, the present invention is hereby proposed. Summary of the Invention

[0009] One objective of this invention is to address the problems in existing static Gaussian allocation strategies, such as rigid allocation rules leading to an excessive number of low-quality samples, inconsistencies between classification and regression task objectives, an excessively large and redundant number of positive samples, and coarse handling of multi-objective conflicts.

[0010] To achieve the above objectives, this invention provides a method for dynamic sample allocation of BEV features based on task alignment, comprising:

[0011] Obtain BEV feature maps and corresponding ground truth bounding boxes, and use a detection head to predict the BEV feature maps to obtain classification prediction results and 3D prediction bounding boxes for each grid point;

[0012] Candidate regions are determined based on the projection range of each real target box in the BEV space, and grid points falling within the candidate regions are selected to form a candidate point set.

[0013] For each candidate point in the candidate point set, obtain its classification prediction score and the localization matching degree between the 3D prediction box and the real target box, and calculate the task alignment score that fuses classification and localization information.

[0014] For each real target bounding box, select several candidate points from its candidate point set according to the task alignment score as positive samples; when the same grid point is selected as a positive sample by multiple real target bounding boxes, assign the grid point to the real target bounding box with the highest localization matching degree, and the grid point not selected by any real target bounding box is a negative sample;

[0015] Based on the allocation results, a classification training objective and a regression training objective are constructed for each grid point; wherein, the classification training objective of positive samples is positively correlated with the task alignment score, and the regression training objective is the encoded value of the assigned true target box;

[0016] The total loss function is calculated based on the classification training objective and the regression training objective to iteratively update the parameters of the detection head.

[0017] Further, the step of predicting the BEV feature map using the detection head to obtain the classification prediction result and 3D prediction box for each grid point includes: acquiring the BEV feature map; wherein the BEV feature map is obtained by encoding sensor data by the BEV backbone network, and the sensor data includes at least LiDAR point cloud data and / or camera image data; acquiring the corresponding ground truth bounding box; wherein the ground truth bounding box is an annotated 3D bounding box containing center coordinates, size, yaw angle, and category information, and organized into a GT list in batches, corresponding one-to-one with the BEV feature map in the batch dimension; inputting the BEV feature map into the detection head, and outputting tensor information through forward inference; wherein the tensor information includes at least a classification score tensor, a center offset tensor, a height tensor, a size tensor, and a rotation angle tensor; for each grid point, decoding the tensor into a classification prediction result and a 3D prediction box; wherein... The step of decoding each grid point into a classification prediction result and a 3D prediction box based on the tensor includes: calculating the center coordinates of the prediction box: correcting the grid point coordinates to the target center point coordinates based on the center offset tensor and the physical resolution of the BEV feature map; calculating the height of the prediction box: directly assigning the target center height based on the height tensor; calculating the size of the prediction box: restoring the log-space encoded size to the physical size by performing an exponential operation on the size tensor; calculating the yaw angle of the prediction box: restoring the target yaw angle by performing an arctangent operation on the sine and cosine values ​​of the rotation angle tensor; assembling the target center coordinates, the size, and the yaw angle obtained above into a 3D prediction box instance according to the dimensions to obtain the 3D prediction box; extracting the prediction probability of the grid point belonging to each category from the classification score tensor S as the classification prediction result of the grid point.

[0018] Further, the step of determining candidate regions by the projection range of each real target box in the BEV space and selecting grid points falling within the candidate regions to form a candidate point set includes: scaling the center point of each real target box from physical world coordinates to BEV feature map coordinates; calculating its projection boundary on the BEV plane based on the scaled center point coordinates and the size of the real target box; applying a preset margin expansion amount around the projection boundary to obtain the boundary of the candidate region; selecting all grid points on the BEV feature map that fall within the boundary of the candidate region, and forming the candidate point set of the corresponding real target box by selecting all the selected grid points.

[0019] Further, the step of obtaining the classification prediction score and the localization matching degree between the 3D prediction box and the real target box for each candidate point in the candidate point set, and calculating the task alignment score that fuses classification and localization information, includes: obtaining the classification prediction result from the network output corresponding to the candidate point, wherein the classification prediction result represents the probability distribution of the candidate point belonging to various target categories; analyzing the classification prediction result and extracting the prediction probability corresponding to the target category as the classification prediction score sp of the candidate point; calculating the 3D intersection-union ratio up between the 3D prediction box of the candidate point and the real target box; wherein the 3D intersection-union ratio is calculated by multiplying the BEV plane intersection-union ratio and the height direction intersection-union ratio; and calculating the task alignment score alignp according to the following formula based on the classification prediction score sp and the 3D intersection-union ratio up: Where α and β are hyperparameters, and β > α > 0; where the height direction intersection-union ratio is... The calculation formula is:

[0020] ;

[0021] in, and These are the top and bottom coordinates of the 3D prediction bounding box in the height direction, respectively. and These are the top and bottom coordinates of the actual target bounding box in the height direction, respectively; The height is the intersection height of the 3D predicted bounding box and the real target bounding box in the height direction.

[0022] Furthermore, the step of selecting a number of candidate points from its candidate point set as positive samples for each real target box according to the task alignment score includes: for each real target box, sorting all candidate points in its candidate point set in descending order according to the task alignment score; determining whether the total number of candidate points in the candidate point set is less than a preset number K; if so, then all candidate points in the candidate point set are used as positive samples of the real target box; if not, then the first K sorted candidate points are selected as positive samples of the real target box.

[0023] Furthermore, the step of assigning the grid point to the real target box with the highest positioning matching degree when the same grid point is selected as a positive sample by multiple real target boxes, and classifying the grid point not selected by any real target box as a negative sample, includes: summarizing the positive sample allocation results of all real target boxes and establishing a global mapping table; traversing all grid points and identifying how many real target boxes selected each grid point as a positive sample; if a grid point is selected as a positive sample by only one real target box, it is directly retained; if a grid point is selected as a positive sample by multiple real target boxes simultaneously, it is assigned to the real target box with the highest positioning matching degree; if a grid point is not selected as a positive sample by any real target box, it is determined as a negative sample.

[0024] Furthermore, the formula for assigning the true target box with the highest positioning matching degree is:

[0025] ;

[0026] Where assign(p) is the allocation result of candidate point p, representing the actual target box to which the candidate point finally belongs; p is the candidate point to be assigned, corresponding to a certain grid or anchor point in the candidate point set; This refers to a real target bounding box in the current scene; To match the true target bounding box A set of candidate points with spatial relationships exists, where the spatial relationships include, but are not limited to, candidate points falling within the BEV range or 3D spatial range of the true target bounding box; p∈ This indicates that candidate point p belongs to the true target box. The associated candidate set, i.e., p and There is a potential matching relationship in space; For candidate point p relative to the ground truth bounding box The localization matching degree, i.e., the 3D intersection-over-union ratio, is used to measure the accuracy of the 3D predicted bounding box of p with the actual location of the target object. The degree of overlap in three-dimensional space.

[0027] Furthermore, the classification training objective for the positive samples is the normalized task alignment score, calculated using the following formula:

[0028] ;

[0029] Where (i,j) is the true target bounding box assigned to the location with the highest matching degree. Grid points; For grid point (i,j), the true bounding box with the highest localization match is... Task alignment score; The 3D predicted bounding box and the ground truth bounding box for grid point (i,j) 3D intersection-union ratio; To match the true target bounding box The highest task alignment score among all associated candidate points; and / or, the regression training objective includes: center offset, target height, logarithmic encoded values ​​of size, sine and cosine values ​​of yaw angle; and, the regression training objective also includes a velocity component.

[0030] The total loss function is obtained by weighted summation of classification loss and regression loss. The classification loss uses the zoom focus loss function, and the regression loss uses the Smooth L1 loss function or CIoU loss function, calculated only for positive samples. The formula for calculating the zoom focus loss (Varifocal Loss) is as follows:

[0031] ;

[0032] Where p is the classification prediction probability, q is the target quality, q is the normalized task alignment score for positive samples, and q is 0 for negative samples; the Smooth L1 loss function is calculated as follows:

[0033] ;

[0034] in, The set of positive samples; is the predicted regression value for the positive sample p; Let p be the target regression value for the positive sample.

[0035] In other embodiments, a task-aligned BEV feature dynamic sample allocation system is provided, capable of executing the task-aligned BEV feature dynamic sample allocation method described above. The system includes: a prediction and decoding module, used to acquire BEV feature maps and corresponding ground truth bounding boxes, perform prediction and decoding using a detection head, and output classification prediction results and 3D prediction boxes for each grid point; a candidate filtering module, used to determine candidate regions for each ground truth bounding box based on its projection range in the BEV space, and filter to obtain a candidate point set; an alignment score calculation module, used to calculate a task alignment score that fuses classification and localization information for each candidate point in the candidate point set; a positive sample selection module, used to select several candidate points from the candidate point set as positive samples for each ground truth bounding box according to its task alignment score; a conflict resolution module, used to allocate the grid point to the ground truth bounding box with the highest localization matching degree when the same grid point is selected as a positive sample by multiple ground truth bounding boxes; and a target construction module, used to construct classification training targets and regression training targets for each grid point based on the allocation results.

[0036] Based on the foregoing description, those skilled in the art will understand that this invention acquires BEV feature maps and ground truth bounding boxes, predicts the classification results and 3D prediction boxes for each grid point using a detection head, determines candidate regions based on the projection range of the ground truth bounding boxes in the BEV space, and filters the candidate point set. For each candidate point, its classification prediction score and localization matching degree are fused to calculate the task alignment score. Then, the Top-K positive samples are selected from the candidate point set according to the task alignment score. When the same grid point is selected as a positive sample by multiple ground truth bounding boxes, it is assigned to the ground truth bounding box with the highest localization matching degree, and the remaining grid points are negative samples. Finally, based on the allocation results, classification training objectives and regression training objectives positively correlated with the task alignment score are constructed, and the detection head parameters are iteratively updated based on the classification loss and regression loss. This invention achieves joint optimization of classification and localization through task alignment score, improving detection accuracy; achieves adaptive control of positive samples through dynamic Top-K selection, reducing redundancy and improving sample balance; improves the allocation rationality in dense scenes through a conflict resolution mechanism driven by localization matching degree; the overall method is compatible with mainstream BEV detection frameworks and supports end-to-end training. Attached Figure Description

[0037] The accompanying drawings, as part of this invention, are used to provide a further understanding of the invention. The illustrative embodiments and descriptions of the invention are used to explain the invention, but do not constitute an undue limitation of the invention. Obviously, the drawings described below are merely some embodiments, and those skilled in the art can obtain other drawings based on these drawings without creative effort. In the drawings:

[0038] Figure 1 This is a flowchart of a BEV feature dynamic sample allocation method based on task alignment in some embodiments of the present invention. Detailed Implementation

[0039] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments will be clearly and completely described below with reference to the accompanying drawings. The following embodiments are used to illustrate the present invention, but are not intended to limit the scope of the present invention.

[0040] Those skilled in the art should understand that the embodiments described below are merely a part of the embodiments of the present invention, and not all of the embodiments of the present invention. These partial embodiments are intended to explain the technical principles of the present invention and are not intended to limit the scope of protection of the present invention. Based on the embodiments provided by the present invention, all other embodiments obtained by those skilled in the art without creative effort should still fall within the scope of protection of the present invention.

[0041] The following reference Figure 1This document will provide a detailed description of the task-aligned dynamic sample allocation method for BEV features in some embodiments of the present invention. Figure 1 This is a flowchart of a BEV feature dynamic sample allocation method based on task alignment in some embodiments of the present invention.

[0042] like Figure 1 As shown, in some embodiments of the present invention, a method for dynamic sample allocation of BEV features based on task alignment is provided, including:

[0043] Step S110 involves acquiring the BEV (Bird's Eye View) feature map and its corresponding ground truth bounding boxes. The BEV feature map is then used by a detection head to predict the classification results and 3D bounding boxes for each grid point. For each training batch, sensor data (point cloud / image) and corresponding annotations are read. The BEV feature map is extracted through BEV encoding and a backbone network. Simultaneously, the ground truth bounding boxes are unified in coordinate system, cropped, filtered, and organized into a structured list of ground truth bounding boxes, aligning them in batch size and spatial extent. This list serves as input for subsequent prediction, candidate region generation, and sample allocation. Specifically, step S110 includes:

[0044] Step S111: Obtain the BEV feature map. The BEV feature map is obtained by encoding sensor data through the BEV backbone network. The sensor data includes LiDAR point cloud data and / or camera image data.

[0045] Specifically, for each frame of the scene, LiDAR point cloud and / or camera image are read, cropped, coordinate unified, and BEV representation constructed (point cloud voxelization / Pillarization or image back-projection fusion) to obtain initial BEV features. , This is a BEV feature map that has not undergone deep backbone processing and is then forward encoded by a BEV backbone network such as CNN / FPN / Transformer, outputting a BEV feature map with rich semantics. , C represents the number of feature channels; Indicates the height of the BEV feature map (number of rows in the y direction); Represents the width of the BEV feature map (number of columns in the x-direction) and records its physical resolution. The sensing range and coordinate system serve as the basic inputs for subsequent detection head prediction and sample allocation, while the corresponding real target boxes are acquired in parallel for training supervision.

[0046] Step S112: Obtain the ground truth bounding boxes corresponding to the BEV feature maps. The ground truth bounding boxes are labeled 3D bounding boxes, which include center coordinates, size, yaw angle and category information, and are organized into a GT list by batch, corresponding one-to-one with the BEV feature maps in the batch dimension.

[0047] Specifically, for each frame of the scene, the original 3D annotations are read, and the center coordinates, size, yaw angle, and category information of the real target boxes are parsed and unified to the same vehicle coordinate system and physical units as the BEV feature map. Boxes that are outside the perception range or have abnormal sizes are filtered out, and the remaining boxes are constructed into structured ground truth representations and organized into a list G by frame. b G b This is the list of ground truth (GT) boxes for frame b. Ultimately, the entire GT list G = {G0, ..., Gb} is generated along the batch dimension. B−1} and BEV feature map This one-to-one correspondence provides accurate and aligned supervision information for subsequent candidate region determination, 3D IoU calculation, task alignment score, and training target construction.

[0048] Step S113: Input the BEV feature map into the detection head, and output tensor information through forward inference.

[0049] The tensor information includes at least the classification score tensor, center offset tensor, height tensor, size tensor, and rotation angle tensor. The classification score tensor S... After Sigmoid activation, the probability values ​​of each grid point belonging to each category are obtained; the center offset tensor O, , used to characterize the offset of a grid point from the target center point; height tensor Z, , used to characterize the height of the target center; size tensor D, The dimensions are used to characterize the length, width, and height of the target, and are encoded in logarithmic space; the rotation angle tensor R, These five tensors are used to characterize the sine and cosine values ​​of the target yaw angle. Aligned in B, H, and W spaces, they are decoded and then assembled dimensionally into 3D prediction boxes for each grid point. This provides a predictive basis for subsequent candidate point sets, task alignment scores, and sample allocation.

[0050] Step S114: For each grid point, decode the tensor into a classification prediction result and a 3D prediction bounding box. Specifically, step S114 includes:

[0051] Step S1141: For each grid point (i,j), calculate the core parameters of the prediction box.

[0052] The core parameters include at least the center coordinates, altitude, dimensions, and yaw angle.

[0053] (1) Calculate the center coordinates: Based on the center offset tensor and the physical resolution of the BEV feature map, correct the grid point coordinates to the target center point coordinates. The calculation process is as follows:

[0054] First, based on the grid index (i,j) and the BEV physical resolution... Calculate the base coordinates (j) of the grid points in physical space. i Then, extract the offset of that point from the center offset tensor. , ), multiply it by The mapping is converted to a physical offset, and finally the two are added together to obtain the corrected target center point coordinates. ).

[0055] Specifically, ; .in, This represents the physical resolution of the BEV feature map, in meters per pixel. and These are the values ​​of the center offset tensor O at grid point (i,j), representing the offset of the grid point towards the target center point.

[0056] Optionally, the physical resolution of the BEV feature map The value ranges from 0.1 m / pixel to 0.5 m / pixel.

[0057] (2) Calculate the prediction box height: The target center height is directly assigned based on the height tensor. The calculation process is as follows:

[0058] For each grid point (i,j) on the BEV feature map, the scalar prediction value of the corresponding position Z[0,i,j] is read from the height tensor Z with a height of 1 channel, and it is directly used as the target center height z of the 3D prediction box corresponding to that grid point. c Without scaling or nonlinear transformation, it provides physically consistent Z-axis center coordinates when assembled with other parameters such as center offset, size, and rotation angle, i.e., the height of the prediction box is... ,in, , The value of the height tensor z at grid point (i,j) is directly used as the height of the target center.

[0059] (3) Calculate the predicted box size: By performing an exponential operation on the size tensor, the size encoded in the logarithmic space is restored to the physical size. The calculation process is as follows:

[0060] For each grid point (i,j) on the BEV feature map, the log-space size value predicted by the network is read from the three channels of the size tensor D. , , Then, the natural exponent operation is applied to each logarithmic dimension to restore it to a positive physical dimension. ), that is, the size of the prediction box is ( ),in, ; , and These are the values ​​of the size tensor D at grid point (i,j), which are then converted back to the physical size through exponential operations in logarithmic space.

[0061] (4) Calculate the yaw angle of the predicted box: The yaw angle of the target is restored by performing arctangent operation on the sine and cosine values ​​of the rotation angle tensor. The calculation process is as follows:

[0062] For each grid point (i,j) on the BEV feature map, read the predicted sine value R[0,i,j]=sinθ from the two channels of the rotation angle tensor R. ij The cosine value R[1,i,j]=cosθ ij Then, these two are passed as input to the four-quadrant arctangent function θpred=atan2(sinθ ij cosθ ij This is used to reduce the target's orientation in the BEV plane to a continuous, full-circumference yaw angle θpred∈(−π,π], which is then assembled with other parameters into a complete 3D prediction bounding box. That is, the yaw angle of the prediction bounding box is... ,in, s and These are the values ​​of the rotation angle tensor R at grid point (i,j), which are then converted to the yaw angle using arctangent operations.

[0063] Ultimately, the parameters involved in the 3D prediction bounding box include at least the center coordinates ( ),size( ), yaw angle .

[0064] Step S1142: Assemble the parameters of all predicted boxes calculated for each grid point (i,j) into a predicted box instance according to its dimensions to obtain the 3D predicted box. This achieves the mapping from BEV feature map grid points to spatial 3D prediction boxes. Specifically, the center coordinates of each grid point (i,j) are calculated... ),size( ), yaw angle Discrete parameters, in a predefined fixed-dimensional order, such as [ , , Fill in the corresponding positions in the prediction box array / tensor sequentially to assemble a complete 3D prediction box instance. This process is repeated for all grid points on the BEV feature map to complete the point-by-point mapping from BEV feature map grid points to spatial 3D prediction boxes, providing a unified prediction box representation for subsequent candidate point set construction, 3D IoU calculation and task alignment score.

[0065] Step S1143: Extract the predicted probabilities of grid points belonging to each category from the classification score tensor S, and use this as the classification prediction result for the grid points. Specifically, for each grid point (i,j) and its corresponding batch b on the BEV feature map, in the classification score tensor S, S∈R B×Ncls×H×W The position S[b,:,i,j] of the grid point (i,j) in batch b is located, where S[b,:,i,j] is the classification score vector of each category for grid point (i,j) in batch b. The length of this position is N. cls The vector is activated by applying Sigmoid activation as needed to obtain the predicted probability p of the grid point belonging to each category. i,j,c =σ(S[b,c,i,j]), where σ represents the Sigmoid activation. Organize this into a classification prediction result vector p. i,j , The classification prediction score used for subsequent extraction of task alignment scores The classification loss calculation and inference stage class determination are performed to achieve the mapping from the classification score tensor to the grid-by-grid point classification prediction result. Represents the true target bounding box Category tags.

[0066] Step S120 involves determining candidate regions based on the projection range of each real target bounding box in the BEV space, and selecting grid points falling within these candidate regions to form a candidate point set; wherein the size of the candidate region is positively correlated with the size of the real target bounding box. Specifically, step S120 includes:

[0067] Step S121: Scale the center point of each ground truth bounding box from physical world coordinates to BEV feature map coordinates. Specifically, for each ground truth bounding box... Project the center point of the ground truth bounding box onto the BEV feature map coordinates:

[0068] ;

[0069] in, This represents the physical resolution of the BEV feature map.

[0070] Step S122: Calculate the projected boundary of the true target box on the BEV plane based on the scaled center point coordinates and the dimensions of the true target box. Specifically, calculate the axis-aligned bounding box (projected boundary) of the true target box on the BEV plane using the following formula:

[0071] ;

[0072] ;

[0073] ;

[0074] ;

[0075] in, , , as well as Together they constitute the projection boundary of the real target box on the BEV plane.

[0076] Step S123: Apply a preset margin expansion amount around the projection boundary to obtain the boundary of the candidate region.

[0077] ;

[0078] ;

[0079] ;

[0080] ;

[0081] Where m is the margin expansion amount, preferably, m can be set to m = 2.0 / .

[0082] by , , as well as As the boundary of the candidate region, it is used to define the candidate point set. , where (i,j) are the grid point indices on the BEV feature map, thereby realizing the spatial determination from the GT projection range to the candidate region.

[0083] Step S124: Filter all grid points on the BEV feature map that fall within the candidate region boundary, and form a candidate point set for the corresponding ground truth bounding boxes from all the filtered grid points. Iterate through all grid points (i,j) on the BEV feature map, and determine the appropriate grid points based on the discrimination criteria. To determine whether a point falls within the candidate region, grid points that meet the conditions are added to the set one by one. After validity verification, a candidate point set corresponding to the true target box g is formed. This enables the mapping from continuous candidate regions to a discrete grid point candidate set, providing a spatially ranged set of candidate points for subsequent task alignment score calculation and positive sample selection.

[0084] Step S130: For each candidate point in the candidate point set, obtain its classification prediction score and the localization matching degree between the 3D predicted bounding box and the real target bounding box, and calculate the task alignment score that fuses classification and localization information. Specifically, step S130 includes:

[0085] Step S131: Obtain the classification prediction result from the network output corresponding to the candidate point. The classification prediction result represents the probability distribution of the candidate point belonging to each type of target. Analyze the classification prediction result and extract the prediction probability corresponding to the target category as the classification prediction score s of the candidate point. p Classification prediction score s p The score is determined based on the classification prediction results of the candidate points. These results are output by the neural network classification head and represent the probability that a candidate point belongs to each target category. The probability value corresponding to the target category is selected as the classification prediction score s. p .

[0086] Specifically, for each candidate point p∈Cg, perform the following operations:

[0087] Obtain the classification score: Obtain the classification prediction score of candidate point p. , ,

[0088] Where S is the classification score tensor output by the detection head; b is the index of the current batch; i,j are the grid coordinates of candidate point p on the BEV feature map;

[0089] Obtain the prediction bounding box: Obtain the 3D prediction bounding box corresponding to the candidate point p. The predicted bounding box is obtained by decoding the BEV feature map by the detection head.

[0090] Step S132: Calculate the 3D intersection-union ratio (u) between the 3D predicted bounding boxes of candidate points and the ground truth bounding boxes. p The 3D crossover ratio is calculated by multiplying the BEV plane crossover ratio by the height crossover ratio.

[0091] Calculate 3D IoU: Calculate the predicted bounding box The 3D intersection-union ratio u between the target bounding box g and the ground truth bounding box g p The calculation formula is:

[0092] ;

[0093] Among them, IoU BEV The intersection-union ratio of the predicted bounding box and the ground truth bounding box on the BEV plane is calculated by the intersection area of ​​the rotation matrix, taking into account the orientation angle θg of the ground truth bounding box.

[0094] Among them, IoU heightThe intersection-union ratio (IUGR) of the predicted bounding box and the ground truth bounding box along the height direction (Z-axis) is calculated using the following formula:

[0095] ;

[0096] in, and These are the top and bottom Z coordinates of the 3D prediction bounding box in the height direction, respectively; and These are the top and bottom Z coordinates of the actual target bounding box in the height direction, respectively; This represents the intersection height of the 3D predicted bounding box and the ground truth bounding box in the height direction.

[0097] In some examples, .

[0098] Step S133, predict the score s based on the classification. p Intersection with 3D and comparison u p Calculate the task alignment score using the following formula: p :

[0099] ;

[0100] Wherein, α and β are hyperparameters, and β > α > 0. Preferably, the hyperparameter α is 1.0 and the hyperparameter β is 6.0.

[0101] Step S140: For each real target box, select several candidate points from its candidate point set according to the task alignment score as positive samples; when the same grid point is selected as a positive sample by multiple real target boxes, assign the grid point to the real target box with the highest localization matching degree, and the grid point not selected by any real target box is a negative sample.

[0102] Specifically, step S140, "for each real target bounding box, select several candidate points from its candidate point set as positive samples based on the task alignment score," includes:

[0103] Step S141: For each ground truth bounding box g, sort all candidate points in its corresponding candidate point set Cg in descending order of task alignment score. (Task alignment score) p The calculation formula is as before, where p∈Cg.

[0104] Step S142: Determine whether the total number of candidate points in the candidate point set Cg is less than the preset number K.

[0105] The range of values ​​for K can be determined based on the operator's experience or through multiple experiments, and no specific limit is set here.

[0106] Step S143: If yes, then all candidate points in the candidate point set Cg are treated as positive samples of the real target box g, forming a positive sample set Pg=Cg.

[0107] Step S144: If not, select the top K candidate points after sorting as positive samples of the true target box g, forming a positive sample set Pg={p1,p2,…,p K}, where p1, p2, ..., p K The top K candidate points are sorted in descending order of task alignment score.

[0108] It should be noted that the number of candidate points in the positive sample set changes dynamically during the model training process: in the early stage of training, positive samples are concentrated in the region with relatively high prediction accuracy, and in the later stage of training, positive samples gradually converge to the region with higher prediction accuracy.

[0109] Specifically, step S140, "when the same grid point is selected as a positive sample by multiple ground truth bounding boxes, the grid point is assigned to the ground truth bounding box with the highest localization matching degree, and the grid point not selected by any ground truth bounding box is a negative sample," includes:

[0110] Step S145: Summarize the positive sample assignment results of all real target boxes and establish a global mapping table;

[0111] Step S146: Traverse all grid points and identify how many real targets have selected each grid point as a positive sample;

[0112] Step S147: If a grid point is selected as a positive sample by only one real target box, it is directly retained;

[0113] Step S148: If a grid point is selected as a positive sample by multiple real target boxes at the same time, then it is assigned to the real target box with the highest localization matching degree.

[0114] The formula for assigning the true target bounding box with the highest localization match is:

[0115] ;

[0116] Where assign(p) is the assignment result of candidate point p, representing the final target box to which the candidate point belongs; p is the candidate point to be assigned, corresponding to a grid or anchor point in the candidate point set; This is a real target bounding box in the current scene; To match the true target bounding box A set of candidate points with spatial relationships, including but not limited to candidate points falling within the BEV range or 3D space range of the ground truth bounding box; p∈ This indicates that candidate point p belongs to the true target box. The associated candidate set, i.e., p and There is a potential matching relationship in space; For candidate point p relative to the ground truth bounding box The localization matching degree, or 3D intersection-union ratio, is used to measure the accuracy of the 3D predicted bounding box of p with the actual localization of p. The degree of overlap in three-dimensional space.

[0117] Step S149: If a grid point is not selected as a positive sample by any real target, then it is determined as a negative sample.

[0118] Step S150: Based on the allocation results, construct classification training objectives and regression training objectives for each grid point; wherein, the classification training objective for positive samples is positively correlated with the task alignment score, and the regression training objective is the encoded value of the assigned true bounding box. Specifically, step S150 includes:

[0119] Step S151: For the real target bounding box assigned to the location with the highest matching degree... For grid points (i,j), construct the classification training target vector and the regression training target vector:

[0120] (1) Constructing the classification training target vector , Assign the normalized alignment score to the dimension corresponding to the category of the true target bounding box, i.e., the score at the 1st position. The value on dimension is the normalized alignment score w. ij The remaining dimensions are zero; the normalized alignment score is determined based on the task alignment score of the grid point relative to the ground truth bounding box, the intersection-union ratio (IoU) of the predicted bounding box and the ground truth bounding box, and the largest task alignment score among all candidate points associated with the ground truth bounding box. That is, the normalized alignment score. The calculation formula is:

[0121] ;

[0122] in, For grid point (i,j) relative to the ground truth bounding box Task alignment score; The 3D predicted bounding box and the ground truth bounding box for grid point (i,j) 3D intersection-union ratio; To match the true target bounding box The highest task alignment score among all associated candidate points.

[0123] (2) Constructing the regression training target vector The regression training target vector contains the geometric difference encoding between grid points and the ground truth bounding boxes. The geometric difference encoding includes at least: position offset, size encoding, orientation encoding, and velocity encoding.

[0124] Specifically, the position offset vector includes the center offset vector and the height, and the formula for calculating the center offset is: (Δx, Δy) = (c x -j,c y -i), where (c x ,c y () represents the true target bounding box The center coordinates; the formula for calculating the height is: That is, the true target bounding box The height value.

[0125] Size coding, also known as size logarithmic coding, is calculated using the following formula: ( , , );in,( , , () represents the true target bounding box The three-dimensional dimensions.

[0126] Orientation encoding, also known as rotation angle encoding, is calculated using the following formula: ;in, For the true target bounding box Yaw angle.

[0127] Velocity encoding, or velocity component, is calculated using the following formula: ;in, For the true target bounding box The velocity vector.

[0128] Step S152: For grid points (i,j) that are not assigned to any real target boxes, all dimensions of their classification training target vector are zero, and they are not included in the regression loss calculation.

[0129] Step S160: Calculate the total loss function based on the classification training objective and the regression training objective to iteratively update the parameters of the detection head.

[0130] Specifically, step S160 includes: the total loss function is obtained by weighted summation of classification loss and regression loss, and the calculation formula is as follows: Preferably, .

[0131] In some specific embodiments, the classification loss uses the zoom focus loss, calculated as follows:

[0132] ;

[0133] Where p is the classification prediction probability, q is the target quality, and the value of q for positive samples is the normalized task alignment score, i.e., w. ijThe q value for negative samples is 0, and α and γ are hyperparameters used to balance the weights of positive and negative samples and easy and difficult samples.

[0134] Preferably, α=0.75 is the hyperparameter for balancing the losses of positive and negative samples; γ=2.0 is the hyperparameter for adjusting the weights of easy and difficult samples.

[0135] The classification loss function described above enables the model to focus on high-quality positive samples and difficult negative samples, and directly optimizes the consistency between classification score and localization quality.

[0136] In some specific embodiments, the regression loss uses the Smooth L1 loss function or the CIoU loss function and is calculated only for positive samples. This embodiment uses the Smooth L1 loss function as an example for illustration:

[0137] The Smooth L1 loss function is calculated based on the difference between the predicted and target regression values. It uses a squared term for small errors and an absolute value term for large errors to balance gradient stability and robustness. The formula for the Smooth L1 loss function is:

[0138] ;

[0139] in, The set of positive samples; is the predicted regression value for the positive sample p; Let p be the target regression value for the positive sample.

[0140] In other embodiments, the regression loss may optionally employ CIoU loss to further optimize the geometric consistency between the predicted bounding box and the ground truth bounding box. CIoU loss, based on the intersection-union ratio loss, comprehensively considers the center distance between the predicted bounding box and the ground truth bounding box, the diagonal length of the minimum bounding box containing both boxes, and the aspect ratio consistency of the two boxes to more comprehensively evaluate the localization quality.

[0141] It should be noted that the method re-executes the steps of candidate region determination, task alignment score calculation, positive sample selection, conflict resolution, and training target construction in each batch of model training, so that the sample allocation results are dynamically optimized as the model's predictive ability improves.

[0142] In other embodiments of the present invention, a task-aligned dynamic sample allocation system for BEV features is provided, which can execute the task-aligned dynamic sample allocation method for BEV features described above.

[0143] The system includes a prediction decoding module, a candidate selection module, an alignment score calculation module, a positive sample selection module, a conflict resolution module, and a target construction module. The prediction decoding module acquires BEV feature maps and corresponding ground truth bounding boxes, performs prediction and decoding using a detection head, and outputs the classification prediction results and 3D prediction boxes for each grid point. The candidate selection module determines candidate regions for each ground truth bounding box based on its projection range in the BEV space and selects a set of candidate points. The alignment score calculation module calculates a task alignment score that fuses classification and localization information for each candidate point in the candidate point set. The positive sample selection module selects several candidate points from the candidate point set for each ground truth bounding box as positive samples based on their task alignment scores. The conflict resolution module assigns the grid point to the ground truth bounding box with the highest localization matching degree when multiple ground truth bounding boxes select it as a positive sample. The target construction module constructs classification and regression training targets for each grid point based on the allocation results, where the classification target of the positive sample is positively correlated with the task alignment score.

[0144] In other embodiments of the present invention, a computer is also provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the relevant steps of the task-aligned BEV feature dynamic sample allocation method described above.

[0145] In other embodiments of the present invention, a computer-readable storage medium is also provided, on which a computer program is stored, the computer program being executed by a processor of the task-aligned BEV feature dynamic sample allocation method described above.

[0146] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, prediction models, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0147] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0148] Those skilled in the art will understand that this invention acquires BEV feature maps and ground truth bounding boxes, predicts the classification results and 3D prediction boxes for each grid point using a detection head, determines candidate regions based on the projection range of the ground truth bounding boxes in the BEV space, and filters the candidate point set. For each candidate point, its classification prediction score and localization matching degree are fused to calculate the task alignment score. Then, the Top-K positive samples are selected from the candidate point set according to the task alignment score. When the same grid point is selected as a positive sample by multiple ground truth bounding boxes, it is assigned to the ground truth bounding box with the highest localization matching degree, and the remaining grid points are negative samples. Finally, based on the allocation results, classification training objectives and regression training objectives positively correlated with the task alignment score are constructed, and the detection head parameters are iteratively updated based on the classification loss and regression loss. This invention effectively solves the problems in existing static Gaussian allocation strategies, such as rigid allocation rules leading to too many low-quality samples, inconsistencies between classification and regression task objectives, excessive and redundant positive samples, and coarse handling of multi-objective conflicts. This method achieves joint optimization of classification and localization through task alignment scores, thereby improving detection accuracy; it achieves adaptive control of positive samples through dynamic Top-K selection, reducing redundancy and improving sample balance; and it improves the rationality of allocation in dense scenes through a conflict resolution mechanism driven by localization matching degree. The overall method is compatible with mainstream BEV detection frameworks and supports end-to-end training.

[0149] Those skilled in the art will understand that modules in the device of the embodiments can be adaptively changed and placed in one or more devices different from that embodiment. Modules, units, or components in the embodiments can be combined into a single module, unit, or component, and further, they can be divided into multiple sub-modules, sub-units, or sub-components. Except where at least some of such features and / or processes or units are mutually exclusive, any combination can be used to combine all features disclosed in this specification (including the accompanying claims, abstract, and drawings) and all processes or units of any method or device so disclosed. Unless expressly stated otherwise, each feature disclosed in this specification (including the accompanying claims, abstract, and drawings) may be replaced by an alternative feature that serves the same, equivalent, or similar purpose.

[0150] Furthermore, those skilled in the art will understand that although some embodiments herein include certain features included in other embodiments but not others, combinations of features from different embodiments are intended to be within the scope of the invention and form different embodiments. For example, in the claims, any of the claimed embodiments can be used in any combination.

[0151] The above are merely preferred embodiments of the present invention and are not intended to limit the present invention in any way. Although the present invention has been disclosed above with reference to preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make some modifications or alterations to the above-described technical content to create equivalent embodiments without departing from the scope of the present invention. The implementation schemes in the above embodiments can be further combined or replaced. Any simple modifications, equivalent changes and alterations made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the scope of the present invention.

Claims

1. A dynamic sample allocation method for BEV features based on task alignment, characterized in that, include: Obtain BEV feature maps and corresponding ground truth bounding boxes, and use a detection head to predict the BEV feature maps to obtain classification prediction results and 3D prediction bounding boxes for each grid point. Candidate regions are determined based on the projection range of each real target box in the BEV space, and grid points falling within the candidate regions are selected to form a candidate point set. For each candidate point in the candidate point set, obtain its classification prediction score and the localization matching degree between the 3D prediction box and the real target box, and calculate the task alignment score that fuses classification and localization information. For each real target bounding box, select several candidate points from its candidate point set according to the task alignment score as positive samples; when the same grid point is selected as a positive sample by multiple real target bounding boxes, assign the grid point to the real target bounding box with the highest localization matching degree, and the grid point not selected by any real target bounding box is a negative sample; Based on the allocation results, a classification training objective and a regression training objective are constructed for each grid point; wherein, the classification training objective of positive samples is positively correlated with the task alignment score, and the regression training objective is the encoded value of the assigned true target box; The total loss function is calculated based on the classification training objective and the regression training objective to iteratively update the parameters of the detection head.

2. The method according to claim 1, characterized in that, The step of predicting the BEV feature map using a detection head to obtain the classification prediction results and 3D prediction boxes for each grid point includes: The BEV feature map is obtained by encoding sensor data by the BEV backbone network, and the sensor data includes at least lidar point cloud data and / or camera image data. Obtain the ground truth bounding box corresponding to the BEV feature map; wherein the ground truth bounding box is an annotated 3D bounding box, which includes center coordinates, size, yaw angle and category information, and is organized into a GT list in batches, corresponding one-to-one with the BEV feature map in the batch dimension; The BEV feature map is input into the detection head, and tensor information is output through forward inference; wherein, the tensor information includes at least the classification score tensor, center offset tensor, height tensor, size tensor, and rotation angle tensor; For each grid point, the tensor is decoded into a classification prediction result and a 3D prediction box. The step of decoding each grid point into a classification prediction result and a 3D prediction box based on the tensor includes: calculating the center coordinates of the prediction box: correcting the grid point coordinates to the target center point coordinates based on the center offset tensor and the physical resolution of the BEV feature map; calculating the height of the prediction box: directly assigning the target center height based on the height tensor; calculating the size of the prediction box: restoring the log-space encoded size to the physical size by performing an exponential operation on the size tensor; calculating the yaw angle of the prediction box: restoring the target yaw angle by performing an arctangent operation on the sine and cosine values ​​of the rotation angle tensor; assembling the target center coordinates, the size, and the yaw angle obtained above into a 3D prediction box instance according to the dimensions to obtain the 3D prediction box; and extracting the prediction probability of the grid point belonging to each category from the classification score tensor S as the classification prediction result of the grid point.

3. The method according to claim 1, characterized in that, The step of determining candidate regions by the projection range of each real target bounding box in the BEV space, and selecting grid points falling within the candidate regions to form a candidate point set, includes: The center point of each real target bounding box is scaled from physical world coordinates to BEV feature map coordinates; Calculate its projected boundary on the BEV plane based on the scaled center point coordinates and the size of the actual target box; A preset margin extension is applied around the projection boundary to obtain the boundary of the candidate region; Filter all grid points on the BEV feature map that fall within the boundary of the candidate region, and form the candidate point set of the corresponding real target box by selecting all the selected grid points.

4. The method according to claim 1, characterized in that, The step of obtaining the classification prediction score and the localization matching degree between the 3D prediction box and the ground truth target box for each candidate point in the candidate point set, and calculating the task alignment score by fusing classification and localization information, includes: The classification prediction result is obtained from the network output corresponding to the candidate point, and the classification prediction result represents the probability distribution of the candidate point belonging to each type of target. The classification prediction result is analyzed, and the prediction probability corresponding to the target category is extracted as the classification prediction score s of the candidate point. p ; Calculate the 3D intersection-over-union ratio (u) between the 3D predicted bounding boxes of the candidate points and the ground truth bounding boxes. p The 3D intersection-to-union ratio is calculated by multiplying the BEV plane intersection-to-union ratio by the height direction intersection-to-union ratio. Based on the classification, the predicted score s p Intersection over union ratio u of the 3D p The task alignment score is calculated according to the following formula: align p : Where α and β are hyperparameters, and β > α > 0; Wherein, the height direction intersection-union ratio The calculation formula is: ; in, and These are the top and bottom coordinates of the 3D prediction bounding box in the height direction, respectively. and These are the top and bottom coordinates of the actual target bounding box in the height direction, respectively; The height is the intersection height of the 3D predicted bounding box and the real target bounding box in the height direction.

5. The method according to claim 1, characterized in that, The step of selecting several candidate points as positive samples from the candidate point set based on the task alignment score for each real target bounding box includes: For each real target bounding box, sort all candidate points in its candidate point set in descending order of task alignment score; Determine whether the total number of candidate points in the candidate point set is less than a preset number K; If so, all candidate points in the candidate point set will be used as positive samples of the real target box; If not, the top K candidate points after sorting are selected as positive samples of the true target bounding boxes.

6. The method according to claim 1, characterized in that, The step of assigning the grid point to the real target box with the highest localization matching degree when the same grid point is selected as a positive sample by multiple real target boxes, and the grid point not selected by any real target box as a negative sample, includes: Summarize the positive sample assignment results of all real target bounding boxes and establish a global mapping table; Traverse all grid points and identify how many real targets have bounded each grid point as a positive sample. If a certain grid point is selected as a positive sample by only one real target, it is directly retained; If a certain grid point is selected as a positive sample by multiple real target boxes at the same time, then it is assigned to the real target box with the highest localization matching degree; If a grid point is not selected as a positive sample by any of the real targets, it is determined as a negative sample.

7. The method according to claim 6, characterized in that, The formula for assigning the true target bounding box with the highest location matching degree is: ; Where assign(p) is the allocation result of candidate point p, representing the actual target box to which the candidate point finally belongs; p is the candidate point to be assigned, corresponding to a certain grid or anchor point in the candidate point set; This refers to a real target bounding box in the current scene; To match the true target bounding box A set of candidate points with spatial relationships exists, where the spatial relationships include, but are not limited to, candidate points falling within the BEV range or 3D spatial range of the true target bounding box; p∈ This indicates that candidate point p belongs to the true target box. The associated candidate set, i.e., p and There is a potential matching relationship in space; For candidate point p relative to the ground truth bounding box The localization matching degree, i.e., the 3D intersection-over-union ratio, is used to measure the accuracy of the 3D predicted bounding box of p with the actual location of the target object. The degree of overlap in three-dimensional space.

8. The method according to claim 1, characterized in that, The classification training objective for the positive samples is the normalized task alignment score, which is calculated using the following formula: ; Where (i,j) is the true target bounding box assigned to the location with the highest matching degree. Grid points; For grid point (i,j), the true bounding box with the highest localization match is... Task alignment score; The 3D predicted bounding box and the ground truth bounding box for grid point (i,j) 3D intersection-union ratio; To match the true target bounding box The highest task alignment score among all associated candidate points; and / or, The regression training objectives include: center offset, target height, logarithmic encoded values ​​of size, sine and cosine values ​​of yaw angle; and the regression training objectives also include velocity components.

9. The method according to claim 1, characterized in that, The total loss function is obtained by weighted summation of classification loss and regression loss. The classification loss adopts the zoom focus loss function, and the regression loss adopts the Smooth L1 loss function or CIoU loss function and is calculated only for positive samples. The formula for calculating the zoom focus loss (Varifocal Loss) is as follows: ; Where p is the classification prediction probability, q is the target quality, q is the normalized task alignment score for positive samples, and q is 0 for negative samples. The formula for calculating the Smooth L1 loss function is as follows: ; in, The set of positive samples; For positive sample p, the predicted regression value is denoted as . Let p be the target regression value for the positive sample.

10. A dynamic sample allocation system for BEV features based on task alignment, characterized in that, The system is capable of executing the task-aligned dynamic sample allocation method for BEV features according to any one of claims 1 to 9; the system comprises: The prediction and decoding module is used to acquire BEV feature maps and corresponding ground truth bounding boxes, perform prediction and decoding through the detection head, and output the classification prediction results and 3D prediction bounding boxes for each grid point. The candidate filtering module is used to determine the candidate region for each real target bounding box based on its projection range in the BEV space, and filter to obtain a set of candidate points. The alignment score calculation module is used to calculate the task alignment score that integrates classification and positioning information for each candidate point in the candidate point set. The positive sample selection module is used to select a number of candidate points as positive samples from the candidate point set of each real target box according to the task alignment score. The conflict resolution module is used to assign the grid point to the real target box with the highest positioning matching degree when the same grid point is selected as a positive sample by multiple real target boxes. The target construction module is used to construct classification training targets and regression training targets for each grid point based on the allocation results.