BEV visual angle 3D target detection method based on self-distillation and storage medium

By building a self-distillation framework on the same model, and using the self-distillation method to perform 3D target detection from the perspective of BEV, the problem of balancing computational resources and accuracy is solved, and efficient feature alignment and real-time detection are achieved, which is suitable for autonomous driving scenarios.

CN120976890APending Publication Date: 2025-11-18DONGFENG MOTOR GRP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511055449.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-30
Publication Date
2025-11-18

AI Technical Summary

Technical Problem

Existing 3D object detection algorithms from a BEV perspective suffer from excessive parameter and computational costs when deployed on real vehicles with limited computing resources. It is difficult to balance accuracy and computational cost, and the cross-modal knowledge distillation process also presents challenges.

Method used

A self-distillation method is adopted to build a dual-branch architecture on the same model. Multi-view images are collected through surround-view cameras, and the image features are converted into BEV feature maps using a visual transformation module. The self-distillation framework is constructed to generate teacher and student BEV features. The model is optimized through the self-distillation loss function, and dynamic soft label fusion and feature alignment are performed by combining LiDAR and image data to achieve knowledge transfer.

Benefits of technology

It reduces the consumption of computing resources, improves detection accuracy and processing speed, meets the real-time requirements of autonomous driving, improves training efficiency, and solves the problem of cross-modal feature alignment in existing technologies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120976890A_ABST
    Figure CN120976890A_ABST
Patent Text Reader

Abstract

A BEV visual angle 3D target detection method based on self-distillation is applied to an automatic driving 3D target detection scene and comprises the steps that S1, a multi-visual angle image is collected through an all-round camera, and image features are converted into a BEV feature map through a visual conversion module; s2, constructing a self-distillation framework, and generating teacher BEV features and student BEV features in parallel on the same model based on a BEV feature map; S3, splicing the teacher BEV features and the student BEV features, inputting the spliced features into an encoder, and calculating BEV feature self-distillation loss LFSD; s4, discretizing 3D bounding box positioning information predicted in the current batch into probability distribution, carrying out KL divergence alignment on the probability distribution and soft label probability distribution generated in the previous batch, and calculating bounding box self-distillation loss LBSD; and S5, combining the LFSD, the LBSD and the target detection original loss Loriginal to form an overall loss function, and carrying out model optimization. According to the method, self-distillation is introduced, and the lightweight of the visual 3D target detection model is realized through a more effective distillation algorithm.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of autonomous driving, and specifically to a BEV-based 3D target detection method and storage medium based on self-distillation. Background Technology

[0002] Thanks to increasingly deeper network layers and larger parameter counts, 3D object detection algorithms from a BEV (Battery Electric Vehicle) perspective are achieving increasingly higher accuracy. However, this usually also means a continuous increase in the number of parameters and computational cost of the algorithm model, making it difficult to deploy on real vehicles with limited computing resources. This requires researchers to make trade-offs between accuracy and computational cost. Knowledge distillation has been introduced to address this problem. Knowledge distillation achieves knowledge transfer by encouraging smaller network models to learn from larger, more accurate network models. Its advantage lies in its ability to utilize pre-trained model resources, using pre-designed distillation methods to guide new training phases with the model's data. Currently, the most common distillation method uses a high-performance LiDAR model as the teacher mode to guide a lower-performance camera model, thereby improving the camera model's accuracy. However, this approach typically presents significant challenges due to the inconsistency in modalities and network model structures between the two methods. Summary of the Invention

[0003] In view of the technical defects and drawbacks existing in the prior art, embodiments of the present invention provide a BEV-based 3D target detection method that overcomes or at least partially solves the above problems, the specific solution of which is as follows:

[0004] As a first aspect of the present invention, a self-distillation-based 3D target detection method based on a BEV (Battery Electric Vehicle) perspective is provided, applicable to 3D target detection scenarios in autonomous driving, comprising the following steps: S1, acquiring multi-view images through a surround-view camera, and using a visual conversion module to convert image features into a bird's-eye view feature map, i.e., a BEV feature map; S2, constructing a self-distillation framework, and generating teacher BEV features and student BEV features in parallel on the same model based on the BEV feature map; S3, concatenating the teacher and student BEV features and inputting them into an encoder, calculating the BEV feature self-distillation loss L. FSD S4, Discretize the 3D bounding box localization information predicted in the current batch into a probability distribution, align it with the KL divergence of the soft label probability distribution generated in the previous batch, and calculate the bounding box self-distillation loss L. BSD S5, combined with L FSD and L BSD and the original loss of target detection L original This constitutes the overall loss function, which is then used for model optimization.

[0005] Furthermore, the BEV feature map is an initial feature representation of the image features converted into a BEV perspective using a visual transformation module. The parallel generation of teacher BEV features and student BEV features on the same model includes:

[0006] For the student branch: Based on the initial feature prediction depth map D, foreground segmentation map S, and contextual features F, the predicted depth map D, foreground segmentation map S, and contextual features F are fused using the foreground-aware SA-BEVPool algorithm to generate student BEV features containing only foreground information. For the teacher branch: fusing ground truth depth maps generated from LiDAR point clouds. And truth segmentation map By combining the student branch predictions D and S, a mixed depth map D in the form of soft labels is generated. T and mixed segmentation map S T Then, the SA-BEVPool algorithm is used to generate teacher BEV features. .

[0007] Furthermore, in the teacher branch, the mixed depth map D T and mixed segmentation map S T Generated through a dynamic soft label fusion mechanism, specifically including: S301, obtaining the ground truth depth map provided by LiDAR point cloud data. And truth segmentation map S302, Receive the depth map D and foreground segmentation map S predicted by the student branch; S303, Generate soft-supervised labels according to the following fusion formula:

[0008] ;

[0009] ;

[0010] Where λ is a configurable balancing factor, the value of which controls the fusion ratio of truth information and student prediction information: when λ=1, the truth information of LiDAR is used completely; when 0<λ<1, a hybrid soft label is formed, which retains the accuracy of the truth and supplements the rich feature representation of the student prediction; the λ is dynamically adjusted during the model training process to achieve progressive knowledge transfer.

[0011] Furthermore, the SA-BEVPool algorithm, acting as a foreground-aware feature extractor, generates BEV features for teacher and student branches through the following differentiation mechanism: S401, Student branch feature generation: Using the foreground segmentation map S predicted by the student branch to mask background noise, only the depth map D of the foreground region and the contextual features F are aggregated. The calculation process is expressed as follows:

[0012]

[0013] Where S serves as a spatial mask, enabling the algorithm to focus on feature extraction of the foreground target; S402, teacher branch feature generation: based on the blended depth map D T and mixed segmentation map S T Enhanced features are generated using the same algorithm architecture:

[0014] .

[0015] Furthermore, the BEV characteristic self-distillation loss L FSD The computation is achieved through a spatially aware feature alignment mechanism, which specifically includes:

[0016] S501, outputs the high-level teacher features from the encoder. With advanced student characteristics Decoupling by spatial coordinates yields and ,in, Let represent the feature vector of the teacher branch at grid position (i,j) in the BEV feature map. The feature vector representing the student branches at the same grid position;

[0017] S502, Normalized Difference Calculation: The difference in the feature vector of each spatial location is standardized by the teacher feature L2 norm.

[0018] ;

[0019] S503, Global Loss Aggregation: Calculates the average difference intensity across the global spatial dimension of the BEV feature map;

[0020] ;

[0021] Where H represents the spatial height dimension of the BEV feature map, corresponding to the vertical resolution; W represents the spatial width dimension of the BEV feature map, corresponding to the horizontal resolution. The normalization operation makes the loss function focus on feature orientation alignment, weakening the interference of feature amplitude differences.

[0022] Furthermore, the self-distillation loss L of the bounding box BSD The calculation is implemented through a time-iterative distillation mechanism, specifically including: S601, historical soft label caching: During model training, the probability distribution of bounding boxes output at each spatial location of the BEV feature map at the (t-1)th iteration is dynamically cached. As a soft label for teachers; S602, current probability distribution generation: in the t-th iteration, the continuous localization information of the 3D bounding box is discretized into a classification probability form to generate the student probability distribution. ;

[0023] S603, and Aligning with KL divergence, we obtain:

[0024] ;

[0025] Where H×W represents the spatial resolution of the BEV feature map, corresponding to the grid partitioning density, and t represents the dynamic temperature coefficient. When t>1, the softening probability distribution promotes knowledge transfer, and when t=1, it degenerates into the original distribution. The mechanism ensures that the current prediction and the historical optimization direction remain consistent in time and space.

[0026] Furthermore, the KL divergence D kL It is used to measure the degree of difference in bounding box localization distribution between the (t-1)th and tth iterations. The larger the value, the greater the difference in probability distribution.

[0027] Furthermore, the temperature coefficient t is scaled using a logarithmic probability output to control the degree of discretization of the bounding box positioning distribution.

[0028] Furthermore, the overall loss function is:

[0029]

[0030] Where α and β are the weighting coefficients of BEV feature self-distillation loss and bounding box self-distillation loss, respectively.

[0031] As a second aspect of the present invention, a computer-readable storage medium is provided, wherein a computer program is stored therein, and when executed by a computer, the computer program causes the computer to perform the self-distilled BEV-view 3D target detection method as described above.

[0032] The present invention has the following beneficial effects:

[0033] 1. Optimization of computing resources: Due to the adoption of a single-model dual-branch architecture, an additional teacher model is avoided, which significantly reduces memory usage and computing overhead.

[0034] 2. Advantages of feature alignment: Distillation is performed within the same model, ensuring consistent feature modes and avoiding cross-modal feature alignment issues.

[0035] Improved real-time performance: The dynamic distillation mechanism improves processing speed and meets the real-time requirements of autonomous driving.

[0036] 3. Improved training efficiency: By utilizing the self-distillation mechanism, the model can iteratively optimize itself, reducing training time. Attached Figure Description

[0037] Figure 1 A flowchart of the method provided in an embodiment of the present invention;

[0038] Figure 2 A self-distillation framework architecture diagram provided for embodiments of the present invention (showing teacher / student branches). Detailed Implementation

[0039] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the present invention, and not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0040] To enable those skilled in the art to better understand the technical solutions of the present invention, exemplary embodiments of the present invention are described below in conjunction with the accompanying drawings, including various details of the embodiments of the present invention to aid understanding. These should be considered merely exemplary. Therefore, those skilled in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present invention. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.

[0041] Where there is no conflict, the various embodiments of the present invention and the features thereof may be combined with each other.

[0042] As used herein, the term “and / or” includes any and all combinations of one or more related enumerated entries.

[0043] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the invention. As used herein, the singular forms “a” and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will also be understood that when the terms “comprising” and / or “made of” are used in this specification, the presence of the stated feature, integral, step, operation, element, and / or component is specified, but the presence or addition of one or more other features, integrals, steps, operations, elements, components, and / or groups thereof is not excluded. Terms such as “connected” or “linked” are not limited to physical or mechanical connections but can include electrical connections, whether direct or indirect.

[0044] Unless otherwise specified, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by one of ordinary skill in the art. It will also be understood that terms such as those defined in commonly used dictionaries should be interpreted as having the meaning consistent with their meaning in the context of the relevant art and the invention, and will not be interpreted as having an idealized or overly formal meaning unless expressly so defined herein.

[0045] In the technical solution of this invention, the collection, storage, use, processing, transmission, provision, and disclosure of user personal information all comply with relevant laws and regulations and do not violate public order and good morals. The use of user data in this technical solution follows relevant national laws and regulations (e.g., the "Information Security Technology - Personal Information Security Specification"). For example: appropriate measures are taken for personal information access control; restrictions are imposed on the display of personal information; the purpose of using personal information does not exceed the scope of direct or reasonable association; and explicit identity targeting is eliminated when using personal information to avoid precisely locating a specific individual.

[0046] To address at least one of the technical problems existing in the aforementioned related technologies, the present invention provides a BEV-based 3D target detection method. Figure 1 This is a flowchart illustrating a BEV-view 3D target detection method based on self-distillation, provided in an embodiment of the present invention. This target detection method is applied to 3D target detection scenarios in autonomous driving and includes the following steps: S1, acquiring multi-view images through a surround-view camera, and converting image features into BEV feature maps using a visual conversion module. The BEV feature map is an initial feature representation of the image features converted from the BEV viewpoint using the visual conversion module. S2, constructing a self-distillation framework, and building a dual-branch structure within the same detection model based on the BEV feature map:

[0047] Student Branch: Based on the initial feature prediction depth map D, foreground segmentation map S, and contextual features F, the predicted depth map D, foreground segmentation map S, and contextual features F are fused using the foreground-aware SA-BEVPool algorithm to generate student BEV features containing only foreground information. The algorithm utilizes S to shield against background interference; Teacher branch: fuses the ground truth depth map generated from LiDAR point clouds. And truth segmentation map By combining the student branch predictions D and S, a mixed depth map D in the form of soft labels is generated. T and mixed segmentation map S T Then, the SA-BEVPool algorithm is used to generate teacher BEV features. .

[0048] refer to Figure 2 The diagram shown is a self-distillation block diagram provided in an embodiment of the present invention. S3: The teacher and student BEV features are concatenated and input into the encoder to calculate the BEV feature self-distillation loss L. FSD S4, Discretize the 3D bounding box localization information predicted in the current batch into a probability distribution, align it with the KL divergence of the soft label probability distribution generated in the previous batch, and calculate the bounding box self-distillation loss L. BSD S5, combined with L FSD and L BSDand the original loss of target detection L original This constitutes the overall loss function, optimizing the detection model parameters.

[0049] Among them, BEV feature map, or Bird's Eye View (BEV), is a technology that observes objects or scenes from above. In the field of autonomous driving environmental perception, visual 3D object detection from the BEV perspective uses surround-view cameras mounted on the vehicle to transform the surround-view image representation into a unified scene representation from the BEV perspective, thereby outputting a unified 3D object detection result. Compared with traditional methods, the BEV perspective offers superior visual effects, ease of processing, and convenient fusion.

[0050] This method proposes a 3D object detection algorithm based on self-distillation from a BEV (Browser-Engineered Entity) perspective. It's a special type of knowledge distillation method whose core idea is to improve model performance through self-supervision and self-learning, without relying on a single teacher model. Self-distillation iteratively optimizes the original model by training the same model multiple times or extracting information from different levels within the model. Compared to traditional knowledge distillation, self-distillation offers greater flexibility and adaptability while reducing computational resource requirements. This method designs the distillation method on the same model based on self-distillation, avoiding problems caused by inconsistent feature modes. Furthermore, this method optimizes the training process by using information from the previous batch to generate soft targets, further improving distillation efficiency and enhancing model performance without introducing additional computational resources.

[0051] In some embodiments, in the teacher branch, the blended depth map D T and mixed segmentation map S T Generated through a dynamic soft label fusion mechanism, specifically including: S301, obtaining the ground truth depth map provided by LiDAR point cloud data. And truth segmentation map S302, Receive the depth map D and foreground segmentation map S predicted by the student branch; S303, Generate soft-supervised labels according to the following fusion formula:

[0052] ;

[0053] ;

[0054] Where λ is a configurable balancing factor, the value of which controls the fusion ratio of truth information and student prediction information: when λ=1, the truth information of LiDAR is used completely, and when 0<λ<1, a hybrid soft label is formed, which retains the accuracy of the truth and supplements the rich feature representation of the student prediction; the λ can be dynamically adjusted during the model training process to achieve progressive knowledge transfer.

[0055] Figure 2A schematic diagram of dual-path input is given, showing the teacher branch receiving the true value of the lidar and the student's prediction.

[0056] In the above embodiment, the student branch generates clean foreground features based on the image prediction depth map D, foreground segmentation map S, and contextual features F through SA-BEVPool; the teacher branch fuses the ground truth value of the LiDAR (… , ) and student predicted values ​​(D, S), generate soft labels (D T ,S T The same algorithm is then used to generate high-precision supervised features. BEV feature distillation is used to standardize and calculate the differences in the high-level features output by the encoder. Bounding box distillation discretizes the bounding boxes into probability distributions, aligning them with historical iteration soft labels using KL divergence. This achieves an 83% reduction in GPU memory usage for a single model architecture (compared to the FrozenDistill solution), with an inference speed of 52ms / frame (Tesla V100), meeting the requirements of Level 4 autonomous driving.

[0057] In some embodiments, the SA-BEVPool algorithm, acting as a foreground-aware feature extractor, generates BEV features for teacher and student branches through the following differential mechanism: S401, Student branch feature generation: Using the foreground segmentation map S predicted by the student branch to mask background noise, only the depth map D and contextual features F of the foreground region are aggregated. The calculation process is expressed as follows:

[0058]

[0059] Where S serves as a spatial mask, enabling the algorithm to focus on feature extraction of the foreground target; S402, teacher branch feature generation: based on the blended depth map D T and mixed segmentation map S T Enhanced features are generated using the same algorithm architecture:

[0060] .

[0061] The differentiated technical effect of the algorithm in the two-branch approach:

[0062] Student branch: Achieve dynamic background suppression and improve feature purity through the predicted S;

[0063] Teacher Branch: Through Soft Label D T / S T Achieve truth-guided feature enhancement and provide high-precision monitoring signals.

[0064] In some embodiments, the BEV characteristic self-distillation loss L FSD The computation is achieved through a spatially aware feature alignment mechanism, which specifically includes:

[0065] S501, outputs the high-level teacher features from the encoder. With advanced student characteristics Decoupling by spatial coordinates yields and ,in, Let represent the feature vector of the teacher branch at grid position (i,j) in the BEV feature map. The feature vector representing the student branches at the same grid position;

[0066] S502, Normalized Difference Calculation: The difference in the feature vector of each spatial location is standardized by the teacher feature L2 norm.

[0067] ;

[0068] S503, Global Loss Aggregation: Calculates the average difference intensity across the global spatial dimension of the BEV feature map;

[0069] ;

[0070] Where H represents the spatial height dimension of the BEV feature map, corresponding to the vertical resolution; W represents the spatial width dimension of the BEV feature map, corresponding to the horizontal resolution. The normalization operation makes the loss function focus on feature orientation alignment, weakening the interference of feature amplitude differences.

[0071] In the above embodiments, by decoupling spatial coordinates, the feature map is decomposed according to grid coordinates (i,j), and further by normalized difference calculation and global aggregation, the feature vector alignment error (comparing cosine similarity) is reduced, thereby reducing the fluctuation of detection accuracy in rainy and foggy weather.

[0072] In some embodiments, the bounding box self-distillation loss L BSD The calculation is implemented through a time-iterative distillation mechanism, specifically including: S601, historical soft label caching: During model training, the probability distribution of bounding boxes output at each spatial location of the BEV feature map at the (t-1)th iteration is dynamically cached. As a soft label for teachers; S602, current probability distribution generation: in the t-th iteration, the continuous localization information of the 3D bounding box is discretized into a classification probability form to generate the student probability distribution. ;

[0073] S603, and Aligning with KL divergence, we obtain:

[0074] ;

[0075] Where H×W represents the spatial resolution of the BEV feature map, corresponding to the grid partitioning density, and t represents the dynamic temperature coefficient. When t>1, the softening probability distribution promotes knowledge transfer, and when t=1, it degenerates into the original distribution. The mechanism ensures that the current prediction and the historical optimization direction remain consistent in time and space.

[0076] Optionally, the method further includes:

[0077] By configurable temperature coefficient t and The two probability distributions are smoothed to obtain the following:

[0078] ;

[0079] ;

[0080] Calculate the KL divergence between each pair of probability distributions;

[0081] ;

[0082] Where k represents the category index after the bounding box is discretized;

[0083] A weighted average is performed across the entire BEV feature map to obtain:

[0084] .

[0085] In some embodiments, the KL divergence DkL is used to measure the degree of difference in the bounding box localization distribution between the (t-1)th and tth iterations, and a larger value indicates a greater difference in probability distribution.

[0086] Specifically, the KL divergence (DKL), as a tool for measuring the difference in bounding box localization distribution, achieves the following technical effects in the time-iterative distillation mechanism: (a) Dynamic error feedback: when the difference in probability distribution between two iterations increases (DKL... KL (a) ↑), indicating that the current bounding box localization prediction deviates from the historical optimization direction, triggering the gradient adjustment mechanism to enhance the model's convergence stability; (b) Localization accuracy quantification: The KL divergence value is directly related to the 3D bounding box center coordinate prediction error (unit: meters), and their mathematical relationship satisfies:

[0087]

[0088] (c) Targeted technical issues: To address the bounding box jitter problem caused by sudden changes in target motion in autonomous driving scenarios, trajectory smoothing is achieved by minimizing KL divergence, which reduces the target displacement variance between adjacent frames by 40%-60%.

[0089] In some embodiments, the temperature coefficient t is scaled by a logarithmic probability output to control the degree of discretization of the bounding box positioning distribution.

[0090] Optionally, the temperature coefficient t serves as a probability distribution smoothing controller, adjusting the bounding box localization discretization process through the following mechanisms: (a) Scaling effect: scaling the log probability output of the bounding box localization prediction with a scaling factor of 1 / t; (b) Working mode switching: when t>1, softening the peak value of the probability distribution to promote knowledge transfer between categories; when t=1, maintaining the original distribution shape as a benchmark reference state; when t<1, sharpening the probability distribution to enhance high-confidence prediction; (c) Dynamic adjustment mechanism: dynamically adjusting the t value based on the bounding box localization accuracy feedback during model training to achieve a progressive optimization strategy that strengthens knowledge transfer in the early training stage (t>1) and focuses on confidence prediction in the later training stage (t≈1).

[0091] In some embodiments, the overall loss function is:

[0092]

[0093] Where α and β are the weighting coefficients of BEV feature self-distillation loss and bounding box self-distillation loss, respectively.

[0094] This invention also provides a computer-readable storage medium, characterized in that the computer-readable storage medium stores a computer program, which, when executed by a computer, causes the computer to perform the self-distilled BEV-view 3D target detection method as described above.

[0095] Those skilled in the art will understand that all or some of the steps, systems, and apparatuses disclosed above, and their functional modules / units, can be implemented as software, firmware, hardware, or suitable combinations thereof. In hardware implementations, the division between functional modules / units mentioned above does not necessarily correspond to the division of physical components; for example, a physical component may have multiple functions, or a function or step may be performed collaboratively by several physical components. Some or all physical components may be implemented as software executed by a processor, such as a central processing unit, digital signal processor, or microprocessor, or as hardware, or as an integrated circuit, such as an application-specific integrated circuit (ASIC). Such software can be distributed on a computer-readable storage medium, which may include computer storage media (or non-transitory media) and communication media (or transient media).

[0096] As is known to those skilled in the art, the term computer storage medium includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information, such as computer-readable program instructions, data structures, program modules, or other data. Computer storage media includes, but is not limited to, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), static random access memory (SRAM), flash memory or other memory technologies, portable compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical disc storage, magnetic cartridges, magnetic tape, disk storage or other magnetic storage devices, or any other medium that can be used to store desired information and is accessible to a computer. Furthermore, it is known to those skilled in the art that communication media typically contain computer-readable program instructions, data structures, program modules, or other data in modulated data signals such as carrier waves or other transmission mechanisms, and may include any information delivery medium.

[0097] The computer-readable program instructions described herein can be downloaded from computer-readable storage media to various computing / processing devices, or downloaded via a network, such as the Internet, local area network, wide area network, and / or wireless network, to an external computer or external storage device. The network may include copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards them to the computer-readable storage media in the respective computing / processing device.

[0098] The computer program instructions used to perform the operations of this invention may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Smalltalk, C++, etc., and conventional procedural programming languages ​​such as the "C" language or similar programming languages. The computer-readable program instructions may be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, electronic circuitry, such as programmable logic circuitry, field-programmable gate arrays (FPGAs), or programmable logic arrays (PLAs), is personalized by utilizing state information from the computer-readable program instructions. This electronic circuitry can execute the computer-readable program instructions to implement various aspects of the invention.

[0099] The computer program product described herein can be implemented specifically through hardware, software, or a combination thereof. In one alternative embodiment, the computer program product is specifically embodied in a computer storage medium; in another alternative embodiment, the computer program product is specifically embodied in a software product, such as a software development kit (SDK), etc.

[0100] Various aspects of the present invention are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.

[0101] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processor of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner; thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.

[0102] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.

[0103] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction, which contains one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than those shown in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

[0104] Example embodiments have been disclosed herein, and while specific terminology has been used, it is for illustrative purposes only and should be construed as such, and is not intended to be limiting. In some instances, it will be apparent to those skilled in the art that features, characteristics, and / or elements described in conjunction with particular embodiments may be used alone, or in combination with features, characteristics, and / or elements described in conjunction with other embodiments, unless otherwise expressly indicated. Therefore, those skilled in the art will understand that various changes in form and detail may be made without departing from the scope of the invention as set forth in the appended claims.

[0105] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A BEV-based 3D target detection method, applied to 3D target detection scenarios in autonomous driving, characterized in that, The process includes the following steps: S1, acquiring multi-view images via a surround-view camera, and using a visual transformation module to convert image features into bird's-eye view feature maps, i.e., BEV feature maps; S2, constructing a self-distillation framework, and generating teacher BEV features and student BEV features in parallel on the same model based on the BEV feature maps; S3, concatenating the teacher and student BEV features and inputting them into the encoder to calculate the BEV feature self-distillation loss L. FSD S4, Discretize the 3D bounding box localization information predicted in the current batch into a probability distribution, align it with the KL divergence of the soft label probability distribution generated in the previous batch, and calculate the bounding box self-distillation loss L. BSD S5, combined with L FSD and L BSD and the original loss of target detection L original This constitutes the overall loss function, optimizing the detection model parameters.

2. The BEV-based 3D target detection method according to claim 1, characterized in that, The BEV feature map is an initial feature representation of the image features converted into a BEV perspective using a visual transformation module. The parallel generation of teacher BEV features and student BEV features on the same model includes: For the student branch: Based on the initial feature prediction depth map D, foreground segmentation map S, and contextual features F, the predicted depth map D, foreground segmentation map S, and contextual features F are fused using the foreground-aware SA-BEVPool algorithm to generate student BEV features containing only foreground information. For the teacher branch: fusing ground truth depth maps generated from LiDAR point clouds. And truth segmentation map By combining the student branch predictions D and S, a mixed depth map D in the form of soft labels is generated. T and mixed segmentation map S T Then, the SA-BEVPool algorithm is used to generate teacher BEV features. .

3. The BEV-based 3D target detection method according to claim 2, characterized in that, In the teacher branch, the blended depth map D T and mixed segmentation map S T Generated through a dynamic soft label fusion mechanism, specifically including: S301, obtaining the ground truth depth map provided by LiDAR point cloud data. And truth segmentation map S302, Receive the depth map D and foreground segmentation map S predicted by the student branch; S303, Generate soft-supervised labels according to the following fusion formula: ; ; Where λ is a configurable balancing factor, the value of which controls the fusion ratio of truth information and student prediction information: when λ=1, the truth information of LiDAR is used completely; when 0<λ<1, a hybrid soft label is formed, which retains the accuracy of the truth and supplements the rich feature representation of the student prediction; the λ is dynamically adjusted during the model training process to achieve progressive knowledge transfer.

4. The BEV-based 3D target detection method according to claim 2 or 3, characterized in that, The SA-BEVPool algorithm, acting as a foreground-aware feature extractor, generates BEV features for teacher and student branches through the following differential mechanism: S401, Student branch feature generation: Using the foreground segmentation map S predicted by the student branch to mask background noise, only the depth map D of the foreground region and the contextual features F are aggregated. The calculation process is expressed as follows: ; in S serves as a spatial mask, enabling the algorithm to focus on feature extraction of the foreground target; S4 02, Teacher Branch Feature Generation: Based on Hybrid Depth Map D T and mixed segmentation map S T Enhanced features are generated using the same algorithm architecture: 。 5. The BEV-based 3D target detection method according to claim 1, characterized in that, The BEV characteristic self-distillation loss L FSD The computation is achieved through a spatially aware feature alignment mechanism, which specifically includes: S501, outputs the high-level teacher features from the encoder. With advanced student characteristics Decoupling by spatial coordinates yields and ,in, Let represent the feature vector of the teacher branch at grid position (i,j) in the BEV feature map. The feature vector representing the student branches at the same grid position; S502, Normalized Difference Calculation: The difference in the feature vector of each spatial location is standardized by the teacher feature L2 norm. ; S503, Global Loss Aggregation: Calculates the average difference intensity across the global spatial dimension of the BEV feature map; ; Where H represents the spatial height dimension of the BEV feature map, corresponding to the vertical resolution; W represents the spatial width dimension of the BEV feature map, corresponding to the horizontal resolution. The normalization operation makes the loss function focus on feature orientation alignment, weakening the interference of feature amplitude differences.

6. The BEV-based 3D target detection method according to claim 1, characterized in that, The self-distillation loss of the bounding box L BSD The calculation is implemented through a time-iterative distillation mechanism, specifically including: S601, historical soft label caching: During model training, the probability distribution of bounding boxes output at each spatial location of the BEV feature map at the (t-1)th iteration is dynamically cached. As a soft label for teachers; S602, current probability distribution generation: in the t-th iteration, the continuous localization information of the 3D bounding box is discretized into a classification probability form to generate the student probability distribution. ; S603, and Aligning with KL divergence, we obtain: ; Where H×W represents the spatial resolution of the BEV feature map, corresponding to the grid partitioning density, and t represents the dynamic temperature coefficient. When t>1, the softening probability distribution promotes knowledge transfer, and when t=1, it degenerates into the original distribution. The mechanism ensures that the current prediction and the historical optimization direction remain consistent in time and space.

7. The BEV-based 3D target detection method according to claim 6, characterized in that, The KL divergence D kL It is used to measure the degree of difference in bounding box localization distribution between the (t-1)th and tth iterations. The larger the value, the greater the difference in probability distribution.

8. The BEV-view 3D target detection method based on self-distillation according to claim 6, characterized in that, The temperature coefficient t is scaled using a logarithmic probability output to control the degree of discretization of the bounding box positioning distribution.

9. The BEV-based 3D target detection method according to claim 1, characterized in that, The overall loss function is: ; Where α and β are the weighting coefficients of BEV feature self-distillation loss and bounding box self-distillation loss, respectively.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a computer, causes the computer to perform the self-distilled BEV-view 3D target detection method as described in any one of claims 1 to 9.