Dynamic shielding target complementing and labeling method and system based on BEV time sequence fusion

By collecting and processing multi-dimensional BEV perception data, setting occlusion completion confidence thresholds and multi-modal feature fusion priorities, constructing a standardized training dataset, and dynamically adjusting the annotation strategy, the problems of low accuracy and low efficiency of pseudo-labels in dynamic occlusion target annotation were solved, achieving efficient and accurate annotation results.

CN121808714APending Publication Date: 2026-04-07SUZHOU KUSHUJU INFORMATION TECHNOLOGY CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-03-10
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

Existing technologies for labeling dynamically occluded targets suffer from problems such as low accuracy in generating pseudo-labels, low labeling efficiency, and a lack of multi-dimensional perceptual data fusion and utilization, resulting in inaccurate labeling results and high costs, making it difficult to meet the needs of large-scale data labeling.

Method used

By collecting multi-dimensional BEV perception data, setting the occlusion completion confidence threshold and multimodal feature fusion priority, performing integrated processing and temporal-spatial alignment of time-series frames, constructing a standardized training dataset, and dynamically adjusting the completion annotation strategy based on the complexity of the occlusion scene and hardware computing power requirements, the mapping relationship between occlusion features and target morphology is learned using the BEV temporal fusion algorithm.

Benefits of technology

It significantly improves the accuracy and efficiency of labeling dynamically occluded targets, reduces human subjective error, reduces manual review costs, provides high-quality data support, and enhances the dynamic adaptability and reliability of the technology.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121808714A_ABST
    Figure CN121808714A_ABST
Patent Text Reader

Abstract

The invention provides a dynamic occlusion target complementing and labeling method and system based on BEV time sequence fusion, and is applied to the technical field of data processing. According to the method, BEV time sequence fusion is taken as a core, multi-dimensional sensing data such as a time sequence image and a three-dimensional coordinate are collected firstly according to dynamic shielding target complementation and labeling requirements, and key parameters such as a complementation confidence threshold value are set in combination with detection precision and labeling specifications; the data quality is optimized through integrated processing, and a standardized training data set is constructed and sorted according to completion contribution degrees. A dynamic completion labeling strategy is determined based on scene complexity, hardware computing power and the like, parameters such as sliding window size and the like are adapted, and data sets are split and then imported into training in parallel. A BEV time sequence fusion algorithm is utilized to learn a mapping relation between shielding features and a target form and a motion rule, multi-modal features are dynamically weighted and fused, and accurate completion and labeling are realized by combining adaptive core parameters such as shielding types and target scales.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing technology, and in particular to a method and system for dynamic occlusion target completion and annotation based on BEV temporal fusion. Background Technology

[0002] Currently, dynamic targets such as vehicles are easily affected by pedestrians, other vehicles, buildings, and road facilities in real-world road scenarios, resulting in partial occlusion, complete occlusion, or cross-occlusion. For the annotation of such dynamically occluded targets, existing technologies mainly rely on manual annotation or simple semi-automatic annotation methods. When faced with occluded targets, annotators often lack effective historical feature references and scientific basis for completion, and can only rely on subjective experience to guess the complete shape and location of the target and draw annotation boxes. This leads to serious data noise in the annotation results, low accuracy in generating pseudo-labels, and an inability to provide high-quality data support for model training.

[0003] Meanwhile, existing annotation methods lack a differentiated annotation mechanism for occluded scenarios. All annotation results must be manually reviewed one by one, regardless of the severity of occlusion or the reliability of the annotation, requiring significant manpower and resulting in low annotation efficiency, failing to meet the practical needs of large-scale data annotation. Furthermore, existing technologies lack effective fusion and utilization of multi-dimensional perceptual data, failing to fully explore the supporting role of temporal features and 3D geometric features in occluded target completion. Completion strategies lack dynamic adaptability, unable to adjust in real time according to occlusion type, target scale, scene dynamic rate, etc., further reducing the accuracy and reliability of completion and annotation.

[0004] It should be noted that the information disclosed in the background section above is only used to enhance the understanding of the background of this disclosure, and therefore may include information that does not constitute prior art known to those skilled in the art. Summary of the Invention

[0005] Other features and advantages of this application will become apparent from the following detailed description, or may be learned in part by practice of the invention.

[0006] According to one aspect of this application, a dynamic occlusion target completion and annotation method based on BEV temporal fusion is provided, comprising: collecting multi-dimensional BEV perception data, including continuous temporal images, high-precision three-dimensional coordinates of the target, refined contours of the occlusion region, dynamic illumination parameters, real-time calibration data of camera intrinsic and extrinsic parameters, and target motion trajectory and attitude records; setting an occlusion completion confidence threshold, multi-modal feature fusion priority weight, and completion fault tolerance range based on target detection recall, precision requirements, and annotation consistency specifications; performing integrated processing on the multi-dimensional BEV perception data to complete image denoising enhancement, temporal frame spatiotemporal alignment, semantic segmentation and feature extraction of the occlusion region, and target motion trend prediction; constructing a standardized training dataset; and reading the target feature dimensions, number of effective temporal frames, multi-modal data types, and quality of the dataset. The system evaluates data and ranks it according to its contribution and relevance to occlusion completion. It determines a dynamic completion annotation strategy based on the complexity of the occlusion scene, the upper limit of hardware computing power, and the real-time requirements of the algorithm. The sliding window size, feature update frequency, and number of completion iterations are set according to the target feature dimension, the number of effective temporal frames, and the frequency of occlusion changes. The dataset is adaptively split according to the strategy to form training batches containing occluded target feature vectors, real-world labels, and scene attributes. Related data of the same type and scene are imported synchronously through a parallel mechanism. The BEV temporal fusion algorithm learns the mapping relationship between occlusion features and the complete shape and motion patterns of the target. The completion parameters are dynamically weighted and optimized by fusion based on pixel semantics, 3D geometry, and temporal motion features. The core parameters and iteration strategy are adapted by combining occlusion type, target scale, scene dynamic rate, and camera motion state.

[0007] Another aspect of this application discloses a dynamic occlusion target completion and annotation system based on BEV temporal fusion, comprising: a multi-dimensional BEV perception data acquisition and standardization module, used to acquire multi-source perception data including continuous temporal images, high-precision three-dimensional coordinates of the target, and refined contours of the occlusion region, and generate a standardized perception dataset containing data modality type, temporal correlation attributes, and quality assessment information through structured acquisition and standardization processing; a completion and annotation core parameter configuration module, used to set the occlusion completion confidence threshold, multi-modal feature fusion priority weight, and completion fault tolerance range based on target detection recall rate, precision requirements, and annotation consistency specifications, and generate parameter configuration standards that meet quality control requirements; and a perception data integrated processing module, used to perform denoising enhancement, spatiotemporal alignment, semantic segmentation, and other processing on the multi-dimensional BEV perception data, extracting occlusion-related features and... The system predicts the target's movement trend, sorts the data according to their contribution to the completion, and outputs an ordered list of training data. A dynamic completion and annotation strategy formulation module is used to combine the complexity of the occlusion scene, the upper limit of hardware computing power, and the real-time requirements of the algorithm, setting parameters such as the sliding window size based on key information such as the target feature dimensions to generate an adaptive dynamic completion and annotation strategy. A training dataset splitting and import module is used to adaptively split the standardized dataset according to the completion and annotation strategy, forming training batches containing occluded target features, real labels, and scene attributes, and synchronously importing related data of the same type and scene through a parallel mechanism. A BEV temporal fusion completion and annotation execution module is used to integrate the output data of the aforementioned modules, learn the mapping relationship between occlusion features and target morphology and movement patterns through the BEV temporal fusion algorithm, dynamically weight and fuse multimodal features and adapt core parameters to complete dynamic occluded target completion and annotation.

[0008] According to another aspect of this application, an electronic device includes: a first processor; and a memory for storing executable instructions of the first processor; wherein the first processor is configured to execute the above-described dynamic occlusion target completion and annotation method based on BEV temporal fusion by executing the executable instructions.

[0009] According to another aspect of this application, a computer-readable storage medium is provided, on which a computer program is stored, which, when executed by a second processor, implements the above-described dynamic occlusion target completion and annotation method based on BEV temporal fusion.

[0010] This application presents a dynamic occluded target completion and annotation method and system based on BEV temporal fusion. With BEV temporal fusion as its core, and addressing the needs of dynamic occluded target completion and annotation, it first collects multi-dimensional BEV perception data, including continuous temporal images and 3D coordinates. Key parameters such as confidence thresholds are set based on detection accuracy and annotation standards. Data quality is optimized through integrated processing, a standardized training dataset is constructed and sorted by completion contribution, and a dynamic completion and annotation strategy is determined based on scene complexity and hardware computing power, adapting parameters such as sliding window size. The dataset is split according to the strategy and imported into parallel for training. The BEV temporal fusion algorithm learns the mapping relationship between occlusion features and target morphology and motion patterns, dynamically weights and fuses multimodal features, and adapts core parameters and iterative strategies to achieve accurate completion and annotation.

[0011] This application significantly reduces human subjective error and improves the accuracy of completion and annotation by using multi-dimensional data fusion and precise parameter adaptation. It solves the problem of low accuracy of pseudo-labels in traditional methods. The data sorting by contribution, parallel import of training and dynamic annotation strategies significantly improve annotation efficiency, reduce manual review costs and meet the needs of large-scale data annotation.

[0012] The system adapts to occlusion types and scene dynamic rates, adjusting the completion strategy in real time to enhance the dynamic adaptability and reliability of the technology, providing high-quality data support for model training. It should be understood that the above general description and the following detailed description are merely illustrative and explanatory, and do not limit this disclosure. Attached Figure Description

[0013] Figure 1 The flowchart illustrates a dynamic occlusion target completion and annotation method based on BEV temporal fusion provided in an embodiment of this application;

[0014] Figure 2 The diagram shows a schematic of the structure of a dynamic occlusion target completion and annotation system based on BEV temporal fusion according to an embodiment of this application. Detailed Implementation

[0015] The preferred embodiments of the present invention will be described below with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are for illustration and explanation only and are not intended to limit the present invention.

[0016] The following is combined with Figure 1 This paper describes a dynamic occlusion target completion and annotation method based on BEV temporal fusion according to an example embodiment of this application. It should be noted that the following application scenarios are shown only to facilitate understanding of the spirit and principles of this application, and the embodiments of this application are not limited in any way. Rather, the embodiments of this application are applicable to any suitable scenario.

[0017] In one implementation, Figure 1 The schematic diagram illustrates a flowchart of a dynamic occlusion target completion and annotation method based on BEV temporal fusion according to an embodiment of this application, including:

[0018] S101 collects multi-dimensional BEV perception data, including continuous time-series images, high-precision 3D coordinates of the target, refined contours of occluded areas, dynamic lighting parameters, real-time calibration data of camera intrinsic and extrinsic parameters, and records of target motion trajectory and attitude.

[0019] In one implementation, focusing on the core needs of automatic annotation and noise reduction in dynamic vehicle occlusion scenarios and reducing manual review costs, the system clarifies the multi-dimensional BEV perception data collection targets. These targets must provide a foundation for storing historical features in a temporal memory, while also supporting occlusion rate calculation and dynamic confidence weighting, ensuring the data is compatible with the business logic of pseudo-label generation and weak supervision signal construction. A high-definition camera adapted to the BEV perspective is used to continuously capture road scene images at a fixed frame rate, ensuring temporal continuity. For example, the camera frame rate is set to 30 frames per second, the acquisition range covers common vehicle driving fields of view, the image resolution is adapted to subsequent feature extraction requirements, and all images are accompanied by precise timestamps, providing raw data support for storing target features from the past N frames in the temporal memory.

[0020] This system integrates LiDAR and visual camera data. LiDAR acquires distance information between the vehicle and surrounding targets, while the pixel coordinates of the camera images are combined to calculate the target's 3D coordinates in the BEV coordinate system. For example, based on feature matching results between LiDAR point cloud data and camera images, the system outputs the target's x, y, and z 3D coordinates with centimeter-level accuracy, providing precise spatial reference for Kalman filtering to predict the location of currently occluded targets. Semantic segmentation technology performs pixel-level processing on the acquired images, identifying occluded and unoccluded areas of the vehicle and surrounding targets. Edge detection algorithms extract the contour information of occluded areas. For example, the semantic segmentation model distinguishes between vehicles, pedestrians, and obstacles, clearly defining the area where the vehicle is occluded by other targets and outputting the pixel coordinate sequence of the contour, providing a direct basis for occlusion rate calculation.

[0021] An integrated light sensor captures real-time parameters such as light intensity and color temperature of the scene, synchronously recording time points of light changes. For example, the light sensor collects data at a frequency of 10 times per second, outputting light intensity values ​​and color temperature ranges to provide parameters for lighting compensation during image feature extraction, preventing target feature distortion caused by lighting changes and ensuring the consistency of historical features in the time-series memory. A pre-set periodic calibration mechanism calculates camera intrinsic parameters (focal length, principal point coordinates) and extrinsic parameters (position, attitude) by photographing a standard calibration board, and updates the calibration data in real time. For example, a calibration process is performed every hour, generating a calibration parameter file to eliminate the impact of parameter drift during equipment use on 3D coordinate calculation and image feature matching, ensuring the consistency of feature space for images from different time series.

[0022] Based on continuous temporal target 3D coordinate data, the system tracks the position changes of vehicles and surrounding targets to form a complete motion trajectory. By analyzing the position changes of key target feature points, it records the target's pitch, rotation, and other attitude information. For example, based on continuous frame 3D coordinate data, it fits the target's motion trajectory curve, extracts parameters such as trajectory velocity and acceleration, and records attitude information through changes in the target's contour shape. This provides a motion pattern reference for Kalman filtering to predict the target's current position and generate pseudo-labels.

[0023] All data acquisition processes across all dimensions are strictly time-series aligned using timestamps, ensuring that multi-source data at the same time point can be correlated and matched. After acquisition, the data is structured and organized, unifying the data format and coordinate system to generate a standardized perception dataset containing data modality types, temporal correlation attributes, and quality assessment information. This dataset provides target feature data that can be directly stored in the temporal memory, and also provides correlation data between the occlusion region contour and the overall target contour for occlusion rate calculation, supporting the implementation of dynamic weighted confidence logic and forming a closed-loop business process of "data acquisition - feature storage - pseudo-label generation - weak supervision signal output".

[0024] S102 sets the occlusion completion confidence threshold, multimodal feature fusion priority weight, and completion error tolerance range based on the target detection recall, precision requirements, and annotation consistency specifications.

[0025] In one implementation, the core objectives are data noise reduction and reduced manual review costs in vehicle dynamic occlusion scenarios. Combining target detection recall and precision requirements with annotation consistency standards, the implementation clarifies the dimensions for setting the occlusion completion confidence threshold, multimodal feature fusion priority weights, and completion tolerance range. Specifically, the confidence threshold is correlated with the pseudo-label credibility assessment in the time-series memory; the fusion priority weights match the differences in the contribution of multimodal features to completion; and the completion tolerance range adapts to the fluctuations in dynamic occlusion scenarios. Through logical integration, recall requirements are mapped to a lower limit constraint of the confidence threshold, precision requirements are transformed into the basis for fusion weight allocation, and annotation consistency standards correspond to the completion tolerance range boundary. This clarifies that each parameter must meet the core constraints of "accurate pseudo-label generation, effective weak supervision signals, and focused manual review."

[0026] If the target detection recall rate is required to be no less than 95%, then the confidence threshold must be set to the lower limit that ensures that the pseudo-labels generated by the temporal memory are not misjudged and removed; if the precision rate is required to be no less than 90%, then the weight ratio of accurate features such as 3D coordinates and motion trajectory must be increased when fusing multimodal features; the annotation consistency specification requires that the annotation deviation of the same type of target does not exceed the preset range, and the error tolerance range of the completion must be limited to this deviation range to ensure that the completion result is consistent with the logic of manual annotation.

[0027] Based on the core requirement of quality control of completion results, it is necessary to establish clear parameter setting standards and definition rules, clarify the functional boundaries of core parameters and auxiliary constraints, and ensure that parameter configuration can support the accuracy of pseudo-label generation in the temporal memory while guaranteeing the effectiveness of dynamic weighted confidence, ultimately reducing the cost of manual review. The definition rules for core parameters and auxiliary constraints are based on "directly affecting completion accuracy and the effectiveness of weak supervision signals." The occlusion completion confidence threshold among the core parameters is a key link between the temporal memory and the weak supervision mechanism. Its value directly determines whether pseudo-labels are included in training or require manual review: pseudo-labels in the high confidence range can be used directly as strong supervision signals for model training, the medium confidence range corresponds to weak supervision signals, and the low confidence range triggers the manual review process. The precise allocation of annotation resources is achieved through signal level classification. The multimodal feature fusion priority weight focuses on the differences in feature contributions during the completion process. By adjusting the weight ratio of different features, the features that are more critical to the generation of pseudo-labels play a leading role. The target motion trajectory feature can reflect the continuous motion pattern of the target, and the three-dimensional coordinate feature can provide accurate spatial positioning. The supporting role of the two in the prediction of pseudo-label position is significantly higher than that of the dynamic lighting parameter feature, which is greatly affected by changes in lighting. Therefore, they are given higher weights to ensure the stability and accuracy of the completion results.

[0028] The allowable range of annotation consistency deviation in auxiliary constraints is a crucial safeguard for ensuring the uniformity of data quality. The tolerance for annotation errors naturally differs for targets of different scales. Small vehicles require higher contour detail, and annotation deviations must be controlled within a smaller pixel range. Large vehicles, due to their wider contour range, can have more relaxed deviation limits. By setting allowable deviation ranges based on target scale differences, annotation chaos caused by uniform standards can be avoided, ensuring consistency between the completion results and manual annotation logic, and providing a high-quality data foundation for model training.

[0029] Considering the real-time changes in dynamic occlusion scenarios, the dynamic optimization rules for core parameters must be scene-adaptable. The dynamic fine-tuning mechanism of the confidence threshold is closely related to changes in the occlusion rate: when the vehicle occlusion rate reaches 80% or other severe occlusion scenarios, the pseudo-labels generated by the temporal memory are the main basis for completion. If the threshold is too high, it is easy to remove valid pseudo-labels. Therefore, it needs to be lowered from the base value to ensure that the core completion data is not lost. When the occlusion rate drops to 30% or other mild occlusion scenarios, the currently visible features can provide more effective information. Raising the threshold can filter out low-confidence completion results generated based on a small number of occlusion features and reduce data noise. The dynamic allocation of priority weights for multimodal feature fusion is based on the historical completion accuracy. Historical features achieve a completion accuracy of 92% in unoccluded scenes, indicating their high reliability in restoring the target shape, and are therefore given a high weight. Kalman filter prediction features achieve a completion accuracy of 88% in moderately occluded scenes, accurately predicting the target position, and are therefore given the second highest weight. Through dynamic adjustment of weights, the advantageous features in different scenes can be fully utilized to improve the completion adaptability.

[0030] Generating the basic data for parameter configuration is a crucial step in implementing parameter settings. It requires integrating parameter constraints, quality control requirements, and optimization rules to form a structured and executable configuration scheme. The parameter types in the data should clearly cover core items such as occlusion completion confidence thresholds, multimodal feature fusion priority weights, and allowable ranges for annotation consistency deviations. Specifications should quantify specific numerical boundaries; for example, confidence thresholds should be divided into high, medium, and low ranges, and fusion weights should clearly define the proportion of each feature within a given range. The correlation logic should clearly label the correspondence between parameters and detection requirements and scene characteristics, such as adjusting the confidence threshold "positively correlated with occlusion rate and constrained by negative correlation with accuracy," clearly identifying the core influencing factors for parameter adjustment. The optimization strategy should refine the triggering conditions and execution methods for dynamic adjustments, such as "for every 10% change in occlusion rate, the threshold should be fine-tuned by 0.05" and "after each batch of training, the weight proportions should be updated based on the completion effect," ensuring that parameter configurations can be continuously optimized with scene changes and algorithm iterations. This foundational data serves as the execution benchmark for subsequent completion and annotation, enabling precise linkage between the generation of pseudo-labels in the time-series memory, dynamic weighting of confidence, and manual review processes, providing stable and reliable parameter support for the entire completion and annotation process.

[0031] S103 performs integrated processing on multi-dimensional BEV perception data, completing image denoising and enhancement, temporal and spatial alignment of time-series frames, semantic segmentation and feature extraction of occluded regions, prediction of target motion trends, constructing a standardized training dataset, reading the target feature dimensions, number of effective time-series frames, multimodal data types and quality assessment information of the dataset, and sorting them according to the contribution and correlation of the data to occlusion completion.

[0032] In one implementation, an integrated data processing technology is employed to process multi-dimensional BEV perception data throughout the entire process, aiming at accurate generation of pseudo-labels and effective output of weak supervision signals in dynamic vehicle occlusion scenarios. Image denoising and enhancement eliminate light and shadow interference in dynamic scenes through filtering techniques, ensuring the clarity of target features; temporal frame spatiotemporal alignment, based on timestamps and camera calibration data, enables accurate matching of continuous frame data in the BEV space, providing a temporal consistency basis for storing historical features in the temporal memory; semantic segmentation of occluded regions accurately divides occluded and unoccluded regions through a semantic model, extracting semantic features such as the boundary and area of ​​occluded regions, providing a core basis for occlusion rate calculation; target motion trend prediction combines historical motion trajectory and attitude data, using trend analysis algorithms to predict the subsequent motion state of the target, assisting Kalman filtering in optimizing the accuracy of pseudo-label positions.

[0033] After denoising and enhancing continuous temporal images, the vehicle outline edges are clearer, avoiding feature blurring caused by changes in illumination; after temporal and spatial alignment of temporal frames, the BEV coordinate deviation of the same vehicle in different frames is controlled within a preset range, ensuring the spatial continuity of historical features stored in the temporal memory; after semantic segmentation of occluded areas, the range and outline shape of the area where the vehicle is occluded by buildings can be clearly identified, directly supporting the calculation of occlusion rate; the target motion trend prediction is based on the uniform linear motion trajectory of the vehicle in the previous 10 frames, predicting the vehicle position in the next frame, providing motion reference for pseudo-label generation.

[0034] By integrating occlusion completion association standards, data contribution ranking rules, and modality classification threshold information, a precise matching relationship between features and completion annotation requirements is established. The occlusion completion association standards clarify the correspondence between various features and pseudo-label generation and confidence weighting. The data contribution ranking rules categorize features according to their importance to completion, and the modality classification threshold information distinguishes between valid and invalid feature data. Through association matching, temporal image features and 3D coordinate features are mapped to pseudo-label location generation requirements, and semantic features of occluded regions are associated with dynamic confidence weighting requirements. Valid features that meet the completion annotation requirements are selected, and a standardized training dataset is constructed to ensure that the dataset possesses the characteristics of "temporally coherent, feature-accurate, and closely correlated" characteristics.

[0035] The occlusion completion association standard stipulates that target motion trajectory features and 3D coordinate features should be prioritized to match the pseudo-label location generation requirements, while occluded region area and contour features should be prioritized to match the confidence weighting requirements. The data contribution ranking rule classifies 3D coordinate features and motion trajectory features as high contribution levels, and dynamic lighting parameters as medium contribution levels. The modality classification threshold information sets data where the confidence level of the semantic features of the occluded region is lower than a preset value, considering it invalid and not included in the dataset. Through the integration of these rules, the constructed standardized training dataset only contains high-value feature data that is effective for completion annotation.

[0036] Core data information is extracted based on a standardized training dataset. Using multi-dimensional BEV perception data as input, occlusion-related features as core parameters, and contribution evaluation rules as the judgment criteria, this method achieves precise reading of target feature dimensions, the number of effective temporal frames, multimodal data types, and quality evaluation information. The core data information extraction uses the standardized training dataset as its core carrier, aiming to select the most valuable data for dynamic occlusion target completion and annotation, providing precise support for subsequent dynamic completion and annotation strategy formulation and parameter optimization. Its core logic is to achieve targeted and efficient data extraction by clearly defining the input dimensions, core parameters, and judgment criteria.

[0037] The input dimension settings comprehensively cover all effective modalities of multi-dimensional BEV perception data, specifically including six categories: continuous temporal images, high-precision 3D coordinates of the target, refined contours of occluded areas, dynamic lighting parameters, real-time calibration data of camera intrinsic and extrinsic parameters, and target motion trajectory and attitude records. The core purpose of this setting is to ensure the comprehensiveness of data extraction and avoid deviations in the formulation of completion and annotation strategies due to the omission of key modal data. For example, continuous temporal images provide a historical feature basis for the temporal memory, real-time calibration data of camera intrinsic and extrinsic parameters ensures the spatial consistency of data across different frames, and dynamic lighting parameters support the illumination compensation processing of image features. All modal data together constitute the complete data support system required for completion and annotation.

[0038] The selection of core parameters focuses on occlusion-related features, eliminating redundant data interference and prioritizing key aspects such as semantic features of occluded regions, target motion features, and 3D coordinate features in the BEV coordinate system. Specifically, semantic features of occluded regions encompass information such as the contour shape, area ratio, and occlusion type of the occluded region, directly supporting occlusion rate calculation and dynamic confidence weighting. Target motion features include parameters such as vehicle speed, acceleration, and trajectory curvature, providing a core basis for Kalman filtering to predict pseudo-label positions. 3D coordinate features accurately reflect the target's spatial position in the BEV space, serving as the fundamental coordinate reference for pseudo-label generation. This focused design of core parameters ensures that the extracted data directly meets the core requirements for completing the annotation, improving the efficiency of subsequent strategy formulation and parameter optimization.

[0039] The contribution evaluation rule serves as the basis for data extraction, providing a clear standard for core data selection. The rule explicitly states that only feature data with a high contribution level will be extracted. The contribution level classification is based on the importance of features to the completion annotation—for example, target 3D coordinate features and motion trajectory features are classified as high contribution level because they directly determine the accuracy of pseudo-label positions; dynamic lighting parameters are classified as medium contribution level because they only affect feature extraction accuracy; and redundant data with weak correlation to occlusion completion is judged as low contribution level and not extracted. The application of this rule effectively filters out invalid data, ensuring that the extracted core data information is highly relevant to the completion annotation requirements, thus improving the quality and usability of the dataset.

[0040] Through the synergistic effect of the above three-fold definition, accurate reading of core data information is achieved: the target feature dimension is clearly defined as 20 dimensions, covering multiple feature dimensions such as vehicle outline, motion parameters, and spatial position, fully supporting the model's feature learning of the target; the effective number of time-series frames is determined to be 1000 frames, which not only meets the needs of the time-series memory to store historical features, but also ensures the efficiency of data processing, ensuring a balance between temporal continuity and computational feasibility; the multimodal data types are selected into 5 effective modalities, including images, coordinates, and trajectories, eliminating redundant modalities and focusing on core data support; the quality assessment information shows that the accuracy of 3D coordinate data reaches 98% and the semantic feature integrity of occluded areas reaches 95%. This data intuitively reflects the reliability of each modal data, providing a key reference for parameter weight allocation and fault tolerance range setting in subsequent strategy formulation, ensuring that the completion and annotation process can be carried out based on high-quality data, and improving the accuracy of the final completion and annotation.

[0041] By combining the contribution strength and correlation priority requirements of data to occlusion completion, the contribution measurement and ranking optimization of the core data information are performed. The correlation weight, temporal attributes, and modal types of various data are clarified, and a training data list ordered by contribution and correlation is generated. Contribution measurement and ranking optimization are key steps to improve the value of training data and ensure the accuracy of dynamic occlusion target completion. The core logic revolves around the two core requirements of "accuracy of pseudo-label generation" and "confidence-weighted effectiveness". The core data information is differentiated and ordered to ensure that high-value data occupies a dominant position in model training, thereby supporting the efficient implementation of pseudo-label generation from the temporal memory and weak supervision signal output.

[0042] Contribution quantification uses the actual effect of data on completion annotation as the core evaluation criterion, and realizes the concrete differentiation of data value by assigning quantitative scores. The quantification process is closely integrated with the core process of completion annotation: in the pseudo-label generation stage, the focus is on evaluating the data's ability to support the restoration of target location and shape; in the confidence weighting stage, the focus is on the data's auxiliary role in determining occlusion rate and dynamically adjusting weights. The target's 3D coordinate data directly provides accurate spatial positioning of the target in BEV space, serving as the core basis for Kalman filtering to predict the location of pseudo-labels and playing a decisive role in the accuracy of pseudo-label generation. Therefore, its quantization score is set to the highest. The semantic features of the occluded region contain key information such as occlusion area and contour shape, directly supporting the calculation of occlusion rate and adjustment of confidence weights. It is an important foundation for the generation of weakly supervised signals, so its quantization score is set to the second highest. Dynamic illumination parameters are only used for illumination compensation during image feature extraction, indirectly affecting feature quality, and have a weak direct contribution to completion annotation. Therefore, their quantization score is set to medium. Although the real-time calibration data of camera intrinsic and extrinsic parameters ensures spatial consistency of data, its role is basic support, and its quantization score is lower than that of the semantic features of the occluded region. Continuous temporal image data provides historical feature sources for the temporal memory and is an important supplement to pseudo-label generation. Its quantization score is slightly higher than that of the dynamic illumination parameters, forming a hierarchical quantization system.

[0043] The priority allocation follows the core principle of "prioritizing the accuracy of pseudo-label location, followed by confidence-weighted effectiveness," further clarifying the application priority of the data. This principle is based on the core objective of completion annotation—ensuring the accuracy of pseudo-label location first, and then optimizing annotation efficiency through confidence-weighted methods, which aligns with the business logic of "accuracy first, efficiency later." The number of effective time-series frames directly determines the total amount of historical features that can be called from the time-series memory. A sufficient number of frames can improve the reliability of pseudo-label generation. The target feature dimension reflects the richness of target features; the more comprehensive the dimension, the better it is for target morphology restoration. Both are directly related to pseudo-label generation and therefore have the highest priority. Auxiliary data such as multimodal data types and quality assessment information are mainly used to judge the reliability and suitability of data, and have a weaker direct impact on pseudo-label generation, so their priority is lower than the former. Although semantic feature-derived data of occluded regions related to confidence-weighted methods are of prominent importance, they are slightly lower in priority than data directly related to pseudo-label generation because they follow the "location accuracy first" principle, forming a clear priority hierarchy.

[0044] The ranking optimization, under the dual constraints of quantization score and association priority, comprehensively ranks the data based on association weights, temporal attributes, and modality types. Association weights are allocated according to quantization scores, with higher quantization scores receiving higher weights to ensure they are prioritized for learning during model training. Regarding temporal attributes, recent time-series frame data contains target state information closer to the current scene, making it more valuable for pseudo-label generation than earlier data; therefore, it is prioritized among data of the same type. Modality types are categorized and arranged according to the logic of "directly impactful data first, followed by indirectly supporting data." The final training data list is arranged in the following order: 3D coordinate data - motion trajectory data - occluded region semantic data - temporal image data - camera intrinsic and extrinsic parameter calibration data - dynamic illumination parameters. 3D coordinate data and motion trajectory data jointly support pseudo-label position and motion trend prediction, occupying the top two positions; occluded region semantic data supports confidence weighting, following closely behind; temporal image data supplements the temporal memory with historical features, ranking fourth; camera intrinsic and extrinsic parameter calibration data and dynamic illumination parameters, as fundamental supporting data, are arranged last. This sorting not only ensures the core requirement of pseudo-label generation, but also takes into account confidence weighting and data consistency, enabling training data to accurately match the core requirements of model training and improving training efficiency and completion accuracy.

[0045] S104 determines the dynamic completion annotation strategy by combining the complexity of the occlusion scene, the upper limit of hardware computing power and the real-time requirements of the algorithm. Based on the target feature dimension, the number of effective time-series frames and the occlusion change frequency, the sliding window size, feature update frequency and completion iteration number are set.

[0046] In one implementation, based on the complexity of the occlusion scene, the upper limit of hardware computing power, and the real-time requirements of the algorithm, the core components, parameter association logic, and adaptation boundaries of the dynamic completion annotation strategy are integrated to clarify the core basis and constraints for strategy formulation. Contribution quantification and ranking optimization are key steps to improve the value of training data and ensure the accuracy of dynamic occlusion target completion. The core logic revolves around the two core requirements of "accuracy of pseudo-label generation" and "confidence-weighted effectiveness," and differentiates and arranges core data information in an orderly manner to ensure that high-value data dominates model training, thereby supporting the efficient implementation of pseudo-label generation in the temporal memory and weak supervision signal output.

[0047] Contribution quantification uses the actual effect of data on completion annotation as the core evaluation criterion, and realizes the concrete differentiation of data value by assigning quantitative scores. The quantification process is closely integrated with the core process of completion annotation: in the pseudo-label generation stage, the focus is on evaluating the data's ability to support the restoration of target location and shape; in the confidence weighting stage, the focus is on the data's auxiliary role in determining occlusion rate and dynamically adjusting weights. The target's 3D coordinate data directly provides accurate spatial positioning of the target in BEV space, serving as the core basis for Kalman filtering to predict the location of pseudo-labels and playing a decisive role in the accuracy of pseudo-label generation. Therefore, its quantization score is set to the highest. The semantic features of the occluded region contain key information such as occlusion area and contour shape, directly supporting the calculation of occlusion rate and adjustment of confidence weights. It is an important foundation for the generation of weakly supervised signals, so its quantization score is set to the second highest. Dynamic illumination parameters are only used for illumination compensation during image feature extraction, indirectly affecting feature quality, and have a weak direct contribution to completion annotation. Therefore, their quantization score is set to medium. Although the real-time calibration data of camera intrinsic and extrinsic parameters ensures spatial consistency of data, its role is basic support, and its quantization score is lower than that of the semantic features of the occluded region. Continuous temporal image data provides historical feature sources for the temporal memory and is an important supplement to pseudo-label generation. Its quantization score is slightly higher than that of the dynamic illumination parameters, forming a hierarchical quantization system.

[0048] The priority allocation follows the core principle of "prioritizing the accuracy of pseudo-label location, followed by confidence-weighted effectiveness," further clarifying the application priority of the data. This principle is based on the core objective of completion annotation—ensuring the accuracy of pseudo-label location first, and then optimizing annotation efficiency through confidence-weighted methods, which aligns with the business logic of "accuracy first, efficiency later." The number of effective time-series frames directly determines the total amount of historical features that can be called from the time-series memory. A sufficient number of frames can improve the reliability of pseudo-label generation. The target feature dimension reflects the richness of target features; the more comprehensive the dimension, the better it is for target morphology restoration. Both are directly related to pseudo-label generation and therefore have the highest priority. Auxiliary data such as multimodal data types and quality assessment information are mainly used to judge the reliability and suitability of data, and have a weaker direct impact on pseudo-label generation, so their priority is lower than the former. Although semantic feature-derived data of occluded regions related to confidence-weighted methods are of prominent importance, they are slightly lower in priority than data directly related to pseudo-label generation because they follow the "location accuracy first" principle, forming a clear priority hierarchy.

[0049] The ranking optimization, under the dual constraints of quantization score and association priority, comprehensively ranks the data based on association weights, temporal attributes, and modality types. Association weights are allocated according to quantization scores, with higher quantization scores receiving higher weights to ensure they are prioritized for learning during model training. Regarding temporal attributes, recent time-series frame data contains target state information closer to the current scene, making it more valuable for pseudo-label generation than earlier data; therefore, it is prioritized among data of the same type. Modality types are categorized and arranged according to the logic of "directly impactful data first, followed by indirectly supporting data." The final training data list is arranged in the following order: 3D coordinate data - motion trajectory data - occluded region semantic data - temporal image data - camera intrinsic and extrinsic parameter calibration data - dynamic illumination parameters. 3D coordinate data and motion trajectory data jointly support pseudo-label position and motion trend prediction, occupying the top two positions; occluded region semantic data supports confidence weighting, following closely behind; temporal image data supplements the temporal memory with historical features, ranking fourth; camera intrinsic and extrinsic parameter calibration data and dynamic illumination parameters, as fundamental supporting data, are arranged last. This sorting not only ensures the core requirement of pseudo-label generation, but also takes into account confidence weighting and data consistency, enabling training data to accurately match the core requirements of model training and improving training efficiency and completion accuracy.

[0050] Based on the need to balance efficiency and accuracy in completion annotation, the strategy parameter setting logic is designed, clarifying the definition criteria for core parameters and related influencing factors. Core parameters include sliding window size, feature update frequency, and number of completion iterations. Related influencing factors include target feature dimension, number of effective temporal frames, and occlusion change frequency. The core objective of the strategy parameter setting logic is to balance completion annotation efficiency and accuracy. By clarifying the definition rules for core parameters and related influencing factors, a dynamic adaptation relationship between the two is established. This ensures that the parameter configuration can both support efficient retrieval of historical features from the temporal memory to generate pseudo-labels and adapt to real-time changes in dynamic occlusion scenarios, while also meeting the requirements of hardware computing power and algorithm real-time performance.

[0051] The core parameters are defined based on their direct impact on the core completion and annotation process, focusing on three key indicators: sliding window size, feature update frequency, and completion iteration count. The sliding window size directly determines the extraction range of historical features from the temporal memory. An excessively large window can lead to redundant feature interference and decreased computational efficiency, while an excessively small window may miss key historical features, affecting the accuracy of pseudo-label location prediction. Its value needs to strike a balance between feature coverage completeness and computational efficiency. The feature update frequency is related to the timeliness of features stored in the temporal memory. Too slow an update will cause historical features to fail to adapt to dynamic changes in the target, while too fast an update will increase computational consumption, failing to keep pace with the dynamic occlusion changes. The completion iteration count directly determines the optimization level of the completion parameters. More iterations result in higher completion accuracy but lower annotation efficiency, and vice versa. It needs to be dynamically adjusted according to scene complexity and accuracy requirements to achieve a balance between accuracy and efficiency.

[0052] The definition of related influencing factors is based on the principle of "indirectly supporting the adaptation of core parameters," covering target feature dimensions, the number of effective time-series frames, and the frequency of occlusion changes. Target feature dimensions reflect the complexity of target features. Complex targets (such as multi-detailed vehicle models) require more feature dimensions to support morphological reconstruction, while simple targets (such as regular sedans) have fewer feature dimensions. Their classification directly provides a basis for adapting the sliding window size. The number of effective time-series frames determines the total amount of historical features stored. Sufficient frames can support higher frequency feature updates, while insufficient frames require a lower update frequency to ensure feature effectiveness. The sufficiency range defines a reasonable range for feature update frequency settings. The frequency of occlusion changes reflects the dynamic characteristics of the occlusion scene. Rapid changes in occlusion require more iterations for a faster response, while slow changes (such as fixed occlusion) can reduce the number of iterations and improve efficiency. The type of change rate provides a direct reference for adjusting the number of iterations.

[0053] The adaptation logic of core parameters and related influencing factors closely revolves around the completion annotation requirements: the sliding window size needs to be precisely matched with the target feature dimension. The higher the target feature dimension, the larger the window needs to be to cover enough effective features to support the reconstruction of complex shapes. For example, a large truck with high-dimensional features needs a larger sliding window than a small car with low-dimensional features. The feature update frequency is deeply linked to the number of effective time-series frames. When there are enough effective time-series frames, the update frequency can be appropriately increased to ensure that the features in the time-series memory can reflect the latest state of the target in a timely manner. For example, the update frequency when there are 1000 effective time-series frames can be higher than that when there are only 500 frames. The number of completion iterations needs to be dynamically adapted to the occlusion change frequency. When the occlusion changes rapidly, the number of iterations needs to be increased to quickly adjust the completion parameters to cope with the dynamic changes in the scene. For example, the number of iterations in a cross-occlusion scene needs to be higher than that in a fixed occlusion scene.

[0054] The refined definition of influencing factors provides a clear basis for setting core parameters: target feature dimensions are divided into three levels of complexity—high, medium, and low—corresponding to complex vehicle models, conventional vehicle models, and simplified vehicle models, respectively; the number of effective time-series frames is set into three intervals—sufficient, medium, and insufficient—clearly defining the upper and lower limits of feature update frequency for different intervals; occlusion change frequency is divided into three categories—fast, medium, and slow—corresponding to cross-occlusion, partial occlusion, and fixed occlusion scenarios. Through this refined definition, the setting of core parameters is no longer a blind selection, but a scientific adaptation based on specific scenario factors, ensuring that parameter configuration meets both completion accuracy requirements and annotation efficiency, providing solid support for the implementation of dynamic completion annotation strategies.

[0055] To address the real-time changes in dynamic occlusion scenarios, parameter optimization rules are established, including adaptive window size adjustment, update frequency adapted to the occlusion rate, and iteration count dynamically adjusted based on feature complexity. This ensures the feasibility of the strategy and the completion effect. The core design logic of the dynamic parameter optimization rules is to adapt to the real-time changes in dynamic occlusion scenarios. By adaptively adjusting the sliding window size, feature update frequency, and completion iteration count, the strategy ensures completion accuracy while also considering computational efficiency and hardware power compatibility, guaranteeing the stability and feasibility of the dynamic completion annotation strategy in complex and ever-changing occlusion scenarios.

[0056] The adaptive adjustment mechanism for the sliding window size uses both target feature dimension and occlusion change frequency as triggering conditions to achieve a dynamic balance between "feature coverage integrity" and "computational efficiency." A higher target feature dimension means a more complex target shape, requiring more historical features to support complete shape reconstruction. In this case, increasing the window size allows for the extraction of richer historical features from the temporal memory, avoiding pseudo-label position deviations due to insufficient features. A higher occlusion change frequency indicates more drastic fluctuations in the target state; increasing the window size covers feature changes across more temporal frames, capturing target motion patterns and improving the stability of pseudo-label tracking. Conversely, when the target feature dimension is low (e.g., simplified vehicle models) and occlusion changes are gradual (e.g., fixed object occlusion), reducing the window size reduces the computational cost of redundant features, improving annotation efficiency. For example, in multi-target cross-occlusion scenarios, where target feature dimensions are complex and occlusion states switch frequently, the window size is increased compared to the base value to ensure sufficient coverage of historical features and dynamic change information. Conversely, when a vehicle is temporarily partially occluded by a fixed roadside guardrail, the window size is reduced to decrease computational pressure while maintaining completion accuracy.

[0057] The logic for adjusting the feature update frequency focuses on balancing the "timeliness" and "computing power consumption" of features in the time-series memory, directly related to the occlusion rate and the stability of the target state. When the occlusion rate is high, the area and degree of target occlusion change rapidly. If the feature update frequency is too low, the features stored in the time-series memory will lag behind the current target state, leading to deviations in pseudo-label generation. Increasing the update frequency allows for real-time synchronization of the latest target features, ensuring the timeliness of the completion basis. When the occlusion rate is slow and the target state is relatively stable, reducing the update frequency can reduce the storage and computation of duplicate features, avoiding wasted computing power. For example, during peak hours on urban roads, multiple targets frequently cross-occlude, resulting in a high occlusion rate. The feature update frequency should be increased accordingly to ensure that the latest target features are always stored in the time-series memory. Conversely, on suburban roads, when vehicles are slowly occluded by distant mountains and the occlusion state is stable, appropriately reducing the update frequency can save computing resources without affecting completion accuracy.

[0058] The dynamic adjustment of the number of completion iterations is based on the complexity of the target features and the difficulty of completion, achieving a precise match between "completion accuracy" and "annotation efficiency." When the target features are complex (such as multi-detailed vehicle models or vehicles with special postures), more parameter iterations are needed during the completion process to restore the complete shape and accurate position of the target. Increasing the number of iterations can gradually correct the completion error and improve the accuracy of pseudo-labels. When the target features are simple (such as regular sedans or vehicles with regular postures), the completion difficulty is low, and reducing the number of iterations can significantly improve annotation efficiency while ensuring accuracy. For example, large trucks have complex body structures and high feature dimensions, so the number of completion iterations is increased compared to the base value to ensure that the completion parameters are fully optimized; small family cars have simple features, so the number of completion iterations is appropriately reduced to quickly output accurate pseudo-labels.

[0059] The dynamic optimization of the three types of parameters does not operate independently, but forms a collaborative mechanism: when the vehicle encounters multiple targets with cross-occlusion (high occlusion change frequency + complex target feature dimensions), the sliding window size is expanded, the feature update frequency is increased, and the number of completion iterations is increased simultaneously, comprehensively strengthening the completion capability and ensuring that the pseudo-label can accurately track the target position; when the vehicle only encounters temporary partial occlusion (gradual occlusion change + simple target feature dimensions), the three types of parameters are adjusted simultaneously, maximizing the reduction of computing power consumption and improving annotation efficiency while ensuring that the completion accuracy meets the requirements, achieving differentiated adaptation of accurate completion in complex scenarios and efficient annotation in simple scenarios.

[0060] The strategy constraints, parameter setting standards, and optimization rules are integrated and processed to generate basic data for dynamic completion annotation configuration, including strategy type, parameter specifications, correlation logic, and optimization strategies. The core design logic of the dynamic parameter optimization rules is to adapt to the real-time changing characteristics of dynamic occlusion scenarios. By adaptively adjusting the sliding window size, feature update frequency, and completion iteration count, the system ensures completion accuracy while also considering computational efficiency and hardware computing power compatibility, thus ensuring the stability and feasibility of the dynamic completion annotation strategy in complex and ever-changing occlusion scenarios.

[0061] The adaptive adjustment mechanism for the sliding window size uses both target feature dimension and occlusion change frequency as triggering conditions to achieve a dynamic balance between "feature coverage integrity" and "computational efficiency." A higher target feature dimension means a more complex target shape, requiring more historical features to support complete shape reconstruction. In this case, increasing the window size allows for the extraction of richer historical features from the temporal memory, avoiding pseudo-label position deviations due to insufficient features. A higher occlusion change frequency indicates more drastic fluctuations in the target state; increasing the window size covers feature changes across more temporal frames, capturing target motion patterns and improving the stability of pseudo-label tracking. Conversely, when the target feature dimension is low (e.g., simplified vehicle models) and occlusion changes are gradual (e.g., fixed object occlusion), reducing the window size reduces the computational cost of redundant features, improving annotation efficiency. For example, in multi-target cross-occlusion scenarios, where target feature dimensions are complex and occlusion states switch frequently, the window size is increased compared to the base value to ensure sufficient coverage of historical features and dynamic change information. Conversely, when a vehicle is temporarily partially occluded by a fixed roadside guardrail, the window size is reduced to decrease computational pressure while maintaining completion accuracy.

[0062] The logic for adjusting the feature update frequency focuses on balancing the "timeliness" and "computing power consumption" of features in the time-series memory, directly related to the occlusion rate and the stability of the target state. When the occlusion rate is high, the area and degree of target occlusion change rapidly. If the feature update frequency is too low, the features stored in the time-series memory will lag behind the current target state, leading to deviations in pseudo-label generation. Increasing the update frequency allows for real-time synchronization of the latest target features, ensuring the timeliness of the completion basis. When the occlusion rate is slow and the target state is relatively stable, reducing the update frequency can reduce the storage and computation of duplicate features, avoiding wasted computing power. For example, during peak hours on urban roads, multiple targets frequently cross-occlude, resulting in a high occlusion rate. The feature update frequency should be increased accordingly to ensure that the latest target features are always stored in the time-series memory. Conversely, on suburban roads, when vehicles are slowly occluded by distant mountains and the occlusion state is stable, appropriately reducing the update frequency can save computing resources without affecting completion accuracy.

[0063] The dynamic adjustment of the number of completion iterations is based on the complexity of the target features and the difficulty of completion, achieving a precise match between "completion accuracy" and "annotation efficiency." When the target features are complex (such as multi-detailed vehicle models or vehicles with special postures), more parameter iterations are needed during the completion process to restore the complete shape and accurate position of the target. Increasing the number of iterations can gradually correct the completion error and improve the accuracy of pseudo-labels. When the target features are simple (such as regular sedans or vehicles with regular postures), the completion difficulty is low, and reducing the number of iterations can significantly improve annotation efficiency while ensuring accuracy. For example, large trucks have complex body structures and high feature dimensions, so the number of completion iterations is increased compared to the base value to ensure that the completion parameters are fully optimized; small family cars have simple features, so the number of completion iterations is appropriately reduced to quickly output accurate pseudo-labels.

[0064] The dynamic optimization of the three types of parameters does not operate independently, but forms a collaborative mechanism: when the vehicle encounters multiple targets with cross-occlusion (high occlusion change frequency + complex target feature dimensions), the sliding window size is expanded, the feature update frequency is increased, and the number of completion iterations is increased simultaneously, comprehensively strengthening the completion capability and ensuring that the pseudo-label can accurately track the target position; when the vehicle only encounters temporary partial occlusion (gradual occlusion change + simple target feature dimensions), the three types of parameters are adjusted simultaneously, maximizing the reduction of computing power consumption and improving annotation efficiency while ensuring that the completion accuracy meets the requirements, achieving differentiated adaptation of accurate completion in complex scenarios and efficient annotation in simple scenarios.

[0065] S105, the dataset is adaptively split according to the strategy to form training batches containing occluded target feature vectors, real labeled labels and scene attributes, and related data of the same scene and type are imported synchronously through a parallel mechanism.

[0066] In one implementation, guided by a dynamic completion annotation strategy, the focus is on accurate pseudo-label training and efficient application of weakly supervised signals in vehicle dynamic occlusion scenarios. The standardized training dataset is adaptively split according to scene attributes, occlusion type, and target feature complexity. The splitting logic must match the feature retrieval patterns of the temporal memory and the dynamic weighting requirements of confidence, ensuring that each training batch contains samples of similar scenes and occlusion types, and that the samples cover cases with different occlusion rates and confidence levels. This provides comprehensive data support for the model to learn occlusion features, pseudo-label generation logic, and confidence weight adjustment rules. During the splitting process, the integrity of the occluded target feature vector must be preserved, and the associated real-world labels serve as supervision for model training. Scene attributes (such as road type and lighting conditions) are added to assist the model in adapting to different application scenarios. For example, if the dynamic completion annotation strategy sets specific parameters for multi-object intersection occlusion scenarios on urban roads, the samples in that scenario are selected to form independent training batches during the splitting process. Each batch contains samples with occlusion rates below 30%, 30%-60%, and above 60%, and each sample is accompanied by a complete vehicle feature vector, a ground truth bounding box, and scene attributes such as urban roads and daytime lighting, ensuring that the model can learn the completion logic for that scenario in a targeted manner.

[0067] Each training batch must clearly define its core components, including the occluded target feature vector, ground truth labels, and scene attributes. The occluded target feature vector needs to integrate multimodal data such as temporal image features, 3D coordinate features, and motion trajectory features to form a unified-dimensional feature vector, adapting to the input requirements of the BEV temporal fusion algorithm. The ground truth labels must include the coordinate and category information of the complete target shape, serving as a benchmark for judging the accuracy of pseudo-label generation. Scene attributes must cover key environmental factors affecting occlusion completion, providing a basis for the model to dynamically adapt core parameters. The batch size needs to be set in conjunction with the hardware computing power limit to ensure that computing resources are fully utilized during training without overload, while ensuring that the number of samples within the batch is sufficient to support iterative optimization of model parameters. For example, when constructing training batches for a single-target transient occlusion scenario on an urban road, the occlusion target feature vector integrates the continuous temporal image features of the vehicle in the scenario, the three-dimensional coordinate sequence in the BEV coordinate system, and the motion trajectory curve features to form a fixed-dimensional vector; the ground truth label includes the coordinate values ​​of the complete outline of the vehicle and the vehicle category identifier; the scenario attribute is labeled as urban road, daytime, and light congestion; the batch size is set to a value that adapts to the hardware computing power to ensure that each batch of training can complete the iteration within a preset time.

[0068] For multimodal data with the same scene, occlusion type, and consistent relevance to the completion task, a parallel mechanism is used to synchronously import the model training process. The core logic of parallel import is to improve data processing efficiency while ensuring that the temporal memory can synchronously acquire historical feature data of similar scenes, supporting cross-sample feature association learning. During the import process, timestamp alignment and data verification mechanisms are used to ensure the temporal consistency and integrity of the imported multimodal data, avoiding model training errors due to data misalignment. The import order must follow the temporal logic of the samples to ensure that the temporal memory can store target features in frame order, providing accurate data support for historical feature retrieval and pseudo-label generation during occlusion. For example, for training data of a multi-target parallel occlusion scene on a highway, the temporal image data, 3D coordinate data, and semantic data of the occluded region of all samples in this scene are grouped by sample and synchronously transmitted to the model training module through a multi-threaded parallel import mechanism. Before importing, timestamp verification is used to ensure the consistency of the multimodal data frame order for the same sample. During the import process, data integrity is monitored in real time. If a set of data is found to be missing, a retransmission mechanism is triggered to ensure that the time-series memory can completely store the historical feature sequence of each sample, providing continuous feature input for Kalman filtering to predict the pseudo-label position.

[0069] S106 learns the mapping relationship between occlusion features and the complete shape and motion law of the target through the BEV temporal fusion algorithm. It dynamically weights and optimizes the completion parameters by fusion based on pixel semantics, 3D geometry and temporal motion features. It also adapts the core parameters and iterative strategies by combining occlusion type, target scale, scene dynamic rate and camera motion state.

[0070] In one implementation, occlusion feature data, target complete shape information, motion law data, and related data such as pixel semantic features, 3D geometric features, and temporal motion features are classified, extracted, and aligned with attributes to generate occlusion feature quantification information, a target shape feature library, motion law adaptation rules, and initial values ​​for multimodal feature weights. Focusing on the need for accurate pseudo-label generation and dynamic weighted confidence in vehicle dynamic occlusion scenarios, multi-dimensional core data is classified and extracted. This includes occlusion feature data (covering occlusion area, contour complexity, etc.), target complete shape information including target 3D dimensions and appearance features, motion law data (involving motion speed, trajectory curvature, etc.), as well as pixel semantic features, 3D geometric features, and temporal motion features. Through attribute alignment processing, data from different sources and in different formats are uniformly mapped to the BEV spatial coordinate system, ensuring data temporal consistency and spatial correlation. Finally, occlusion feature quantification information is generated, such as specific occlusion rate values, contour regularity index, target morphology feature library (storing standard morphology templates for various types of vehicles), motion law adaptation rules (clarifying the feature matching logic under different motion states), and initial values ​​of multimodal feature weights (providing a basis for subsequent weighted fusion).

[0071] Feature data of a vehicle partially obscured by a building was extracted, and the occlusion rate was quantified to be 45% and the contour regularity was 0.8. The standard three-dimensional dimensions and appearance feature templates of the vehicle model were matched from the target morphology feature library. Based on the vehicle's continuous frame motion data, the regularity adaptation rules of uniform linear motion were extracted. Initial weights were assigned to pixel semantic features, three-dimensional geometric features, and temporal motion features, with the initial weight of three-dimensional geometric features being higher than that of other features because they are more critical for position prediction.

[0072] Based on the goal of precise optimization of completion parameters, this method associates occlusion feature quantification information with the target morphological feature library and motion law adaptation rules. It integrates the influence weights and collaborative information of various features in the completion process using dynamic weighting logic to generate a basis for completion parameter optimization. With the goal of precise optimization of completion parameters, a relationship is established between occlusion feature quantification information, the target morphological feature library, and motion law adaptation rules. Through dynamic weighting logic, the influence weights and collaborative information of various features in the completion process are integrated: occlusion feature quantification information determines the completion priority (the higher the occlusion rate, the more reliant it is on historical feature completion); the target morphological feature library provides the morphological benchmark for completion; and the motion law adaptation rules guide the position prediction of completion. The dynamic weighting logic adjusts the weights of each feature in real time according to the occlusion scenario. For example, in cases of severe occlusion, the weights of temporal motion features and historical morphological features are increased; in cases of mild occlusion, the weight ratio of currently visible features is enhanced. Ultimately, a basis for optimizing completion parameters, including feature weight allocation schemes, morphological matching criteria, and position prediction logic, is generated.

[0073] When the vehicle occlusion rate reaches 70% (severe occlusion), the dynamic weighting logic increases the weight of historical motion trajectory features and target shape features in the temporal memory and decreases the weight of semantic features of currently visible pixels. Based on the motion law adaptation rules, the current vehicle position is predicted based on historical uniform motion data. Combined with the standard size in the target shape feature library, the basis for optimizing the completion parameters is formed, and it is clear that the completion should prioritize matching historical motion trends and standard shapes.

[0074] Based on the core parameters and iterative strategy adaptation targets, occlusion type, target scale, scene dynamic rate, and camera motion state data are respectively mapped to parameter adaptation models and iterative strategy adjustment models to generate initial adaptation information for core parameters and iterative strategies. The core logic of the initial adaptation of core parameters and iterative strategies is based on the precise matching of scene characteristics and completion requirements. Through the collaborative operation of the parameter adaptation model and the iterative strategy adjustment model, initial parameters and strategy schemes that fit the current scene are generated, laying the foundation for subsequent completion parameter optimization. This ensures that the BEV temporal fusion algorithm can quickly adapt to different dynamic occlusion scenes, balancing completion accuracy and computational efficiency.

[0075] First, the adaptation goals are clearly defined: the core parameters must meet the requirements of accuracy in pseudo-label generation and effectiveness in confidence-weighted calculation under the current scene; the iterative strategy must balance completion accuracy and annotation efficiency; and it must also adapt to the hardware computing power and algorithm real-time requirements. Under these goals, four key scene data categories—occlusion type, target scale, scene dynamic rate, and camera motion state—are respectively mapped to the parameter adaptation model and the iterative strategy adjustment model. These two models each perform their respective functions and work in deep collaboration to jointly support the generation of initial adaptation information.

[0076] Occlusion types are categorized into partial occlusion, complete occlusion, and cross-occlusion based on the degree and form of occlusion, directly affecting the dependence of completion on historical features: partial occlusion requires consideration of both currently visible features and historical features; complete occlusion relies entirely on historical features and motion pattern predictions from the temporal memory; and cross-occlusion needs to handle frequent switching of occlusion states among multiple targets. Target scale is categorized into small cars, medium cars, and large cars based on size, affecting the complexity of target morphological features and the difficulty of completion: large cars have higher feature dimensions and are more difficult to complete, while small cars have relatively simpler features and are more efficient to complete. Scene dynamic rate is categorized into low-speed congestion, medium-speed driving, and high-speed traffic based on vehicle driving status, reflecting the intensity of scene dynamic changes: in high-speed traffic scenarios, target movement speed is fast and occlusion states change frequently, requiring higher real-time performance of completion parameters and faster response speed of iterative strategies; low-speed congestion scenarios are relatively smooth. Camera movement is categorized by speed of motion into stationary, slow, and fast movement, which affects the stability of data acquisition. Fast camera movement can easily cause data offset, requiring parameter adjustments to improve the anti-interference capability of data acquisition.

[0077] The core function of the parameter adaptation model is to output the initial values ​​of the completion parameters. The input consists of the four types of scene data mentioned above, and the output includes core parameters such as multimodal feature fusion weights and completion accuracy thresholds. The allocation logic of feature fusion weights is strongly correlated with the scene data: in fully occluded scenes, the weight of temporal motion features (from the temporal memory bank) is significantly higher than that of pixel semantic features (currently visible features); in cross-occlusion scenes, the weight of 3D geometric features needs to be increased to ensure the distinguishability of multiple target positions; in high-speed traffic scenes, the weight of temporal motion features needs to be further increased to adapt to the position prediction requirements of rapid target movement; when the camera moves rapidly, the weight of camera intrinsic and extrinsic parameter calibration-related features needs to be increased to offset the impact of data offset. The setting of the completion accuracy threshold is combined with the target scale and occlusion type: the completion accuracy threshold for large vehicles is more stringent, the threshold for fully occluded scenes can be appropriately relaxed to balance recall and precision, and the threshold for partially occluded scenes needs to be strict to filter out low-confidence completion results.

[0078] The core function of the iterative strategy adjustment model is to determine the initial number of iterations and the convergence condition. The input is also data from four different scenarios, and the output must adapt to the scenario's completion difficulty and real-time requirements. The initial number of iterations is dynamically adjusted based on the completion difficulty: in scenarios with complete occlusion, large vehicles, and high-speed traffic, the completion difficulty is higher, requiring more iterations to fully optimize the completion parameters and ensure pseudo-label accuracy; in scenarios with partial occlusion, small vehicles, and low-speed congestion, the completion difficulty is lower, and the number of iterations can be reduced to improve annotation efficiency. The convergence condition is set based on the dynamic characteristics of the scenario and accuracy requirements: in scenarios with high-speed traffic and rapid camera movement, the convergence condition can be appropriately relaxed to meet real-time requirements and avoid delays caused by excessive iteration; in low-speed stationary and completely occluded scenarios, the convergence condition needs to be more stringent to ensure the accuracy of pseudo-label generation and provide a reliable basis for dynamic weighted confidence.

[0079] The collaborative operation process of the two types of models is clear: First, the scene data is synchronously input into the two types of models. The parameter adaptation model outputs initial parameters such as feature fusion weights and completion accuracy thresholds based on the quantitative analysis of the scene data. The iterative strategy adjustment model outputs the initial number of iterations and convergence conditions based on the scene completion difficulty assessment. Finally, the outputs of the two types of models are integrated to form the core parameters and initial adaptation information of the iterative strategy. For example, when the input data consists of a small car completely occluded, a high-speed traffic scene, and a camera moving rapidly: the parameter adaptation model analysis suggests that complete occlusion requires relying on historical features, high-speed traffic requires strengthening motion trend prediction, and rapid camera movement requires improving anti-interference capabilities. Therefore, it outputs a higher weight for temporal motion features (with a higher proportion than pixel semantics and 3D geometric features) and a strict completion accuracy threshold (to ensure that the pseudo-label position error is controllable). The iterative strategy adjustment model evaluation suggests that complete occlusion is difficult, but high-speed scenes have real-time requirements. Therefore, it sets a large number of initial iterations (to ensure sufficient parameter optimization) and relatively flexible convergence conditions (to avoid delays). The final generated initial adaptation information can accurately adapt to the completion requirements of the scene, providing a high-quality starting point for subsequent parameter optimization based on the BEV temporal fusion algorithm.

[0080] Based on the completion parameter optimization criteria and initial adaptation information, key attributes (such as occlusion rate, target size, motion speed, and camera movement rate) in the core data extracted for classification are marked and weighted according to dynamic weighting logic. Combining the historical feature retrieval results from the temporal memory with the position prediction output of the Kalman filter, core parameters such as multimodal feature weights, completion accuracy threshold, and number of iterations are adjusted. Simultaneously, the iteration strategy is optimized, clarifying the iteration termination conditions and parameter update frequency under different occlusion scenarios. Finally, the optimized completion parameter configuration (including dynamically adjusted feature weights and allowable completion error range) and dynamic iteration strategy scheme (specifying the iteration process and adjustment logic under different scenarios) are generated to ensure accurate completion results and adaptability to the needs of weakly supervised signal generation.

[0081] The key attributes such as the occlusion rate of the marker (60%), the small size of the target, the movement speed (80 km / h), and the camera movement speed (5 m / s) are weighted and calculated to increase the weight of temporal motion features and 3D geometric features, thereby narrowing the allowable range of completion error. The dynamic iteration strategy stipulates that if the completion error still exceeds the threshold after 3 iterations, more historical frame features from the temporal memory are called to recalculate until the error requirement is met or the maximum number of iterations is reached. The confidence of the generated pseudo-label is dynamically adjusted according to the completion error, providing a basis for weak supervision signals.

[0082] In one implementation, such as Figure 2 As shown, this application also provides a dynamic occlusion target completion and annotation system based on BEV temporal fusion, including:

[0083] The multi-dimensional BEV perception data acquisition and standardization module 201 is used to acquire multi-source perception data including continuous temporal images, high-precision three-dimensional coordinates of the target, and fine contours of occluded areas. Through structured acquisition and standardization processing, it generates a standardized perception dataset containing data modality type, temporal correlation attributes, and quality assessment information.

[0084] The completion annotation core parameter configuration module 202 is used to set the occlusion completion confidence threshold, multimodal feature fusion priority weight and completion fault tolerance range based on the target detection recall rate, precision requirements and annotation consistency specifications, and generate parameter configuration standards that meet quality control requirements.

[0085] The integrated processing module 203 for perceptual data is used to perform denoising enhancement, spatiotemporal alignment, semantic segmentation and other processing on multi-dimensional BEV perceptual data, extract occlusion-related features and predict the target motion trend, sort the data according to their contribution to the completion, and output an ordered training data list.

[0086] The dynamic completion annotation strategy formulation module 204 is used to combine the complexity of the occlusion scene, the upper limit of hardware computing power and the real-time requirements of the algorithm, and set parameters such as the sliding window size based on key information such as the target feature dimension to generate an adaptive dynamic completion annotation strategy.

[0087] The training dataset splitting and import module 205 is used to adaptively split the standardized dataset according to the completion annotation strategy to form training batches containing occluded target features, real labels and scene attributes, and synchronously import related data of the same type in the same scene through a parallel mechanism.

[0088] The BEV temporal fusion completion and annotation execution module 206 is used to integrate the output data of the aforementioned modules, learn the mapping relationship between occlusion features and target shape and motion law through the BEV temporal fusion algorithm, dynamically weight and fuse multimodal features and adapt core parameters to complete dynamic occlusion target completion and annotation.

[0089] The computer-readable storage medium provided in the above embodiments of this application and the dynamic occlusion target completion and annotation method based on BEV temporal fusion provided in the embodiments of this application are based on the same inventive concept and have the same beneficial effects as the methods adopted, run or implemented by the application stored therein.

Claims

1. A dynamic occlusion target completion and annotation method based on BEV temporal fusion, characterized in that, include: Collect multi-dimensional BEV perception data, including continuous time-series images, high-precision three-dimensional coordinates of the target, refined contours of occluded areas, dynamic lighting parameters, real-time calibration data of camera intrinsic and extrinsic parameters, and records of target motion trajectory and attitude. Based on the requirements for target detection recall and precision, as well as the standard for consistent annotation, we set the confidence threshold for occlusion completion, the priority weight for multimodal feature fusion, and the tolerance range for completion. Multi-dimensional BEV perception data is processed in an integrated manner to complete image denoising and enhancement, temporal and spatial alignment of time-series frames, semantic segmentation and feature extraction of occluded regions, and prediction of target motion trends. A standardized training dataset is constructed, and the target feature dimensions, number of effective time-series frames, multimodal data types and quality assessment information of the dataset are read. The data is sorted according to its contribution and correlation to occlusion completion. The dynamic completion annotation strategy is determined by combining the complexity of the occlusion scene, the upper limit of hardware computing power and the real-time requirements of the algorithm. Based on the target feature dimension, the number of effective time frames and the occlusion change frequency, the sliding window size, feature update frequency and completion iteration number are set. The dataset is adaptively split according to the strategy to form training batches containing occluded target feature vectors, real labeled labels and scene attributes. Related data of the same type in the same scene are imported synchronously through a parallel mechanism. The BEV temporal fusion algorithm learns the mapping relationship between occlusion features and the complete shape and motion law of the target. The parameters are dynamically weighted and optimized by pixel semantics, 3D geometry and temporal motion features. The core parameters and iterative strategies are adapted by combining occlusion type, target scale, scene dynamic rate and camera motion state.

2. The method as described in claim 1, characterized in that, Based on the requirements for object detection recall and precision, as well as annotation consistency standards, we set the occlusion completion confidence threshold, multimodal feature fusion priority weights, and completion error tolerance range, including: Based on the requirements for target detection recall and precision, as well as the standard for consistent annotation, the dimensions and related logic for setting the confidence threshold for occlusion completion, the priority weight for multimodal feature fusion, and the tolerance range for completion are integrated to clarify the core constraints and adaptation requirements of each parameter. Based on the quality control requirements of the completion results, the parameter setting standard was designed, and the definition rules of core parameters and auxiliary constraints were clarified. The core parameters include the occlusion completion confidence threshold and the priority weight of multimodal feature fusion, and the auxiliary constraints include the allowable range of annotation consistency deviation. In combination with the completion stability requirements of dynamic occlusion scenarios, parameter optimization rules are set to dynamically fine-tune the confidence threshold and allocate the fusion weight according to the feature contribution, so as to ensure that the completion results meet the detection and annotation requirements; The parameter constraints, quality control requirements, and parameter optimization rules are integrated and processed to generate complete parameter configuration basic data that includes parameter types, setting specifications, correlation logic, and optimization strategies.

3. The method as described in claim 1, characterized in that, Multi-dimensional BEV perception data is processed in an integrated manner to complete image denoising and enhancement, spatiotemporal alignment of time-series frames, semantic segmentation and feature extraction of occluded regions, and prediction of target motion trends. A standardized training dataset is constructed, and the target feature dimensions, number of effective time-series frames, multimodal data types, and quality assessment information of the dataset are read. The data is ranked according to its contribution and correlation to occlusion completion, including: An integrated data processing technology is used to process multi-dimensional BEV perception data to generate image denoising and enhancement results, time-series frame spatiotemporal alignment data, semantic feature set of occluded areas, and target motion trend prediction results. The feature set is integrated with the occlusion completion association criteria, data contribution ranking rules and modality classification threshold information to establish a precise matching relationship between features and completion annotation requirements, and generate a standardized training dataset. Based on the standardized training dataset, core data information is extracted. Multi-dimensional BEV perception data is used as the input dimension, occlusion-related features are used as the core parameters, and contribution evaluation rules are used as the judgment basis to achieve accurate reading of target feature dimensions, number of effective time-series frames, multi-modal data types, and quality evaluation information. By combining the contribution intensity and correlation priority requirements of the data to occlusion completion, the contribution measurement and sorting optimization of the core data information are performed, the correlation weight, temporal attributes and modal types of various data are clarified, and a training data list is generated in order of contribution and correlation.

4. The method as described in claim 1, characterized in that, A dynamic completion annotation strategy is determined by considering the complexity of the occlusion scene, the upper limit of hardware computing power, and the real-time requirements of the algorithm. Based on the target feature dimension, the number of effective time-series frames, and the frequency of occlusion changes, the sliding window size, feature update frequency, and number of completion iterations are set, including: Based on the complexity of occlusion scenarios, the upper limit of hardware computing power, and the real-time requirements of algorithms, this paper integrates the core components, parameter association logic, and adaptation boundaries of dynamic completion annotation strategies, and clarifies the core basis and constraints for strategy formulation. Based on the need to balance the efficiency and accuracy of completion annotation, the logic for setting strategy parameters is designed, and the definition criteria of core parameters and related influencing factors are clarified. Among them, the core parameters include sliding window size, feature update frequency, and completion iteration number, while the related influencing factors include target feature dimension, number of effective temporal frames, and occlusion change frequency. Based on the real-time changing characteristics of dynamic occlusion scenarios, parameter optimization rules are set to adaptively adjust window size, adapt update frequency to occlusion rate, and dynamically increase or decrease the number of iterations according to feature complexity, so as to ensure the feasibility of the strategy and the completion effect. The system integrates and processes strategy constraints, parameter setting standards, and optimization rules to generate dynamic completion annotation configuration base data that includes strategy type, parameter specifications, correlation logic, and optimization strategies.

5. The method as described in claim 4, characterized in that, Through the BEV temporal fusion algorithm, the mapping relationship between occlusion features and the complete shape and motion patterns of the target is learned. The parameters are dynamically weighted and optimized by fusion based on pixel semantics, 3D geometry, and temporal motion features. The core parameters and iterative strategies are adapted by combining occlusion type, target scale, scene dynamic rate, and camera motion state, including: Classification, extraction, and attribute alignment processing are performed on occlusion feature data, target complete shape information, motion law data, pixel semantic features, three-dimensional geometric features, and temporal motion feature related data to generate occlusion feature quantification information, target shape feature library, motion law adaptation rules, and initial values ​​of multimodal feature weights. Based on the precise optimization of completion parameters, the quantitative information of occlusion features is associated with the target morphology feature library and motion law adaptation rules. Combined with dynamic weighting logic, the influence weights and collaborative information of various features in the completion process are integrated to generate the basis for completion parameter optimization. Based on the core parameters and iterative strategies, the occlusion type, target scale, scene dynamic rate, and camera motion state data are respectively assigned to the parameter adaptation model and iterative strategy adjustment model to generate the initial adaptation information of the core parameters and iterative strategies. Based on the optimization criteria for completion parameters and the initial adaptation information of core parameters and iterative strategies, key attributes in the core data extracted by classification are marked and weighted to generate optimized completion parameter configurations and dynamic iterative strategy schemes.

6. A dynamic occlusion target completion and annotation system based on BEV temporal fusion, characterized in that, The system includes: The multi-dimensional BEV perception data acquisition and standardization module is used to acquire multi-source perception data including continuous temporal images, high-precision three-dimensional coordinates of the target, and fine contours of occluded areas. Through structured acquisition and standardization processing, it generates a standardized perception dataset containing data modality type, temporal correlation attributes, and quality assessment information. The completion annotation core parameter configuration module is used to set the occlusion completion confidence threshold, multimodal feature fusion priority weight and completion error tolerance range based on the target detection recall rate, precision requirements and annotation consistency specifications, and generate parameter configuration standards that meet quality control requirements. The integrated processing module for perception data is used to perform noise reduction and enhancement, spatiotemporal alignment, and semantic segmentation on multi-dimensional BEV perception data, extract occlusion-related features and predict target motion trends, sort the data according to their contribution to the completion, and output an ordered training data list. The dynamic completion annotation strategy formulation module is used to combine the complexity of the occlusion scene, the upper limit of hardware computing power and the real-time requirements of the algorithm, and set the sliding window size parameter according to the target feature dimension to generate an adaptive dynamic completion annotation strategy. The training dataset splitting and importing module is used to adaptively split the standardized dataset according to the completion annotation strategy to form training batches containing occluded target features, real labels and scene attributes, and synchronously import related data of the same type in the same scene through a parallel mechanism; The BEV temporal fusion completion and annotation execution module is used to integrate the output data of the aforementioned modules. It learns the mapping relationship between occlusion features and target shape and motion law through the BEV temporal fusion algorithm, dynamically weights and fuses multimodal features and adapts core parameters to complete dynamic occlusion target completion and annotation.

7. An electronic device, characterized in that, include: First processor; and memory for storing executable instructions of the first processor; The first processor is configured to execute the dynamic occlusion target completion and annotation method based on BEV temporal fusion as described in any one of claims 1 to 5 by executing the executable instructions.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the second processor, it implements the dynamic occlusion target completion and annotation method based on BEV temporal fusion as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Transform-based high-precision map real-time prediction method and system

    CN116071721A

  • Traffic target detection method and system based on cross-modal cross attention mechanism

    CN117173399A

  • Map construction method and device, vehicle, storage medium and computer program product

    CN118443005A

  • BEV perception optimization method based on semantic projection compensation and edge structure enhancement

    CN120656158A

  • System and Method for Multi-Modal Hyperspectral Image Generation with Cross-Modal Attention and Adaptive Quality Assurance

    US20250315932A1