A robust collaborative perception method and system based on temporal enhancement and impaired feature repair
Patent Information
- Application Number
- CN202610794730.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-03
- Publication Date
- 2026-09-25
AI Technical Summary
[0006]针对现有协同感知方法在非理想通信条件下易受传输时延与局部丢包影响,导致协作特征时效性不足、完整性受损以及多源特征融合稳定性下降的问题,本发明提供了一种基于时序增强与受损特征修复的鲁棒协同感知方法及系统
(1)本发明将自车端和协作端提取的当前鸟瞰图特征进行时序增强与受损特征修复,其中,通过对自车历史特征进行筛选并结合局部动态时序建模,得到包含历史上下文信息的时序增强特征;通过利用协作方历史特征对当前受损区域进行粗补偿,并结合受损感知引导的细修复网络恢复缺失区域的结构与语义信息,得到修复后的协作特征;将所述修复后的协作特征与自车当前特征输入进行可变形跨车聚合,得到空间补充特征;再对自车当前特征、空间补充特征以及时序增强特征进行融合,得到鲁棒性强且信息完整的融合特征;基于融合特征生成三维目标检测结果。本发明通过时序增强有效提升了自车特征对目标运动状态、场景连续结构及局部变化的表达能力,增强了当前感知特征的稳定性;本发明通过受损特征修复和时空区域自适应融合机制则实现了对受损协作信息的恢复以及对多源特征的稳健利用,显著提高了系统在非理想通信条件下的检测精度与鲁棒性。
Smart Images

Figure CN122821079A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of 3D target detection technology for autonomous driving based on computer vision, and in particular to a robust collaborative perception method and system based on temporal enhancement and damaged feature repair. Background Technology
[0002] In recent years, with the continuous development of autonomous driving technology, vehicle-to-everything (V2X) communication technology, and intelligent transportation infrastructure, environmental perception, as a key component of autonomous driving systems, undertakes the important task of identifying, locating, and understanding surrounding traffic participants, obstacles, and the road environment. Its performance directly affects the safety and reliability of subsequent planning and control. Three-dimensional target detection, as one of the core technologies in environmental perception, can output information such as the position, size, and category of targets, playing a crucial role in autonomous driving systems.
[0003] Traditional single-vehicle perception mainly relies on sensors such as onboard LiDAR, cameras, and millimeter-wave radar for environmental detection. However, due to limitations such as sensor installation location, detection range, and field-of-view obstruction, single-vehicle perception is prone to problems such as blind spots, missed detections, and inaccurate localization in scenarios such as curves, large vehicle obstructions, complex intersections, and long-distance target detection. To overcome the physical limitations of single-vehicle perception, collaborative perception achieves complementary multi-view and multi-spatial location perception information through information sharing between vehicles and between vehicles and roadside infrastructure, thereby effectively expanding the perception range and improving detection capabilities in complex traffic scenarios. Existing collaborative sensing methods can generally be categorized into three types: early-stage collaborative sensing, mid-stage collaborative sensing, and late-stage collaborative sensing. Early-stage collaborative sensing directly shares raw sensor data, preserving complete information, but incurs significant communication overhead. Late-stage collaborative sensing shares detection results, reducing communication burden, but suffers from greater loss of intermediate representation information, resulting in limited improvement in sensing performance. Mid-stage collaborative sensing achieves a better balance between sensing performance and communication costs by sharing intermediate features, thus becoming the mainstream technical approach in current collaborative sensing research. Existing mid-stage collaborative sensing methods mostly focus on feature extraction, feature compression, communication selection, or feature fusion, and have improved collaborative sensing performance to some extent. However, most existing cooperative sensing methods are designed based on ideal communication conditions. In real-world deployment environments, wireless links are susceptible to factors such as network congestion, bandwidth fluctuations, channel interference, and transmission delays, leading to insufficient timeliness and compromised integrity of cooperative features during transmission. On one hand, transmission delays disrupt the spatiotemporal correspondence between cooperative features and current vehicle features, making it difficult for the receiving end to reflect the current scene status in a timely manner. On the other hand, local packet loss can cause continuous gaps, incomplete structures, and semantic incoherence in cooperative features. If such unstable cooperative information is directly used for subsequent fusion, outdated or damaged information can be introduced, thereby reducing overall detection performance.
[0004] Furthermore, when collaborative information is affected by latency and packet loss, relying solely on the vehicle's features in a single frame at the current moment is often insufficient to support a complete perception of the scene. Due to the inherent sparsity of single-frame observations by LiDAR, in situations involving long distances, occlusion, or rapid movement of dynamic targets, the current vehicle features have limited ability to express the target's motion state, the continuous structure of the scene, and local changes. Simultaneously, for collaborative features that are partially lost during transmission, relying solely on the currently remaining features is usually insufficient to recover the complete target outline and contextual structure.
[0005] Therefore, existing technologies still lack a collaborative perception scheme that can address non-ideal communication conditions while simultaneously compensating for the timeliness of collaborative information, recovering damaged collaborative features, and robustly fusing multi-source features. How to improve the availability of collaborative information under conditions of communication latency and local packet loss, and how to more stably and effectively utilize the vehicle's current features, historical time-series information, and supplementary collaborative information during the fusion phase, thereby enhancing the accuracy and robustness of 3D target detection in complex traffic scenarios, has become a pressing technical problem to be solved in this field. Summary of the Invention
[0006] To address the issues of existing cooperative sensing methods being susceptible to transmission delays and local packet loss under non-ideal communication conditions, leading to insufficient timeliness and integrity of cooperative features, and decreased stability of multi-source feature fusion, this invention provides a robust cooperative sensing method and system based on temporal enhancement and damaged feature repair. This method introduces a temporal enhancement module, a damaged feature repair module, a deformable cross-vehicle aggregation module, and a spatiotemporal adaptive fusion module to achieve joint modeling of the vehicle's historical information, cooperative historical information, and current multi-source features, thereby improving the accuracy and robustness of 3D target detection in complex communication environments.
[0007] To achieve the objective of this invention, the present invention provides a robust collaborative sensing method based on temporal enhancement and damaged feature repair, comprising the following steps: Step 1: Multiple intelligent agents acquire their own sensor data and convert the sensor data to a unified coordinate system to extract their respective initial features; Step 2: Input the current features and historical frame features of the vehicle into the temporal enhancement module for filtering and enhancement to obtain temporal enhanced features; Step 3: Perform packet loss modeling on the collaborative features, construct damaged features during transmission, and perform coarse compensation and fine repair on the damaged features to obtain the repaired collaborative features; Step 4: Input the vehicle features and the repaired collaborative features into the cross-vehicle aggregation module to perform feature alignment and spatial interaction to obtain spatial supplementary features; Step 5: Based on the perception reliability of different spatial regions, perform spatiotemporal region adaptive fusion of the vehicle features, the spatial supplementary features, and the temporal enhancement features to obtain the final fused features; Step 6: Input the final fused features into the detection head and output the three-dimensional target detection result.
[0008] The above technical solution first maps the LiDAR point clouds collected by the autonomous vehicle and the collaborating end to a unified reference coordinate system, and extracts the corresponding features through a point cloud coding network. Then, it uses the historical frame information of the autonomous vehicle to perform temporal enhancement on the current autonomous vehicle features, improving the ability of the current features to express the target motion state, continuous structure of the scene, and local changes. At the same time, to address the problem of local continuous area loss during the transmission of collaborative features, it uses the historical information of the collaborating party to perform coarse compensation for the damaged areas, and combines a fine repair network guided by damage perception to restore the structural and semantic information of the missing areas. On this basis, it performs deformable cross-vehicle aggregation on the repaired collaborative features to obtain spatial supplementary features.
[0009] Optionally, in step 1, the sensor data is lidar point cloud data; based on shared pose information, the point cloud data of the vehicle and the collaborating end are transformed to a unified reference coordinate system, and the corresponding features are extracted through the point cloud feature encoding network and the backbone network, thereby providing a unified spatial representation basis for subsequent temporal compensation, damage recovery and cross-agent fusion.
[0010] Optionally, in step 2, the temporal enhancement includes: firstly, performing temporal dimension aggregation on the vehicle's historical features and generating a historical filtering weight map in combination with the current features to filter out historical context regions that are more relevant to the current detection task; then, inputting the filtered historical enhanced features and the current features into a local dynamic temporal LSTM module, and jointly modeling the local dynamic change relationship and stable neighborhood structure information between consecutive moments through a local neighborhood attention branch and a local convolutional context branch, thereby obtaining the temporal enhanced features.
[0011] Optionally, in step 3, the construction of the damaged features includes: applying a block-shaped random mask to the current features of the collaborating end to simulate the loss of local continuous regions during wireless link transmission; the coarse compensation includes using historical features of the collaborating end to perform cross-attention compensation on the currently missing region to restore the basic structure and contextual clues of the missing location; the fine repair includes combining the damaged perception guidance information and using a convolutional encoder-decoder network to further refine the coarse compensation results to improve the structural continuity, semantic consistency and local detail integrity of the collaborating features.
[0012] Optionally, in step 4, the cross-vehicle aggregation module is a deformable cross-vehicle aggregation module. The module performs local sparse sampling, alignment and spatial interaction on the repaired collaborative features through learnable sampling positions to alleviate the problem of cross-agent feature inconsistency caused by perspective differences, scale changes and local offsets, and generates the spatial supplementary features.
[0013] Optionally, in step 5, the spatiotemporal adaptive fusion includes: dividing the spatial region into a strong perception region and a weak perception region based on the perception reliability of the vehicle's current features at different spatial locations; in the strong perception region, fusion is performed primarily based on the vehicle's current features to maintain the representation stability of the reliable region; in the weak perception region, importance estimation is performed on the vehicle's current features, spatial supplementary features, and temporal enhancement features, and adaptive weighted fusion is performed to enhance the information completion capability in the weak perception region, thereby obtaining the final fused features. Optionally, in step 6, the detection head includes a classification branch and a regression branch, which are used to output the target category probability and the target bounding box parameters, respectively; the target bounding box parameters include at least the target center position, size and orientation information.
[0014] The present invention also provides a robust collaborative sensing system based on temporal enhancement and damaged feature repair.
[0015] The present invention also provides a computer device.
[0016] The present invention also provides a computer-readable storage medium.
[0017] Compared with the prior art, the present invention can achieve at least the following technical effects: (1) This invention performs temporal enhancement and damaged feature repair on the current bird's-eye view features extracted from the vehicle and the collaborating end. Specifically, by filtering the historical features of the vehicle and combining them with local dynamic temporal modeling, temporal enhancement features containing historical context information are obtained. By using the historical features of the collaborating party to perform coarse compensation on the current damaged area and combining it with a fine repair network guided by damage perception to restore the structure and semantic information of the missing area, repaired collaborative features are obtained. The repaired collaborative features are combined with the current features of the vehicle to perform deformable cross-vehicle aggregation to obtain spatial supplementary features. The current features of the vehicle, spatial supplementary features and temporal enhancement features are then fused to obtain robust and information-complete fused features. Based on the fused features, three-dimensional target detection results are generated. This invention effectively improves the ability of vehicle features to express the target motion state, continuous structure of the scene and local changes through temporal enhancement, and enhances the stability of the current perception features. This invention realizes the recovery of damaged collaborative information and robust utilization of multi-source features through damaged feature repair and spatiotemporal adaptive fusion mechanism, which significantly improves the detection accuracy and robustness of the system under non-ideal communication conditions.
[0018] (2) This invention can achieve more robust compensation, repair and fusion of multi-source features under non-ideal communication environments. In order to address the problems of insufficient timeliness of cooperative information caused by transmission delay in cooperative perception, as well as feature gaps, structural damage and semantic incoherence caused by local packet loss, the invention strengthens the support role of the vehicle's historical context for current perception by designing a temporal enhancement module, and restores the structural continuity and semantic integrity of cooperative features by a damaged feature repair module. Furthermore, by combining a deformable cross-vehicle aggregation module and a spatiotemporal adaptive fusion module, the invention achieves differentiated utilization of the vehicle's current features, spatial supplementary features and temporal enhancement features. While maintaining the stability of reliable regional features, the invention fully explores the cooperative complementary information in the weak perception areas, and finally achieves higher fusion quality and stronger environmental perception robustness, resulting in better three-dimensional target detection performance. Attached Figure Description
[0019] Figure 1 The flowchart illustrates a robust collaborative sensing method based on temporal enhancement and damaged feature repair, provided in an embodiment of the present invention.
[0020] Figure 2 This is a schematic diagram of the overall framework of a robust collaborative sensing method based on temporal enhancement and damaged feature repair provided in an embodiment of the present invention.
[0021] Figure 3 A schematic diagram of a timing enhancement module provided in an embodiment of the present invention.
[0022] Figure 4 A schematic diagram of a damaged feature repair module provided in an embodiment of the present invention. Detailed Implementation
[0023] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0024] This invention proposes a robust cooperative sensing method based on temporal enhancement and damaged feature repair to address the problems of insufficient timeliness and integrity of cooperative information, as well as decreased stability of subsequent fusion, caused by transmission delay and local packet loss in real wireless communication environments in existing cooperative sensing methods. This method focuses on the degradation phenomenon of cooperative sensing under non-ideal communication conditions, constructing an overall technical solution for temporal enhancement, damaged feature repair, deformable cross-vehicle aggregation, and spatiotemporal adaptive fusion. Within a unified bird's-eye view space, it completes temporal compensation of multi-source features, damaged feature recovery, cross-agent interaction, and detection output.
[0025] Example 1 See Figure 1 and Figure 2 The present invention provides a robust collaborative sensing method based on temporal enhancement and damaged feature repair, comprising the following steps: Step 1: Multiple intelligent agents acquire their own sensor data and convert the sensor data to a unified coordinate system to extract their respective initial features.
[0026] In this step, a smart agent is selected as the autonomous vehicle, and a communication connection is established with at least one cooperative smart agent within its communication range; the LiDAR point cloud data collected by the autonomous vehicle and each cooperative smart agent at the current moment is acquired, and the corresponding historical point cloud data is also acquired; based on shared pose information, the point cloud data is uniformly transformed to the coordinate system of the autonomous vehicle; the point cloud data of the autonomous vehicle and the cooperative smart agents are processed using a grid-based feature encoding network to generate corresponding initial bird's-eye view features. The current features of the autonomous vehicle can be represented as follows: The current characteristics of the collaborating party can be represented as: The historical feature sequence of a vehicle can be represented as The historical feature sequence of the collaborating parties can be represented as , Indicates the index of the vehicle. An index representing the collaborating parties that establish communication connections with the vehicle. Indicates the current moment. This indicates the number of historical time steps traced backward. Indicates a bicycle Before the current moment, the Historical features extracted from historical moments Represents a collaborative intelligent agent Before the current moment, the Historical features extracted from historical moments.
[0027] Step 2: Perform temporal enhancement on the current features and historical features of the vehicle to obtain rich historical context information and obtain temporally enhanced features.
[0028] By filtering the vehicle's historical feature sequences and combining them with local dynamic temporal modeling, historical context information more relevant to the current detection task is extracted, thereby enhancing the expressive power of the vehicle's current features. The overall process can be represented as follows:
[0029] in, This represents the vehicle features after time-enhanced processing. This represents the temporal augmentation mapping. Through this step, the vehicle's current features are transformed from a single-moment observation representation into an augmented representation supported by historical context, thereby improving its ability to characterize the target's motion state, the continuous structure of the scene, and local changes.
[0030] Step 3: Perform packet loss modeling on the current characteristics of the collaborating parties, construct damaged collaborative features during transmission, and perform coarse compensation and fine repair on the damaged collaborative features based on the damaged collaborative features and the historical feature sequence of the collaborating parties to obtain the repaired collaborative features.
[0031] In this embodiment, a block-based random masking strategy is used to model packet loss based on the current features of the collaborating parties. The damaged collaborative features can be represented as follows:
[0032] in, Represents a binary mask. This represents element-wise multiplication. This indicates a damaged collaborative characteristic.
[0033] Furthermore, during the repair of damaged features, the historical features of collaborating parties are first used to perform coarse compensation for the currently missing region. Then, a fine repair network guided by damage perception is used to refine the coarse compensation result. The overall process can be represented as follows:
[0034] in, This indicates the cooperative features after the repair. This indicates the repair process of damaged features.
[0035] Step 4: Perform cross-agent local alignment and spatial interaction between the current features of the vehicle and the repaired collaborative features to transform cross-vehicle aggregation and obtain spatial supplementary features.
[0036] The deformable cross-vehicle aggregation mitigates the inconsistency of cross-agent features caused by viewpoint differences, scale variations, and local offsets by locally sparsely sampling and aligning the repaired collaborative features using learnable sampling locations. The overall process of this step can be represented as follows:
[0037] in, Indicates spatial supplementary features, Represents a deformable cross-vehicle aggregation mapping. This indicates the number of collaborating parties involved in the collaboration.
[0038] Step 5: Based on the perception reliability of different spatial regions, the current features of the vehicle, the spatial supplementary features, and the temporal enhancement features are spatiotemporally adaptively fused to obtain the final fused features.
[0039] In this embodiment, the spatiotemporal adaptive fusion performs differentiated fusion of feature maps based on the perception reliability of different spatial regions. In regions with reliable perception, conservative fusion is performed, primarily based on the vehicle's current features. In regions with weak perception, adaptive weighted fusion is performed on the vehicle's current features, spatial supplementary features, and temporal enhancement features. The overall process can be represented as follows:
[0040] in, Indicates the final fusion characteristics, This indicates spatiotemporal region adaptive fusion.
[0041] The adaptive fusion of three features in the weakly perceived region can be further expressed as:
[0042] in, , and Let represent the fusion weights corresponding to the vehicle branch, spatial branch, and temporal branch, respectively, and satisfy the following:
[0043] Finally, the outputs from the strong-sensing region and the weak-sensing region are combined to obtain the final fused feature.
[0044] Step 6: Decode the final fused features and output the 3D target detection results.
[0045] In this step, the final fusion features will be... The data is input to a 3D target detection head. Through the classification and regression branches in the detection head, the target category probability distribution and the geometric parameters of the bounding box in 3D space are output in parallel. The geometric parameters include at least the center point coordinates, size and orientation, thereby completing the collaborative perception task.
[0046] Example 2 This embodiment further refines the key steps in Embodiment 1; see [link to Embodiment 1]. Figure 3 and Figure 4 A robust collaborative sensing method based on temporal enhancement and damaged feature repair includes the following steps: Step 1: Data Alignment and Feature Extraction. LiDAR, a key sensor in autonomous driving perception systems, generates point cloud data containing 3D coordinates and reflection intensity by emitting laser beams and receiving echoes, providing accurate spatial geometric information for 3D target detection. The purpose of this step is to transform the raw, unordered point cloud data into structured bird's-eye view features in a unified coordinate system, providing a unified input for subsequent temporal modeling and cross-agent feature fusion. Specifically, this includes: Step 1.1, Point Cloud Preprocessing and Spatial Discretization: The vehicle and the cooperative agent first convert the current and historical point clouds they collected into the vehicle's reference coordinate system based on shared pose information. Then, the three-dimensional space is discretized within a preset spatial range, and the point clouds are divided into corresponding voxel or columnar voxel units according to their three-dimensional coordinates, retaining only non-empty voxel units that contain valid point cloud data.
[0047] Step 1.2, Voxel Point Cloud Normalization: For each non-empty acceleration unit, a point count threshold T can be set. When the number of points within a voxel is greater than T, it is randomly downsampled; when the number of points within a voxel is less than T, it is zero-padded to ensure that all voxels have a uniform data organization form, thereby facilitating batch parallel processing.
[0048] Step 1.3, Voxel Feature Encoding: For T points within a voxel, calculate their relative coordinates to the voxel center and combine them with the point's absolute coordinates and reflection intensity to form a rich initial local feature set. Input the initial local features into a point cloud feature encoding network for feature learning. Aggregate the features of all points within the voxel using max pooling to obtain a fixed-dimensional feature vector. This feature vector robustly represents the overall geometric structure and semantic information of the voxel unit. Map the final feature vectors of all non-empty voxels to their original spatial positions on the XY plane to form a structured initial bird's-eye view feature map (structured feature representation). Input this initial bird's-eye view feature map into the backbone network for further feature extraction to obtain the initial features corresponding to the autonomous vehicle agent and the cooperative agent.
[0049] Step 2, timing enhancement.
[0050] Step 2.1, Historical Feature Aggregation. First, the vehicle's historical features are accumulated over time to obtain a historical feature representation. :
[0051] in, This indicates the number of historical frames. This historical feature represents the contextual information retained for integrating continuous observations from the vehicle. This represents the time index in the historical frame sequence.
[0052] Step 2.2, Historical Information Filtering. To extract the most relevant and useful regions from the vehicle's historical feature representations for the current moment, the vehicle's current features are filtered... and historical characteristics Perform average pooling and max pooling, and generate spatial statistical responses through convolution:
[0053]
[0054] in, Indicates channel splicing. This represents the current spatial response map generated from the vehicle's current features, used to characterize the information distribution of the current features in the spatial dimension. This represents a historical spatial statistical response map generated from historical features, used to describe the effectiveness of historical context information at different spatial locations.
[0055] Subsequently, the current spatial statistical response is merged with the historical spatial statistical response to generate a historical screening weight map. :
[0056] in, This represents the Sigmoid activation function.
[0057] Based on the historical filtering weight map, an adaptive weighted fusion of the vehicle's current features and historical feature representations is performed to obtain the filtered historical enhanced features. :
[0058] in, This indicates element-wise multiplication.
[0059] This method can adaptively balance the contributions of current features and historical features in spatial location, thereby suppressing redundant background and spatiotemporal mismatch information. Step 2.3: Local Dynamic Temporal Modeling. The current features of the vehicle, the filtered historical enhanced features, and the hidden state from the previous recursive time step are input into the local dynamic temporal LSTM module for fusion to obtain intermediate features. :
[0060] To enhance the modeling capability for local dynamic changes, query features are constructed within a fixed local window. Key features Sum value characteristics :
[0061] in These represent query mapping, key mapping, and value mapping, respectively, used to map input features to query features, key features, and value features.
[0062] The local dynamic temporal LSTM module includes a local neighborhood attention branch and a local convolutional context branch, used to simultaneously model local dynamic changes and local stable context information. In the local neighborhood attention branch, for spatial location... Let its local neighborhood be Then the neighborhood attention weight for:
[0063] in, This represents a neighborhood location within the local neighborhood. This indicates the spatial location of traversing the local neighborhood during normalized calculations. Indicates the key feature in the neighborhood position eigenvectors at location This represents the channel dimension of the feature vector. Indicates the spatial location of the query feature The feature vector at that location, This represents the vector transpose operation.
[0064] From this, we can obtain the spatial location. Local dynamic features :
[0065] in, Indication feature in neighborhood location The eigenvector at that location.
[0066] Meanwhile, to reduce the sensitivity of local attention to instantaneous noise and local misalignment, a local convolutional context branch is introduced to extract local contextual features:
[0067] in This represents the local context features extracted by the local convolutional context branch, which are used to supplement the relatively stable spatial structure information within the local neighborhood.
[0068] Then, the two sets of features are fused:
[0069] in This represents the enhanced input feature obtained by fusing local dynamic features and local contextual features. This represents the local dynamic features of the output of the local neighborhood attention branch.
[0070] And it performs recursive updates based on the LSTM gating mechanism:
[0071]
[0072]
[0073]
[0074]
[0075]
[0076] Finally, the hidden states are output as temporal augmentation features:
[0077] in This represents the forget gate, used to control the information that needs to be retained from the cell's state at the previous time step; This represents the input gate, used to control the proportion of the current candidate state written into the cell state; This indicates the current state of the candidate cells; This represents the output gate, used to control the output content in the current hidden state; Indicates the current state of the cell; , , , These represent the learnable convolutional weights corresponding to the forget gate, input gate, candidate state, and output gate, respectively. , , , These represent the corresponding bias terms.
[0078] Through the above steps, historical information can be stably accumulated over consecutive time intervals, and the ability to represent local dynamic changes, neighborhood structural relationships, and sparse observation scenarios can be enhanced.
[0079] Step 3: Repairing Damaged Features In real-world wireless links, cooperative features are prone to spatial continuity holes during transmission due to bandwidth fluctuations, channel interference, and local transmission failures, resulting in structural damage and semantic incoherence. To address this, a two-stage strategy of coarse historical compensation and fine-grained damage awareness repair is employed to gradually restore damaged cooperative features. The specific workflow of this module is as follows: Step 3.1, Packet Loss Modeling: For the first... Features of the cooperative agents at the current moment Construct a binary mask .in, Indicates position Successfully received. To indicate packet loss, a mask value of 1 indicates successful reception of the corresponding spatial block, while a mask value of 0 indicates that the corresponding spatial block has been lost. Therefore, the characteristics of a damaged cooperation can be represented as follows:
[0080] The binary mask is generated using a block-based random masking strategy, which divides the feature map into several spatial blocks of fixed size, and then randomly masks these local units according to a preset packet loss rate, thereby forming a spatially continuous missing region.
[0081] Step 3.2, Historical Coarse Compensation. For damaged cooperative features, the missing regions are usually continuously distributed, and it is difficult to recover the complete target contour relying solely on the current residual features. Therefore, historical frame information of the cooperative parties is introduced as a recovery reference. Specifically, firstly, the missing regions in the current damaged cooperative features are determined based on the binary mask, and query features are generated using the damaged cooperative features and their missing region information; simultaneously, the historical feature sequences of the cooperative parties, after historical filtering, are mapped to key features and value features, respectively. Thus, the coarse compensation module can initiate cross-attention queries only for the missing regions, retrieve temporal reference information related to the current missing regions from the historical frame feature sequences of the cooperative parties, and write the extracted temporal reference information into the missing regions to obtain the coarse compensation features. Let the query, key, and value features be respectively... and combined with relative time coding Let represent the time interval between historical frames and the current moment. Then, under the effect of cross-attention, the coarse compensation result can be expressed as:
[0082] in, This represents the coarse-compensated features obtained from the historical frame features of the collaborating party. Indicates the first The time interval between historical frames and the current moment. By introducing time coding, the timeliness of different historical frames can be distinguished, thereby reducing the interference of outdated historical information on the current compensation. After coarse compensation, the originally blank missing areas can be transformed into a coarsely restored state with basic contours and semantic clues.
[0083] Step 3.3, Damage-Aware Fine-Finish Repair. Since the coarse compensation result may still have issues such as blurred boundaries and uneven transitions in high-frequency structures like boundary details, local textures, and target contours, a damage-aware fine-finish repair module is further designed. This module first constructs a missing region indication map based on a binary mask, and generates location-related guiding weights through a damage-aware guiding branch to characterize the degree of damage and repair needs at different spatial locations. Subsequently, the guiding weights are used to modulate the coarse compensation features to obtain the coarse compensation features after damage-aware guidance. Let the fine-finish repair network be denoted as... The detailed repair result can be expressed as:
[0084] in, This represents the coarse compensation features after being guided by the perception of damage. This represents a convolutional encoder-decoder fine-tuning network. This represents the fine-repair result output by the fine-repair network. In each decoding stage, low-resolution features are first upsampled, then concatenated with the corresponding layer features from the encoding stage, and finally feature fusion and reconstruction are completed through two 3×3 convolutions. At the end of the network, a 1×1 convolution is used to complete channel mapping, projecting the restored result back to the original cooperative feature space.
[0085] The convolutional encoder-decoder fine-tuning network includes an encoder, a bottleneck layer, a decoder, and an output mapping layer. At the encoder, multi-level semantic features are extracted through step-by-step downsampling. At the decoder, spatial resolution and local details are restored through step-by-step upsampling combined with skip connections, resulting in a fine-tuning result. Only missing regions are selectively replaced, while the original features of successfully received regions are preserved, thus obtaining the repaired collaborative features.
[0086] The bottleneck layer is located between the encoder and decoder, and is used to integrate high-level semantic information and expand the feature perception range at a lower spatial resolution. The output mapping layer is used to perform channel transformation on the output features of the decoder so that its dimension is consistent with the original collaborative features, thereby facilitating subsequent replacement of missing regions and cross-vehicle aggregation.
[0087] Step 3.4, Selective Write-Back. Considering that the goal of this module is to finely complete missing regions rather than regenerate the entire feature map, only missing regions are replaced during the output stage, while the original features of successfully received regions remain unchanged. The final repaired collaborative features are:
[0088] Through the above processing, the structural continuity, semantic consistency, and local boundary integrity of the damaged collaborative features are improved.
[0089] Step 4, Deformable Cross-Vehicle Aggregation: After completing the cooperative feature repair, the current features of the vehicle and the repaired cooperative features are input into the deformable cross-vehicle aggregation module to perform deformable cross-vehicle aggregation, so as to alleviate the problem of cross-agent feature inconsistency caused by perspective differences, scale changes and local spatial offset.
[0090] Using vehicle features as a unified query basis, deformable sampling and information aggregation are performed within the local neighborhood of collaborative features to obtain spatially supplementary features. Specifically, for any spatial location in the vehicle feature map... First, the vehicle features at this spatial location are extracted and combined with the location code to construct a query representation. :
[0091] in, Represents a linear mapping matrix. Indicates position code, This represents the feature vector of the vehicle's current features at spatial location q. Subsequently, sparse sampling and local alignment of the collaborative features are performed using learnable sampling locations, including: first, predicting several sampling offsets from the query representation, and generating corresponding deformable sampling locations in the collaborative feature map. :
[0092] in, Indicates the first Individual attention is focused on spatial location For the first The first collaborator predicted the first Each sampling offset.
[0093] Subsequently, based on the deformable sampling location and the corresponding attention weight, local information aggregation is performed within the cooperative feature neighborhood to obtain the spatial location. Spatial supplementary features:
[0094] in, Indicates the number of heads of attention. This represents the number of sampling points corresponding to each attention head. Indicates the number of collaborating parties. Indicates the first The output mapping matrix of each attention head This represents the corresponding attention weight. This indicates that in the collaborative features repaired by the j-th collaborator, according to the deformable sampling position... The extracted collaborative features.
[0095] The above method eliminates the need for strict positional correspondence and enables adaptive searching within the collaborative feature neighborhood around the vehicle's current location. This allows for more flexible extraction of supplementary information related to the current query and provides a more robust spatial feature representation for subsequent spatiotemporal adaptive fusion.
[0096] Step 5: Spatiotemporal region adaptive fusion Step 5.1, Division of Strong and Weak Perception Regions. First, based on the reliability of the vehicle's current characteristics at different spatial locations, the spatial region is divided into strong perception regions. and weak perception area Among them, the strong perception area corresponds to the location where the vehicle's perception is relatively sufficient, while the weak perception area corresponds to the location that needs to be supplemented with collaborative information and temporal information.
[0097] Step 5.2, Strong Perception Region Fusion. In the strong perception region, a conservative fusion strategy dominated by the vehicle's current features is adopted, allowing only spatial and temporal branches to participate in a restricted manner, that is, limiting the participation intensity of the spatial supplementary features and the temporal enhancement features, specifically expressed as follows:
[0098] in, and These represent the upper limit coefficients for the participation of spatial supplementary information and temporal enhancement information in the highly perceptible region, respectively. This design can supplement situations where local boundaries are blurred or slightly missing, while maintaining the dominance of vehicle features.
[0099] Step 5.3, Weak Perception Region Fusion. In the weak perception region, importance scores are calculated for the vehicle branch, spatial branch, and temporal branch (corresponding to vehicle features, spatial supplementary features, and temporal enhancement features, respectively):
[0100]
[0101]
[0102] in, This represents a confidence map generated from the current features of the vehicle. This represents a confidence map generated from spatially supplemented features. This represents a confidence plot generated from temporal enhancement features. This represents a location prior map. These represent the importance estimation functions corresponding to the three branches.
[0103] Subsequently, the fusion weights are obtained through Softmax normalization:
[0104] and satisfy
[0105] Therefore, the fusion result of the weakly sensed regions is as follows:
[0106] This method enables adaptive decisions in areas with weak perception to retain more vehicle information at the current location, as well as to introduce more spatial supplementary information and temporal enhancement information.
[0107] Step 5.4, Final Fusion Output. Finally, the outputs of the strong sensing region and the weak sensing region are combined to obtain the final fused feature:
[0108] Therefore, while maintaining the reliable regional perception stability of the vehicle, it can make fuller use of the complementary spatial information from the collaborative perspective and the historical temporal context information to form a more robust and complete spatiotemporal fusion representation.
[0109] Step 6: Feature Decoding and Target Detection Final fusion features A 3D target detection head is fed in to generate detection results. Specifically, this includes: Step 6.1, Classification and Regression: A series of 3D anchor boxes of different sizes and orientations are predefined at each spatial location on the feature map. For each anchor box, the detection head performs two tasks in parallel. The classification branch predicts the probability that each anchor box contains a specific class of object (e.g., vehicle, pedestrian, cyclist) using convolutional layers. The regression branch predicts the parametric offset of each anchor box relative to its matched ground truth bounding box using convolutional layers; these parameters typically include the 3D coordinates of the object's center point. ,size and orientation angle .
[0110] Step 6.2, Post-processing: Using the non-maximum suppression (NMS) algorithm, the large number of overlapping detection boxes generated by the prediction are filtered to remove redundant boxes, and finally a series of refined 3D bounding boxes and their class labels are output as the final result of collaborative perception.
[0111] The collaborative perception model proposed in this invention (including a temporal enhancement module, a damaged feature repair module, a deformable cross-vehicle aggregation module, and a spatiotemporal region adaptive fusion module) has an overall loss function during training that includes classification loss and regression loss, in the following form: in, For classification loss, focus loss is used to address the class imbalance problem between foreground and background anchor boxes; For the regression loss, Smooth L1 Loss is used to regress the bounding box parameters; and Weighting coefficients to balance the two types of losses.
[0112] Experiments were conducted to verify the technical effectiveness of the methods described in the aforementioned embodiments.
[0113] This experiment validated the method on the publicly available cooperative sensing datasets OPV2V, V2XSet, and DAIR-V2X. The OPV2V dataset was used to verify the model's detection capability in simulated V2V scenarios; the V2XSet dataset was used to further verify the model's adaptability in general V2X multi-agent cooperative environments; and the DAIR-V2X dataset was used to evaluate the detection performance of the proposed method in real-world vehicle-to-infrastructure (V2I) scenarios. By conducting experiments on these three datasets, the effectiveness and generalization ability of the proposed method in simulated scenarios, general cooperative scenarios, and real-world scenarios can be comprehensively verified.
[0114] The method in this embodiment of the invention is implemented based on the OpenCOOD framework. The backbone network uses PointPillars as the point cloud feature extraction network. The width and length of the voxels are set to 0.4 meters, and the height is set to 4 meters. The input point cloud range is set to x∈[-100.8 m, 100.8 m], y∈[-40 m, 40 m], z∈[ [3 m, 1 m]. During the experiment, communication latency, noise interference, and packet loss issues from collaborative sensing were introduced to simulate non-ideal communication scenarios. The Adam optimizer was used during training, with an initial learning rate set to [value missing]. The learning rate was decayed every 10 epochs with a decay factor of 0.1. All models were trained on eight NVIDIA GeForce RTX 3090 GPUs. The experimental environment also included an Intel Core i9-10980XE processor, Ubuntu 22.04 operating system, CUDA 11.3, and PyTorch 1.12 deep learning framework.
[0115] The model evaluation uses Average Precision (AP), a widely used metric in 3D object detection, as the core indicator, and evaluates it at Intersection over Union (IoU) thresholds of 0.5 and 0.7 (AP@0.5 and AP@0.7). To fully verify the superior performance of the method in this embodiment, it is compared with various baseline algorithms and representative cooperative sensing models, including No Collaboration, Late Fusion, When2com, V2VNet, V2X-ViT, DiscoNet, and Where2comm. To verify the robustness of the method in this embodiment under non-ideal communication conditions, comparative experiments are further conducted under joint non-ideal conditions of 100 ms communication delay, 0.2m positioning error, 0.2° heading angle error, and 20% packet loss rate. The experimental comparison results are shown in Table 1.
[0116] Table 1 presents a comparison of experimental results between the method proposed in this embodiment and other representative collaborative sensing methods in the field. As can be seen from the data in Table 1, the method proposed in this embodiment significantly outperforms the comparative methods in AP@0.5 and AP@0.7 on the OPV2V, V2XSet, and DAIR-V2X datasets, demonstrating its excellent detection performance and generalization ability.
[0117] Table 1 Performance Evaluation Table of the Method of the Present Invention and Existing Models in the Field
[0118] Example 3 This embodiment provides a robust collaborative sensing system based on temporal enhancement and damaged feature repair, used to implement the method described in the foregoing embodiments. The system includes: The data acquisition module is used by multiple agents to acquire their own sensor data, transform the sensor data into a unified coordinate system, and then extract their respective initial features, including current features and historical features. The temporal enhancement module is used to filter and enhance the current features and historical features of the vehicle to obtain temporal enhanced features; The damaged feature repair module is used to model the current damaged features of the collaborator, and combine the collaborator's historical features to perform coarse compensation and fine repair to obtain the repaired collaborative features. The cross-vehicle aggregation module is used to align and spatially interact with the current features of the vehicle and the repaired collaborative features to obtain spatial supplementary features. The spatiotemporal adaptive fusion module is used to adaptively fuse the vehicle's current features, spatial supplementary features, and temporal enhancement features to obtain the final fused features. The decoding module is used to decode the final fused features and output the 3D target detection results.
[0119] This invention introduces a temporal enhancement module, a damaged feature repair module, a deformable cross-vehicle aggregation module, and a spatiotemporal adaptive fusion module into the collaborative perception process under non-ideal communication conditions. The temporal enhancement module enhances the ability of current features to express the target's motion state, continuous scene structure, and local changes by filtering historical features of the vehicle and performing local dynamic temporal modeling. The damaged feature repair module effectively restores the structural continuity and semantic integrity of collaborative features under local packet loss conditions through a two-stage strategy combining coarse historical compensation and fine damaged perception repair. Furthermore, the deformable cross-vehicle aggregation module alleviates feature inconsistencies caused by perspective differences, scale variations, and local offsets through local deformable sampling and cross-agent information interaction. The spatiotemporal adaptive fusion module differentiates and fuses the vehicle's current features, spatial supplementary features, and temporal enhancement features according to the perception reliability of different spatial regions, maintaining the stability of features in reliable regions while fully utilizing complementary collaborative information and historical temporal information in weakly perceived regions. The synergistic effect of these modules mitigates the adverse effects of latency, noise disturbances, and local packet loss on collaborative perception performance under non-ideal communication conditions, thereby achieving superior 3D target detection results.
[0120] Example 4 This embodiment provides a computer device, including a memory, a processor, and a computer program stored in the memory. When the computer program is executed by the processor, it implements a robust collaborative perception method based on temporal enhancement and damaged feature repair as described in the foregoing embodiment.
[0121] Example 5 This embodiment provides a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, it implements a robust collaborative sensing method based on temporal enhancement and damaged feature repair as described in the foregoing embodiment.
[0122] The system, device, and medium described herein have the same technical effects as those achieved by the methods described in the foregoing embodiments.
[0123] This document uses specific examples to illustrate the principles and implementation methods of the present invention. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of the present invention. Furthermore, those skilled in the art will recognize that, based on the ideas of the present invention, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of the present invention.
Claims
1. A robust collaborative sensing method based on temporal enhancement and damaged feature repair, characterized in that, Includes the following steps: Multiple intelligent agents acquire their own sensor data and transform the sensor data into a unified coordinate system, thereby extracting their respective initial features, including current features and historical features; The current and historical features of the vehicle are filtered and enhanced to obtain temporally enhanced features; Model packet loss based on the current characteristics of the collaborating party, and construct damaged features during transmission; The damaged features are coarsely compensated and finely repaired to obtain the repaired cooperative features; The current features of the vehicle are aligned and spatially interacted with the repaired collaborative features across intelligent agents to obtain spatial supplementary features. Based on the perception reliability of different spatial regions, the current features of the vehicle, spatial supplementary features, and temporal enhancement features are adaptively fused in a spatiotemporal region to obtain the final fused features. The final fused features are decoded to output the 3D target detection results.
2. The robust collaborative sensing method based on temporal enhancement and damaged feature repair according to claim 1, characterized in that, The sensor data is transformed to a unified coordinate system, and then their respective initial features are extracted, including: The lidar point cloud data collected by the vehicle and its collaborators will be transformed into a unified reference coordinate system based on shared pose information. The transformed point cloud space is discretized, and the point cloud is mapped to the corresponding voxel unit; For non-empty cells, the spatial coordinate information and reflection intensity information of their internal points are aggregated to generate structured feature representations; The structured feature representation is input into the backbone network to obtain the corresponding features.
3. The robust collaborative sensing method based on temporal enhancement and damaged feature repair according to claim 1, characterized in that, The process of filtering and enhancing the current and historical features of the vehicle to obtain time-enhanced features includes: By performing time-dimensional aggregation on the historical features of the vehicle, a representation of the historical features is obtained; Average pooling and max pooling are performed on the current and historical feature representations of the vehicle respectively to extract spatial statistical information; A historical screening weight map is generated based on spatial statistical information, and the current features and historical features of the vehicle are adaptively weighted and fused based on the historical screening weight map to obtain the screened historical enhanced features. The current features of the vehicle, the filtered historical enhanced features, and the hidden state of the previous recursive time step are input into the local dynamic temporal LSTM module to obtain temporal enhanced features. The local dynamic temporal LSTM module includes a local neighborhood attention branch and a local convolutional context branch, which are used to simultaneously model local dynamic change relationships and local stable context information.
4. The robust collaborative sensing method based on temporal enhancement and damaged feature repair according to claim 1, characterized in that, The step of modeling packet loss based on the current characteristics of the collaborating party and constructing damaged features during transmission includes: The current features of the collaborating party are divided into multiple spatial blocks; a binary mask corresponding to the spatial blocks is generated based on a preset packet loss rate; the binary mask is applied to the current features of the collaborating party; and the damaged collaborative features are obtained based on the binary mask and the current features of the collaborating party.
5. A robust collaborative sensing method based on temporal enhancement and damaged feature repair according to claim 4, characterized in that, The coarse compensation includes: Query features are generated using damaged collaboration features, with collaborators' historical features serving as key and value features; missing regions are determined based on binary masks, and cross-attention queries are initiated only for these missing regions to extract temporal reference information corresponding to the missing regions from the collaborators' historical frame feature sequences; the extracted temporal reference information is written into the missing regions to obtain coarse compensation features; and / or, The detailed repair specifically includes: Based on the binary mask, damage perception guidance information is constructed, and the damage perception guidance information is used to modulate the coarse compensation features. The modulated coarse compensation features are input into a convolutional encoder-decoder fine repair network. The convolutional encoder-decoder fine repair network extracts multi-level semantic features by downsampling at the encoding end, and restores spatial resolution and local details by upsampling at the decoding end and combining skip connections, thus obtaining the fine repair result. Only the missing regions are selectively replaced, while the original features of the successfully received regions remain unchanged, thereby obtaining the repaired cooperative features.
6. The robust collaborative sensing method based on temporal enhancement and damaged feature repair according to claim 1, characterized in that, Based on the vehicle's current features and the repaired collaborative features, spatial supplementary features are obtained by sparse sampling and local alignment of the collaborative features using learnable sampling locations; and / or, Based on the perception reliability of the vehicle's current features at different spatial locations, the spatial region is divided into a strong perception region and a weak perception region. In the strong perception region, the vehicle's current features are the primary focus of fusion, while limiting the participation intensity of the spatial supplementary features and the temporal enhancement features. In the weak perception region, the importance weights of the vehicle's current features, the spatial supplementary features, and the temporal enhancement features are estimated respectively, and the three features are adaptively weighted and fused according to the importance weights. The final fused feature is obtained by combining the fusion results of the strong sensing region and the weak sensing region.
7. A robust collaborative sensing method based on temporal enhancement and damaged feature repair according to any one of claims 1-6, characterized in that, The decoding of the final fused features and the output of the 3D target detection result include: The final fused features are input into the 3D target detection head, and the target category probability is output through the classification branch and the target bounding box parameters are output through the regression branch; the target bounding box parameters include at least the target center position, size and orientation information.
8. A robust collaborative sensing system based on temporal enhancement and damaged feature repair, characterized in that, The system for implementing the method according to any one of claims 1-7 comprises: The data acquisition module is used by multiple agents to acquire their own sensor data, transform the sensor data into a unified coordinate system, and then extract their respective initial features, including current features and historical features. The temporal enhancement module is used to filter and enhance the current features and historical features of the vehicle to obtain temporal enhanced features; The damaged feature repair module is used to model the current damaged features of the collaborator, and combine the collaborator's historical features to perform coarse compensation and fine repair to obtain the repaired collaborative features. The cross-vehicle aggregation module is used to align and spatially interact with the current features of the vehicle and the repaired collaborative features to obtain spatial supplementary features. The spatiotemporal adaptive fusion module is used to adaptively fuse the vehicle's current features, spatial supplementary features, and temporal enhancement features to obtain the final fused features. The decoding module is used to decode the final fused features and output the 3D target detection results.
9. A computer device, characterized in that, The system includes a memory, a processor, and a computer program stored in the memory. When the computer program is executed by the processor, it implements a robust collaborative sensing method based on temporal enhancement and damaged feature repair as described in any one of claims 1 to 7.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements a robust collaborative perception method based on temporal enhancement and damaged feature repair as described in any one of claims 1-7.