A collaborative perception interference management method based on hierarchical reinforcement learning

By adopting a hierarchical reinforcement learning algorithm in a collaborative perception system to dynamically allocate spectrum and transmission power, the problems of resource allocation complexity and reduced perception accuracy in wireless environments are solved, achieving higher perception accuracy and resource utilization efficiency.

CN119012396BActive Publication Date: 2025-09-16SOUTHEAST UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411092209.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-09
Publication Date
2025-09-16
Estimated Expiration
2044-08-09

AI Technical Summary

Technical Problem

Existing collaborative sensing systems have difficulty in achieving complete information transmission in complex dynamic wireless environments with limited radio resources, resulting in reduced perception accuracy, complex resource allocation strategies, and high communication overhead.

Method used

Using a hierarchical reinforcement learning approach, a centralized deep reinforcement learning algorithm is designed to allocate resources and optimize perception performance in a collaborative sensing interference management scenario. Using state observations, action spaces, and reward functions, the algorithm trains an intelligent agent to dynamically allocate spectrum and transmission power, improving perception accuracy and reducing bandwidth requirements.

Benefits of technology

By learning resource allocation strategies, vehicles can upload more and more useful features to roadside units, improving the recognition accuracy of the overall scene while achieving a balance between perception and communication.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119012396B_ABST
    Figure CN119012396B_ABST
Patent Text Reader

Abstract

The present invention discloses a collaborative perception interference management method based on hierarchical reinforcement learning: The present invention studies the interference management problem in vehicle-to-infrastructure (V2I) communications in a collaborative perception scenario, models the spectrum allocation and power control problems as a Markov decision process, uses the spatial confidence of vehicle-side features and the channel state information of the communication link as the observed state space, designs a reward constraint model that simultaneously considers the communication rate and perception accuracy, and adopts an interference management strategy driven by a hierarchical reinforcement learning model. The present invention strikes a balance between communication rate and perception accuracy, effectively improving perception performance compared to random selection and maximum rate selection methods, while approaching the upper limit of perception performance under high bandwidth conditions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of combining wireless communication and machine learning in Internet of Vehicles (IoV), involves deep learning and IoV technology, and is a method for collaborative perception of transmission perception features in IoV communication. Background Art

[0002] Collaborative perception technology facilitates the exchange of perception information between connected autonomous vehicles (CAVs) and roadside units (RSUs), overcoming the limitations of independent perception and improving the accuracy of the collaborative system's environmental perception. Accurate environmental perception is essential for autonomous driving, ensuring the system effectively interprets its surroundings. However, independent perception systems are limited by their limited perception range and obstructed vision at long distances, exposing potential risks. These limitations can lead to serious safety issues, highlighting the performance limitations of independent perception systems. To address this challenge, the industry has turned to the development of collaborative perception systems, leveraging vehicle-to-vehicle (V2V) and vehicle-to-infrastructure (V2I) communications to achieve joint perception of the surrounding environment. This facilitates data exchange and information integration between different vehicles and roadside units (RSUs), aiming to significantly improve the accuracy and reliability of the overall perception system. This approach leverages collaborative perception to broaden the perception field, reduce blind spots, and enhance the safety and efficiency of autonomous driving.

[0003] Collaborative perception can be categorized into three main types: early-stage collaboration, mid-stage collaboration, and late-stage collaboration. Early-stage collaboration occurs at the raw sensor data level, integrating raw data inputs from other vehicles and RSUs, such as point clouds and RGB images. Early-stage collaboration faces the greatest challenge in practical deployments due to its high demand for communication resources. In contrast, late-stage collaboration focuses on exchanging perception results. While significantly reducing the demand for communication bandwidth, it is susceptible to positioning errors and data transmission errors. Therefore, collaborative perception often considers mid-stage collaboration, which strikes a balance between communication overhead and perception performance through the exchange of feature-level information.

[0004] Existing methods rely on the ideal assumption of complete information transmission between CAV vehicles and RSUs. In complex dynamic wireless environments with limited radio resources, this assumption is difficult to achieve, thereby reducing the feasibility of practical applications. Therefore, effectively allocating wireless resources to improve the perception accuracy of collaborators becomes a key issue. It is very difficult to directly find the best solution to this problem because it requires frequent exchange of rapidly changing feature information between RSUs and vehicles, which incurs huge communication overhead and makes the problem very complex. To this end, based on deep learning technology, the present invention develops an algorithm based on hierarchical reinforcement learning (RL), which is based on centralized deep reinforcement learning (DRL) and is used to allocate resources to optimize perception performance in collaborative perception interference management scenarios. Summary of the Invention

[0005] Technical problem: In response to the shortcomings of the existing technology, the present invention provides a collaborative perception interference management method based on hierarchical reinforcement learning. By utilizing the self-learning ability of reinforcement learning, a resource allocation strategy is learned to help vehicles upload more and more useful features to roadside units for feature fusion, thereby improving the recognition accuracy of the overall scene.

[0006] Technical solution: The present invention adopts a collaborative perception interference management method based on hierarchical reinforcement learning, which includes the following steps:

[0007] Spectrum allocation and power control for collaborative sensing scenarios consider communication rate and perception accuracy. The spectrum allocation and power control problems are modeled as Markov decision processes, and the state observation space, action space, and reward function are designed. A hierarchical reinforcement learning-based intelligent agent is trained to dynamically allocate spectrum and transmission power to complete collaborative sensing, thereby improving the perception accuracy of roadside units and reducing the demand for transmission bandwidth. The interference management method is carried out in the following steps:

[0008] Step 1: The roadside unit (RSU) and all the intelligent communication vehicles (CAV) participating in the collaborative perception collect the original point cloud data, which is represented by X m and X r , where X m represents the point cloud features of the mth intelligent communication vehicle, X r Point cloud features representing roadside units.

[0009] Step 2: CAV aligns its own point cloud data with the coordinate system of RSU and uses the backbone network (such as PointPillars) to align the point cloud data X of CAV with the coordinate system of RSU. m In the feature encoder, features are extracted, F m =Φ(X m ),in C represents the number of channels, H×W represents the spatial resolution, which are the height and width of the feature tensor respectively. In addition, the roadside unit also uses the backbone network to obtain the original point cloud data X r Get the feature tensor F r , expressed as F r =Φ(X r ).

[0010] Step 3: After generating the feature tensor, vehicle m converts the feature tensor F m Input the neural network-based detection head P(·) to obtain the spatial confidence map C m , similarly, the spatial confidence map C at the roadside unit r By F r Input into P(·) to obtain, where C m ∈[0,1]H×W 、C r ∈[0,1] H×W .

[0011] Step 4: RSU broadcasts the request graph R before each round of transmission begins. r To each CAV, where R r =1-C r After receiving the request map, each CAV will send its local confidence map Perform element-wise multiplication with the request graph to generate the feature selection matrix Recorded as Where ⊙ represents element-wise multiplication, represents the top-k feature selection strategy and Φ mask ∈[0,1] H×W represents the feature transfer mask, if Φ mask (h, w) = 0 means that the feature of the corresponding coordinate has been transmitted. At the same time, if the vehicle transmits some features in this round of transmission, the transmission mask at the corresponding feature will be set to 0, otherwise it will be set to 1.

[0012] Step 5: After feature selection is completed, the vehicle uploads features through the infrastructure communication (V2I) link m, where the set of all V2I links is represented as M is the maximum number of vehicles and V2I links in the scenario, and each CAV uses only one V2I link m. The uplink V2I link m uploads the corresponding CAV’s features via the allocated resource block (RB) k, where the set of all resource blocks (RBs) is used For the V2I link m at this time, its signal to interference and noise ratio is expressed as where σ 2 is the noise power, Indicates whether the V2I link m uses RBk to transmit characteristics. If Then the V2I link m uses RBk for transmission, otherwise it is 0, and link m can only use one RB at the same time, that is, P m and are transmit power and channel gain respectively, where α m is the frequency-independent large-scale fading, including path loss and shadowing, represents small-scale fading. The transmission rate of link m using RBk to upload features at time t is W is the bandwidth of each RB.

[0013] Step 6: After this round of feature transmission is completed, the cumulative features received by RSU from CAVm are RSU integrates it with its own BEV characteristics, and the fused characteristics are Where Γ(·) is the attention network. After completing the feature fusion of all vehicles in the current time slot, the prediction head will generate bounding box proposals and their corresponding confidence scores. Detection Network The generated detection output is represented as in It consists of regression results and classification results.

[0014] The action space is Where η and p are RBs allocation and power level selection, respectively, including three power levels: [23, 10.5, -100] dBm.

[0015] The reward function is in is the detection loss for the object detection task.

[0016] The observation state space of the agent is and in is the value of CAV to RSU, given by Calculated.

[0017] The training hierarchical agents are responsible for RBs allocation and power level selection. Specifically, after the RSU receives the status information S1(t), the agent responsible for RBs allocation outputs an action After obtaining the state information S2(t), the agent responsible for power level selection outputs the action Finally, we get action a t , complete the upload of CAV features.

[0018] Beneficial effects: The main innovation of the present invention is that it considers the interference management problem of the V2I link in the actual collaborative perception scenario. Based on the measurement of channel state information and feature effectiveness, a collaborative perception interference management method based on a hierarchical reinforcement learning model is proposed. A resource allocation strategy is learned through the reinforcement learning method to help vehicles upload more and more useful features to roadside units for feature fusion, so as to improve the recognition accuracy of the overall scene and achieve a balance between perception and communication. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 This is a collaborative sensing workflow diagram in this collaborative sensing interference management method.

[0020] Figure 2 This is the architecture diagram of reinforcement learning for interference management in this collaborative perception interference management method.

[0021] Figure 3This is a comparison chart of the target recognition task in the V2X-Sim dataset in this collaborative perception interference management method and the conventional method. The performance evaluation standard is AP@0.5.

[0022] Figure 4 This is a comparison chart of the target recognition task in the V2X-Sim dataset in this collaborative perception interference management method and the conventional method. The performance evaluation standard is AP@0.7. DETAILED DESCRIPTION

[0023] The present invention will be further described below with reference to the accompanying drawings and specific embodiments:

[0024] like Figure 1 As shown, the specific implementation steps of the present invention are as follows:

[0025] 1) The roadside unit (RSU) and all the intelligent communication vehicles (CAV) participating in the collaborative perception collect the original point cloud data, which is represented by X m and X r , where X m Represents the point cloud features of the mth intelligent communication vehicle. CAV aligns its own point cloud data with the coordinate system of RSU and uses a backbone network (such as PointPillars) to extract the point cloud data X of CAV. m In the feature encoder, features are extracted, F m =Φ(X m ),in C represents the number of channels, H×W represents the spatial resolution, which are the height and width of the feature tensor respectively. In addition, the roadside unit also uses the backbone network to obtain the original point cloud data X r Get the feature tensor F r , expressed as F r =Φ(X r ). After generating the feature tensor, vehicle m will generate the feature tensor F m Input the neural network-based detection head P(·) to obtain the spatial confidence map C m , similarly, the spatial confidence map C at the roadside unit r By F r Input into P(·) to obtain, where C m ∈[0,1] H×W 、C r ∈[0,1] H×W .

[0026] 2) RSU broadcasts the request graph R before each round of transmission begins r To each CAV, where R r =1-C r After receiving the request map, each CAV will send its local confidence map Perform element-wise multiplication with the request graph to generate the feature selection matrix Recorded as Where ⊙ represents element-wise multiplication, represents the top-k feature selection strategy and Φ mask ∈[0,1] H×W represents the feature transfer mask, if Φ mask (h, w) = 0 means that the feature of the corresponding coordinate has been transmitted. At the same time, if the vehicle transmits some features in this round of transmission, the transmission mask at the corresponding feature will be set to 0, otherwise it will be set to 1.

[0027] 3) The set of all V2I links is represented as M is the maximum number of vehicles and V2I links in the scenario, and each CAV uses only one V2I link m. The uplink V2I link m uploads the corresponding CAV’s features via the allocated resource block (RB) k, where the set of all resource blocks (RBs) is used For the V2I link m at this time, its signal to interference and noise ratio is expressed as where σ 2 is the noise power, Indicates whether the V2I link m uses RBk to transmit characteristics. If Then the V2I link m uses RBk for transmission, otherwise it is 0, and link m can only use one RB at the same time, that is, P m and are transmit power and channel gain respectively, where α m is the frequency-independent large-scale fading, including path loss and shadowing, represents small-scale fading. The transmission rate of link m using RBk to upload features at time t is W is the bandwidth of each RB.

[0028] The CAV vehicle packages its own value and link information and sends it to the RSU server. As an element of the agent's state observation value, the overall observation value is

[0029] 4) The present invention adopts Proximal Policy Optimization (PPO) as the reinforcement learning model, and receives observations at the RSU end server Then, S1(t) is input into the agent PPO1 to output the resource block allocation action. Then S1(t) and Splicing to obtain the observation value of agent PPO2 And output power control action Finally, the combined action Resource block allocation and power control for each V2I link.

[0030] 5) CAV vehicles upload features in this time slot, according to the top-k feature selection strategy and feature transmission mask Φ mask ∈[0,1] H×W The top-k feature selection strategy is based on the actual rate of the current V2I link. to transmit the first k features, and the feature transmission mask is used to filter the features that have been transmitted.

[0031] 6) After this round of feature transmission is completed, the cumulative features received by RSU from CAVm are RSU integrates it with its own BEV characteristics, and the fused characteristics are Where Γ(·) is the attention network. After completing the feature fusion of all vehicles in the current time slot, the prediction head will generate bounding box proposals and their corresponding confidence scores. Detection Network The generated detection output is represented as in It consists of regression results and classification results. The agent calculates the reward value R obtained after this round of transmission t , and obtain new observation values ​​S1(t+1) and S2(t+1) according to the method in step 4, and collect trajectories respectively Update the network parameters of PPO1 and PPO2 respectively according to the trajectory.

[0032] All simulation results below are based on the object recognition task on the V2X-Sim dataset. The evaluation metrics used are AP@0.5 and AP@0.7. The training scenario is an intersection scenario from the V2X-Sim dataset with one deployed RSU and four CAVs, with a training set to test set ratio of 8:2. The sensor streams in the V2X-Sim dataset are recorded at a 5Hz frequency, so the time slot T is set to 0.2 seconds. The communication environment is set with an update period of 200ms for large-scale channel fading and 1ms for small-scale channel fading. The decision interval for interference management is set to 20ms, resulting in 10 decision intervals per set for feature transmission. The number of RBs is 2, the carrier frequency is 5.9GHz, and the bandwidth ranges from 2MHz to 5MHz with a step size of 0.5MHz. The comparison methods are random access and maximum sum rate. Random access randomly allocates RBs according to different V2I links and transmits features to RSU at random power levels. The maximum rate method selects the two V2I links with the best channel state and the highest total rate each time and transmits features to RSU at the maximum power level.

[0033] Figure 3This figure compares the target recognition task of this collaborative perception interference management method with the conventional method on the V2X-Sim dataset. The performance evaluation standard is AP@0.5. Figure 4 This is a comparison chart of the target recognition task in the V2X-Sim dataset in this collaborative perception interference management method and the conventional method. The performance evaluation standard is AP@0.7. Figure 3 and Figure 4 In a comparison of two methods in the 2MHz to 5MHz bandwidth spectrum, the proposed hierarchical reinforcement learning method demonstrated superior performance compared to random and maximum rate methods at both AP@0.50 and AP@0.70. Furthermore, as bandwidth decreases, the proposed method exhibits a significant advantage over random access and maximum rate. The proposed method estimates the eigenvalues ​​of each CAV for the RSU based on feature importance. By using a hierarchical reinforcement learning model, the intelligent agent learns an effective strategy for balancing communication and perception. It achieves near-ceiling performance in high-bandwidth scenarios while improving perception accuracy in low-bandwidth situations.

[0034] It should be noted that those skilled in the art can make several improvements and modifications without departing from the principles of the present invention, and such improvements and modifications should also be considered within the scope of protection of the present invention. Components not specified in this embodiment can be implemented using existing technologies.

Claims

1. A collaborative perception interference management method based on hierarchical reinforcement learning, characterized in that: The steps include: Step 1: The CAV vehicle packages its own value and link information and sends it to the RSU server within a time slot; Step 2: The RSU server models the spectrum allocation and power control problem as a Markov decision process, and designs the state observation space, action space, and reward function. The CAV vehicle's own value and link information are used as an element of the agent's state observation value. The hierarchical reinforcement learning-based agent is trained to dynamically allocate spectrum and transmission power for each V2I link in the time slot. The RSU server uses proximal policy optimization as the reinforcement learning model. After receiving the state observation value S1(t), the RSU server inputs S1(t) into the agent PPO1 to output the resource block allocation action. Then S1(t) and The observation value S2(t) of the agent PPO2 is obtained by splicing, and the power control action is output Finally, the combined action Resource block allocation and power control for each V2I link, where η and p are resource block allocation and power level selection, respectively, and M is the maximum number of vehicles and V2I links in the scenario. Each CAV vehicle only uses one V2I link m; Step 3: The CAV vehicle uploads features via its corresponding V2I link within the time slot; Step 4: After the current round of feature transmission is completed, the RSU server will fuse the accumulated features received from each CAV vehicle with its own BEV features in sequence; after completing the feature fusion of all vehicles in the current time slot, the recognition result is obtained through the detector.

2. The collaborative perception interference management method based on hierarchical reinforcement learning according to claim 1 is characterized in that: In step 3, the CAV vehicle follows the top-k feature selection strategy and the feature transfer mask Φ mask ∈[0,1] H×W The top-k feature selection strategy is based on the actual rate of the current V2I link. To transmit the top k features with higher confidence, the feature transmission mask is used to filter the features that have been transmitted, and H×W represents the spatial resolution.

3. The collaborative perception interference management method based on hierarchical reinforcement learning according to claim 1 is characterized in that: Step 4 specifically includes: After the current round of feature transmission is completed, the cumulative features received by RSU from CAVm are RSU integrates it with its own BEV characteristics, and the fused characteristics are Where Γ(·) is the attention network, t represents the time slot, is the feature selection matrix, ⊙ represents element-by-element multiplication, Indicates whether V2I link m uses resource block k to transmit characteristics. If Then the V2I link m uses RBk for transmission, otherwise it is 0, and link m can only use one resource block at the same time, that is, After completing the feature fusion of all vehicles in the current time slot, the prediction head will generate bounding box proposals and their corresponding confidence scores; the detection network The generated detection output is represented as in It consists of regression results and classification results; the agent calculates the reward value R obtained after this round of transmission t , and obtain new observation values ​​S1(t+1) and S2(t+1), and collect trajectories respectively and Update the network parameters of PPO1 and PPO2 respectively according to the trajectory.

4. The collaborative perception interference management method based on hierarchical reinforcement learning according to claim 1 is characterized in that: Step 1 specifically includes: Step 11: RSU and all CAV vehicles participating in collaborative perception collect original point cloud data, which are represented as X m and X r , where X m Represents the point cloud features of the mth CAV vehicle; Step 12: CAV aligns its own point cloud data with the coordinate system of RSU, and uses the backbone network to align the point cloud data X of CAV. m In the feature encoder, features are extracted, F m =Φ(X m ),in C represents the number of channels, H×W represents the spatial resolution, which are the height and width of the feature tensor respectively; RSU also uses the backbone network to obtain the original point cloud data X r Get the feature tensor F r , expressed as F r =Φ(X r ); Step 13: After generating the feature tensor, vehicle m converts the feature tensor F m Input the neural network-based detection head P(·) to obtain the spatial confidence map C m , similarly, the spatial confidence graph C of RSU r By F r Input into P(·) to obtain, where C m ∈[0,1] H×W 、C r ∈[0,1] H×W ; Step 14: RSU broadcasts the request graph R before each round of transmission begins. r To each CAV, where R r =1-C r After receiving the request map, each CAV will send its local confidence map Perform element-wise multiplication with the request graph to generate the feature selection matrix Recorded as Where ⊙ represents element-wise multiplication, represents the top-k feature selection strategy and Φ mask ∈[0,1] H×W represents the feature transfer mask, if Φ mask (h,w)=0 means that the feature of the corresponding coordinate has been transmitted. At the same time, if the vehicle transmits some features in this round of transmission, the transmission mask at the corresponding feature is set to 0, otherwise it is set to 1; the feature selection matrix The sum of the elements in the string is the value of the string itself.

5. The collaborative perception interference management method based on hierarchical reinforcement learning according to claim 1 is characterized in that: The reward function is in, is the detection loss of the target detection task, λ1 and λ2 represent weights, λ1+λ2=1, and K represents the set of all resource blocks.

6. The collaborative perception interference management method based on hierarchical reinforcement learning according to claim 1 is characterized in that: The observation state space of the agent is and in, is the feedback information of the channel, is the value of CAV to RSU, given by Calculated.

7. The collaborative perception interference management method based on hierarchical reinforcement learning according to claim 1 is characterized in that: The training hierarchical agents are responsible for RBs allocation and power level selection respectively. Specifically, after the RSU receives the state information S1(t), the agent responsible for RBs allocation outputs an action After obtaining the state information S2(t), the agent responsible for power level selection outputs the action Finally, we get action a t , complete the upload of CAV features.