Sensor data processing and transmission decision method for vehicle-to-everything complementary perception
Through an intelligently guided, reward-free reinforcement learning method, a system model is constructed and sensor data processing and transmission decisions are optimized, which solves the shortcomings of the Internet of Vehicles collaborative perception system in multi-source data integration and resource allocation, realizes efficient and flexible perception data processing and transmission, and improves the system's adaptability and accuracy.
Patent Information
- Application Number
- CN202510118719.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2045-01-24
AI Technical Summary
The existing collaborative perception system of the Internet of Vehicles fails to effectively integrate various sensor information when processing multi-source perception data, especially in the dynamic allocation of perception tasks and resource utilization. In addition, the traditional reward mechanism is difficult to respond flexibly in a changing traffic environment, affecting the real-time and accuracy of perception data.
An intelligently guided non-reward reinforcement learning method is used to build a system model. By balancing the detection accuracy and delay objective functions, non-reward reinforcement learning is used to optimize sensor data processing and transmission decisions, dynamically adjust resource allocation, and optimize the collaborative efficiency of communication, perception, and computing.
It improves the system's adaptability to complex traffic environments, optimizes resource utilization efficiency, significantly reduces data processing and transmission delays, and improves the efficiency and accuracy of perception data processing. It is suitable for vehicle networking environments of various sizes and types.
Smart Images

Figure CN119996965B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of vehicle networking, and specifically relates to a sensor data processing and transmission decision method for complementary perception in vehicle networking. Background Art
[0002] In connected vehicle environments, collaborative perception technology is crucial to the safety and efficiency of autonomous vehicles. By enabling vehicles to share and receive perception data in real time, this technology significantly improves their responsiveness to complex traffic conditions. However, while collaborative perception technology provides vehicles with a more comprehensive view of their surroundings, it also places increased demands on communication and computing resources, especially in dense urban environments.
[0003] Collaborative perception technology requires efficient processing and transmission of large amounts of data between autonomous vehicles and intelligent transportation infrastructure. This places high demands on communication bandwidth and computing resources, which can lead to resource bottlenecks, especially during peak traffic periods, impacting the timeliness and accuracy of perception data. Therefore, developing technologies that can efficiently utilize limited resources while ensuring data transmission quality and processing speed is crucial for achieving efficient collaborative perception systems in connected vehicles. Combining perception data from different sources and optimizing data transmission and processing to reduce overall resource demands and improve overall system performance are key technical directions for enhancing collaborative perception capabilities in future connected vehicle systems.
[0004] Current research focuses on improving the performance of cooperative perception systems through advanced data processing algorithms and network optimization techniques. For example, some studies focus on reducing network burden through data compression and selective information sharing. Li et al. (Y. Li, F.R. Yu, and M. Wang, “Data Compression and Selective Sharing for Cooperative Perception in Vehicular Networks”) proposed a data compression and selective sharing mechanism that uses an intelligent algorithm to determine which data should be compressed or transmitted first, effectively reducing network pressure. Chen et al. (X. Chen, H. Zhang, and L. Liu, “Resource-Efficient Data Processing for Cooperative Vehicular Networks”) developed a resource-efficient data processing framework designed to optimize the processing of perception data and the allocation of communication resources, reducing latency and increasing data processing speed.
[0005] A large body of research is also devoted to improving resource allocation mechanisms and optimizing data processing and transmission efficiency. For example, Wang et al. (Z. Wang, J. Liu, and S. Lee, “Optimized Resource Allocation for Integrated Communication and Computation in Cooperative Vehicular Networks”) proposed a resource allocation optimization method for integrated communication and computing. This method uses a deep learning algorithm to dynamically adjust resource allocation to accommodate varying traffic and network conditions, significantly improving the system's response speed and data processing capabilities.
[0006] However, existing technologies often fail to effectively integrate diverse sensor information when processing multi-source perception data, particularly when it comes to dynamically allocating perception tasks and resource utilization. Furthermore, existing collaborative perception systems rely on traditional reward mechanisms to optimize data processing and transmission decisions. This often results in insufficient adaptability in volatile traffic environments, making it difficult to flexibly respond to environmental changes, thus impacting the real-time and accuracy of perception data. Summary of the Invention
[0007] In order to solve the above problems existing in the prior art, the present invention provides a sensor data processing and transmission decision-making method for complementary perception in the Internet of Vehicles. The technical problem to be solved by the present invention is achieved through the following technical solutions:
[0008] An embodiment of the present invention provides a sensor data processing and transmission decision-making method for complementary perception in an Internet of Vehicles (IoV), the method comprising:
[0009] Modeling a system model including CAVs, RAUs, and sensor ensembles;
[0010] Based on the system model, a detection accuracy objective function and a delay objective function are constructed, and an overall objective optimization function consisting of the detection accuracy objective function and the delay objective function is constructed by balancing weights;
[0011] Construct a state set including road and traffic state, sensor data state, vehicle dynamic state, and vehicle perception demand, a decision set including sensor management decision, data processing decision, and communication decision, and an observation value set including road and traffic observation values, sensor data observation values, vehicle dynamic observation values, vehicle perception demand observation values, and delay observation values, and based on the state set, use an intelligently guided reward-free reinforcement learning method to obtain the optimal decision set for sensor data transmission, so as to solve the overall objective optimization function through the optimal decision set.
[0012] Beneficial effects of the present invention:
[0013] The sensor data processing and transmission decision-making method for complementary perception of the Internet of Vehicles proposed in the present invention utilizes an intelligently guided non-reward reinforcement learning method, which has significant advantages over the existing technology: traditional collaborative perception systems usually rely on fixed reward mechanisms to guide the processing and transmission of perception data. This method is often unable to flexibly respond to various situations in the dynamically changing Internet of Vehicles environment, resulting in low efficiency in the collaboration between perception, communication, and computing. The present invention, through an intelligently guided non-reward reinforcement learning method, can dynamically adjust resource allocation and perception tasks according to real-time environmental data and vehicle status, thereby optimizing the efficiency of collaboration between communication, perception, and computing. At the same time, this method not only improves the efficiency of resource utilization, but also significantly reduces the delay in data processing and transmission. The intelligently guided non-reward reinforcement learning method of the present invention improves the system's adaptability to complex traffic environments. In addition, the method can be flexibly adjusted according to different vehicles and road conditions, has good scalability, and is suitable for Internet of Vehicles environments of various sizes and types. In summary, the proposed intelligently guided, unrewarded reinforcement learning method aims to optimize the collaborative perception and communication efficiency of CAVs during multi-source sensor data processing and transmission. This approach enables the system to significantly improve the efficiency and accuracy of the system's perception data processing while ensuring real-time response, providing safer and more reliable driving support for the Internet of Vehicles.
[0014] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] Figure 1 This is a flow chart of a sensor data processing and transmission decision-making method for complementary perception in an Internet of Vehicles, provided by an embodiment of the present invention;
[0016] Figure 2 Schematic diagram of the comparison results between three traditional methods and the method proposed in the present invention in terms of perceived data delay;
[0017] Figure 3 This is a schematic diagram of the comparison results between three traditional methods and the method proposed in this invention in terms of perception fusion accuracy. DETAILED DESCRIPTION
[0018] The present invention will be further described in detail below with reference to specific examples, but the embodiments of the present invention are not limited thereto.
[0019] See Figure 1 The embodiment of the present invention provides a sensor data processing and transmission decision method for complementary perception in an Internet of Vehicles, which specifically includes the following steps:
[0020] S10. Modeling a system model including CAVs, RAUs, and sensor collection.
[0021] The system modeling of the embodiment of the present invention includes the description of the scene and the modeling of perception characteristics, specifically including the sensor deployment position and height, perception distance, horizontal viewing angle, vertical viewing angle, obstacle impact, calculation and processing modeling, communication modeling and problem statement.
[0022] For example, the research scenario is set in a complex urban traffic environment, which is equipped with an advanced system model RSNs, which includes CAVs as well as ordinary vehicles (called CAVs with level 0 intelligence), sensors and roadside assistance units (RAUs). The embodiment of the present invention envisions an advanced urban traffic environment equipped with multiple sensors, RAUs and CAVs, as well as ordinary vehicles. Each road is divided into multiple sections, each section is equipped with an RAU, and the RAUs are strategically deployed on the road to ensure optimal coverage and functionality. RAUs are used to aggregate perception data from sensors deployed on the roadside infrastructure and transmit the processed perception information to CAVs with different perception needs. Assume that the group of CAVs served by a specific RAU is {1,2,…,V}, where V represents the number of CAVs in the system model. Similarly, the sensor set is defined as {1,2,…,S}, where S represents the total number of deployed sensors. CAVs have different requirements for perception data, such as detection accuracy and latency. Accuracy level (P v ) represents the level of detail of the data required for the vth CAV, ranging from the original data (P v =1, representing the highest accuracy) to abstract information (P v =0, indicating lower accuracy). In addition, CAVs have different tolerances to latency in sensing data, with key operations requiring low latency (L v →0, while regular tasks can tolerate higher delays L v .
[0023] Sensors are strategically deployed on both sides of the road or above it to ensure that normal traffic flow is not disturbed. The location coordinates of each sensor are expressed as Where k is the deployment location number and l is the sensor number. The deployment height of the sensor is expressed by the following formula:
[0024]
[0025] Among them, H k,l represents the installation height of the lth sensor deployed at the deployment location number k from the ground. Due to budget and operational constraints, not all potential locations can deploy sensors. The number of sensors at each location is given by Range is 0 to the maximum number of sensors allowed at this location
[0026] S20, based on the system model, a detection accuracy objective function is constructed, and a delay objective function is constructed, and a total target optimization function composed of the detection accuracy objective function and the delay objective function is constructed by balancing weights.
[0027] The embodiment of the application constructs a detection accuracy objective function, which includes: respectively constructing a distance perception detection probability function, a horizontal perception detection probability function, a vertical perception detection probability function and an obstacle perception detection probability function of a single sensor; and constructing a detection accuracy objective function from the distance perception detection probability function, the horizontal perception detection probability function, the vertical perception detection probability function and the obstacle perception detection probability function. More specifically:
[0028] The perception distance refers to the maximum distance at which the sensor can effectively detect the target. The distance perception detection probability function of the sensor R k,l constructed by the embodiment of the application is expressed by the formula:
[0029]
[0030] Among them, represents the distance perception detection probability function, d kl,j represents the distance between the sensor R k,l and the jth target O j , j takes a value of 1 to J, J represents the number of targets, here the target refers to the object to be detected, such as a car, a person, a building, etc., R k,l represents the lth sensor deployed at the deployment position number k, represents the maximum effective distance of the sensor R k,l , Λ represents the sensitivity adjustment parameter of R , Λ∈[0,10], Γ represents the threshold adjustment parameter of R
[0031] , Γ∈[0,100]; this formula describes the detection ability of the sensor with distance, reflecting how the detection probability changes with the increase of distance.
[0032]
[0033] Among them, represents the horizontal perception detection probability function, φ kl,j represents the horizontal deviation angle of the sensor R k,l and the jth target O j from the sensor axis, Ω represents The slope adjustment parameter, Ω∈[0,1], Π represents The maximum effective horizontal angle is Π∈[0,100]; this formula reflects the impact of horizontal deviation on the detection ability of the sensor, indicating that the detection probability of the sensor will decrease when the detection angle deviates from the central axis.
[0034] The vertical viewing angle refers to the sensor's ability to detect targets in the vertical direction. The vertical perception detection probability function constructed in the embodiment of the present invention is expressed as follows:
[0035]
[0036] in, represents the vertical perception detection probability function, θ kl,j Indicates sensor R k,l With the j-th target O j The vertical deviation angle of the sensor axis, Δ represents The slope adjustment parameter, Δ∈[0,1], Φ represents The maximum effective vertical angle is Φ∈[0,100]; this formula shows the impact of vertical deviation on the detection probability of the sensor, and reflects the change in the detection probability of the sensor when the detection angle deviates from the central axis.
[0037] Obstacles have a significant impact on the sensor's detection capabilities. The obstacle perception detection probability function formula constructed in this embodiment of the present invention is expressed as:
[0038]
[0039] in, represents the obstacle perception detection probability function, P obs Indicates the location information of the obstacle, G obs (P obs ) indicates P obs The obstacle characteristic factor of the obstacle at Indicates sensor R k,l The position information of the object, α represents the rate at which the perception probability changes with distance, α∈[2,4]. This formula quantifies the change in detection probability due to the presence of obstacles and explains how obstacles affect the detection performance of the sensor.
[0040] Finally, the detection accuracy objective function constructed by formulas (2) to (5) is expressed as:
[0041]
[0042] Furthermore, the system model modeled in the embodiments of the present invention also includes an edge computing model. The edge computing model includes several edge nodes, each equipped with computing resources. The edge computing model processes data from various sensors by introducing edge nodes. The edge computing nodes are responsible for communicating with CAVs and providing perception assistance to CAVs based on currently available sensors. Each edge node is equipped with computing resources for processing sensor data.
[0043] The process of constructing a delay target function in an embodiment of the present invention includes:
[0044] When CAVs send sensor perception data to RAUs, the RAUs processing capacity is calculated based on the computing resources and the size of the perception data. When RAUs send sensor perception data to CAVs, the CAVs processing capacity is calculated based on the computing resources and the size of the perception data. The perception data includes the size of the perception data, the vehicle processing density, and the processing decision weight. The RAUs processing time is calculated based on the RAUs processing capacity and the perception data, and the CAVs processing time is calculated based on the CAVs processing capacity and the perception data. Based on the characteristics of the wireless channel for communication between CAVs and RAUs, the uplink data transmission rate from CAVs to RAUs and the downlink data transmission rate from RAUs to CAVs are calculated respectively, and the uplink transmission time and the downlink transmission time are calculated based on the uplink data transmission rate and the downlink data transmission rate. The delay objective function is constructed by the CAVs processing time, the RAUs processing time, the uplink transmission time, and the downlink transmission time. More specifically:
[0045] The embodiment of the present invention calculates the RAUs processing capacity using the formula:
[0046]
[0047] in, represents the RAUs processing capability when processing the nth sensing data at time t, represents the available computing resources found by RAUs when CAVs sends sensor perception data to RAUs at time t, D n (t) represents the size of the nth sensing data sent by CAVs to RAUs at time t, and N represents the number of sensing data sent by CAVs to RAUs;
[0048] The embodiment of the present invention calculates the CAVs processing capability, and the formula is expressed as follows:
[0049]
[0050] in, represents the CAVs processing capability when processing the n1th perception data at time t, It means that when RAUs send sensor perception data to CAVs at time t, CAVs find available computing resources. It represents the size of the n1th sensing data sent by RAUs to CAVs at time t, and N′ represents the amount of sensing data sent by RAUs to CAVs.
[0051] Through formulas (7) and (8), computing resources are distributed among all sensor data, effectively balancing the computing load and ensuring efficient data processing with minimized latency.
[0052] The embodiment of the present invention calculates the RAUs processing time, and the formula is expressed as:
[0053]
[0054] in, Indicates RAUs processing time, a n (t) represents the decision processing weight of RAUs when processing the nth perception data at time t, a n (t) ranges from [0,1], λ n (t) represents the processing complexity of RAUs when processing the nth sensor data at time t. Formula (9) describes the processing delay of data on the RAU side, ensuring efficient execution of tasks and minimizing delays.
[0055] The embodiment of the present invention calculates the CAVs processing time, and the formula is expressed as:
[0056]
[0057] in, represents the CAVs processing time, represents the decision processing weight of CAVs when processing the n1th perception data at time t, The value range is [0,1], represents the processing complexity of CAVs when processing the n1th sensory data at time t. Formula (10) describes the processing delay of data on the CAV side, ensuring efficient execution of tasks and minimizing delays.
[0058] The communication model focuses on the interaction between CAVs, RAUs, and sensors within the RSNs system model, addressing the efficient transmission of sensor data and information exchange to optimize perception and decision-making. Communication primarily occurs between CAVs and RAUs over wireless channels, including uplink (from CAV to RAU) and downlink (from RAU to CAV) data transmission.
[0059] The embodiment of the present invention calculates the uplink transmission time using the formula:
[0060]
[0061] Among them, T up Indicates the uplink transmission time, Indicates the uplink bandwidth of the wireless channel. Indicates the uplink transmission power of the wireless channel, Indicates the uplink channel gain of the wireless channel, N represents the uplink interference power of the wireless channel, and N0 represents the noise power of the wireless channel.
[0062] The embodiment of the present invention calculates the downlink transmission time, and the formula is expressed as follows:
[0063]
[0064] Among them, T dn Indicates the downlink transmission time, Indicates the downlink bandwidth of the wireless channel, Indicates the downlink transmission power of the wireless channel, Indicates the downlink channel gain of the wireless channel, Indicates the downlink interference power of the wireless channel.
[0065] The parameters in formulas (11) and (12) jointly determine the efficiency and reliability of data transmission in the Internet of Vehicles. Delay is a key factor in the system model, especially for safety-critical applications in CAVs. The total delay constructed in the embodiment of the present invention includes CAV processing time, uplink transmission time, RAU processing time, and downlink transmission time. The delay objective function constructed by formulas (9) to (12) is expressed as follows:
[0066]
[0067] Furthermore, the communication model is seamlessly integrated with the perception and computation models, which allows the sensor data transmission strategy to be dynamically adjusted based on data availability, network conditions, and vehicle computing capabilities. The main goal of the system model of the embodiment of the present invention is to maximize the perception detection probability of CAVs while minimizing the delay caused by communication and task processing, so that the final overall objective optimization function is expressed as follows:
[0068]
[0069] Among them, R represents the overall objective optimization function, Θ total For the optimization goal of detection accuracy, T total For optimization purposes related to latency, including processing latency and communication delay T up +T dn , ω represents the balance weight, V represents the number of CAVs in the system model, represents the perceptual detection probability of the vth CAV, Θ min represents the minimum perceptual detection threshold, T total represents the total transmission delay, which is calculated by formula (13), T max Indicates the maximum delay threshold, Indicates the uplink bandwidth of the wireless channel. Represents the downlink bandwidth of the wireless channel, W total represents the total bandwidth of the wireless channel, Indicates the computing processing capacity of the i-th RAU, i ranges from 1 to I, and I represents the number of RAUs in the system model. Indicates the maximum computing processing capacity of RAU, represents the number of sensors deployed at the deployment location numbered k, Represents the maximum number of sensors deployed. Among them, in the constraints: The first constraint ensures that the perception detection probability of each CAV is not lower than the minimum detection threshold Θ required by the system min To ensure the basic perceptual performance requirements; the second constraint limits the total delay of each CAV to not exceed the maximum delay threshold T allowed by the system max , ensuring the timeliness of data; the third constraint indicates that the total bandwidth allocated for uplink and downlink cannot exceed the total bandwidth W of the wireless channel total ; The fourth constraint ensures that each roadside unit RAU i The computing resource usage of the system does not exceed its maximum computing capacity; the fifth constraint limits the total number of sensors deployed in the system model to not exceed the maximum number of sensors allowed by the system. These constraints together ensure the reasonable allocation of system resources and the satisfaction of performance requirements.
[0070] The embodiments of the present invention integrate multi-source sensor data in an Internet of Vehicles environment and optimize data preprocessing and analysis through edge computing technology to ensure the real-time and reliability of the data.
[0071] S30. Construct a state set including road and traffic status, sensor data status, vehicle dynamic status, and vehicle perception needs, a decision set including sensor management decisions, data processing decisions, and communication decisions, and an observation value set including road and traffic observation values, sensor data observation values, vehicle dynamic observation values, vehicle perception need observation values, and delay observation values. Based on the state set, decision set, and observation value set, use the intelligent-guided reward-free reinforcement learning method to obtain the optimal decision set for sensor data transmission, so as to solve the overall objective optimization function through the optimal decision set.
[0072] To address the limitations of traditional methods that rely on explicit reward functions, embodiments of the present invention introduce a non-rewarded intelligently guided reinforcement learning method, replacing the traditional reinforcement learning method of defining reward functions based on task types. This non-rewarded intelligently guided reinforcement learning method optimizes the processing and transmission decisions of perceptual information in collaborative perception of the Internet of Vehicles. This method does not rely on traditional reward mechanisms, but instead uses intelligence as a guide based on understanding the environment to dynamically adjust resource allocation and optimize data processing flows. A detailed description of the non-rewarded intelligently guided reinforcement learning method is as follows:
[0073] The environment state of reinforcement learning is a multidimensional structure that contains various aspects that affect the dynamics of CAVs operation. At any given time t, the environment state is represented by S t , S t include:
[0074] Road and traffic conditions reflects dynamic and static road conditions, including traffic density, road construction, and weather conditions;
[0075] Sensor data status Aggregates a large amount of data from various vehicle-mounted and roadside sensors;
[0076] Vehicle dynamic status Covers key parameters such as vehicle speed, direction and other relevant dynamic states;
[0077] Vehicle perception needs It is used to represent the perception needs of CAVs and can also reflect the intelligence level of the vehicle.
[0078] Therefore, the state set S at time t t It can be expressed as:
[0079]
[0080] In the reinforcement learning framework, the observation set O t It can be expressed as:
[0081]
[0082] in, represents road and traffic observation values, represents the sensor data observation value, represents the vehicle dynamic observation value, represents the vehicle perception demand observation value, Represents a delayed observation.
[0083] Agents usually cannot directly access the real state of the environment, but obtain state information through observation. Therefore, the embodiment of the present invention considers preference as a unique observation mode. Given a state set S t , observation set O t , the joint prior preference distribution of the reinforcement learning model parameters θ can be expressed as p Ψ (S t ,O t ,θ).
[0084] Decision Set (A t ) represents the strategic response of the system at time t, which is affected by the current state. t Includes: Sensor Management Decisions Involves decisions related to sensor selection and tuning; data processing decisions Methods involving sensor data processing; communication decision making Determines the information exchange strategy between CAVs and RAUs. Therefore, the decision set A at time t is t Expressed as:
[0085]
[0086] The core of the data transfer decision lies in evaluating whether to process the data locally or transfer it to RAUs for more comprehensive processing. This decision is represented by the binary variable Where 1 represents transmission to RAU and 0 represents local processing, as shown in the following formula:
[0087]
[0088] in, Represents an edge node (E i ) of the load, Represents an edge node (E i ) data processing requirements, θ trans Indicates the transmission threshold.
[0089] The dynamics of the system include the transition of states and the impact of decisions. The state transition is expressed by the following formula:
[0090] S t =F(S t ,A t ) (19);
[0091] Where F() represents the dynamics of the CAV environment that integrates the current state and decision. In the framework of reward-free reinforcement learning, reward-free guidance is achieved through a set of abstract environmental states and goals, bypassing the need for explicit reward signals. The decision-making algorithm is based on active inference, which uses these guidance to make wise choices that are consistent with the overall goal of the system, namely to improve the operational safety and efficiency of CAVs in RSNs. In addition, when higher-level cognition is developed, there is a difference between the predicted state and the actual state, and this difference changes through the learning process. In the study, the embodiment of the present invention uses the basic concept of intelligence to evaluate the change of this difference. Intelligence, as an advanced indicator for quantifying learning effects, borrows the concepts of energy and information. It is a comparative measure that evaluates the change in information distribution over time due to learning or the degree of information dispersion relative to the initial state.
[0092] Furthermore, the embodiment of the present invention uses an intelligently guided, non-rewarded reinforcement learning method based on a state set, a decision set, and an observation set to obtain an optimal decision set for sensor data transmission, so as to solve the overall objective optimization function through the optimal decision set, including:
[0093] Given a first probability distribution of the state set predicted by the current policy, and a second probability distribution of the state set predicted by the decision set, the KL divergence between the first probability distribution and the second probability distribution is used as the policy optimization function to obtain the optimal policy by minimizing the policy optimization function; given a third probability distribution of the state set predicted by the optimal policy, the observation value set, and the reinforcement learning model parameters, the joint prior preference distribution of the given state set, observation value set, and reinforcement learning model parameters is calculated, and the KL divergence between the third probability distribution and the joint prior preference distribution is used as the expected free energy function; the optimization problem of minimizing the expected free energy function is converted into the optimization problem of maximizing the negative expected free energy function; the negative expected free energy function is decomposed into a first term that captures the expected information gain and a second term that captures the extrinsic value; the approximate posterior distribution of the first term and the second term is solved respectively to obtain the optimal policy distribution, and the optimal decision set for sensor data transmission is obtained based on the optimal policy distribution to solve the overall objective optimization function through the optimal decision set. More specifically:
[0094] The basic concept of the decision framework is to optimize decisions based on predicting their impact on the future state of the environment. It involves maintaining a probabilistic model of the environment and predicting the outcomes of different decisions. Therefore, the embodiment of the present invention predicts a first probability distribution of the set of states given the current policy, and predicts a second probability distribution of the set of states given the decision set. The KL divergence between the first and second probability distributions is solved as a policy optimization function, with the goal of minimizing the policy optimization function to obtain the optimal policy. The optimal policy obtained by minimizing the policy optimization function is expressed as follows:
[0095]
[0096] Among them, π * represents the optimal strategy, A t represents the decision set, S t represents the state set, π represents the current strategy, q(S t |π) represents the predicted state set S given the current policy π t The first probability distribution, P(S t |A t ) represents a given decision set A t Prediction state set S t The second probability distribution, D KL (·||·) represents the Kullback-Leibler divergence function. The model is continuously updated based on the decision set to ensure that it accurately reflects the changing dynamics of the CAVs environment:
[0097] P(S t+1 |S t ,A t )=f(S t ,A t ,θ) (21);
[0098] Among them, f() is a function describing the state transition probability, θ represents the reinforcement learning model parameters, P(S t+1 |S t ,A t ) represents a given decision set A t , state set S t Prediction state set S t+1 The probability distribution of .
[0099] The decision-making process of the embodiment of the present invention aims to minimize the expected future free energy, which represents the difference between the predicted and target state distributions. By continuously adjusting the optimal strategy π * To minimize the KL divergence, we can make optimal decisions under uncertainty and enhance adaptability. The strategy selection process is achieved through the expected free energy function The model is modeled by minimization, where the expected free energy function formula is expressed as:
[0100]
[0101] in, represents the expected free energy function, S t Represents a state set, O t represents the set of observations, π * represents the optimal strategy, θ represents the reinforcement learning model parameters, q(S t ,O t ,θ|π * ) represents the given optimal strategy π * Prediction state set S t , observation set O t , the third probability distribution of the reinforcement learning model parameter θ, p Ψ (S t ,O t ,θ) represents the state set S t , observation set O t , the joint prior preference distribution of reinforcement learning model parameters θ, D KL (·||·) represents the Kullback-Leibler divergence function.
[0102] The optimal policy distribution q(π * ) can be minimized by the following expected free energy function formula:
[0103]
[0104] therefore, It shows that the strategy that minimizes the expected free energy is more likely to be selected, which promotes a unified framework for exploration and exploitation. σ represents the Sigmoid function, which is used to convert the input value into a value with a probability between 0 and 1. The goal of active inference is to minimize the expected free energy. However, by converting to maximize the negative expected free energy We can transform the problem into a standard maximization problem. This is more consistent with traditional optimization theory and makes the problem easier to handle and understand. The balance between exploration and exploitation is achieved by The decomposition into a first term that captures the expected information gain and a second term that captures the extrinsic value is an approximation because it ignores some higher-order terms and cross-terms. However, in practice, this approximation is usually accurate enough and greatly simplifies calculations and interpretations. The negative expected free energy function formula of the embodiment of the present invention is expressed as:
[0105]
[0106] in, represents the negative expected free energy function, A t represents the decision set, S t Represents a state set, O t represents the set of observations, π * represents the optimal strategy, θ represents the reinforcement learning model parameters, q(A t |π * ) represents the given optimal strategy π * Prediction decision set A t The probability distribution of Denotes a given optimal policy π * The expectation of the decision set, q(S t ,θ|O t ,π * ) represents the given optimal strategy π * , observation set O t Prediction state set S t , the probability distribution of the reinforcement learning model parameters θ, q(S t ,θ|π * ) represents the given optimal strategy π * Prediction state set S t , the probability distribution of the reinforcement learning model parameters θ, q(S t |π * ) represents the given optimal strategy π * Prediction state set S t The probability distribution of Denotes a given optimal policy π * Next state set S t The expectation, q(O t ,θ|S t ,π * ) represents the given optimal strategy π * , state set S t Predicted observation set O t , the probability distribution of reinforcement learning model parameters θ, p Ψ (O t ,θ|π * ) represents the given optimal strategy π * Predicted observation set O t , the prior preference distribution of reinforcement learning model parameters θ, D KL (·||·) represents the Kullback-Leibler divergence function.
[0107] Formula (24) naturally integrates exploration behavior (through information gain) and the use of known strategies (through extrinsic value). Based on the variational inference principle. First, As a variational lower bound, it is necessary to maximize it. By applying variational inference, this maximization problem can be transformed into finding the best approximate posterior distribution q. At each time t, the decision set reflects the decisions of sensor management, data processing and communication strategies. In order to dynamically optimize the strategy distribution q(π * ), where the first term is the expectation of the log-likelihood and the second term is the entropy of the approximate posterior. By maximizing formula (24), the optimal approximate posterior q distribution can be obtained:
[0108]
[0109] Among them, S τ:T represents the set of all states from time τ to time T, O τ:T represents the set of all observations from time τ to time T, q(S τ:T ,O τ:T ,θ|π * ) is the optimal approximate posterior q distribution, that is, given the optimal strategy π * Prediction state set S τ:T , observation set O τ:T , the probability distribution of reinforcement learning model parameters θ, q(O τ |S τ ,θ,π * ) represents a given state set S τ 、Optimal strategy π * , reinforcement learning model parameters θ predict observation value set O τ The probability distribution of q(S τ |S τ-1 ,θ,π * ) represents a given state set S τ-1 、Optimal strategy π * , reinforcement learning model parameters θ predict state set S τ The probability distribution of p Ψ (O τ |S τ ) represents a given state set S τ Predicted observation set O τ The prior preference distribution of Denotes a given optimal policy π * , the state set S under the reinforcement learning model parameter θ τ expectations, Represents a given state set S τ-1 、Optimal strategy π * , reinforcement learning model parameters θ predict state set S τ The prior preference distribution of Denotes a given optimal policy π* , the expected value of the state set S under the reinforcement learning model parameter θ. τ-1
[0110] This decomposition form of formula (25) reflects the dynamic characteristics of the system, where the state and observation at each time depend on the state at the previous time. q(π * ) is converted into a diagonal Gaussian distribution to simplify the calculation, and the Exponential Change of Measure (ECM) algorithm is used to optimize q(π * ) to be consistent with the minimization of J(π ). The goal of the ECM algorithm is to determine the optimal strategy for sensor data transmission based on the current state of the CAVs and RSN environment. Ultimately, the total objective function of formula (14) can be optimized through the action of the optimal strategy.
[0111] The embodiment of the present application adopts an innovative intelligent guided reward-free reinforcement learning method, which does not rely on traditional reward mechanisms, but guides decision-making through environmental understanding. This method can dynamically adjust resource allocation and optimize the processing and transmission process of perception information, thereby adapting to high dynamic and variable traffic environments.
[0112] In order to verify the effectiveness of the sensor data processing and transmission decision method for Internet of Vehicles complementary perception provided by the embodiment of the present application, through comparison with traditional methods, the superiority of the present application in terms of perception accuracy and delay in multiple actual driving scenarios is demonstrated. The numerical results clearly show that the proposed method (Our Proposal) has significant improvement in improving the efficiency of perception, communication and calculation cooperation and the efficiency of resource allocation compared with three traditional methods (Rainbow DRL, SoftActor-Critic RL, Policy Optimization). Please refer to Figure 2 : The proposed method of the present application ensures the detection delay under different scenarios. Compared with the other three traditional methods, the proposed method of the present application has a peak delay of 76.38 milliseconds at medium density, and then slightly decreases at the highest density. It can be seen that the proposed method of the present application can reduce the detection delay under high traffic density, and shows excellent delay management capability. Please refer to Figure 3 As the number of roadside sensors increases, the overall detection accuracy of the proposed method rapidly increases from 63.91% to 94.56%. The other three traditional methods struggle to achieve the same detection level with the same number of available sensors. Comprehensive simulation results demonstrate that the proposed intelligently guided, reward-free reinforcement learning method excels in optimizing detection accuracy, reducing detection latency, and maintaining packet delivery rates, demonstrating its significant potential in the RSN system model.
[0113] In summary, the sensor data processing and transmission decision-making method for complementary perception of the Internet of Vehicles proposed in the embodiment of the present invention utilizes an intelligently guided non-reward reinforcement learning method, which has significant advantages over the existing technology: traditional collaborative perception systems usually rely on a fixed reward mechanism to guide the processing and transmission of perception data. This method is often unable to flexibly respond to various situations in a dynamically changing Internet of Vehicles environment, resulting in low efficiency in the collaboration between perception, communication, and computing. The present invention, through an intelligently guided non-reward reinforcement learning method, can dynamically adjust resource allocation and perception tasks according to real-time environmental data and vehicle status, thereby optimizing the efficiency of collaboration between communication, perception, and computing. At the same time, this method not only improves the efficiency of resource utilization, but also significantly reduces the delay in data processing and transmission. The intelligently guided non-reward reinforcement learning method of the present invention improves the adaptability of the system to complex traffic environments. In addition, the method can be flexibly adjusted according to different vehicles and road conditions, has good scalability, and is suitable for Internet of Vehicles environments of various sizes and types. In summary, the intelligently guided, unrewarded reinforcement learning method proposed in the embodiments of the present invention aims to optimize the collaborative perception and communication efficiency of CAVs during multi-source sensor data processing and transmission. This method enables the system to significantly improve the efficiency and accuracy of the system's perception data processing while ensuring real-time response, providing safer and more reliable driving support for the Internet of Vehicles.
[0114] In the description of the present invention, it should be understood that the terms "first" and "second" are used for descriptive purposes only and should not be understood to indicate or imply relative importance or implicitly specify the number of the technical features indicated. Therefore, a feature specified as "first" or "second" may explicitly or implicitly include one or more of the features. In the description of the present invention, "plurality" means two or more, unless otherwise specifically defined.
[0115] Although the present invention is described herein in conjunction with various embodiments, those skilled in the art may understand and implement other variations of the disclosed embodiments by reviewing the specification and accompanying drawings in the process of implementing the claimed invention. In the specification, the word "comprising" does not exclude other components or steps, and "a" or "an" does not exclude multiple components or steps. The fact that certain measures are described in different embodiments does not mean that these measures cannot be combined to produce good results.
[0116] The above is a further detailed description of the present invention in conjunction with specific preferred embodiments, and the specific implementation of the present invention should not be considered to be limited to these descriptions. For those skilled in the art of the present invention, without departing from the concept of the present invention, several simple deductions or substitutions can be made, which should be considered to fall within the scope of protection of the present invention.
Claims
1. A sensor data processing and transmission decision method for complementary perception in Internet of Vehicles, characterized in that: The method comprises: Modeling a system model including CAVs, RAUs, and sensor ensembles; Based on the system model, a detection accuracy objective function is constructed according to the distance perception detection probability function, the horizontal perception detection probability function, the vertical perception detection probability function, and the obstacle perception detection probability function of a single sensor, and a delay objective function is constructed, and an overall objective optimization function composed of the detection accuracy objective function and the delay objective function is constructed by balancing weights; Constructing a state set including road and traffic states, sensor data states, vehicle dynamic states, and vehicle perception requirements, a decision set including sensor management decisions, data processing decisions, and communication decisions, and an observation set including road and traffic observations, sensor data observations, vehicle dynamics observations, vehicle perception requirement observations, and delay observations, and using an intelligently guided, reward-free reinforcement learning method to obtain an optimal decision set for sensor data transmission based on the state set, the decision set, and the observation set, so as to solve the overall objective optimization function using the optimal decision set; Among them, the constructed distance perception detection probability function is expressed as follows: ; in, represents the distance perception detection probability function, Indicates sensor With the j goals The distance between j The value is 1~ J , J represents the number of targets, Indicates that the deployment location number is k The deployed l sensors, Indicates sensor The maximum effective distance, express Sensitivity adjustment parameter, express Threshold adjustment parameters; The horizontal perception detection probability function formula is expressed as: ; in, represents the horizontal perception detection probability function, Indicates sensor With the j goals Horizontal deviation angle at the sensor axis, express The slope adjustment parameter, express The maximum effective horizontal angle; The vertical perception detection probability function formula is expressed as: ; in, represents the vertical perception detection probability function, Indicates sensor With the j goals The vertical deviation angle of the sensor axis, express The slope adjustment parameter, express The maximum effective vertical angle; The obstacle perception detection probability function formula is expressed as: ; in, represents the obstacle perception detection probability function, Indicates the location information of the obstacle. express The obstacle characteristic factor of the obstacle at Indicates sensor location information, represents the rate at which the perception probability changes with distance; The modeled system model also includes an edge computing model; the edge computing model includes a number of edge nodes, each edge node is equipped with computing resources; and the corresponding delay objective function is constructed, including: When CAVs send sensor perception data to RAUs, the RAUs processing capacity is calculated based on the computing resources and the size of the perception data; when RAUs send sensor perception data to CAVs, the CAVs processing capacity is calculated based on the computing resources and the size of the perception data; wherein the perception data includes the size of the perception data, processing complexity, and decision processing weight; the RAUs processing time is calculated based on the RAUs processing capacity and the perception data, and the CAVs processing time is calculated based on the CAVs processing capacity and the perception data; based on the characteristics of the wireless channel for communication between CAVs and RAUs, the uplink data transmission rate from CAVs to RAUs and the downlink data transmission rate from RAUs to CAVs are calculated respectively, and the uplink transmission time and the downlink transmission time are calculated based on the uplink data transmission rate and the downlink data transmission rate; a delay objective function is constructed based on the CAVs processing time, the RAUs processing time, the uplink transmission time, and the downlink transmission time; The formula for calculating RAUs processing capacity is: ; in, express Time processing n RAUs processing capability when sensing data, express The available computing resources found by RAUs when CAVs send sensor perception data to RAUs at the moment, express The first time CAVs sends to RAUs n The size of the perception data, It represents the amount of sensing data sent by CAVs to RAUs; The formula for calculating CAVs processing capacity is: ; in, express Time processing n CAVs processing capability when sensing data, express When RAUs send sensor perception data to CAVs, CAVs find available computing resources. express At this moment, RAUs sends the first n The size of 1 perception data, It represents the amount of sensing data sent by RAUs to CAVs; The formula for calculating RAUs processing time is: ; in, Indicates RAUs processing time, express Time processing n The decision processing weight of RAUs when sensing data, The value range is [0,1], express Time processing n The processing complexity of RAUs when sensing data; Calculate the CAVs processing time, the formula is expressed as: ; in, represents the CAVs processing time, express Time processing n 1 perception data, the decision processing weight of CAVs, The value range is [0,1], express Time processing n The processing complexity of CAVs when there is 1 sensory data; Calculate the uplink transmission time, the formula is expressed as: ; in, Indicates the uplink transmission time, Indicates the uplink bandwidth of the wireless channel. Indicates the uplink transmission power of the wireless channel, Indicates the uplink channel gain of the wireless channel, Indicates the uplink interference power of the wireless channel, represents the noise power of the wireless channel; Calculate the downlink transmission time, the formula is expressed as: ; in, Indicates the downlink transmission time, Indicates the downlink bandwidth of the wireless channel, Indicates the downlink transmission power of the wireless channel, Indicates the downlink channel gain of the wireless channel, Indicates the downlink interference power of the wireless channel; The constructed detection accuracy objective function is expressed as follows: ; The delay objective function constructed is expressed as follows: ; The constructed total objective optimization function is expressed as follows: ; in, represents the overall objective optimization function, represents the balance weight, V represents the number of CAVs in the system model, Indicates the v The perception detection probability of a CAV, represents the minimum perceptual detection threshold, represents the total transmission delay, Indicates the maximum delay threshold, Indicates the uplink bandwidth of the wireless channel. Indicates the downlink bandwidth of the wireless channel, represents the total bandwidth of the wireless channel, Indicates the i The computing power of each RAU, i The value is 1~ I , I represents the number of RAUs in the system model, Indicates the maximum computing processing capacity of RAU, Indicates that the deployment location number is k The number of sensors deployed, Indicates the maximum number of sensors deployed.
2. The sensor data processing and transmission decision method for complementary perception in Internet of Vehicles according to claim 1, characterized in that: Based on the state set, the decision set, and the observation value set, an optimal decision set for sensor data transmission is obtained using an intelligently guided non-rewarded reinforcement learning method, so as to solve the overall objective optimization function using the optimal decision set, including: Given a current policy predicting a first probability distribution of the state set, given a decision set predicting a second probability distribution of the state set, solving the KL divergence between the first probability distribution and the second probability distribution as a policy optimization function, and solving the optimal policy by minimizing the policy optimization function; Predicting a third probability distribution of the set of states, the set of observations, and reinforcement learning model parameters based on the optimal strategy, calculating a joint prior preference distribution for the set of states, the set of observations, and the reinforcement learning model parameters, and solving a KL divergence between the third probability distribution and the joint prior preference distribution as an expected free energy function; Convert the optimization problem of minimizing the expected free energy function into the optimization problem of maximizing the negative expected free energy function; Decomposing the negative expected free energy function into a first term capturing expected information gain and a second term capturing extrinsic value; Approximate posterior distributions are respectively solved for the first item and the second item to obtain an optimal strategy distribution, and an optimal decision set for sensor data transmission is obtained according to the optimal strategy distribution, so as to solve the overall objective optimization function through the optimal decision set.
3. The sensor data processing and transmission decision method for complementary perception in Internet of Vehicles according to claim 2, characterized in that: The optimal strategy is obtained by minimizing the strategy optimization function to solve the goal. The formula is expressed as: ; in, represents the optimal strategy, represents the decision set, Represents a state set, Indicates the current strategy, Indicates that the current strategy is given Predict the state set The first probability distribution of Represents a given decision set Predict the state set The second probability distribution of represents the Kullback-Leibler divergence function.
4. The sensor data processing and transmission decision method for complementary perception in Internet of Vehicles according to claim 2, characterized in that: The expected free energy function formula is expressed as: ; in, represents the expected free energy function, Represents a state set, represents a set of observations, represents the optimal strategy, represents the reinforcement learning model parameters, Represents a given optimal strategy Prediction status set , observation set , reinforcement learning model parameters The third probability distribution, Represents a state set , observation set , reinforcement learning model parameters The joint prior preference distribution of represents the Kullback-Leibler divergence function.
5. The sensor data processing and transmission decision method for complementary perception in Internet of Vehicles according to claim 2, characterized in that: The negative expected free energy function formula is expressed as: ; in, represents the negative expected free energy function, represents the decision set, Represents a state set, represents a set of observations, represents the optimal strategy, represents the reinforcement learning model parameters, Represents a given optimal strategy Prediction decision set The probability distribution of Represents a given optimal strategy Next decision set expectations, Represents a given optimal strategy , observation set Prediction status set , reinforcement learning model parameters The probability distribution of Represents a given optimal strategy Prediction status set , reinforcement learning model parameters The probability distribution of Represents a given optimal strategy Prediction status set The probability distribution of Represents a given optimal strategy Next state set expectations, Represents a given optimal strategy , state collection Predicted observation set , reinforcement learning model parameters The probability distribution of Represents a given optimal strategy Predicted observation set , reinforcement learning model parameters The prior preference distribution of represents the Kullback-Leibler divergence function.
Citation Information
Patent Citations
Multi-sensor over-the-horizon ad hoc network method based on traffic semantics and game theory
CN112437501A
CAV speed guidance system and method based on deep reinforcement learning
CN117612396A