Sensor data processing and transmission decision-making method for complementary perception of Internet of Vehicles
Through the intelligently guided reward-free reinforcement learning method, dynamically adjusting resource allocation and perception tasks, the problem of insufficient resource integration and resource utilization in the existing technology is solved, efficient perceived data processing and transmission is achieved, and the adaptability and accuracy of the system is improved.
Patent Information
- Application Number
- CN202510118719.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-01-24
Smart Images

Figure CN119996965A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of Internet of Vehicles, and specifically relates to a sensor data processing and transmission decision method for complementary perception in Internet of Vehicles. Background Art
[0002] In the connected vehicle environment, collaborative perception technology is essential for the safety and efficiency of autonomous vehicles. By enabling vehicles to share and receive perception data in real time, this technology greatly improves the vehicle's responsiveness to complex traffic environments. However, although collaborative perception technology provides vehicles with a more comprehensive environmental perception view, its requirements for communication and computing resources also increase accordingly, especially in dense urban environments.
[0003] Collaborative perception technology requires efficient processing and transmission of large amounts of data between autonomous vehicles and intelligent transportation infrastructure. This places high demands on communication bandwidth and computing resources, especially during peak traffic hours, which may lead to resource bottlenecks, affecting the timeliness and accuracy of perception data. Therefore, developing technologies that can efficiently utilize limited resources while ensuring data transmission quality and processing speed is crucial to achieving an efficient collaborative perception system for the Internet of Vehicles. Combining perception data from different sources and optimizing data transmission and processing processes to reduce the overall demand for resources and improve overall system performance are key technical directions for improving collaborative perception capabilities in future Internet of Vehicles systems.
[0004] Current research focuses on improving the performance of cooperative perception systems through advanced data processing algorithms and network optimization techniques. For example, some studies focus on reducing network burden through data compression and selective information sharing. Li et al. proposed a data compression and selective sharing mechanism in the document Y. Li, FR Yu, and M. Wang, "Data Compression and Selective Sharing for Cooperative Perception in Vehicular Networks", which effectively reduces network pressure by using intelligent algorithms to determine which data should be compressed or transmitted first; Chen et al. developed a resource-efficient data processing framework in the document X. Chen, H. Zhang, and L. Liu, "Resource-Efficient Data Processing for Cooperative Vehicular Networks", which aims to optimize the processing of perception data and the allocation of communication resources, reduce latency and increase data processing speed.
[0005] A large number of studies are also devoted to improving resource allocation mechanisms and optimizing data processing and transmission efficiency. For example, Wang et al. proposed a resource allocation optimization method for integrated communication and computing in the document Z. Wang, J. Liu, and S. Lee, "Optimized Resource Allocation for Integrated Communication and Computation in Cooperative Vehicular Networks". This method dynamically adjusts resource allocation through deep learning algorithms to cope with different traffic and network conditions, significantly improving the system's response speed and data processing capabilities.
[0006] However, existing technologies often fail to effectively integrate various sensor information when processing multi-source perception data, especially in terms of dynamic allocation of perception tasks and resource utilization. In addition, existing collaborative perception systems rely on traditional reward mechanisms to optimize data processing and transmission decisions, which often leads to insufficient adaptability in a changing traffic environment and makes it difficult to flexibly respond to environmental changes, thus affecting the real-time and accuracy of perception data. Summary of the invention
[0007] In order to solve the above problems existing in the prior art, the present invention provides a sensor data processing and transmission decision method for complementary perception of Internet of Vehicles. The technical problem to be solved by the present invention is achieved through the following technical solutions:
[0008] An embodiment of the present invention provides a sensor data processing and transmission decision method for complementary perception in an Internet of Vehicles, the method comprising:
[0009] Modeling a system model including CAVs, RAUs, and sensor ensembles;
[0010] Based on the system model, a detection accuracy objective function and a delay objective function are constructed, and a total objective optimization function composed of the detection accuracy objective function and the delay objective function is constructed by balancing weights;
[0011] Construct a state set including road and traffic state, sensor data state, vehicle dynamic state, and vehicle perception demand, a decision set including sensor management decision, data processing decision, and communication decision, and an observation value set including road and traffic observation values, sensor data observation values, vehicle dynamic observation values, vehicle perception demand observation values, and delay observation values, and based on the state set, the decision set, and the observation value set, use an intelligently guided reward-free reinforcement learning method to obtain the optimal decision set for sensor data transmission, so as to solve the overall objective optimization function through the optimal decision set.
[0012] Beneficial effects of the present invention:
[0013] The sensor data processing and transmission decision method for complementary perception of the Internet of Vehicles proposed in the present invention uses an intelligently guided non-reward reinforcement learning method, which has significant advantages over the prior art: the traditional collaborative perception system usually relies on a fixed reward mechanism to guide the processing and transmission of perception data. This method often cannot flexibly respond to various situations in a dynamically changing Internet of Vehicles environment, resulting in low efficiency in the collaboration between perception, communication and computing. The present invention uses an intelligently guided non-reward reinforcement learning method to dynamically adjust resource allocation and perception tasks according to real-time environmental data and vehicle status, thereby optimizing the collaboration efficiency between communication, perception and computing. At the same time, this method not only improves the efficiency of resource utilization, but also significantly reduces the delay of data processing and transmission. The intelligently guided non-reward reinforcement learning method of the present invention improves the adaptability of the system to complex traffic environments. In addition, the method can be flexibly adjusted according to different vehicles and road conditions, has good scalability, and is suitable for Internet of Vehicles environments of various sizes and types. In general, the reward-free reinforcement learning method based on intelligent guidance proposed in this invention aims to optimize the collaborative perception and communication efficiency of CAVs in the process of multi-source sensor data processing and transmission. This method enables the system to greatly improve the system's perception data processing efficiency and accuracy while ensuring real-time response, providing safer and more reliable driving support for the Internet of Vehicles.
[0014] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] Figure 1 It is a flow chart of a sensor data processing and transmission decision method for complementary perception of Internet of Vehicles provided by an embodiment of the present invention;
[0016] Figure 2 is a schematic diagram of the comparison results of three traditional methods and the method proposed in the present invention in terms of sensing data delay;
[0017] Figure 3 It is a schematic diagram of the comparison results of three traditional methods and the method proposed in the present invention in terms of perception fusion accuracy. DETAILED DESCRIPTION
[0018] The present invention is further described in detail below with reference to specific embodiments, but the embodiments of the present invention are not limited thereto.
[0019] See also Figure 1 The embodiment of the present invention provides a sensor data processing and transmission decision method for complementary perception of Internet of Vehicles, which specifically includes the following steps:
[0020] S10. Modeling a system model including CAVs, RAUs, and sensor collection.
[0021] The system modeling of the embodiment of the present invention includes the description of the scene and the modeling of perception characteristics, specifically including the sensor deployment position and height, perception distance, horizontal viewing angle, vertical viewing angle, obstacle impact, calculation and processing modeling, communication modeling and problem statement.
[0022] For example, the research scenario is set in a complex urban traffic environment equipped with an advanced system model RSNs, which includes CAVs as well as ordinary vehicles (referred to as CAVs with Level 0 intelligence), sensors and roadside assistance units (RAUs). The embodiment of the present invention envisions an advanced urban traffic environment equipped with multiple sensors, RAUs and CAVs, as well as ordinary vehicles. Each road is divided into multiple sections, each section is equipped with a RAU, and the RAUs are strategically deployed on the road to ensure optimal coverage and functionality. RAUs are used to aggregate perception data from sensors deployed on the roadside infrastructure and transmit the processed perception information to CAVs with different perception requirements. Assume that the group of CAVs served by a specific RAU is {1,2,…,V}, where V represents the number of CAVs in the system model. Similarly, the sensor set is defined as {1,2,…,S}, where S represents the total number of deployed sensors. CAVs have different requirements for perception data, such as detection accuracy and latency. Accuracy level (P v ) represents the level of detail of the data required for the vth CAV, ranging from the original data (P v =1, representing the highest accuracy) to abstract information (P v =0, indicating lower accuracy). In addition, CAVs have different tolerances to latency in sensory data, with critical operations requiring low latency (L v →0, while regular tasks can tolerate higher delays L v .
[0023] Sensors are strategically deployed on both sides or above the road to ensure that normal traffic flow is not disturbed. The location coordinates of each sensor are expressed as Where k is the deployment location number and l is the sensor number. The deployment height of the sensor is expressed by the following formula:
[0024]
[0025] Among them, H k,l represents the installation height of the lth sensor deployed at the deployment location numbered k from the ground. Due to budget and operational constraints, not all potential locations can deploy sensors. The number of sensors at each location is given by Range is 0 to the maximum number of sensors allowed at this location.
[0026] S20. Based on the system model, a detection accuracy objective function and a delay objective function are constructed, and an overall objective optimization function consisting of the detection accuracy objective function and the delay objective function is constructed by balancing weights.
[0027] The embodiment of the present invention constructs a detection accuracy objective function, including: constructing a distance perception detection probability function, a horizontal perception detection probability function, a vertical perception detection probability function, and an obstacle perception detection probability function of a single sensor respectively; and constructing a detection accuracy objective function by using the distance perception detection probability function, the horizontal perception detection probability function, the vertical perception detection probability function, and the obstacle perception detection probability function. More specifically:
[0028] The sensing distance refers to the maximum distance at which the sensor can effectively detect the target. The sensor R constructed in the embodiment of the present invention k,l The distance perception detection probability function is expressed as follows:
[0029]
[0030] in, represents the distance-aware detection probability function, d kl,j Indicates sensor R k,l With the jth target O j The distance between them, j ranges from 1 to J, J represents the number of targets, where the target refers to the object to be detected, such as a car, a person, a building, etc., R k,l represents the lth sensor deployed at the deployment location numbered k, Indicates sensor R k,l The maximum effective distance, Λ represents The sensitivity adjustment parameter, Λ∈[0,10], Γ represents The threshold adjustment parameter is Γ∈[0,100]; this formula describes the detection capability of the sensor as the distance changes, reflecting how the detection probability changes with increasing distance.
[0031] The horizontal viewing angle refers to the ability of the sensor to detect targets in the horizontal direction. The horizontal perception detection probability function formula constructed in the embodiment of the present invention is expressed as:
[0032]
[0033] in, represents the horizontal perception detection probability function, φ kl,j Indicates sensor R k,l With the jth target O j The horizontal deviation angle of the sensor axis, Ω represents The slope adjustment parameter, Ω∈[0,1], Π represents The maximum effective horizontal angle, Π∈[0,100]; this formula reflects the impact of horizontal deviation on the detection ability of the sensor, indicating that when the detection angle of the sensor deviates from the central axis, the detection probability will decrease.
[0034] The vertical viewing angle refers to the ability of the sensor to detect targets in the vertical direction. The vertical perception detection probability function formula constructed in the embodiment of the present invention is expressed as:
[0035]
[0036] in, represents the vertical perception detection probability function, θ kl,j Indicates sensor R k,l With the jth target O j The vertical deviation angle of the sensor axis, Δ, represents The slope adjustment parameter, Δ∈[0,1], Φ represents The maximum effective vertical angle is Φ∈[0,100]; this formula shows the impact of vertical deviation on the detection probability of the sensor, and reflects the change in the detection probability of the sensor when the detection angle deviates from the central axis.
[0037] Obstacles have a significant impact on the detection capability of sensors. The obstacle perception detection probability function formula constructed in the embodiment of the present invention is expressed as:
[0038]
[0039] in, represents the obstacle perception detection probability function, P obs Indicates the location information of the obstacle, G obs (P obs ) indicates P obs The obstacle characteristic factor of the obstacle at Indicates sensor R k,l The position information of the object, α represents the rate at which the perception probability changes with distance, α∈[2,4]. This formula quantifies the change in detection probability due to the presence of obstacles and explains how obstacles affect the detection performance of the sensor.
[0040] Finally, the detection accuracy objective function constructed by formulas (2) to (5) is expressed as:
[0041]
[0042] Furthermore, the system model modeled by the embodiment of the present invention also includes an edge computing model; the edge computing model includes a number of edge nodes, each edge node is equipped with computing resources; the edge computing model processes data from various sensors by introducing edge nodes, and the edge computing nodes are responsible for communicating with CAVs and providing perception assistance to CAVs based on currently available sensors. Each edge node is equipped with computing resources for processing sensor data.
[0043] The process of constructing a delay objective function in an embodiment of the present invention includes:
[0044] When CAVs send sensor perception data to RAUs, the RAUs processing capacity is calculated based on the computing resources and the size of the perception data; when RAUs send sensor perception data to CAVs, the CAVs processing capacity is calculated based on the computing resources and the size of the perception data; wherein the perception data includes the size of the perception data, the vehicle processing density, and the processing decision weight; the RAUs processing time is calculated based on the RAUs processing capacity and the perception data, and the CAVs processing time is calculated based on the CAVs processing capacity and the perception data; according to the characteristics of the wireless channel for communication between CAVs and RAUs, the uplink data transmission rate from CAVs to RAUs and the downlink data transmission rate from RAUs to CAVs are calculated respectively, and the uplink transmission time and the downlink transmission time are calculated based on the uplink data transmission rate and the downlink data transmission rate; the delay objective function is constructed by the CAVs processing time, the RAUs processing time, the uplink transmission time, and the downlink transmission time. More specifically:
[0045] The embodiment of the present invention calculates the RAUs processing capacity, and the formula is expressed as:
[0046]
[0047] in, represents the RAUs processing capability when processing the nth sensing data at time t, represents the available computing resources found by RAUs when CAVs send sensor perception data to RAUs at time t, D n (t) represents the size of the nth sensing data sent by CAVs to RAUs at time t, and N represents the number of sensing data sent by CAVs to RAUs;
[0048] The embodiment of the present invention calculates the CAVs processing capability, and the formula is expressed as:
[0049]
[0050] in, represents the CAVs processing capability when processing the n1th perception data at time t, It means that when RAUs send sensor perception data to CAVs at time t, CAVs find available computing resources. It represents the size of the n1th perception data sent by RAUs to CAVs at time t, and N′ represents the amount of perception data sent by RAUs to CAVs.
[0051] Through formulas (7) and (8), computing resources are distributed among all sensor data, effectively balancing the computing load and ensuring efficient data processing with minimized latency.
[0052] The embodiment of the present invention calculates the RAUs processing time, and the formula is expressed as:
[0053]
[0054] in, Indicates the RAUs processing time, a n (t) represents the decision processing weight of RAUs when processing the nth perception data at time t, a n (t) ranges from [0,1], λ n (t) represents the processing complexity of RAUs when processing the nth sensor data at time t. Formula (9) describes the processing delay of data on the RAU side, ensuring efficient execution of tasks and minimizing delays.
[0055] The embodiment of the present invention calculates the CAVs processing time, and the formula is expressed as:
[0056]
[0057] in, represents the CAVs processing time, represents the decision processing weight of CAVs when processing the n1th perception data at time t, The value range is [0,1]. It represents the processing complexity of CAVs when processing the n1th sensory data at time t. Formula (10) describes the processing delay of data on the CAV side, ensuring efficient execution of tasks and minimizing delays.
[0058] The communication model focuses on the interaction of CAVs, RAUs, and sensors within the system model RSNs, dealing with the effective transmission of sensor data and information exchange to optimize perception and decision-making. Communication is mainly carried out between CAVs and RAUs through wireless channels, including uplink (from CAV to RAU) and downlink (from RAU to CAV) data transmission.
[0059] The embodiment of the present invention calculates the uplink transmission time, and the formula is expressed as:
[0060]
[0061] Among them, T up Indicates the uplink transmission time. Indicates the uplink bandwidth of the wireless channel. represents the uplink transmission power of the wireless channel, represents the uplink channel gain of the wireless channel, represents the uplink interference power of the wireless channel, and N0 represents the noise power of the wireless channel.
[0062] The embodiment of the present invention calculates the downlink transmission time, and the formula is expressed as:
[0063]
[0064] Among them, T dn Indicates the downlink transmission time, Indicates the downlink bandwidth of the wireless channel. represents the downlink transmission power of the wireless channel, represents the downlink channel gain of the wireless channel, Indicates the downlink interference power of the wireless channel.
[0065] The parameters in formulas (11) and (12) jointly determine the efficiency and reliability of data transmission in the Internet of Vehicles. Delay is a key factor in the system model, especially for safety-critical applications in CAVs. The total delay constructed in the embodiment of the present invention includes CAV processing time, uplink transmission time, RAU processing time and downlink transmission time. The delay objective function constructed by formulas (9) to (12) is expressed as follows:
[0066]
[0067] Furthermore, the communication model is seamlessly integrated with the perception and computation models, which allows the transmission strategy of sensor data to be dynamically adjusted based on data availability, network conditions, and vehicle computing capabilities. The main goal of the system model of the embodiment of the present invention is to maximize the perception detection probability of CAVs while minimizing the delay caused by communication and task processing, so that the final constructed overall objective optimization function is expressed as follows:
[0068]
[0069] Among them, R represents the overall objective optimization function, Θ total For the optimization of detection accuracy, T total For optimization purposes, latency goals include processing latency and communication delay T up +T dn , ω represents the balance weight, V represents the number of CAVs in the system model, represents the perceptual detection probability of the vth CAV, Θ min represents the minimum perceptual detection threshold, T total represents the total transmission delay, which is calculated by formula (13), T max Indicates the maximum delay threshold, Indicates the uplink bandwidth of the wireless channel. represents the downlink bandwidth of the wireless channel, W total represents the total bandwidth of the wireless channel, It represents the computing processing capacity of the i-th RAU, i ranges from 1 to I, and I represents the number of RAUs in the system model. Indicates the maximum computing processing capacity of RAU, represents the number of sensors deployed at the deployment location numbered k, Represents the maximum number of sensors deployed. Among them, in the constraints: The first constraint ensures that the perception detection probability of each CAV is not less than the minimum detection threshold Θ required by the system min , to ensure the basic perceptual performance requirements; the second constraint limits the total delay of each CAV to not exceed the maximum delay threshold T allowed by the system max , ensuring the timeliness of the data; the third constraint indicates that the sum of the bandwidth allocated for uplink and downlink cannot exceed the total bandwidth W of the wireless channel total ; The fourth constraint ensures that each roadside unit RAU i The computing resource usage of the system does not exceed its maximum computing capacity; the fifth constraint limits the total number of sensors deployed in the system model to not exceed the maximum number of sensors allowed by the system. These constraints together ensure the reasonable allocation of system resources and the satisfaction of performance requirements.
[0070] The embodiment of the present invention integrates multi-source sensor data in a vehicle networking environment, and optimizes data preprocessing and analysis through edge computing technology to ensure the real-time and reliability of the data.
[0071] S30. Construct a state set including road and traffic states, sensor data states, vehicle dynamic states, and vehicle perception needs, a decision set including sensor management decisions, data processing decisions, and communication decisions, and an observation value set including road and traffic observation values, sensor data observation values, vehicle dynamic observation values, vehicle perception need observation values, and delay observation values, and based on the state set, decision set, and observation value set, use an intelligently guided reward-free reinforcement learning method to obtain the optimal decision set for sensor data transmission, so as to solve the overall objective optimization function through the optimal decision set.
[0072] In view of the limitation of traditional methods that rely on explicit reward functions, the embodiments of the present invention introduce a non-reward intelligently guided reinforcement learning method to replace the traditional reinforcement learning method of defining reward functions according to task types. The non-reward intelligently guided reinforcement learning method optimizes the perception information processing and transmission decision of the collaborative perception of the Internet of Vehicles, so that the present invention does not rely on the traditional reward mechanism, but uses the understanding of the environment as a guide to dynamically adjust resource allocation and optimize the data processing process. The specific introduction of the non-reward intelligently guided reinforcement learning method is as follows:
[0073] The environment state of reinforcement learning is a multidimensional structure that contains various aspects that affect the operating dynamics of CAVs. At any given time t, the environment state is represented by S t , S t include:
[0074] Road and traffic conditions reflects dynamic and static road conditions, including traffic density, road construction and weather conditions;
[0075] Sensor data status Aggregates a large amount of data from different on-board and roadside sensors;
[0076] Vehicle dynamic status Covers key parameters such as vehicle speed, direction and other relevant dynamic states;
[0077] Vehicle Sensing Requirements It is used to represent the perception needs of CAVs and can also reflect the intelligence level of the vehicle.
[0078] Therefore, the state set S at time t t It can be expressed as:
[0079]
[0080] In the reinforcement learning framework, the observation set O t It can be expressed as:
[0081]
[0082] in, represents road and traffic observation values, represents the sensor data observation value, represents the vehicle dynamic observation value, represents the vehicle perceived demand observation value, Represents a delayed observation.
[0083] Agents usually cannot directly access the real state of the environment, but obtain state information through observation. Therefore, the embodiment of the present invention considers preference as a unique observation mode. Given a state set S t , observation set O t , the joint prior preference distribution of the reinforcement learning model parameters θ can be expressed as p Ψ (S t ,O t ,θ).
[0084] Decision set (A t ) represents the strategic response of the system at time t, which is affected by the current state. t Includes: Sensor Management Decisions Involves decisions related to sensor selection and tuning; data processing decisions Methods involving sensor data processing; communication decision making Determines the information exchange strategy between CAVs and RAUs. Therefore, the decision set A at time t is t It is expressed as:
[0085]
[0086] The core of the data transfer decision lies in evaluating whether to process the data locally or transfer it to the RAUs for more comprehensive processing. This decision is represented by the binary variable Where 1 represents transmission to RAU and 0 represents local processing, as shown in the following formula:
[0087]
[0088] in, Represents an edge node (E i ) of the load, Represents an edge node (E i ) data processing requirements, θ trans Indicates the transmission threshold.
[0089] The dynamics of the system include the transition of states and the impact of decisions. The state transition is expressed by the following formula:
[0090] S t =F(S t ,A t ) (19);
[0091] Where F() represents the dynamics of the CAV environment that integrates the current state and decision. In the framework of reward-free reinforcement learning, reward-free guidance is achieved through a set of abstract environmental states and goals, bypassing the need for explicit reward signals. The decision algorithm is based on active inference, which uses these guidance to make wise choices that are consistent with the overall goal of the system, namely to improve the operating safety and efficiency of CAVs in RSNs. In addition, when higher-level cognition develops, there is a difference between the predicted state and the actual state, which changes through the learning process. In the study, the embodiments of the present invention used the basic concepts of intelligence to evaluate the changes in this difference. Intelligence, as an advanced indicator for quantifying learning effects, borrows the concepts of energy and information. As a comparative measure, it evaluates the changes in information distribution over time due to learning or the degree of information dispersion relative to the initial state.
[0092] Furthermore, the embodiment of the present invention uses an intelligently guided non-rewarded reinforcement learning method based on a state set, a decision set, and an observation value set to obtain an optimal decision set for sensor data transmission, so as to solve the overall objective optimization function through the optimal decision set, including:
[0093] Given the first probability distribution of the state set predicted by the current policy, and the second probability distribution of the state set predicted by the decision set, the KL divergence between the first probability distribution and the second probability distribution is used as the policy optimization function to solve the optimal policy by minimizing the policy optimization function; given the third probability distribution of the state set predicted by the optimal policy, the observation value set, and the reinforcement learning model parameters, the joint prior preference distribution of the given state set, the observation value set, and the reinforcement learning model parameters is calculated, and the KL divergence between the third probability distribution and the joint prior preference distribution is used as the expected free energy function; the optimization problem of minimizing the expected free energy function is converted into the optimization problem of maximizing the negative expected free energy function; the negative expected free energy function is decomposed into a first term that captures the expected information gain and a second term that captures the extrinsic value; the first term and the second term are respectively approximated by the posterior distribution to obtain the optimal policy distribution, and the optimal decision set for sensor data transmission is obtained according to the optimal policy distribution, so as to solve the overall objective optimization function through the optimal decision set. More specifically:
[0094] The basic concept of the decision framework is to optimize decisions based on predicting the impact of decisions on the future state of the environment. It includes maintaining a probabilistic model of the environment and predicting the results of different decisions. Therefore, the embodiment of the present invention predicts a first probability distribution of the state set given the current strategy, and predicts a second probability distribution of the state set given the decision set, and solves the KL divergence between the first probability distribution and the second probability distribution as a policy optimization function, and solves the optimal policy by minimizing the policy optimization function as the goal; the optimal policy is solved by minimizing the policy optimization function as the goal, and the formula is expressed as:
[0095]
[0096] Among them, π * represents the optimal strategy, A t represents the decision set, S t represents the state set, π represents the current strategy, q(S t |π) represents the predicted state set S given the current strategy π t The first probability distribution, P(S t |A t ) represents a given decision set A t Prediction state set S t The second probability distribution, D KL (·||·) represents the Kullback-Leibler divergence function. The model is continuously updated based on the decision set to ensure accurate reflection of the dynamic changes in the CAVs environment:
[0097] P(S t+1 |S t ,A t )=f(S t ,A t ,θ) (21);
[0098] Among them, f() is a function describing the probability of state transition, θ represents the reinforcement learning model parameters, P(S t+1 |S t ,A t ) represents a given decision set A t , state set S t Prediction state set S t+1 The probability distribution of .
[0099] The decision-making process of the embodiment of the present invention aims to minimize the expected future free energy, which represents the difference between the predicted and target state distributions. By continuously adjusting the optimal strategy π * To minimize the KL divergence, we can make optimal decisions under uncertainty and enhance adaptability. The strategy selection process is achieved through the expected free energy function The minimization of is used to model, where the expected free energy function formula is expressed as:
[0100]
[0101] in, represents the expected free energy function, S t Represents a state set, O t represents the set of observations, π * represents the optimal strategy, θ represents the reinforcement learning model parameters, q(S t ,O t ,θ|π * ) represents the given optimal strategy π * Prediction state set S t , observation set O t , the third probability distribution of the reinforcement learning model parameter θ, p Ψ (S t ,O t ,θ) represents the state set S t , observation set O t , the joint prior preference distribution of reinforcement learning model parameters θ, D KL (·||·) denotes the Kullback-Leibler divergence function.
[0102] The optimal strategy distribution q(π * ) can be minimized by the following expected free energy function formula:
[0103]
[0104] therefore, It shows that the strategy that minimizes the expected free energy is more likely to be selected, promoting a unified framework for exploration and exploitation. σ represents the Sigmoid function, which is used to convert the input value into a value with probability between 0 and 1. The goal of active inference is to minimize the expected free energy. However, by converting to maximize the negative expected free energy We can transform the problem into a standard maximization problem. This is more consistent with traditional optimization theory and makes the problem easier to handle and understand. The balance between exploration and exploitation is Decomposed into a first term that captures the expected information gain and a second term that captures the extrinsic value, this decomposition is an approximation because it ignores some higher-order terms and cross terms. However, in practice, this approximation is usually accurate enough and greatly simplifies calculations and interpretations. The negative expected free energy function formula of the embodiment of the present invention is expressed as:
[0105]
[0106] in, represents the negative expected free energy function, A t represents the decision set, S t Represents a state set, O t represents the set of observations, π * represents the optimal strategy, θ represents the reinforcement learning model parameters, q(A t |π * ) represents the given optimal strategy π * Prediction decision set A t The probability distribution of Denotes the given optimal strategy π * The expectation of the decision set under t ,θ|O t ,π * ) represents the given optimal strategy π * , observation set O t Prediction state set S t , the probability distribution of the reinforcement learning model parameter θ, q(S t ,θ|π * ) represents the given optimal strategy π * Prediction state set S t , the probability distribution of the reinforcement learning model parameter θ, q(S t |π * ) represents the given optimal strategy π * Prediction state set S t The probability distribution of Denotes the given optimal strategy π * Next, the state set S t The expectation, q(O t ,θ|S t ,π * ) represents the given optimal strategy π * , state set S t Predicted observation set O t , the probability distribution of reinforcement learning model parameters θ, p Ψ (O t ,θ|π * ) represents the given optimal strategy π * Predicted observation set O t , the prior preference distribution of the reinforcement learning model parameter θ, D KL (·||·) denotes the Kullback-Leibler divergence function.
[0107] Formula (24) naturally integrates exploration behavior (through information gain) and the use of known strategies (through extrinsic value). Based on the variational inference principle, is considered as a variational lower bound that needs to be maximized. By applying variational inference, this maximization problem can be transformed into finding the best approximate posterior distribution q. At each time t, the decision set reflects the decisions of sensor management, data processing and communication strategies. In order to dynamically optimize the strategy distribution q(π * ), where the first term is the expectation of the log-likelihood and the second term is the entropy of the approximate posterior. By maximizing formula (24), the optimal approximate posterior q distribution can be obtained:
[0108]
[0109] Among them, S τ:T represents the set of all states from time τ to time T, O τ:T represents the set of all observations from time τ to time T, q(S τ:T ,O τ:T ,θ|π * ) is the optimal approximate posterior q distribution, that is, given the optimal strategy π * Prediction state set S τ:T , observation set O τ:T , the probability distribution of the reinforcement learning model parameter θ, q(O τ |S τ ,θ,π * ) represents a given state set S τ 、Optimal strategy π * , reinforcement learning model parameters θ predict observation set O τ The probability distribution of q(S τ |S τ-1 ,θ,π * ) represents a given state set S τ-1 、Optimal strategy π * , reinforcement learning model parameters θ predict state set S τ The probability distribution of p Ψ (O τ |S τ ) represents a given state set S τ Predicted observation set O τ The prior preference distribution of Denotes the given optimal strategy π * , the state set S under the reinforcement learning model parameter θ τ expectations, Represents a given state set S τ-1 、Optimal strategy π * , reinforcement learning model parameters θ predict state set S τ The prior preference distribution of Denotes the given optimal strategy π* , the state set S under the reinforcement learning model parameter θ τ-1 expectations.
[0110] The decomposition form of formula (25) reflects the dynamic characteristics of the system, in which the state and observation at each moment depend on the state at the previous moment. * ) is transformed into a diagonal Gaussian distribution to simplify the calculation, and the Exponential Change of Measure (ECM) algorithm is used to optimize q(π * ) to make it as close as possible to The goal of the ECM algorithm is to determine the best strategy for sensor data transmission based on the current state of the CAVs and RSN environment. Ultimately, the overall objective function of formula (14) can be optimized through the action of the optimal strategy.
[0111] The embodiment of the present invention adopts an innovative intelligent-guided non-reward reinforcement learning method, which does not rely on traditional reward mechanisms, but guides decision-making through environmental understanding. This method can dynamically adjust resource allocation and optimize the processing and transmission of perception information, thereby adapting to highly dynamic and changeable traffic environments.
[0112] In order to verify the effectiveness of the sensor data processing and transmission decision method for complementary perception of the Internet of Vehicles provided by the embodiment of the present invention, the superiority of the present invention in terms of perception accuracy and latency in multiple actual driving scenarios is demonstrated by comparison with traditional methods. The numerical results clearly show that the proposed method (Our Proposal) has significant improvements over the three traditional methods (Rainbow DRL, SoftActor-Critic RL, Policy Optimization) in improving the efficiency of perception, communication and computing collaboration, as well as resource allocation. See Figure 2 :The method proposed in the present invention ensures the detection delay in different scenarios. Compared with the other three traditional methods, when the traffic density increases, the delay of the method proposed in the present invention reaches a peak of 76.38 milliseconds at medium density, and then decreases slightly at the highest density. It can be seen that the method proposed in the present invention can reduce the detection delay under high traffic density and show excellent delay management ability. Please refer to Figure 3When the number of roadside sensors increases, the overall detection accuracy of the proposed method increases rapidly from 63.91% to 94.56%, while the other three traditional methods are difficult to reach the detection level of the proposed method under the same number of available sensors. Comprehensive simulation results show that the proposed intelligent guidance-based non-reward reinforcement learning method performs well in optimizing detection accuracy, reducing detection delay and maintaining data packet delivery rate, demonstrating great potential in the system model RSNs.
[0113] In summary, the sensor data processing and transmission decision method for complementary perception of the Internet of Vehicles proposed in the embodiment of the present invention, using the intelligently guided non-reward reinforcement learning method, has significant advantages over the prior art: the traditional collaborative perception system usually relies on a fixed reward mechanism to guide the processing and transmission of perception data. This method is often unable to flexibly respond to various situations in a dynamically changing Internet of Vehicles environment, resulting in low efficiency in the collaboration between perception, communication and computing. The present invention, through the intelligently guided non-reward reinforcement learning method, can dynamically adjust resource allocation and perception tasks according to real-time environmental data and vehicle status, thereby optimizing the collaboration efficiency between communication, perception and computing. At the same time, this method not only improves the efficiency of resource utilization, but also significantly reduces the delay of data processing and transmission. The intelligently guided non-reward reinforcement learning method of the present invention improves the adaptability of the system to complex traffic environments. In addition, the method can be flexibly adjusted according to different vehicles and road conditions, has good scalability, and is suitable for Internet of Vehicles environments of various sizes and types. In general, the non-rewarded reinforcement learning method based on intelligent guidance proposed in the embodiment of the present invention aims to optimize the collaborative perception and communication efficiency of CAVs in the process of multi-source sensor data processing and transmission. This method enables the system to greatly improve the system's perception data processing efficiency and accuracy while ensuring real-time response, providing safer and more reliable driving support for the Internet of Vehicles.
[0114] In the description of the present invention, it should be understood that the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the features. In the description of the present invention, the meaning of "plurality" is two or more, unless otherwise clearly and specifically defined.
[0115] Although the present invention is described herein in conjunction with various embodiments, in the process of implementing the claimed invention, those skilled in the art may understand and implement other variations of the disclosed embodiments by viewing the specification and its drawings. In the specification, the word "comprising" does not exclude other components or steps, and "a" or "one" does not exclude multiple situations. Certain measures are recorded in different embodiments, but this does not mean that these measures cannot be combined to produce good results.
[0116] The above contents are further detailed descriptions of the present invention in combination with specific preferred embodiments, and it cannot be determined that the specific implementation of the present invention is limited to these descriptions. For ordinary technicians in the technical field to which the present invention belongs, several simple deductions or substitutions can be made without departing from the concept of the present invention, which should be regarded as falling within the protection scope of the present invention.
Claims
1. A sensor data processing and transmission decision method for complementary perception in Internet of Vehicles, characterized in that: The method comprises: Modeling a system model including CAVs, RAUs, and sensor ensembles; Based on the system model, a detection accuracy objective function and a delay objective function are constructed, and a total objective optimization function composed of the detection accuracy objective function and the delay objective function is constructed by balancing weights; Construct a state set including road and traffic state, sensor data state, vehicle dynamic state, and vehicle perception demand, a decision set including sensor management decision, data processing decision, and communication decision, and an observation value set including road and traffic observation values, sensor data observation values, vehicle dynamic observation values, vehicle perception demand observation values, and delay observation values, and based on the state set, the decision set, and the observation value set, use an intelligently guided reward-free reinforcement learning method to obtain the optimal decision set for sensor data transmission, so as to solve the overall objective optimization function through the optimal decision set.
2. The sensor data processing and transmission decision method for complementary perception of Internet of Vehicles according to claim 1 is characterized in that: Construct the detection accuracy objective function, including: Construct the distance perception detection probability function, horizontal perception detection probability function, vertical perception detection probability function and obstacle perception detection probability function of a single sensor respectively; The detection accuracy target function is jointly constructed by the distance perception detection probability function, the horizontal perception detection probability function, the vertical perception detection probability function and the obstacle perception detection probability function.
3. The sensor data processing and transmission decision method for complementary perception of Internet of Vehicles according to claim 2 is characterized in that: The distance perception detection probability function constructed is expressed as follows: in, represents the distance-aware detection probability function, d kl,j Indicates sensor R k,l With the jth target O j The distance between them, j ranges from 1 to J, J represents the number of targets, R k,l represents the lth sensor deployed at the deployment location numbered k, Indicates sensor R k,l The maximum effective distance, Λ represents The sensitivity adjustment parameter, Γ represents Threshold adjustment parameters of The horizontal perception detection probability function formula is expressed as: in, represents the horizontal perception detection probability function, φ kl,j Indicates sensor R k,l With the jth target O j The horizontal deviation angle of the sensor axis, Ω represents The slope adjustment parameter, Π represents The maximum effective horizontal angle; The vertical perception detection probability function formula is expressed as: in, represents the vertical perception detection probability function, θ kl,j Indicates sensor R k,l With the jth target O j The vertical deviation angle of the sensor axis, Δ, represents The slope adjustment parameter, Φ, represents The maximum effective vertical angle; The obstacle perception detection probability function formula is expressed as: in, represents the obstacle perception detection probability function, P obs Indicates the location information of the obstacle, G obs (P obs ) indicates P obs The obstacle characteristic factor of the obstacle at Indicates sensor R k,l The position information of the sensor is represented by α, and α represents the rate at which the perception probability changes with distance.
4. The sensor data processing and transmission decision method for complementary perception of Internet of Vehicles according to claim 3 is characterized in that: The modeled system model also includes an edge computing model; the edge computing model includes a number of edge nodes, each edge node is equipped with computing resources; correspondingly constructing a delay objective function, including: When CAVs send sensor perception data to RAUs, the RAUs processing capacity is calculated according to the computing resources and the size of the perception data; when RAUs send sensor perception data to CAVs, the CAVs processing capacity is calculated according to the computing resources and the size of the perception data; wherein the perception data includes the size of the perception data, the processing complexity, and the decision processing weight; calculating RAUs processing time according to the RAUs processing capability and the sensing data, and calculating CAVs processing time according to the CAVs processing capability and the sensing data; According to the characteristics of the wireless channel for communication between the CAVs and the RAUs, respectively calculating an uplink data transmission rate from the CAVs to the RAUs and a downlink data transmission rate from the RAUs to the CAVs, and calculating an uplink transmission time and a downlink transmission time according to the uplink data transmission rate and the downlink data transmission rate; A delay objective function is constructed by the CAVs processing time, the RAUs processing time, the uplink transmission time and the downlink transmission time.
5. The sensor data processing and transmission decision method for complementary perception of Internet of Vehicles according to claim 4 is characterized in that: Calculate the RAUs processing capacity, the formula is expressed as: in, represents the RAUs processing capability when processing the nth sensing data at time t, represents the available computing resources found by RAUs when CAVs send sensor perception data to RAUs at time t, D n (t) represents the size of the nth sensing data sent by CAVs to RAUs at time t, and N represents the number of sensing data sent by CAVs to RAUs; Calculate the CAVs processing capacity, the formula is expressed as: in, represents the CAVs processing capability when processing the n1th perception data at time t, It means that when RAUs send sensor perception data to CAVs at time t, CAVs find available computing resources. represents the size of the n1th sensing data sent by RAUs to CAVs at time t, and N' represents the number of sensing data sent by RAUs to CAVs; Calculate the RAUs processing time, the formula is expressed as: in, Indicates the RAUs processing time, a n (t) represents the decision processing weight of RAUs when processing the nth perception data at time t, a n (t) ranges from [0,1], λ n (t) represents the processing complexity of RAUs when processing the nth sensing data at time t; Calculate the CAVs processing time, the formula is expressed as: in, represents the CAVs processing time, represents the decision processing weight of CAVs when processing the n1th perception data at time t, The value range is [0,1]. represents the processing complexity of CAVs when processing the n1th perception data at time t; Calculate the uplink transmission time, the formula is: Among them, T up Indicates the uplink transmission time. Indicates the uplink bandwidth of the wireless channel. represents the uplink transmission power of the wireless channel, represents the uplink channel gain of the wireless channel, represents the uplink interference power of the wireless channel, and N0 represents the noise power of the wireless channel; Calculate the downlink transmission time, the formula is: Among them, T dn Indicates the downlink transmission time, Indicates the downlink bandwidth of the wireless channel. represents the downlink transmission power of the wireless channel, represents the downlink channel gain of the wireless channel, Indicates the downlink interference power of the wireless channel.
6. The sensor data processing and transmission decision method for complementary perception of Internet of Vehicles according to claim 5 is characterized in that: The constructed detection accuracy objective function is expressed as follows: The delay objective function constructed is expressed as follows: The total objective optimization function constructed is expressed as follows: Where R represents the overall objective optimization function, ω represents the balance weight, V represents the number of CAVs in the system model, represents the perceptual detection probability of the vth CAV, Θ min represents the minimum perceptual detection threshold, T total represents the total transmission delay, T max Indicates the maximum delay threshold, Indicates the uplink bandwidth of the wireless channel. represents the downlink bandwidth of the wireless channel, W total represents the total bandwidth of the wireless channel, It represents the computing processing capacity of the i-th RAU, i ranges from 1 to I, and I represents the number of RAUs in the system model. Indicates the maximum computing processing capacity of RAU, represents the number of sensors deployed at the deployment location numbered k, Indicates the maximum number of sensors deployed.
7. The sensor data processing and transmission decision method for complementary perception of Internet of Vehicles according to claim 1 is characterized in that: Based on the state set, the decision set and the observation value set, an optimal decision set for sensor data transmission is obtained by using an intelligently guided non-rewarded reinforcement learning method, so as to solve the overall objective optimization function through the optimal decision set, including: Given the current strategy predicting a first probability distribution of the state set, given the decision set predicting a second probability distribution of the state set, solving the KL divergence between the first probability distribution and the second probability distribution as a strategy optimization function, and solving the optimal strategy by minimizing the strategy optimization function as a goal; Given the optimal strategy, predicting a third probability distribution of the state set, the observation set, and the reinforcement learning model parameters, calculating a joint prior preference distribution of the state set, the observation set, and the reinforcement learning model parameters, and solving the KL divergence between the third probability distribution and the joint prior preference distribution as the expected free energy function; The optimization problem of minimizing the expected free energy function is converted into the optimization problem of maximizing the negative expected free energy function; Decomposing the negative expected free energy function into a first term capturing expected information gain and a second term capturing extrinsic value; The approximate posterior distribution of the first item and the second item is solved respectively to obtain the optimal strategy distribution, and the optimal decision set for sensor data transmission is obtained according to the optimal strategy distribution, so as to solve the overall objective optimization function through the optimal decision set.
8. The sensor data processing and transmission decision method for complementary perception of Internet of Vehicles according to claim 7 is characterized in that: The optimal strategy is obtained by minimizing the strategy optimization function as the goal, and the formula is expressed as: Among them, π * represents the optimal strategy, A t represents the decision set, S t represents the state set, π represents the current strategy, q(S t |π) represents the state set S predicted by the current strategy π t The first probability distribution, P(S t |A t ) represents a given decision set A t Predict the state set S t The second probability distribution, D KL (·||·) denotes the Kullback-Leibler divergence function.
9. The sensor data processing and transmission decision method for complementary perception of Internet of Vehicles according to claim 7, characterized in that: The expected free energy function formula is expressed as: in, represents the expected free energy function, S t Represents a state set, O t represents the set of observations, π * represents the optimal strategy, θ represents the reinforcement learning model parameters, q(S t ,O t ,θ|π * ) represents the given optimal strategy π * Prediction state set S t , observation set O t , the third probability distribution of the reinforcement learning model parameter θ, p Ψ (S t ,O t ,θ) represents the state set S t , observation set O t , the joint prior preference distribution of reinforcement learning model parameters θ, D KL (·||·) denotes the Kullback-Leibler divergence function.
10. The sensor data processing and transmission decision method for complementary perception of Internet of Vehicles according to claim 7, characterized in that: The negative expected free energy function formula is expressed as: in, represents the negative expected free energy function, A t represents the decision set, S t Represents a state set, O t represents the set of observations, π * represents the optimal strategy, θ represents the reinforcement learning model parameters, q(A t |π * ) represents the given optimal strategy π * Prediction decision set A t The probability distribution of Denotes the given optimal strategy π * Next, the decision set A t The expectation, q(S t ,θ|O t ,π * ) represents the given optimal strategy π * , observation set O t Prediction state set S t , the probability distribution of the reinforcement learning model parameter θ, q(S t ,θ|π * ) represents the given optimal strategy π * Prediction state set S t , the probability distribution of the reinforcement learning model parameter θ, q(S t |π * ) represents the given optimal strategy π * Prediction state set S t The probability distribution of Denotes the given optimal strategy π * Next, the state set S t The expectation, q(O t ,θ|S t ,π * ) represents the given optimal strategy π * , state set S t Predicted observation set O t , the probability distribution of reinforcement learning model parameters θ, p Ψ (O t ,θ|π * ) represents the given optimal strategy π * Predicted observation set O t , the prior preference distribution of the reinforcement learning model parameter θ, D KL (·||·) denotes the Kullback-Leibler divergence function.
Citation Information
Patent Citations
Multi-sensor over-the-horizon ad hoc network method based on traffic semantics and game theory
CN112437501A
CAV speed guidance system and method based on deep reinforcement learning
CN117612396A
Segmentable task unloading and service caching method in cloud-side collaborative Internet of Vehicles scene
CN118540743A
Three-dimensional roadside sensor deployment method for collaborative awareness of Internet of Vehicles
CN119211868A
Unmanned aerial vehicle body sensing network deployment method based on entropy feedback control
CN119233275A
Cited By
Self-adaptive multi-unmanned aerial vehicle information transmission system and method based on channel estimation
CN120811471A
Adaptive multi-uav information transmission system and method based on channel estimation
CN120811471B
A perception decision control closed-loop compression method and system for internet of vehicles
CN122747946A