Method for sensor data processing and transmission decision for complementary perception in internet of vehicles

By using an intelligently guided, reward-free reinforcement learning method, the sensor data processing and transmission decision-making of the vehicle-to-everything (V2X) system are optimized. This solves the problems of insufficient integration of multi-source perception data and resource utilization in traditional systems, achieving efficient and flexible perception and communication, and improving the system's adaptability and data processing accuracy.

WO2026157158A1PCT designated stage Publication Date: 2026-07-30XIDIAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
XIDIAN UNIV
Filing Date
2025-07-24
Publication Date
2026-07-30

AI Technical Summary

Technical Problem

Existing vehicle-to-everything (V2X) cooperative perception systems fail to effectively integrate information from various sensors when processing multi-source perception data. In particular, they are inadequate in dynamically allocating perception tasks and utilizing resources. Furthermore, traditional reward mechanisms are not adaptable enough to changing traffic environments, affecting the real-time performance and accuracy of perception data.

Method used

A system model is constructed using an intelligently guided, rewardless reinforcement learning method. By balancing the objective functions of detection accuracy and delay, and combining sensor management, data processing, and communication decisions, the system optimizes resource allocation and perception tasks using rewardless reinforcement learning, and dynamically adjusts resource allocation to adapt to complex traffic environments.

Benefits of technology

It improves the system's adaptability to complex traffic environments, optimizes the collaborative efficiency of communication, perception, and computing, reduces data processing and transmission latency, and enhances the efficiency and accuracy of perception data processing, making it suitable for vehicle-to-everything (V2X) environments of various sizes and types.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025110315_30072026_PF_FP_ABST
    Figure CN2025110315_30072026_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed in the present invention is a method for sensor data processing and transmission decision for complementary perception in Internet of Vehicles, comprising: modeling a system model comprising CAVs, RAUs, and a sensor set; on the basis of the system model, constructing a detection accuracy objective function and constructing a latency objective function, and constructing, by means of a balance weight, a total objective optimization function composed of the detection accuracy objective function and the latency objective function; and constructing a state set, a decision set, and an observation value set, on the basis of the state set, the decision set, and the observation value set, using an intelligently guided reward-free reinforcement learning method to obtain an optimal decision set for sensor data transmission, and solving the total objective optimization function by means of the optimal decision set. In the present invention, by means of the intelligently guided reward-free reinforcement learning method, resource allocation and sensing tasks can be dynamically adjusted on the basis of real-time environmental data and a vehicle state, thereby optimizing the collaboration efficiency among communication, sensing, and computing, improving the resource utilization efficiency, and reducing the data processing and transmission latency.
Need to check novelty before this filing date? Find Prior Art

Description

Sensor data processing and transmission decision-making methods for complementary perception in vehicle-to-everything (V2X) networks Technical Field

[0001] This invention belongs to the field of vehicle networking, specifically relating to a sensor data processing and transmission decision-making method for complementary perception in vehicle networking. Background Technology

[0002] In the connected vehicle environment, cooperative perception technology is crucial for the safety and efficiency of autonomous vehicles. By enabling vehicles to share and receive perception data in real time, this technology greatly improves their responsiveness to complex traffic environments. However, while cooperative perception technology provides vehicles with a more comprehensive view of their environment, it also increases the demands on communication and computing resources, especially in densely populated urban environments.

[0003] Cooperative perception technology requires the efficient processing and transmission of large amounts of data between autonomous vehicles and intelligent transportation infrastructure. This places high demands on communication bandwidth and computing resources, especially during peak traffic periods, potentially leading to resource bottlenecks and affecting the timeliness and accuracy of perceived data. Therefore, developing technologies that can efficiently utilize limited resources while ensuring data transmission quality and processing speed is crucial for achieving efficient vehicle-to-everything (V2X) cooperative perception systems. Combining perception data from different sources and optimizing data transmission and processing to reduce overall resource requirements and improve overall system performance is a key technological direction for enhancing the cooperative perception capabilities of future V2X systems.

[0004] Current research focuses on improving the performance of cooperative perception systems through advanced data processing algorithms and network optimization techniques. For example, some studies focus on reducing network load through data compression and selective information sharing. Li et al., in their paper "Data Compression and Selective Sharing for Cooperative Perception in Vehicular Networks," proposed a data compression and selective sharing mechanism that uses intelligent algorithms to determine which data should be compressed or prioritized for transmission, effectively alleviating network pressure. Chen et al., in their paper "Resource-Efficient Data Processing for Cooperative Vehicular Networks," developed a resource-efficient data processing framework aimed at optimizing the processing of perception data and the allocation of communication resources, reducing latency and improving data processing speed.

[0005] Numerous studies have also focused on improving resource allocation mechanisms and optimizing data processing and transmission efficiency. For example, Wang et al., in their paper "Optimized Resource Allocation for Integrated Communication and Computation in Cooperative Vehicular Networks," proposed a resource allocation optimization method that integrates communication and computation. This method dynamically adjusts resource allocation through deep learning algorithms to cope with different traffic and network conditions, significantly improving the system's response speed and data processing capabilities.

[0006] However, existing technologies often fail to effectively integrate information from various sensors when processing multi-source sensing data, particularly in dynamically allocating sensing tasks and utilizing resources. Furthermore, existing collaborative sensing systems rely on traditional reward mechanisms to optimize data processing and transmission decisions, which often leads to insufficient adaptability in dynamic traffic environments, making it difficult to flexibly respond to environmental changes and thus affecting the real-time performance and accuracy of the sensing data. Summary of the Invention

[0007] To address the aforementioned problems in the existing technology, this invention provides a sensor data processing and transmission decision-making method for complementary sensing in vehicle-to-everything (V2X) networks. The technical problem to be solved by this invention is achieved through the following technical solution:

[0008] This invention provides a sensor data processing and transmission decision-making method for complementary perception in vehicle-to-everything (V2X) networks, the method comprising:

[0009] Modeling includes system models of CAVs, RAUs, and sensor assemblies;

[0010] Based on the system model, a detection accuracy objective function and a delay objective function are constructed. A total objective optimization function composed of the detection accuracy objective function and the delay objective function is constructed by balancing the weights.

[0011] A set of states is constructed, including road and traffic conditions, sensor data conditions, vehicle dynamic conditions, and vehicle perception needs; a set of decisions is constructed, including sensor management decisions, data processing decisions, and communication decisions; and a set of observations is constructed, including road and traffic observations, sensor data observations, vehicle dynamic observations, vehicle perception needs observations, and delay observations. Based on the set of states, the set of decisions, and the set of observations, an optimal set of decisions for sensor data transmission is obtained using an intelligently guided, reward-free reinforcement learning method. The overall objective function is then solved using the optimal set of decisions.

[0012] The beneficial effects of this invention are:

[0013] This invention proposes a sensor data processing and transmission decision-making method for complementary perception in vehicle-to-everything (V2X) networks. Utilizing an intelligently guided, reward-free reinforcement learning approach, it offers significant advantages over existing technologies. Traditional cooperative perception systems typically rely on fixed reward mechanisms to guide the processing and transmission of perception data. This approach often fails to adapt flexibly to various situations in dynamically changing V2X environments, leading to low efficiency in the collaboration between perception, communication, and computation. In contrast, this invention, through its intelligently guided, reward-free reinforcement learning method, dynamically adjusts resource allocation and perception tasks based on real-time environmental data and vehicle status, thereby optimizing the collaboration efficiency between communication, perception, and computation. Furthermore, this method not only improves resource utilization efficiency but also significantly reduces data processing and transmission latency. The intelligently guided, reward-free reinforcement learning method of this invention enhances the system's adaptability to complex traffic environments. Moreover, this method can be flexibly adjusted according to different vehicle and road conditions, exhibiting good scalability and applicability to V2X environments of various sizes and types. In summary, the intelligent-guided, reward-free reinforcement learning method proposed in this invention aims to optimize the collaborative perception and communication efficiency of CAVs in the process of multi-source sensor data processing and transmission. This method enables the system to significantly improve the efficiency and accuracy of perception data processing while ensuring real-time response, providing safer and more reliable driving support for vehicle networking.

[0014] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Attached Figure Description

[0015] Figure 1 is a flowchart illustrating a sensor data processing and transmission decision-making method for complementary perception in vehicle networking provided by an embodiment of the present invention;

[0016] Figure 2 is a schematic diagram showing the comparison results of three traditional methods and the method proposed in this invention in terms of perceived data latency.

[0017] Figure 3 is a schematic diagram showing the comparison results of the three traditional methods and the method proposed in this invention in terms of perceptual fusion accuracy. Detailed Implementation

[0018] The present invention will be further described in detail below with reference to specific embodiments, but the implementation of the present invention is not limited thereto.

[0019] Please refer to Figure 1. This embodiment of the invention provides a sensor data processing and transmission decision-making method for complementary perception in vehicle-to-everything (V2X) networks, specifically including the following steps:

[0020] S10. Model a system model that includes CAVs, RAUs, and sensor ensembles.

[0021] The system modeling in this embodiment of the invention includes scene description, perception characteristic modeling, specifically including sensor deployment location and height, perception distance, horizontal viewpoint, vertical viewpoint, obstacle influence, computation and processing modeling, communication modeling, and problem description.

[0022] For example, the research scenario is set in a complex urban traffic environment equipped with advanced system models (RSNs), including vehicle-mounted vehicles (CAVs) as well as ordinary vehicles (referred to as CAVs with Level 0 intelligence), sensors, and roadside assistance units (RAUs). This invention envisions an advanced urban traffic environment equipped with multiple sensors, RAUs, CAVs, and ordinary vehicles. Each road is divided into multiple segments, each segment equipped with one RAU, which are strategically deployed on the road to ensure optimal coverage and functionality. RAUs are used to aggregate perception data from sensors deployed on roadside infrastructure and transmit the processed perception information to CAVs with different perception needs. Assume the group of CAVs served by a specific RAU is {1,2,…,V}, where V represents the number of CAVs in the system model. Similarly, the sensor set is defined as {1,2,…,S}, where S represents the total number of deployed sensors. CAVs have different requirements for perception data, such as detection accuracy and latency. Accuracy level (P... v ) indicates the level of detail required for the v-th CAV, ranging from the original data (P) v =1, representing the highest precision) to abstract information (P) v =0, representing lower precision). Furthermore, CAVs have varying tolerances for latency in perceived data; critical operations require low latency (L...). v →0, while regular tasks can tolerate higher latency L. v .

[0023] Sensors are strategically deployed on or above roadsides to ensure uninterrupted normal traffic flow. The location coordinates of each sensor are represented as follows: Where k represents the deployment location number and l represents the sensor number. The deployment height of the sensor is expressed by the following formula:

[0024] Among them, H k,l This represents the installation height of the l-th sensor deployed at deployment location k above the ground. Due to budget and operational constraints, not all potential locations can be used to deploy sensors. The number of sensors at each location is determined by... The range is from 0 to the maximum number of sensors allowed at that location.

[0025] S20. Based on the system model, construct the detection accuracy objective function and the delay objective function, and construct the overall objective optimization function composed of the detection accuracy objective function and the delay objective function by balancing the weights.

[0026] The present invention constructs a detection accuracy target function, including: constructing a distance perception detection probability function, a horizontal perception detection probability function, a vertical perception detection probability function, and an obstacle perception detection probability function for a single sensor respectively; and constructing the detection accuracy target function together from the distance perception detection probability function, the horizontal perception detection probability function, the vertical perception detection probability function, and the obstacle perception detection probability function.

[0027] More specifically:

[0028] Sensing distance refers to the maximum distance at which a sensor can effectively detect a target. The distance sensing detection probability function of the sensor Rk,l constructed in this embodiment of the invention is expressed by the formula:

[0029] in, Let d represent the distance-aware detection probability function. kl,j Indicates sensor R k,l With the j-th target O j The distance between them, j takes values ​​from 1 to J, where J represents the number of targets. Here, a target refers to the object being detected, such as a car, a person, or a building. R k,l This indicates the l-th sensor deployed at deployment location k. Indicates sensor R k,l The maximum effective distance, Λ represents The sensitivity adjustment parameters, Λ∈[0,10], Γ represents The threshold adjustment parameter, Γ∈[0,100], describes the sensor's detection capability as distance changes, reflecting how the detection probability changes with increasing distance.

[0030] Horizontal field of view refers to the sensor's ability to detect targets in the horizontal direction. The horizontal sensing detection probability function constructed in this embodiment of the invention is expressed as follows:

[0031] in, Let φ represent the horizontal sensing detection probability function. kl,j Indicates sensor R k,l With the j-th target O j The horizontal deviation angle of the sensor axis, Ω represents... The slope adjustment parameter, Ω∈[0,1], Π represents The maximum effective horizontal angle, Π∈[0,100]; this formula reflects the influence of horizontal deviation on the sensor's detection capability, indicating that the detection probability will decrease when the sensor's detection angle deviates from the central axis.

[0032] Vertical perspective refers to the sensor's ability to detect targets in the vertical direction. The vertical sensing detection probability function constructed in this embodiment of the invention is expressed as follows:

[0033] in, Let θ represent the vertical sensing detection probability function. kl,j Indicates sensor R k,l With the j-th target O j The vertical deviation angle of the sensor axis, Δ represents... The slope adjustment parameter, Δ∈[0,1], Φ represents The maximum effective vertical angle, Φ∈[0,100]; this formula shows the influence of vertical deviation on the sensor detection probability, reflecting the change in detection probability when the sensor's detection angle deviates from the central axis.

[0034] Obstacles have a significant impact on the sensor's detection capability. The obstacle perception and detection probability function constructed in this embodiment of the invention is expressed as follows:

[0035] in, Let P represent the obstacle perception and detection probability function. obs G represents the location information of the obstacle. obs (P obs ) represents P obs Obstacle characteristic factor of obstacles, Indicates sensor R k,l The location information is given by α, which represents the rate at which the perception probability changes with distance, α∈[2,4]. This formula quantifies the change in detection probability due to the presence of obstacles, illustrating how obstacles affect the sensor's detection performance.

[0036] Finally, the detection accuracy objective function constructed using formulas (2) to (5) is expressed as follows:

[0037] Furthermore, the system modeling in this embodiment of the invention also includes an edge computing model; the edge computing model includes several edge nodes, each equipped with computing resources; the edge computing model processes data from various sensors by introducing edge nodes, which are responsible for communicating with CAVs and providing perception assistance to CAVs based on currently available sensors. Each edge node is equipped with computing resources for processing sensor data.

[0038] The process of constructing the delay objective function in this embodiment of the invention includes:

[0039] When CAVs send sensor data to RAUs, the RAUs processing capacity is calculated based on computing resources and the size of the sensor data; when RAUs send sensor data to CAVs, the CAVs processing capacity is calculated based on computing resources and the size of the sensor data. The sensor data includes its size, vehicle processing density, and processing decision weights. The RAUs processing time and CAVs processing time are calculated based on the RAUs processing capacity and the sensor data. Based on the characteristics of wireless channel communication between CAVs and RAUs, the uplink data transmission rate from CAVs to RAUs and the downlink data transmission rate from RAUs to CAVs are calculated respectively, and the uplink and downlink transmission times are calculated based on these rates. A delay objective function is constructed from the CAVs processing time, RAUs processing time, uplink transmission time, and downlink transmission time. More specifically:

[0040] The RAUs processing capacity is calculated according to the following formula in this embodiment of the invention:

[0041] in, This represents the RAUs processing capacity when processing the nth sensing data at time t. D represents the available computing resources found by RAUs when CAVs send sensor data to RAUs at time t. n (t) represents the size of the nth sensing data sent by CAVs to RAUs at time t, and N represents the number of sensing data sent by CAVs to RAUs;

[0042] The CAVs processing capacity is calculated according to the following formula in this embodiment of the invention:

[0043] in, This represents the processing capacity of CAVs when processing the n1th sensing data at time t. This indicates that at time t, when RAUs sends sensor data to CAVs, CAVs find available computing resources. N represents the size of the n1th sensing data sent by RAUs to CAVs at time t, and N′ represents the number of sensing data sent by RAUs to CAVs.

[0044] By using formulas (7) and (8), computational resources are allocated among all sensor data, effectively balancing the computational load and ensuring efficient data processing with minimal latency.

[0045] The processing time for RAUs in this embodiment of the invention is calculated using the following formula:

[0046] in, Indicates the processing time of RAUs, a n (t) represents the decision processing weight of RAUs when processing the nth sensing data at time t, a n (t) takes values ​​in the range [0,1], λ n (t) represents the processing complexity of RAUs when processing the nth sensing data at time t. Equation (9) describes the processing latency of data on the RAU side, ensuring efficient task execution and minimizing latency.

[0047] The CAVs processing time is calculated according to the following formula in this embodiment of the invention:

[0048] in, Indicates the processing time of CAVs. This represents the decision processing weights of CAVs when processing the n1th sensing data at time t. The value range is [0,1]. This represents the processing complexity of CAVs when processing the n1th sensing data at time t. Equation (10) describes the processing latency of data on the CAV side, ensuring efficient task execution and minimizing latency.

[0049] The communication model focuses on the interaction between CAVs, RAUs, and sensors within the system model RSNs, handling the efficient transmission and information exchange of sensor data to optimize perception and decision-making. Communication primarily occurs between CAVs and RAUs via wireless channels, including uplink (from CAV to RAU) and downlink (from RAU to CAV) data transmission.

[0050] The uplink transmission time is calculated according to the following formula in this embodiment of the invention:

[0051] Among them, T up Indicates the uplink transmission time. This indicates the uplink bandwidth of the wireless channel. Indicates the uplink transmission power of the wireless channel. This represents the uplink channel gain of the wireless channel. N represents the uplink interference power of the wireless channel, and N0 represents the noise power of the wireless channel.

[0052] The downlink transmission time is calculated according to the following formula in this embodiment of the invention:

[0053] Among them, T dn Indicates downlink transmission time. This indicates the downlink bandwidth of the wireless channel. Indicates the downlink transmission power of the wireless channel. This represents the downlink channel gain of the wireless channel. This indicates the downlink interference power of the wireless channel.

[0054] The parameters in formulas (11) and (12) together determine the efficiency and reliability of data transmission in the Internet of Vehicles. Latency is a key factor in the system model, especially for safety-critical applications in CAVs. The total latency constructed in this embodiment of the invention includes CAV processing time, uplink transmission time, RAU processing time, and downlink transmission time. Finally, the latency objective function constructed through formulas (9) to (12) is expressed as follows:

[0055] Furthermore, the communication model is seamlessly integrated with the perception and computing models, enabling the sensor data transmission strategy to be dynamically adjusted based on data availability, network conditions, and vehicle computing capabilities. The primary objective of the system model in this embodiment is to maximize the perception and detection probability of CAVs while minimizing latency caused by communication and task processing. The final overall objective optimization function is expressed as:

[0056] Where R represents the overall objective function, Θ total To optimize the target regarding detection accuracy, T total To optimize latency-related objectives, including latency handling and communication delay T up +T dn ω represents the balancing weights, and V represents the number of CAVs in the system model. Let Θ represent the perceptual detection probability of the v-th CAV. min T represents the minimum perceptual detection threshold. total T represents the total transmission delay, which is calculated using formula (13). max Indicates the maximum delay threshold. This indicates the uplink bandwidth of the wireless channel. W represents the downlink bandwidth of the wireless channel. total This represents the total bandwidth of the wireless channel. This represents the computational processing power of the i-th RAU, where i ranges from 1 to I, and I represents the number of RAUs in the system model. This indicates the maximum computing power of the RAU. This indicates the number of sensors deployed at deployment location number k. This represents the maximum number of sensors deployed. Among the constraints: the first constraint guarantees that the detection probability of each CAV is not lower than the minimum detection threshold required by the system. min This ensures basic sensing performance requirements are met; the second constraint limits the total latency of each CAV to no more than the maximum allowable latency threshold T of the system. max This ensures data timeliness; the third constraint states that the total bandwidth allocated for uplink and downlink cannot exceed the total bandwidth W of the wireless channel. total The fourth constraint guarantees that each roadside unit RAU i The fifth constraint limits the total number of sensors deployed in the system model to no more than the maximum number of sensors allowed by the system. These constraints together ensure the rational allocation of system resources and the satisfaction of performance requirements.

[0057] This invention integrates multi-source sensor data in a vehicle-to-everything (V2X) environment and optimizes data preprocessing and analysis through edge computing technology to ensure data real-time performance and reliability.

[0058] S30. Construct a set of states including road and traffic conditions, sensor data conditions, vehicle dynamic conditions, and vehicle perception needs; a set of decisions including sensor management decisions, data processing decisions, and communication decisions; and a set of observations including road and traffic observations, sensor data observations, vehicle dynamic observations, vehicle perception needs observations, and delay observations. Based on the set of states, the set of decisions, and the set of observations, use an intelligently guided, reward-free reinforcement learning method to obtain the optimal set of decisions for sensor data transmission, and solve the overall objective function using the optimal set of decisions.

[0059] To address the limitations of traditional methods that rely on explicit reward functions, this invention introduces a reward-free, intelligently guided reinforcement learning method, replacing the traditional approach of defining reward functions based on task type. This reward-free, intelligently guided reinforcement learning method optimizes the processing and transmission decisions of sensory information in vehicle-to-everything (V2X) collaborative perception. By leveraging an understanding of the environment and using intelligence as guidance, this invention dynamically adjusts resource allocation and optimizes data processing flows without relying on traditional reward mechanisms. A detailed description of the reward-free, intelligently guided reinforcement learning method is as follows:

[0060] The environment state in reinforcement learning is a multi-dimensional structure containing various aspects that influence the operational dynamics of CAVs. At any given time t, the environment state is represented as S. t S t include:

[0061] Road and traffic conditions It reflects dynamic and static road conditions, including traffic density, road construction, and weather conditions;

[0062] Sensor data status It compiles a large amount of data from various vehicle-mounted and roadside sensors;

[0063] Vehicle dynamic status It covers key parameters such as vehicle speed, direction and other relevant dynamic states;

[0064] Vehicle perception requirements It can be used to represent the perception needs of CAVs and can also reflect the intelligence level of the vehicle.

[0065] Therefore, the state set S at time t t It can be represented as:

[0066] In the reinforcement learning framework, the set of observations O t It can be represented as:

[0067] in, Represents road and traffic observations. Represents sensor data observations. Represents vehicle dynamic observation values. This represents the observed values ​​of vehicle perception needs. This indicates a delayed observation.

[0068] Agents typically cannot directly access the true state of the environment; instead, they obtain state information through observation. Therefore, this embodiment of the invention considers preference as a unique observation mode, given a set of states S... t Observation set O t The joint prior preference distribution of the reinforcement learning model parameters θ can be represented as p Ψ (S t O t ,θ).

[0069] Decision set (A) t The decision set A represents the system's strategic response at time t, influenced by the current state. t Includes: sensor management decisions Decisions involving sensor selection and tuning; data processing decisions Methods involving sensor data processing; communication decision-making This determines the strategy for information exchange between CAVs and RAUs. Therefore, the decision set A at time t is... t Represented as:

[0070] The core of data transfer decisions lies in assessing whether to process the data locally or transfer it to RAUs for more comprehensive processing. This decision uses a binary variable... Where 1 represents transmission to RAU and 0 represents local processing, as shown in the following formula:

[0071] in, Represents edge nodes (E) i ) load, Represents edge nodes (E) i ) data processing requirements, θ trans This indicates the transmission threshold.

[0072] The dynamics of a system include state transitions and the impact of decisions. State transitions are represented by the following formula: S t =F(S) t A t (19);

[0073] Here, F() represents the CAV environment dynamics that integrate the current state and decisions. In the framework of rewardless reinforcement learning, rewardless guidance is achieved through a set of abstract environmental states and goals, bypassing the need for explicit reward signals. The decision-making algorithm is based on active inference, utilizing these guidelines to make informed choices that align with the overall system goal, i.e., improving the operational safety and efficiency of CAVs in RSNs. Furthermore, as higher levels of cognition develop, a discrepancy exists between the predicted state and the actual state, and this discrepancy changes through the learning process. In this study, embodiments of the invention utilize the fundamental concept of intelligence to assess the change in this discrepancy. Intelligence, as a high-level indicator for quantifying learning effectiveness, draws on the concepts of energy and information; it serves as a comparative measure, assessing the change in information distribution over time due to learning, or the degree of information dispersion relative to the initial state.

[0074] Furthermore, in this embodiment of the invention, based on a state set, a decision set, and an observation set, an intelligently guided rewardless reinforcement learning method is used to obtain the optimal decision set for sensor data transmission, so as to solve the overall objective optimization function through the optimal decision set, including:

[0075] Given a first probability distribution of the current policy prediction state set and a second probability distribution of the decision set prediction state set, the KL divergence between the first and second probability distributions is used as the policy optimization function. The optimal policy is obtained by minimizing the policy optimization function. Given a third probability distribution of the optimal policy prediction state set, observation set, and reinforcement learning model parameters, the joint prior preference distribution of the given state set, observation set, and reinforcement learning model parameters is calculated. The KL divergence between the third probability distribution and the joint prior preference distribution is used as the expected free energy function. The optimization problem of minimizing the expected free energy function is transformed into an optimization problem of maximizing the negative expected free energy function. The negative expected free energy function is decomposed into a first term for capturing expected information gain and a second term for capturing extrinsic value. The optimal policy distribution is obtained by approximating the posterior distribution of the first and second terms respectively. The optimal decision set for sensor data transmission is obtained based on the optimal policy distribution. The overall objective optimization function is then solved using the optimal decision set. More specifically:

[0076] The basic concept of a decision-making framework is to optimize decisions based on predicting the impact of decisions on the future state of the environment. It includes maintaining a probabilistic model of the environment and predicting the outcomes of different decisions. Therefore, in this embodiment of the invention, given a current policy, a first probability distribution predicting the set of states is used; given a set of decisions, a second probability distribution predicting the set of states is used; the KL divergence between the first and second probability distributions is used as the policy optimization function; and the optimal policy is obtained by minimizing this policy optimization function. The formula for obtaining the optimal policy by minimizing the policy optimization function is expressed as:

[0077] Where, π * Let A represent the optimal strategy. t Let S represent the decision set. t Let q(S) represent the set of states, π represent the current policy, and q(S) represent the set of states. t |π) represents the set of predicted states S given the current policy π. t The first probability distribution, P(S t |A t ) represents a given decision set A t Predicted state set S t The second probability distribution, D KL (·||·) denotes the Kullback-Leibler divergence function. This model is based on a continuously updated decision set, ensuring an accurate reflection of the dynamic changes in the CAVs environment: P(S t+1 |S t A t )=f(S t A t ,θ) (21);

[0078] Where f() is a function describing the state transition probability, θ represents the reinforcement learning model parameters, and P(S t+1 |S t A t ) represents a given decision set A t State set S t Predicted state set S t+1 The probability distribution.

[0079] The decision-making process in this invention aims to minimize the desired future free energy, which represents the difference between the predicted and target state distributions. This is achieved by continuously adjusting the optimal policy π. * To minimize the KL divergence and make optimal decisions under uncertainty, thus enhancing adaptability, the policy selection process is conducted via the expected free energy function. The model is based on minimizing , where the expected free energy function is expressed as:

[0080] in, Let S represent the expected free energy function. t Represents a set of states, O t Represents the set of observations, π * Let θ represent the optimal policy, q(S) represent the reinforcement learning model parameters, and q(S) represent the optimal policy. t O t ,θ|π * ) represents a given optimal policy π * Predicted state set S t Observation set O t The third probability distribution of the reinforcement learning model parameters θ, p Ψ (S t O t ,θ) represents the set of states S t Observation set O t The joint prior preference distribution of the reinforcement learning model parameters θ, D KL (·||·) denotes the Kullback-Leibler divergence function.

[0081] The optimal policy distribution q(π) * The desired free energy function can be minimized as follows. formula:

[0082] therefore, This demonstrates that strategies minimizing the expected free energy are more likely to be chosen, promoting a unified framework for exploration and exploitation. σ represents the sigmoid function used to transform input values ​​into probabilities between 0 and 1. The goal of active inference is to minimize the expected free energy. However, this is achieved by transforming it into maximizing the negative expected free energy. We can transform the problem into a standard maximization problem. This aligns better with traditional optimization theory, making the problem easier to handle and understand. The balance between exploration and exploitation will be achieved through… The decomposition into a first term capturing the expected information gain and a second term capturing the extrinsic value is an approximation because it ignores some higher-order and cross-terms. However, in practice, this approximation is often accurate enough and greatly simplifies calculation and interpretation. The negative expected free energy function formula of this invention is expressed as:

[0083] in, Let A represent the negative expected free energy function. t Let S represent the decision set. t Represents a set of states, O t Represents the set of observations, π * Let θ represent the optimal policy, θ represent the parameters of the reinforcement learning model, and q(A) represent the parameters of the reinforcement learning model. t |π * ) represents a given optimal policy π * Predictive decision set A t The probability distribution, Given the optimal policy π * The expected value of the decision set, q(S) t ,θ|O t ,π * ) represents a given optimal policy π * Observation set O t Predicted state set S t The probability distribution of the reinforcement learning model parameters θ, q(S t ,θ|π * ) represents a given optimal policy π * Predicted state set S t The probability distribution of the reinforcement learning model parameters θ, q(S t |π * ) represents a given optimal policy π * Predicted state set S t The probability distribution, Given the optimal policy π * The following pair of states S t The expected value, q(O) t ,θ|S t ,π *) represents a given optimal policy π * State set S t Predicted observation set O t The probability distribution of the reinforcement learning model parameters θ, p Ψ (O t ,θ|π * ) represents a given optimal policy π * Predicted observation set O t The prior preference distribution of the reinforcement learning model parameter θ, D KL (·||·) denotes the Kullback-Leibler divergence function.

[0084] Formula (24) naturally integrates exploratory behavior (through information gain) and the utilization of known strategies (through extrinsic value). This is based on the principle of variational inference. First, [the following is a partial translation of the original text, which is incomplete and requires further context]. This can be considered a variational lower bound, which needs to be maximized. By applying variational inference, this maximization problem can be transformed into finding the optimal approximate posterior distribution q. At each time t, the decision set... This reflects the decisions made regarding sensor management, data processing, and communication strategies. To dynamically optimize the strategy distribution q(π)... * The first term is the expectation of the log-likelihood, and the second term is the entropy of the approximate posterior. By maximizing formula (24), the optimal approximate posterior q-distribution can be obtained:

[0085] Among them, S τ:T O represents the set of all states from time τ to time T. τ:T Let q(S) represent the set of all observations from time τ to time T. τ:T O τ:T ,θ|π * The optimal approximate posterior q-distribution is given by the optimal policy π. * Predicted state set S τ:T Observation set O τ:T The probability distribution of the reinforcement learning model parameters θ, q(O τ |S τ ,θ,π * ) represents a given set of states S τ Optimal strategy π * The reinforcement learning model parameters θ predict the set of observations O. τ The probability distribution, q(S) τ |S τ-1 ,θ,π * ) represents a given set of states S τ-1 Optimal strategy π * The reinforcement learning model parameters θ predict the set of states S. τThe probability distribution, Represents a given set of states S τ Predicted observation set O τ The prior preference distribution, Given the optimal policy π * , reinforcement learning model parameters θ for state set S τ Expectations Represents a given set of states S τ-1 Optimal strategy π * The reinforcement learning model parameters θ predict the set of states S. τ The prior preference distribution, Given the optimal policy π * , reinforcement learning model parameters θ for state set S τ-1 The expectation.

[0086] Formula (25) reflects the dynamic characteristics of the system, where the state and observation at each moment depend on the state at the previous moment, and q(π) * The distribution is transformed into a diagonal Gaussian distribution to simplify calculations, and the Exponential Change of Measure (ECM) algorithm is used to optimize q(π). * ), to make it as close as possible to The goal of the ECM algorithm is to determine the optimal strategy for sensor data transmission based on the current state of the CAVs and RSN environment. Ultimately, the overall objective function of equation (14) can be optimized through the action of the optimal strategy.

[0087] This invention employs an innovative intelligent-guided, reward-free reinforcement learning method. This method does not rely on traditional reward mechanisms but instead guides decision-making through environmental understanding. This approach can dynamically adjust resource allocation and optimize the processing and transmission of perceived information, thereby adapting to highly dynamic and changing traffic environments.

[0088] To verify the effectiveness of the sensor data processing and transmission decision-making method for complementary perception in vehicle-to-everything (V2X) provided in this embodiment of the invention, a comparison with traditional methods demonstrates the superiority of this invention in terms of perception accuracy and latency in multiple real-world driving scenarios. Numerical results clearly show that the proposed method (Our Proposal) significantly improves perception, communication, and computational collaboration efficiency, as well as resource allocation efficiency, compared to three traditional methods (Rainbow DRL, Soft Actor-Critic RL, and Policy Optimization). See Figure 2: The proposed method ensures detection latency under different scenarios. Compared to the other three traditional methods, the proposed method achieves a peak latency of 76.38 milliseconds at medium traffic density, which then slightly decreases at the highest density. This demonstrates that the proposed method can reduce detection latency under high traffic density, exhibiting excellent latency management capabilities. Please refer to Figure 3. As the number of roadside sensors increases, the overall detection accuracy of the proposed method rapidly increases from 63.91% to 94.56%, while the other three traditional methods struggle to achieve the same detection level with the same number of available sensors. Comprehensive simulation results show that the intelligently guided, reward-free reinforcement learning method proposed in this invention performs excellently in optimizing detection accuracy, reducing detection latency, and maintaining packet delivery rate, demonstrating its significant potential in system models (RSNs).

[0089] In summary, the sensor data processing and transmission decision-making method for complementary perception in vehicle-to-everything (V2X) proposed in this invention utilizes an intelligently guided, reward-free reinforcement learning method. Compared with existing technologies, this method has significant advantages: traditional cooperative perception systems typically rely on fixed reward mechanisms to guide the processing and transmission of perception data. This approach often cannot flexibly adapt to various situations in dynamically changing V2X environments, leading to low collaboration efficiency among perception, communication, and computing. In contrast, this invention, through its intelligently guided, reward-free reinforcement learning method, can dynamically adjust resource allocation and perception tasks based on real-time environmental data and vehicle status, thereby optimizing the collaboration efficiency among communication, perception, and computing. Furthermore, this method not only improves resource utilization efficiency but also significantly reduces data processing and transmission latency. The intelligently guided, reward-free reinforcement learning method of this invention enhances the system's adaptability to complex traffic environments. In addition, this method can be flexibly adjusted according to different vehicle and road conditions, exhibiting good scalability and applicability to V2X environments of various sizes and types. In summary, the intelligent-guided, rewardless reinforcement learning method proposed in this invention aims to optimize the collaborative perception and communication efficiency of CAVs in the process of multi-source sensor data processing and transmission. This method enables the system to significantly improve the efficiency and accuracy of perception data processing while ensuring real-time response, providing safer and more reliable driving support for vehicle networking.

[0090] In the description of this invention, it should be understood that the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this invention, "a plurality of" means two or more, unless otherwise explicitly specified.

[0091] Although the invention has been described herein in conjunction with various embodiments, those skilled in the art, by reviewing the specification and accompanying drawings, will understand and implement other variations of the disclosed embodiments in carrying out the claimed invention. In the specification, the word "comprising" does not exclude other components or steps, and "a" or "an" does not exclude a plurality. While certain measures are described in different embodiments, this does not mean that these measures cannot be combined to produce good results.

[0092] The above description, in conjunction with specific preferred embodiments, provides a further detailed explanation of the present invention. It should not be construed that the specific implementation of the present invention is limited to these descriptions. For those skilled in the art, various simple deductions or substitutions can be made without departing from the concept of the present invention, and all such modifications and substitutions should be considered within the scope of protection of the present invention.