Multi-source sensing fusion and dynamic resource allocation method based on vehicle infrastructure cooperation
By building a vehicle-road collaborative network architecture and a multi-source perception fusion model, the problem of complementarity optimization of multi-source heterogeneous data and insufficient adaptability to dynamic scenarios is solved, efficient perception fusion and resource allocation are achieved, and the safety and efficiency of autonomous driving are improved.
Patent Information
- Application Number
- CN202510490229.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-18
- Publication Date
- 2025-07-22
AI Technical Summary
In the existing Internet of Vehicles technology, the complementarity optimization of multi-source heterogeneous data and insufficient adaptability to dynamic scenarios lead to difficulties in space-time alignment of heterogeneous data, large fusion errors, real-time conflicts with computing resources, and overload load of on-board computing platforms, high response delays, affecting the safety of autonomous driving.
A multi-layer network architecture with vehicle-road collaboration is built, and a collaborative whole-domain perception model, multi-source perception fusion model and resource allocation algorithm are adopted. Through cross-modal spatiotemporal alignment and confidence weighting, combined with improved discrete squid group optimization algorithm and hybrid action space depth deterministic strategy gradient algorithm, the perceived fusion quality and resource allocation are optimized.
It improves the accuracy of occlusion detection, reduces perception error and delay, improves perception fusion efficiency, meets the safety requirements of autonomous driving, and reduces the risk of computing resource consumption and data leakage.
Smart Images

Figure CN120358519A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of mobile communication, and relates to a multi-source perception fusion and dynamic resource allocation method based on vehicle-road cooperation. Background Art
[0002] With the rapid development of artificial intelligence and multi-modal perception fusion technology, the vehicle-road cooperative intelligent transportation system (ITS) integrates advanced communication technologies, promoting the evolution of traditional Internet of Vehicles (IoVs) towards intelligent environmental perception, and significantly improving driving safety and urban traffic management efficiency. As the core carrier for realizing road-side safe driving, CAV constructs a global environmental data capture network covering "human-vehicle-road-cloud" through a multi-source heterogeneous sensor collaborative perception mechanism, providing technical support for real-time traffic situation analysis and decision-making control. However, the public's trust in autonomous driving technology still faces severe challenges.
[0003] As the core pillar of ITS, the core contradiction of environmental perception technology lies in the lack of complementary optimization of multi-source heterogeneous data and dynamic scenario adaptation ability. Current technologies mainly rely on the heterogeneous fusion of lidar (LiDAR), vision sensors, and millimeter-wave radars. Although they can construct a high-precision three-dimensional environmental model, there are still the following bottlenecks: Firstly, it is the problem of spatio-temporal alignment of heterogeneous data. Differences in sampling frequencies, coordinate systems, and accuracies of different sensors lead to fusion errors. Especially in harsh environments such as rain, snow, haze, etc., point cloud data distortion and image blurring will cause cross-modal feature mismatches. Secondly, there is a conflict between real-time performance and computing resources. Although deep neural networks (such as 3D object detection models) can improve perception accuracy, the number of model parameters is huge (for example, the Tesla FSD system needs to process 230 billion operations per second), resulting in overloading of in-vehicle computing platforms, causing data synchronization delays and decision-making lags. Research shows that the average response delay of existing CAVs in sudden obstacle scenarios is as high as 320 ms, far exceeding the safety threshold (≤100 ms), directly causing a 13% increase in the traffic accident rate.
[0004] Based on the above problems, the present invention designs a multi-source perception fusion and dynamic resource allocation method based on vehicle-road cooperation. First, a two-layer vehicle-road network architecture based on perception fusion is constructed, and a cooperative global perception model is proposed based on this network architecture to provide real-time perception information for CAVs. Second, a multi-source perception fusion model is proposed, which describes the perception fusion of multi-source occluders to improve the accuracy and robustness of perception. Then, based on the COP model and the multi-source perception fusion model, a perception fusion optimization algorithm based on IDSS is proposed to maximize the perception fusion quality of the vehicle ROI. Finally, a resource allocation algorithm based on HDDPG is proposed to optimize the ROI fusion resource allocation problem and minimize the ROI perception fusion delay. Summary of the Invention
[0005] In view of this, the object of the present invention is to provide a multi-source perception fusion and dynamic resource allocation method based on vehicle-road cooperation.
[0006] To achieve the above object, the present invention provides the following technical solutions:
[0007] In the first aspect, according to complex traffic scenarios, embodiments of the present invention provide an environmental perception optimization and edge computing resource scheduling method to improve the perception fusion quality and reduce the processing delay. The method includes the following steps:
[0008] S1: Propose a two-layer vehicle-road network architecture based on perception fusion;
[0009] S2: Construct a cooperative global perception model;
[0010] S3: Construct a multi-source perception fusion model;
[0011] S4: Model the multi-source perception fusion and resource allocation optimization problem;
[0012] S5: Design a perception fusion optimization algorithm based on IDSS;
[0013] S6: Design a resource allocation algorithm based on HDDPG.
[0014] In the second aspect, in S1 of the embodiments of the present invention, a two-layer vehicle-road network architecture based on perception fusion is constructed, which consists of a CAV layer and an RSU layer. Among them, the CAV layer integrates multi-modal sensors and lightweight computing units, performs local perception data acquisition and feature compression, and realizes data sharing through the V2V / V2I communication protocol. A high-performance edge server is deployed at the RSU, supporting dynamic resource allocation and cooperative computing task offloading, and realizing low-latency data backhaul through the Orthogonal Frequency Division Multiplexing (OFDM) communication mechanism.
[0015] In a third aspect, in S2 of the embodiments of the present invention, a Collaborative Omni-Perception (COP) model is established. Through the dual-role dynamic switching mechanism of CAVs, distributed collaboration for environmental perception is achieved with the assistance of RSUs. At the D-CAV, when an occluder in the ROI is detected, a collaborative perception requirement instruction is first generated and broadcast to nearby CAVs, and then the perception information received from multiple R-CAVs is fused with the assistance of the RSU to reconstruct the global ROI. At the R-CAV, when the collaborative perception requirement sent by the D-CAV is received, the original perception information needs to be perceptually compressed first. Then, to ensure low-latency and low-power consumption transmission, the R-CAV needs to perform ROI extraction on the compression result, and finally send the ROI extraction information to the D-CAV that initiated the perception requirement.
[0016] In a fourth aspect, in S3 of the embodiments of the present invention, a multi-source perception fusion model is established to quantify the fusion quality on the basis of the COP model through cross-modal spatio-temporal alignment and confidence weighting. First, the D-CAV will perform operations such as normalization of heterogeneous data, timestamp synchronization, coordinate system conversion, and three-dimensional reconstruction, so as to obtain unified description information of the occluder. Secondly, the present invention defines the description of the ROI perception fusion quality as a quantization index based on multi-dimensional joint evaluation, including observation distance, computing resources, and perception confidence. Finally, the D-CAV will select a description for each occluder from the perception fusion information for fusion.
[0017] In a fifth aspect, in S4 of the embodiments of the present invention, based on the COP model and considering multi-source perception fusion, the multi-source perception fusion and resource allocation optimization problem is modeled. In order to obtain the best real-time ROI perception fusion quality under the COP model, the modeled optimization problem is divided into two sub-optimization problems: the ROI perception fusion quality optimization problem and the ROI fusion resource allocation optimization problem.
[0018] In a sixth aspect, in S5 of the embodiments of the present invention, a perception fusion optimization algorithm based on IDSS is designed. Through the chain collaborative search mechanism and the dynamic parameter adjustment strategy, the global optimal solution that maximizes the ROI perception fusion quality is efficiently solved in the discrete solution space. First, the population is initialized to define a discrete search space. Secondly, the position of the leader (Leader Salp, LS) is dynamically updated according to the food position, and the position update of the followers (Follow Salps, FSs) is constrained by the LS and the individuals before and after. Then, the fusion quality dynamically adjusts the new position and retains the optimal result. Finally, the occluder description scheme that meets the maximum ROI fusion quality is output.
[0019] In a seventh aspect, in step S6 of the embodiments of the present invention, a resource allocation algorithm based on HDDPG is designed to solve the non-convex characteristic of the ROI fusion resource allocation optimization problem and overcome the defect that traditional gradient descent methods are prone to falling into local optima. First, the ROI fusion resource allocation optimization problem is modeled as a Parameterized ActionSpace Markov Decision Process (PAMDPs), and then, by combining the Actor-Critic algorithm with a mixed action space and the Deep Q-Network (DQN) algorithm, an HDDPG algorithm is proposed to optimize the ROI fusion resource allocation problem.
[0020] The beneficial effects of the present invention are as follows:
[0021] (1) By constructing a vehicle-road double-layer collaborative sensing network architecture (CAV layer and RSU layer), a cross-modal spatio-temporal alignment and confidence-weighted fusion mechanism for multi-source sensor data is realized. The dynamic region detection algorithm is used to screen key ROI information, and combined with three-dimensional reconstruction and Kalman filtering technologies, the time drift and spatial deviation of heterogeneous data are effectively eliminated. Experiments show that this method improves the accuracy by 15%-20% in the occluder detection scenario and significantly reduces the sensing error caused by data mismatch.
[0022] (2) An improved discrete salp swarm optimization algorithm (IDSS) is designed, and the convergence factor and random perturbation parameters are dynamically adjusted through a chain collaborative search strategy. Compared with the traditional genetic algorithm, the global search efficiency of this algorithm in the discrete solution space is increased by 30%, and at the same time, the risk of local optima is reduced, and the computational resource consumption when the ROI fusion quality is maximized is reduced by 22%.
[0023] (3) A hybrid action space deep deterministic policy gradient algorithm (HDDPG) is developed, and the resource allocation is modeled as a parameterized Markov decision process. Through the joint optimization of the discrete offloading decision, continuous offloading rate, and RSU resource allocation by a three-branch Actor network, edge-cloud collaborative computing is realized in complex traffic scenarios. Measured data shows that the average system response delay is reduced from 320 ms to 85 ms, meeting the 100 ms safety threshold requirement.
[0024] (4) The proposed collaborative omnidirectional sensing model (COP) supports the dynamic role switching of CAVs between demand vehicles (D-CAVs) and response vehicles (R-CAVs) and is compatible with V2V / V2I multi-mode communication protocols. This architecture can flexibly adapt to different scales of roadside unit deployments and support a single RSU node to process the collaborative requests of 50+ CAVs simultaneously while maintaining a 95% sensing accuracy.
[0025] (5) Ensure the anti-interference ability of V2X communication through the OFDM (Orthogonal Frequency Division Multiplexing) mechanism and the encrypted transmission protocol. Combined with the edge computing task offloading strategy, the sensitive data processing is completed entirely within the secure boundary of the RSU, reducing the data leakage risk by 60% compared to the pure cloud solution, meeting the requirements of the third-level information security protection for the vehicle networking.
[0026] Other advantages, objectives, and features of the present invention will be described to some extent in the subsequent specification, and to some extent, will be obvious to those skilled in the art based on the study of the following text, or can be taught from the practice of the present invention. The objectives and other advantages of the present invention can be achieved and obtained through the following specification. Brief Description of the Drawings
[0027] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be described in detail preferably with reference to the accompanying drawings, where:
[0028] Figure 1 is a double-layer network architecture diagram based on the COP model;
[0029] Figure 2 is a training flowchart of the resource allocation algorithm based on HDDPG;
[0030] Figure 3 is an execution flowchart of the multi-source perception fusion and dynamic resource allocation method based on vehicle-road cooperation. Detailed Embodiments
[0031] The following illustrates the embodiments of the present invention through specific examples. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through different specific embodiments. The details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the diagrams provided in the following embodiments only illustrate the basic concept of the present invention schematically. Without conflict, the following embodiments and the features in the embodiments can be combined with each other.
[0032] Among them, the drawings are only for illustrative purposes, showing only schematic diagrams, not physical diagrams, and should not be construed as a limitation to the present invention; in order to better illustrate the embodiments of the present invention, some components in the drawings will be omitted, enlarged, or reduced, which do not represent the dimensions of the actual product; for those skilled in the art, it is understandable that some well-known structures and their descriptions in the drawings may be omitted.
[0033] In the accompanying drawings of the embodiments of the present invention, the same or similar reference numerals correspond to the same or similar components; in the description of the present invention, it should be understood that if there are terms such as "upper", "lower", "left", "right", "front", "rear", etc. indicating the orientation or positional relationship, they are based on the orientation or positional relationship shown in the accompanying drawings. This is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, the terms describing the positional relationship in the accompanying drawings are only for illustrative purposes and should not be construed as a limitation of the present invention. For those of ordinary skill in the art, the specific meanings of the above terms can be understood according to specific circumstances.
[0034] Figure 1 Fig. shows a possible structural schematic diagram of the communication system involved in the embodiments of the present invention. As Figure 1 shown, this network architecture considers a two-layer network, which consists of N CAVs and R RSUs, and their sets are respectively and where each CAV is equipped with computing resources Let and represent the sets of D-CAV and R-CAV respectively. There are two roles among CAVs: R-CAV and D-CAV. In the COP framework, a CAV can be both a D-CAV and an R-CAV. Among them, an R-CAV will go through three steps: perception compression, ROI extraction, and information transmission; a D-CAV also goes through three steps: demand broadcast, information reception, and ROI fusion to reconstruct a high-quality global perception scenario. In addition, in order to reduce the CAV computing load and achieve efficient perception fusion, the present invention uses an RSU to assist the CAV to complete the processing of ROI fusion.
[0035] 1. Vehicle-road two-layer network architecture based on perception fusion
[0036] The present invention considers a vehicle-road two-layer network architecture based on perception fusion, including CAVs and RSUs.
[0037] The transmission and exchange of perception information between CAVs and between CAVs and RSUs can be supported by communication technologies such as V2V and V2I. The present invention only considers the uplink offloading delay and ignores the downlink result feedback delay of V2V and V2I. The communication link in the system adopts the OFDM scheme, and there is no channel interference in the transmission of perception information between vehicles. Therefore, at time slot t, when CAV i connects to RSU m through the V2I link for information transmission, the uplink transmission rate is:
[0038]
[0039] where, W RtLet \(B\) be the transmission channel bandwidth. Rt Define \(W\) as the bandwidth of the V2I link. Therefore, when multiple CAVs unload fusion information to the edge RSU simultaneously, the maximum number of vehicles that the RSU can connect to is: \(K = B\) Rt / W Rt ; Let \(P_i\) be the transmission power of CAV \(i\); Let \(h_{V2I}\) be the V2I channel gain; Let \(l_{i,m}\) denote the path loss from CAV \(i\) to RSU \(m\); \(\sigma^2\) 2 denotes the power of Gaussian channel noise.
[0040] According to the Shannon formula, the transmission rate of the V2V link for R-CAV \(i\) to transmit sensing information to D-CAV \(j\) at time slot \(t\) is:
[0041]
[0042] where, Let \(h_{V2V}\) be the V2V channel gain; Let \(l_{i,j}\) denote the path loss between R-CAV \(i\) and D-CAV \(j\).
[0043] 2. Cooperative Global Sensing Model
[0044] When a CAV is driving on the road, it can send and receive sensing information at any time and play a dual role according to the COP model, and the sensing information can be processed in parallel.
[0045] A. Responding Vehicle Role: In this scenario, the CAV, as an R-CAV, will perform the following three steps:
[0046] 1) Sensing Compression. The amount of original sensing information is large, and directly transmitting the original sensing information will result in too large a transmission delay. Therefore, the sensing compression of R-CAV \(i\) transforms the original sensing information into low-dimensional and high-semantic sensing information through feature extraction and coding techniques, while retaining the key features of cooperative sensing and significantly reducing the redundancy of sensing information and communication overhead. Therefore, the execution delay of sensing compression for R-CAV \(i\) is:
[0047]
[0048] where \(L\) represents the size of the original sensing information; Let \(C_i\) denote the computational intensity of R-CAV \(i\) for performing sensing compression; Let \(r_i\) denote the computing power ratio allocated by R-CAV \(i\) to sensing compression.
[0049] 2) ROI extraction. To ensure that the system transmits ROI-related perception information to D-CAVj with low latency and low energy consumption, the perception compression result of R-CAV i needs to be further extracted. The ROI extraction of R-CAV i is performed through a dynamic region detection algorithm (such as YOLOv7) to filter out the ROI perception information that is strongly related to the driving of D-CAV. Therefore, the execution latency of ROI extraction of R-CAV i for D-CAVj is:
[0050]
[0051] where ζ i→j represents the ROI extraction ratio of R-CAV i relative to D-CAVj; is the computational intensity required for R-CAV i to perform ROI extraction; represents the computing power ratio allocated by R-CAV i for ROI extraction.
[0052] 3) Information transmission. In data transmission, R-CAV packs the ROI extraction result and sends it to D-CAV. The transmitted data will include the perception information of the ROI occluder, that is, the transmission latency of the perception information of R-CAV i to D-CAVj is:
[0053]
[0054] Finally, R-CAV i completes collaborative perception through perception compression, ROI extraction, and information transmission, thereby determining the latency for D-CAVj to receive the perception information transmitted by R-CAV i, which can be expressed as:
[0055]
[0056] B. Role of the required vehicle: In this scenario, CAV, as D-CAV, will perform the following three steps:
[0057] 1) Required broadcast. When D-CAV i perceives that there is an occluder in the ROI, it will generate a collaborative perception requirement instruction, then encapsulate the perception requirement of the occluder, and broadcast it to the nearby CAVs and RSUs. The perception requirement includes the requirement instruction, ROI range, and location.
[0058] 2) Information reception. D-CAV fuses all the received perception information to reconstruct the global ROI. The latency for D-CAV i to receive all the perception information is:
[0059]
[0060] where τ is the tolerable data reception latency of D-CAV, and the perception data packets that time out will be discarded.
[0061] 3) ROI Fusion. After the D-CAV receives the perception information from all R-CAVs, it first performs multi-source data unpacking and spatio-temporal alignment to eliminate the time drift and spatial deviation of heterogeneous perception fusion, and constructs a consistent spatio-temporal benchmark for subsequent multi-source perception fusion. Then, based on the perception information of the occluder and the perception information of the D-CAV itself, an optimal description is selected for ROI fusion to reconstruct a high-quality global perception scene. The above ROI fusion process requires high computing resources for CAVs. However, the current computing resources of CAVs cannot support the complete local processing of ROI fusion. Therefore, to obtain high-quality ROI perception fusion, some perception fusion information needs to be offloaded to the RSU to assist the CAV in completing ROI fusion. Therefore, ROI fusion includes local and edge computing, and the total delay of its ROI fusion can be expressed as:
[0062]
[0063] Among them, is the transmission delay from CAV i to RSU m, which can be obtained by the following formula:
[0064]
[0065] where α i,m is the offloading ratio, indicating the ratio of the data offloaded from D-CAV i to RSU m.
[0066] and respectively represent the local computing and edge computing of ROI fusion, which are represented by Equation (9) and Equation (10) respectively:
[0067]
[0068]
[0069] where M f represents the computing intensity of ROI fusion.
[0070] So far, D-CAV i makes a reasonable offloading decision, and the total delay for CAV i to obtain the global ROI is:
[0071] T i t = T i rt + T i ft (11)
[0072] 3. Multi-source Perception Fusion Model
[0073] In the COP model, the D-CAV perception fusion information Including its own perception information P i and R-CAV perception information P j→i , the perception fusion information can be expressed as:
[0074]
[0075] Define O u as the relevant status information of the occluder u in the perception fusion information . O u includes: the occluder type ty u (vehicle, pedestrian, non-motor vehicle driver, and road construction area), the position information loc of the occluder relative to the CAV coordinate system u , the perception distance d u , the required computing resources rs u and the three-dimensional detection bounding box feature box u and other isomorphic information, which can be expressed as:
[0076] O u ={ty u ,loc u ,d u ,rs u ,box u} (13)
[0077] Among them, loc u =(x u ,y u ) is the two-dimensional coordinate information of the occluder u; box u =(l u ,h u ,w u ) is the three-dimensional detection bounding box of the occluder u. The CAV obtains the type and position information of the occluder u through ty u ,loc u , and evaluates the quality of the occluder u according to d u ,rs u ,box u . In order to more clearly quantify the effect of ROI fusion, the description of the ROI perception fusion quality is defined as a quantitative index based on multi-dimensional joint evaluation, including the allocation of computing resources, the observation distance, the geometric accuracy, and the perception confidence, that is, O u The description is defined as:
[0078]
[0079] Among them, ω1, ω2, and ω3 represent the weights of the observation distance, computing resources, and perception confidence, respectively, and ω1 + ω2 + ω3 = 1. Different weights directly affect the ROI perception fusion quality of the system by adjusting the priorities of the observation distance, computing resources, and perception confidence. s u is the perception confidence, which is obtained from pc u where the CAV assigns confidence scores for different occluders ty u . The description of O u is affected by the observation distance, fusion computing resources, perception confidence, and their corresponding weights. Specifically, the closer the observation distance, the more information the CAV perceives, and the better the quality of O u ; the higher the computing resources required for the CAV to process the occluder, the more information it perceives, and the better the quality of O u ; the higher the confidence score and the larger the volume of the occluder u, the greater the impact on the safety of the system, and the better the quality of O u .
[0080] Considering the perception fusion information from multiple R-CAVs, there may be multiple repeated descriptions of the same occluder. Therefore, the D-CAV i needs to classify the perception fusion information u according to ty u , loc and select one description for each occluder u for fusion. Assuming contains U occluders, the formula (12) can be rewritten as where each O u has K descriptions, that is Finally, the D-CAV i selects one description from each O u to achieve perception fusion, thus presenting the global ROI.
[0081] 4. Modeling the Multi-Source Perception Fusion and Resource Allocation Problem
[0082] Based on the COP model, considering multi-source perception fusion, in order to improve the accuracy and robustness of perception fusion and minimize the system perception fusion delay, the multi-source perception fusion and resource allocation optimization problem is modeled as:
[0083]
[0084] Among them, \(T\) is the total number of time slots, and the constraint conditions \(C1\), \(C2\), \(C3\) and \(C4\) are restrictions on decision parameters; the constraint condition \(C4\) is the impact of CAV computing resources on the quality of ROI perception fusion; the constraint condition \(C5\) is the restriction on RSU computing resources, which needs to ensure that the sum of the computing resources allocated by the RSU to all associated CAVs does not exceed its maximum computing resources; the constraint condition \(C6\) is to ensure that the ROI perception information of the R-CAV can arrive on time at \(\tau\).
[0085] To obtain the best and real-time ROI perception fusion quality under the COP model, it is necessary to maximize the ROI perception fusion quality and also minimize the ROI perception fusion delay. Therefore, the optimization problem \(P1\) is divided into two sub-optimization problems \(P2\) and \(P3\).
[0086] 1) Sub-problem one: The ROI perception fusion quality optimization problem \(P2\) can be expressed as:
[0087]
[0088] where \(\Omega\) i represents the optimal solution of the ROI perception fusion quality of CAV \(i\), which includes the best description solution for each type of occluder.
[0089] 2) Sub-problem two: The ROI fusion resource allocation optimization problem \(P3\) can be expressed as:
[0090]
[0091] 5. Design of the perception fusion optimization algorithm based on IDSS
[0092] Since the optimization problem \(P2\) is an NP-hard problem, existing solutions include exact techniques and genetics, etc. The calculation time of exact techniques increases significantly with the increase in the number of occluders. Although high ROI perception fusion quality is obtained, it faces the problem of high delay. To solve this problem, the present invention proposes a perception fusion optimization algorithm based on IDSS, aiming to obtain high ROI perception fusion quality and prevent falling into local optimal solutions.
[0093] In the proposed IDSS algorithm, all salps are composed of a leader \(LS\) and multiple followers \(FSs\). The \(LS\) is the first salp on the chain, guiding the \(FSs\) to search for food in a chain-like behavior. All salps are connected end to end by individuals to form a "chain" and move following one another in sequence. The \(LS\) moves towards the food and guides the movement of the \(FSs\). The movement of the \(FSs\) follows a strict "hierarchical" system, that is, it is only affected by the previous salp. During the movement process, the \(LS\) conducts global exploration, while the \(FSs\) conduct full local exploration, greatly reducing the situation of falling into local optima. During the chain-like movement and foraging process of all salps, the position update of the \(LS\) is:
[0094]
[0095] Among them, represents the position of LS in the u-th dimension; F u represents the position of food, which usually represents the optimal solution of the current situation; ub u and lb u are the corresponding upper and lower bounds, which are K (the number of description types of objects) and 1 respectively; is the convergence factor, which plays a role in balancing global exploration and local exploitation, and gradually decreases during the iteration process. z and Z are the current iteration and the maximum number of iterations respectively. c2 and c3 are random variables drawn from the interval [0, 1]. The above formula shows that the position update of LS is only related to the position of food. When c3 ≥ 0.5, the new position of LS moves in the direction of the food position F u , increasing the random factor; when c3 < 0.5, the new position of LS deviates from the direction of the food position F u , increasing the exploration factor.
[0096] The position update of FSs is affected by LS and the individuals before and after. Therefore, the position update of FSs is:
[0097]
[0098] Among them, and are the positions of the updated FSs and the FSs before update in the u-th dimension respectively.
[0099] Therefore, the overall process of the perception fusion optimization algorithm based on IDSS is as follows: First, initialize the population, define the search space as the Euclidean space of H × U, where H is the space dimension and U is the population size. The positions of all salps in the space are stored in Among them represents the description number of the selected occluder u in plan h. For each plan solution, the discrete variable is Each row in X represents a solution Ω i . After Z iterations, the row with the maximum sum of quality is selected as the optimal solution
[0100] 6. Design of Resource Allocation Algorithm Based on HDDPG
[0101] Since the objective function and constraints of the ROI fusion resource allocation optimization problem P3 are non-convex, traditional gradient descent algorithms are trapped in local optima and it is difficult to guarantee the global optimality of the solution. Therefore, the present invention introduces a reinforcement learning (RL) method based on a hybrid action space to solve the ROI fusion resource allocation optimization problem.
[0102] Regarding the high-dimensional complexity of the hybrid action space, the present invention simplifies it into a hierarchically parameter-driven hybrid action space, which can be expressed as Specifically, the hierarchically parameter-driven hybrid action space needs to select a higher-level discrete action k from the action set and then select a continuous action x from the continuous action set associated with the discrete action k to execute together with the action k, so as to form the hybrid action space of the deep reinforcement learning (DRL) algorithm. Taking the hierarchically parameter-driven hybrid action space of the CAV in the ROI fusion resource allocation optimization problem as an example: at time slot t, CAV i needs to select a discrete action from the discrete action space to decide to offload the ROI perception fusion information to RSU m. Subsequently, the RSU assigns continuous actions and for the perception fusion information of CAV i, which respectively represent the offloading rate of CAV i and the computing resources allocated by RSU m for CAV i, and together constitute the discrete-continuous hybrid action space and will be used in the subsequent DRL algorithm.
[0103] Based on the COP model, the ROI fusion resource allocation optimization problem is modeled as PAMDPs, that is where S represents the state space, A represents the action space, R is the reward value, and π is the policy. This process mainly includes three parts:
[0104] 1) State space: At time slot t, the computing resources f of CAV i t , the computing resources of RSU and the perception fusion information P of CAV i i t . Therefore, the state is defined as the state s at time slot t t ∈S.
[0105] 2) Action space: At time slot t, the action space includes the offloading decision (determine that CAV i will offload to the server at time slot t ) Unloading rate Computing resources allocated by RSU m to CAV i
[0106] 3) Reward function: designed as:
[0107]
[0108] Therefore, through the DRL framework, the RSU aims to maximize the long-term cumulative reward:
[0109]
[0110] Using the discount factor γ ∈ [0, 1] to balance short-term and long-term benefits. The policy π θ Learns the state-to-action mapping through a trial-and-error mechanism and finally obtains the optimal policy:
[0111]
[0112] Based on the above Markov decision process, the solution process of the resource allocation algorithm based on HDDPG proposed by the present invention is as Figure 2 shown. As Figure 2 shown, the HDDPG algorithm is equipped with an Actor network with a three-branch structure and a global Critic network, that is, the Actor network contains three network branches, mainly two different types of networks: a discrete policy network and a continuous policy network. The discrete policy network learns the policy to select the discrete action m, and the continuous policy network learns the policy to select the continuous action of resource allocation, including the unloading rate α and the RSU computing resource f. The discrete policy network outputs i values M = {m1,..., m i}, and randomly samples through the sigmod(m) distribution to obtain the discrete action m to be taken. The continuous policy network generates the policy of the continuous action by calculating the mean and variance of each parameter Gaussian distribution Therefore, the discrete action m selected by the discrete policy network and the continuous actions (α and f) selected by the continuous policy network together constitute the hybrid action
[0113] The calculation formulas for the discrete action m and the continuous actions (α, f) are as follows:
[0114]
[0115]
[0116] where S represents the state space; θ d and θ crespectively represent the discrete policy network parameters and the continuous policy network parameters, jointly constituting the current network and the target network θ of the Actor network π and θ π' .
[0117] Therefore, the HDDPG algorithm proposed in the present invention can be regarded as a combination of the Actor-Critic algorithm based on the hybrid action space and the DQN algorithm. Among them, the Actor network updates the policy network parameters in the gradient direction of the action value function; the Critic network uses a differentiable approximation function to approximate the action value function. Therefore, the action value function Q(·) is defined to approximate the long-term cumulative reward Use Q(s t , a t ) to evaluate and improve the optimal policy of the Actor network. The action value function can be expressed as:
[0118]
[0119] From this, it can be obtained that the action value function of the Actor current network can be given by the Bellman equation:
[0120]
[0121] Among them, R t (s t , a t ) represents the instantaneous reward obtained by the agent when taking the action a t under the given state s t ; represents the Q-value estimation for the next moment obtained based on the policy π.
[0122] At the same time, the Critic current network in the HDDPG algorithm is the network in the DQN model, and the Q-value is an evaluation of the quality of the policy of the Actor current network. Assuming that the policy of the Actor network is deterministic, when the Actor network generates an action, the Q-value function of the Critic network can be derived as:
[0123]
[0124] Similar to Q-learning and DQN, the Critic network in the HDDPG algorithm uses the TD error to update the Q-value so that the RSU can search for the optimal policy from the Q-table. Finally, Q(s t , a t ) is updated by the following formula:
[0125] Q q (s t+1 , a t+1 ) = Q q (s t , at ) + ηδ t (28)
[0126] Among them, η represents the learning rate; δ t is the TD error and can be expressed as:
[0127]
[0128] Among them, γ represents the discount factor; Q(s', a') represents the historical Q value. Theoretically, it can be iterated through Equation (28) until |Q(s', a') - Q(s, a)| < ξ to obtain an approximately optimal Q value, where ξ is a very small positive number.
[0129] To stabilize the training process, the HDDPG algorithm requires the target network to calculate the final loss function. Therefore, based on the loss function of DQN, the loss in the current state is:
[0130]
[0131] Among them, y t represents the Q value obtained by the current Critic network and can be expressed as:
[0132] y t = R t (s t , a t ) + γQ(s t+1 , π(s t+1 )|θ q ) (31)
[0133] Among them, π(·) and Q(·) represent the network parameters of the current Actor network and the current Critic network respectively.
[0134] According to the loss function of the target network Loss(θ q ) to update the current Critic network parameter θ q . Finally, using the gradient ascent algorithm, the current Actor network updates the current Actor network parameter θ π according to the policy gradient, and the loss gradient of its network is:
[0135]
[0136] To reduce the impact of data correlation on the training accuracy of the neural network, the RSU collects the environmental information of the IoVs scenario and stores it in the experience replay pool, where each piece of data can be expressed as (s t , a t , R t , s t+1) When the experience pool is full of data, the earliest stored experience needs to be replaced with new experience to ensure that the latest samples are always retained in the pool. Each time training is performed, a batch of historical samples will be randomly drawn from the experience replay pool for update. At the same time, the introduced target network mechanism copies the parameters of the current Actor network and the current Critic network to the target Actor network and the target Critic network, and uses the target network to obtain y t , that is, y t can be obtained from the target Actor network θ π' and the target Critic network θ q' . Therefore, the y t in Equation (31) can be changed to:
[0137] y t = R t (s t , a t ) + γQ'(s t+1 , π'(s t+1 )|θ q' ) (33)
[0138] Finally, the HDDPG algorithm updates the current Actor network and the current Critic parameters through the policy gradient algorithm. Then, the parameters of the target network are updated using the following formula:
[0139] θ π' ← τθ π + (1 - τ)θ π' (34)
[0140] θ q' ← τθ q + (1 - τ)θ q' (35)
[0141] Among them, τ is the soft update parameter, and its main purpose is to weight-average the current network parameters and the new target network parameters and assign them to the target network. Due to the use of soft update in the above network parameter update method, the network output is more stable, further improving the stability of the learning process of the HDDPG algorithm.
[0142] Since the deterministic policy may limit the interaction between the agent and the environment, resulting in a limited exploration range for the agent and making it difficult to try enough actions to obtain effective learning signals, thus affecting the full exploration of the environment. Therefore, random noise is introduced to enable the agent to effectively explore in the mixed action space, that is, a random noise is added to the deterministic policy π(s t |θ π ) The policy of the HDDPG algorithm is as follows:
[0143]
[0144] In summary, the resource allocation algorithm based on HDDPG is as follows: First, randomly initialize the Actor-Critic network and the target network, and clear the experience pool; then select actions through the policy and noise, and store the state-action-reward data after execution; finally, sample the experience to batch update the Critic and the Actor, and softly update the target network parameters. Ultimately, achieve dynamic optimization of resource allocation through hybrid actions, minimize the system delay, and improve the reward value.
[0145] 7. System flow diagram
[0146] Figure 3 The following is the execution flow chart of the multi-source perception fusion and dynamic resource allocation method based on vehicle-road cooperation in a complex traffic environment, and the specific steps are as follows:
[0147] S701 - S703: System initialization. The CAV starts the multi-modal sensor and calibrates the parameters, initializes the local computing unit, and loads the lightweight perception model (such as YOLOv7). The RSU starts the edge server, configures the computing resource pool, initializes the dynamic resource allocation module, and sets the maximum allocable resource ( in Equation (15)). The CAV establishes a communication link with the RSU, configures the OFDM parameters, and allocates channel resources;
[0148] S704: When the local perception confidence of the D-CAV is lower than the threshold, it is considered that an obstacle is sensed in the ROI;
[0149] S705: The D-CAV encapsulates the requirement instructions (ROI range, location, timestamp) and broadcasts them to neighboring CAVs through V2V, and notifies the RSU through V2I;
[0150] S706 - S708: Response and data processing of the R-CAV. Collect the original perception data in the form of point clouds and images, then perform feature compression of the original data and ROI extraction of the compressed data, and send the data packet containing the occluder description after spatio-temporal alignment to the D-CAV;
[0151] S709 - S710: The D-CAV triggers the IDSS algorithm. First, initialize the population and define the discrete search space. Secondly, update the positions of the LS and FS through chain collaborative search. Then, perform dynamic parameter adjustment under the control of the convergence factor. Finally, output the optimal fusion scheme;
[0152] S711: According to the fusion quality output by the IDSS algorithm and the local resource margin, if the output fusion quality is greater than the threshold and the local resource margin is greater than the minimum value, select local processing and execute S712. Otherwise, trigger offloading and execute S713;
[0153] S712: Perform local ROI fusion for D-CAV;
[0154] S713 - S716: Perform the HDDPG resource allocation process. After the RSU triggers the HDDPG algorithm, it selects and executes discrete and continuous hybrid actions, stores experiences, samples batch data to update the Critic network, updates the Actor network with policy gradients, then performs a soft update on the target network, and finally feeds back the resource allocation result;
[0155] S717: Allocate computing resources according to Equation (10) and transmit the fusion result back through OFDM;
[0156] S718 - S719: D-CAV offloads the data fusion calculation task to the RSU according to the resource allocation strategy, and the RSU performs the fusion of sensing data;
[0157] S720: D-CAV receives the fusion result from the RSU, aligns the local fusion result and the edge fusion result spatiotemporally, and finally updates the global ROI model. In addition, Kalman filtering is used to eliminate fusion noise and improve the robustness of the model;
[0158] S721: The RSU collects the interaction data (state, action, reward) between the CAV and the environment, stores it in the experience replay pool for HDDPG training, and at the same time softly updates the parameters of the Critic and Actor target networks to ensure the stability of training.
[0159] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the purpose and scope of the present technical solution, and they should all be covered within the scope of the claims of the present invention.
Claims
1. A multi-source perception fusion and dynamic resource allocation method based on vehicle-road cooperation, characterized in that: The method includes the following steps: S1: Construct a vehicle-road double-layer network architecture based on perception fusion, including a Connected Automated Vehicle (CAV) layer and a Road Side Unit (RSU) layer. The CAV layer realizes data sharing through Vehicle-to-Vehicle (V2V) and Vehicle-to-Infrastructure (V2I) communications, and the RSU layer assists the CAV to complete data fusion through edge computing resources; S2: Based on the vehicle-road double-layer network architecture, establish a Collaborative Omni-Perception (COP) model, and realize distributed environmental perception through the dual-role dynamic switching mechanism of the CAV between the Demanding Connected Automated Vehicle (D-CAV) and the Responding Connected Automated Vehicle (R-CAV); S3: Construct a multi-source perception fusion model, fuse the local perception information of the D-CAV and the heterogeneous perception information of the R-CAV through cross-modal spatio-temporal alignment and confidence weighting, and quantify the perception fusion quality; S4: Model the joint optimization problem of multi-source perception fusion and resource allocation, and divide it into a sub-problem of maximizing the ROI perception fusion quality and a sub-problem of minimizing the global processing delay; S5: Design a perception fusion optimization algorithm based on the Improved Discrete Salp Swarm (IDSS), and solve the sub-problem of maximizing the ROI perception fusion quality through a chain collaborative search mechanism and a dynamic parameter adjustment strategy; S6: Design a resource allocation algorithm based on the Hybrid Action Space Based Deep Deterministic Policy Gradient (HDDPG), and solve the sub-problem of minimizing the global processing delay through a hierarchical hybrid action space modeling and a reinforcement learning framework.
2. The multi-source perception fusion and dynamic resource allocation method based on vehicle-road cooperation according to claim 1, wherein: In the above S1, the construction of the vehicle-road double-layer network architecture includes: The CAV layer performs local perception data acquisition, feature compression, and ROI extraction, and transmits the ROI data to the RSU layer through the V2V / V2I protocol; The RSU layer receives the CAV data through the Orthogonal Frequency Division Multiplexing (OFDM) communication mechanism, and assists the CAV to complete the ROI data fusion based on the edge computing resource dynamic allocation strategy.
3. The multi-source perception fusion and dynamic resource allocation method based on vehicle-road cooperation according to claim 1, characterized in that: In the above S2, the implementation of the collaborative global perception model includes: When the D-CAV detects an obstacle, it broadcasts a collaborative perception demand and receives the ROI extraction information fed back by the R-CAV; In response to the requirements, R-CAV sequentially performs perception compression, ROI extraction, and information transmission, and transmits the processed data back to D-CAV with the assistance of RSU. RSU dynamically allocates edge resources according to the computing resource load of CAV to optimize the ROI fusion delay.
4. The multi-source perception fusion and dynamic resource allocation method based on vehicle-road cooperation according to claim 1, characterized in that: In step S3, the construction of the multi-source perception fusion model includes: Perform timestamp synchronization, coordinate system transformation, and three-dimensional reconstruction on the heterogeneous perception data of D-CAV and R-CAV to generate unified description information. Quantify the ROI perception fusion quality based on multi-dimensional joint evaluation of observation distance, computing resource allocation, and perception confidence.
5. The multi-source perception fusion and dynamic resource allocation method based on vehicle-road cooperation according to claim 1, characterized in that: In step S5, the execution of the IDSS algorithm includes: Initialize the population of the discrete solution space, and define the chain collaborative search mechanism of the leader (Leader Salp, LS) and the followers (Follower Salps, FSs). Balance global exploration and local exploitation by dynamically adjusting the convergence factor and random perturbation parameters, and iteratively solve the optimal solution of the occluder description scheme.
6. The multi-source perception fusion and dynamic resource allocation method based on vehicle-road cooperation according to claim 1, characterized in that: In step S6, the implementation of the HDDPG algorithm includes: Model the resource allocation problem as a Parameterized Action Space Markov Decision Process (PAMDPs). Output the discrete offloading decision, continuous offloading rate, and RSU resource allocation action through the three-branch Actor network respectively, and optimize the long-term cumulative reward by combining the Q-value evaluation strategy of the Critic network.
7. The multi-source perception fusion and dynamic resource allocation method based on vehicle-road cooperation according to claim 2, wherein: The resource allocation strategy of the RSU layer includes: Dynamically decide whether to offload the fusion task to the RSU based on the local computing margin of CAV and the ROI fusion quality threshold. Transmit the fusion result back to D-CAV through the OFDM channel, and use Kalman filtering to eliminate the spatio-temporal alignment noise.
8. The multi-source perception fusion and dynamic resource allocation method based on vehicle-road collaboration according to claim 3, characterized in that: The specific ROI extraction of R-CAV includes: Adopt a dynamic region detection algorithm to screen the ROI perception information strongly related to the driving of D-CAV. Discard the timeout data packets to ensure real-time performance by the transmission delay constraint of the compressed sensing data.
9. The multi-source perception fusion and dynamic resource allocation method based on vehicle-road cooperation according to claim 5, characterized in that: In the IDSS algorithm, the position update formula of the leader LS is: Among them, represents the position of LS in the u-th dimension; F u represents the position of the food, usually representing the optimal solution for the current situation; ub u and lb u are the corresponding upper and lower bounds, which are the number of description types K and 1 of the object respectively; c1 = 2e -(4zZ)2 is the convergence factor, which plays a role in balancing global exploration and local exploitation, and gradually decreases during the iteration process. z and Z are the current iteration and the maximum number of iterations respectively; c2 and c3 are random variables drawn from the interval [0, 1].
10. The multi-source perception fusion and dynamic resource allocation method based on vehicle-road cooperation according to claim 6, characterized in that: The reward function of the HDDPG algorithm is designed as: Among them, At time slot t, the intelligent agent RSU is in state s t Takes action a t The immediate reward value obtained after that; T i local Is the total delay for CAV i to fully locally process ROI fusion; T i t Is the actual total processing delay of CAV i in time slot t; N is the total number of connected and autonomous vehicles CAV in the system.
Citation Information
Cited By
BEV-based multi-source sensing data fusion and vehicle and road resource allocation method
CN120708406A
A method for multi-source perception data fusion and road-vehicle resource allocation based on BEV
CN120708406B
Multi-agent cooperative sensing method and system for Internet of Vehicles
CN121837878A
An internet of vehicles multi-agent cooperative perception method and system
CN121837878B