Communication method based on communication-sensing-reflection cooperation of vehicle-road cloud integration

By using an integrated perception and communication system assisted by an intelligent reflective surface, combined with alternating optimization and multi-agent reinforcement learning, the problems of communication interruption and perception failure in complex urban scenarios have been solved. This has achieved a joint improvement in communication rate and radar detection performance, thereby enhancing the reliability and safety of the autonomous driving system.

CN121485726BActive Publication Date: 2026-04-24JILIN UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
JILIN UNIVERSITY
Filing Date
2026-01-08
Publication Date
2026-04-24

AI Technical Summary

Technical Problem

In complex urban scenarios, obstacles can disrupt the line-of-sight link between base stations and users/targets, leading to communication interruptions and perception failures. Existing technologies struggle to achieve a dynamic balance between maximizing communication rates and maintaining radar beam pattern fidelity in vehicle-road-cloud systems, and multi-IRS collaborative control faces challenges of dimensional explosion and local optima.

Method used

By constructing an integrated sensing and communication system assisted by an intelligent reflector, and combining alternating optimization strategies and multi-agent reinforcement learning, the active beam of the base station and the passive beam of the IRS are optimized, achieving a joint improvement in communication rate and radar detection performance without the need for precise channel state information.

Benefits of technology

It significantly improves signal coverage in non-line-of-sight blind spots in complex urban scenarios, enhances the reliability and safety of autonomous driving systems, reduces hardware costs and deployment complexity, and is suitable for low-cost, low-power applications on large-scale roadside infrastructure.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121485726B_ABST
    Figure CN121485726B_ABST
Patent Text Reader

Abstract

The application belongs to the technical field of intelligent transportation, and particularly relates to a vehicle-road cloud integrated communication method based on communication-sensing-reflection cooperation. The method comprises the following steps: step one, constructing an intelligent reflective surface assisted integrated sensing and communication system model; step two, base station active beamforming optimization based on equivalent conversion; step three, adopting multi-agent reinforcement learning to optimize IRS passive beamforming; and step four, alternating optimization and system deployment. The application is used in the scene where there are buildings, large vehicles and other obstacles on urban roads, and solves the problems of communication interruption and sensing failure in non-line-of-sight areas through the intelligent reflective surface assisted integrated sensing and communication system. Through alternating optimization of base station active beam and passive beam, the method realizes the joint improvement of communication rate and radar detection performance without accurate channel state information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of intelligent transportation technology, specifically a vehicle-road-cloud integrated communication method based on communication-sensing-reflection coordination. Background Technology

[0002] With the development of 6G and intelligent connected vehicles, vehicle-road-cloud integrated systems face stringent demands for both highly reliable communication and high-precision environmental perception. Integrated Sensing and Communication (ISAC) technology achieves "one-to-two" transmission and reception by sharing hardware and spectrum resources, significantly improving spectrum and energy efficiency. However, in complex urban scenarios, obstacles such as large vehicles and buildings can easily disrupt the line-of-sight (LoS) link between the base station (BS) and the user / target, leading to communication interruptions and perception failures, severely restricting the safety and reliability of autonomous driving.

[0003] To address the challenge of blind zone coverage, Intelligent Reflectors (IRS) have been introduced into ISAC systems. IRSs consist of numerous low-cost, passive, programmable reflective units. By dynamically adjusting the reflection phase of each unit, they can reconstruct the wireless propagation environment, construct a "virtual line-of-sight" link, enhance communication signal coverage, and improve sensing performance. However, this technology faces three core scientific challenges: First, IRSs typically lack radio frequency links and cannot actively estimate Channel State Information (CSI), making traditional optimization methods relying on precise CSI difficult to deploy. Second, there is an inherent conflict between maximizing communication rate and radar beammap fidelity; communication tends to concentrate energy in the user direction, while sensing requires beam coverage across a wide area, necessitating a performance balance in dynamic environments. Third, multi-IRS collaborative control faces the dilemma of "dimensional explosion" and local optima; centralized algorithms have poor scalability, and distributed methods lack effective collaboration mechanisms. Existing technologies either only optimize communication, ignore IRS deployment constraints, or rely on ideal CSI, making it difficult to achieve efficient coordination of sensing and reflection in real vehicle-road-cloud scenarios. Therefore, there is an urgent need for a new collaborative optimization framework to achieve efficient collaboration of "cloud control optimization + roadside execution + IRS reflection" under the vehicle-road-cloud architecture. Summary of the Invention

[0004] This invention provides a vehicle-road-cloud integrated communication method based on communication-sensing-reflection coordination. It addresses communication interruptions and sensing failures in non-line-of-sight (NLoS) areas by integrating a sensing and communication (ISAC) system with the aid of an intelligent reflector surface (IRS) in urban road scenarios with obstacles such as buildings and large vehicles. This method achieves a combined improvement in communication rate and radar detection performance without requiring precise channel state information (CSI) by alternately optimizing the active beam of the base station and the passive beam of the IRS.

[0005] The technical solution of this invention is described below in conjunction with the accompanying drawings:

[0006] This invention provides a vehicle-road-cloud integrated communication method based on communication-sensing-reflection coordination, comprising the following steps:

[0007] S1. Construct an integrated sensing and communication system model assisted by intelligent reflective surfaces;

[0008] S11. Deploy the system;

[0009] S12. Optimize target definition to maximize the total vehicle communication rate while ensuring that the radar beam pattern error does not exceed the threshold ε.

[0010] S2. Optimization of active beamforming for base stations based on equivalent transformation;

[0011] S21. With the phase of the fixed reflector, the non-convex communication rate target is equivalently transformed through fractional programming.

[0012] S22. Construct a convex upper bound surrogate function for beammap error;

[0013] S23. Iteratively solve the convex sub-protrusion problem to optimize the active beam of the base station;

[0014] S3. Employ multi-agent reinforcement learning to optimize passive beamforming of intelligent reflectors;

[0015] S31. Model each reflective surface as an intelligent agent and construct a local state based on the vehicle's position and its own phase.

[0016] S32. The MADDPG algorithm is used for centralized training, with the total communication rate of the system as the global reward, and the optimal reflection strategy is learned collaboratively.

[0017] S4, Alternating Optimization and System Deployment;

[0018] S41. S2 and S3 are executed alternately to form a closed loop of "optimizing base station beam - updating reflector phase";

[0019] S42. Ultimately, the collaborative deployment of "cloud / roadside periodic optimization and autonomous response of lightweight model at the reflector end" is achieved in the vehicle-road-cloud architecture.

[0020] Furthermore, the specific method of S11 is as follows:

[0021] A multi-antenna base station is deployed on the roadside unit; the antenna base station performs two tasks simultaneously: communication function: transmitting data to multiple networked vehicles; sensing function: detecting multiple moving targets; multiple IRSs are deployed on the infrastructure on both sides of the road, each IRS consisting of dozens to hundreds of passive reflective elements that can independently adjust the phase;

[0022] The specific method of S12 is as follows:

[0023] The system optimization objective is to maximize the total communication rate of all vehicles while ensuring radar detection accuracy. Specifically, radar performance is measured by beam pattern error, which is the mean square error between the actual transmitted radar beam and the ideal beam; this error must be below a preset threshold ε. The base station transmit power cannot exceed the maximum value P. max The IRS phase must be within the range of [0, 2π]. The formula for calculating the beammap error constraint is as follows:

[0024] ;

[0025] In the formula, The beammap error cost function; For base station transmission matrix; The guide vector describes the array at an angle. The response below; The ideal beam pattern is the desired beam gain. These are the weighting coefficients; It is the conjugate transpose; This is the actual beam pattern response; This represents the total number of angle sampling points; For the first Each sampling direction angle; The maximum permissible beammap error threshold.

[0026] Furthermore, the specific method of S21 is as follows:

[0027] The original optimization problem aims to maximize the communication rate, and is in the form of:

[0028] ;

[0029] In the formula, This is the channel vector; Noise power;

[0030] The objective function is transformed into an equivalent convex optimization form using fractional programming techniques. An auxiliary variable t is introduced, and we let:

[0031] ;

[0032] The specific method of S22 is as follows:

[0033] Perform a local approximation within the neighborhood of the current solution and construct a convex upper bound surrogate function. It satisfies the following conditions: 1. It is tangent to the original function at the current iteration point; 2. For all , 3. yes The convex function; using a first-order Taylor expansion approximation, the result is:

[0034] ;

[0035] In the formula, This is the current optimal solution at the k-th iteration; for The conjugate transpose of ; for The conjugate transpose of; This represents the total number of sampling directions.

[0036] The specific method of S23 is as follows:

[0037] In each iteration, solve the following convex optimization subproblem:

[0038] ;

[0039] ;

[0040] In the formula, The objective function for communication rate is convexized; For the convex upper bound surrogate function constructed; For transmit power constraints;

[0041] After obtaining a new solution, the beam matrix is ​​updated, and the approximate constraints are reconstructed based on the new points to enter the next iteration.

[0042] Furthermore, the specific method of S31 is as follows:

[0043] Each IRS is treated as an independent agent, and a multi-agent decision-making system is constructed; the multi-agent system has the following components:

[0044] 1. Intelligent Agent; Assume there are M IRSs in the system, each IRS is equipped with There are 1 programming reflection unit, and the reflection coefficient matrix is:

[0045] ;

[0046] In the formula, Let be the reflection coefficient matrix of the m-th IRS; Let n be the reflection phase of the nth unit, which is the variable to be optimized. The imaginary unit is used to construct phase adjustments in complex exponential form; The total number of reflective elements in the m-th IRS; This refers to the IRS number in the system;

[0047] 2. State space; each agent The local state input is defined as:

[0048] ;

[0049] In the formula, For the first Two-dimensional location of each vehicle user; This is the current phase vector of the m-th IRS;

[0050] 3. Action Space: The action of each agent is defined as the adjustment amount of the phase of its own reflection unit; at each decision moment, the agent outputs a continuous phase adjustment vector according to the current environmental state. The dimension of the phase adjustment vector is consistent with the number of reflection units controlled by the agent; each component represents the phase change amplitude of the corresponding reflection unit.

[0051] 4. Reward Function; The system adopts a global reward mechanism, where the reward value is jointly determined by the communication quality of all vehicle users; the global reward is the total communication rate reported by all user receivers, mathematically expressed as follows:

[0052] ;

[0053] In the formula, The global reward represents the total communication rate of the system. This represents the total number of vehicle users in the system. From the base station to the Direct connection path channel status between users; This is the base station transmission matrix, which controls the direction of signal transmission. For from the first The intelligent reflective surface to the first Reflection path channel for each user; For the first The reflection coefficient matrix of each intelligent reflective surface is determined by the phase of each unit; From the base station to the An incident channel for a smart reflective surface; The additive white Gaussian noise power at the receiving end; This represents the total number of IRSs deployed in the system.

[0054] The specific method of S32 is as follows:

[0055] A multi-agent deep deterministic strategy gradient algorithm is used to achieve efficient collaboration among multiple intelligent reflective surfaces;

[0056] During training, the system constructs a shared learning platform where the state and actions of all agents are monitored by a central controller. A global commentator network is responsible for evaluating the long-term benefits of the joint behavior of all agents, providing precise value feedback by analyzing the global state and the actions of each agent. Meanwhile, each agent has an independent executor network for generating its own control policy. The network outputs specific phase adjustment instructions based on local observations and updates its policy using global value signals.

[0057] During training, the system continuously interacts with the environment, collects state transition, action execution, and reward feedback data, and stores them in a shared replay buffer. The data is randomly sampled for iterative optimization of network parameters. In addition, the algorithm introduces a target network mechanism to update the target network parameters.

[0058] After training is completed, each intelligent reflector only needs to deploy an executor network to operate independently in the real environment. At this time, each agent no longer depends on the information or communication coordination of other agents, but makes autonomous decisions based on its own perceived local state to achieve distributed control.

[0059] Furthermore, the specific method of S41 is as follows:

[0060] An alternating optimization strategy is adopted, and S2 and S3 are executed cyclically to form a collaborative closed loop of communication-sensing-reflection. First, under the premise of fixed IRS phase, the base station beam is optimized based on equivalent transformation to ensure that the radar sensing performance meets the requirements. Then, based on the updated beam, a multi-agent reinforcement learning framework is launched to collaboratively optimize the reflection phase of each IRS.

[0061] The specific method of S42 is as follows:

[0062] The system adopts a layered architecture of "cloud control optimization, roadside execution, and terminal response"; the equivalent conversion and stepwise optimization algorithms run periodically in the cloud or roadside units, and update the base station beam in combination with global information to ensure the coordination of perception and communication; a lightweight neural network model is deployed at the IRS end to receive vehicle position and communication rate feedback, that is, to autonomously adjust the reflection phase according to the local state.

[0063] The beneficial effects of this invention are as follows:

[0064] 1. This invention effectively solves the signal coverage problem in non-line-of-sight blind spots in complex urban scenarios by constructing a collaborative optimization framework of communication, perception, and reflection. By reconstructing the electromagnetic environment using intelligent reflective surfaces and combining it with an alternating optimization strategy, it significantly improves vehicle communication quality and radar perception accuracy without requiring precise channel information, thereby enhancing the reliability and safety of autonomous driving systems.

[0065] 2. This invention employs a multi-agent reinforcement learning mechanism of "centralized training and distributed execution," avoiding the dependence on channel state information in traditional methods. IRS only requires vehicle position and rate feedback to make autonomous decisions, eliminating the need for radio frequency links and inter-vehicle communication, significantly reducing hardware costs and deployment complexity, and making it suitable for low-cost, low-power applications in large-scale roadside infrastructure.

[0066] 3. This invention supports integrated vehicle-road-cloud layered deployment, with base station beams periodically optimized by the cloud or roadside, and a lightweight IRS model responding to dynamic environments in real time. This architecture balances global performance with local flexibility, enabling efficient sharing of sensing resources and providing a scalable and easily deployable technical solution for future intelligent transportation and 6G converged networks. Attached Figure Description

[0067] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as a limitation on the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0068] Figure 1 This is a diagram of the architecture of the present invention;

[0069] Figure 2 This is a schematic diagram illustrating the change in the total communication rate of the system after each alternating optimization iteration;

[0070] Figure 3 The reward curve for the MADDPG training process;

[0071] Figure 4 The curve represents the phase evolution process of the IRS.

[0072] Figure 5 This is the evolution curve of the global Q value;

[0073] Figure 6 This is a diagram illustrating the distribution of training rewards.

[0074] Figure 7 This is a schematic diagram of the vehicle's movement trajectory;

[0075] Figure 8 This is a schematic diagram of the direct path.

[0076] Figure 9 A schematic diagram of the layout of a visual vehicle-to-everything (V2X) system;

[0077] Figure 10 This is a schematic diagram of the obstacle's movement path. Detailed Implementation

[0078] The present invention will now be described in further detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and not intended to limit it. Furthermore, it should be noted that, for ease of description, the accompanying drawings show only the parts relevant to the present invention, and not all of the structures.

[0079] See Figures 1-10 This embodiment provides a vehicle-road-cloud integrated communication method based on communication-sensing-reflection coordination, including the following steps:

[0080] S1. Construct an integrated sensing and communication system model assisted by an intelligent reflective surface. The specific method is as follows:

[0081] S11. Deploy the system;

[0082] A multi-antenna base station (BS) is deployed in the roadside unit (RSU). This base station performs two tasks simultaneously: communication function: transmitting data to multiple networked vehicles (Communication Users, CUs); and sensing function: detecting multiple moving targets (TGs), such as pedestrians and non-motorized vehicles. Multiple IRSs are deployed on infrastructure such as streetlights and billboards on both sides of the road. Each IRS consists of dozens to hundreds of passive reflective elements with independently adjustable phases.

[0083] S12. Optimize target definition to maximize the total vehicle communication rate while ensuring that the radar beam pattern error does not exceed the threshold ε.

[0084] The system optimization objective is to maximize the total communication rate of all vehicles while ensuring radar detection accuracy. Specifically, radar performance is measured by "beam pattern error," which is the mean square error between the actual transmitted radar beam and the ideal beam; this error must be below a preset threshold ε. The base station transmit power cannot exceed the maximum value P. max The IRS phase must be within the range of [0, 2π] to reflect its passive and low-power characteristics. The formula for calculating the beammap error constraint is as follows:

[0085] ;

[0086] In the formula, The beammap error cost function; For base station transmission matrix; The guide vector describes the array at an angle. The response below; The ideal beam pattern is the desired beam gain. These are the weighting coefficients; It is the conjugate transpose; This is the actual beam pattern response; This represents the total number of angle sampling points; For the first Each sampling direction angle; The maximum allowable beammap error threshold; the core idea expressed by this formula is: the actual beammap (derived from the transmission matrix) The desired beam pattern should be as close as possible to the ideal beam pattern. At all sampling angles The average error does not exceed the threshold. .

[0087] S2. Optimization of active beamforming for base stations based on equivalent conversion, the specific method is as follows:

[0088] With the phase configuration of the intelligent reflector (IRS) fixed, the goal of this phase is to optimize the base station's transmit beam matrix. This is to satisfy the radar beammap error constraints and provide a high-quality initial solution for the subsequent joint optimization of the IRS. Since the original optimization problem contains non-convex constraints, it is difficult to solve directly. Therefore, this step adopts a processing method based on equivalent transformation and stepwise optimization, which transforms the original problem into a series of solvable subproblems and approximates the feasible solution through iteration.

[0089] S21. With the phase of the fixed reflector, the non-convex communication rate target is equivalently transformed through fractional programming.

[0090] The original optimization problem aims to maximize the communication rate, and is in the form of:

[0091] ;

[0092] In the formula, This is the channel vector; This represents noise power.

[0093] The objective function is non-concave and involves logarithmic terms, making it difficult to handle directly. Therefore, fractional programming is used to transform it into an equivalent convex optimization form. An auxiliary variable t is introduced, let:

[0094] ;

[0095] At this point, the objective becomes maximizing t, and the constraint becomes a quadratic inequality, which facilitates subsequent processing.

[0096] S22. Construct a convex upper bound surrogate function for beammap error;

[0097] The core difficulty of step two lies in handling the non-convex beammap error constraint in S12, because... It is non-convex because Since it is a quadratic term, the overall solution is neither differentiable nor convex. Therefore, this invention performs a local approximation within the neighborhood of the current solution, constructing a convex upper bound surrogate function. It satisfies the following conditions: 1. It is tangent to the original function at the current iteration point; 2. For all , 3. yes The convex function. Using a first-order Taylor expansion approximation, the result is:

[0098] ;

[0099] In the formula, This is the current optimal solution at the k-th iteration; for The conjugate transpose of ; for The conjugate transpose of; This represents the total number of sampling directions.

[0100] The function is tangent to the original function at the current iteration point and is always greater than or equal to the original function, thus transforming the original problem into a standard convex optimization problem, and thus the entire optimization problem into a standard solvable model;

[0101] S23. Iteratively solve the convex sub-protrusion problem to optimize the active beam of the base station;

[0102] In each iteration, solve the following convex optimization subproblem:

[0103] ;

[0104] ;

[0105] In the formula, The objective function for communication rate is convexized; For the convex upper bound surrogate function constructed; The transmit power constraint is used. This subproblem is a standard solvable optimization model, which can be solved using general algorithms such as the interior-point method. After obtaining a new solution, the beam matrix is ​​updated, and the approximate constraints are reconstructed based on the new points, entering the next iteration. This process is repeated until the change in system performance (such as communication rate and beam error) is less than a preset threshold, at which point the algorithm is considered to have converged.

[0106] S3. Multi-agent reinforcement learning is used to optimize IRS passive beamforming. The specific method is as follows:

[0107] After obtaining the base station beam P, this step optimizes the reflection phase of each IRS. Since IRSs typically lack radio frequency links and cannot obtain accurate channel information (CSI), traditional model-based optimization methods (such as convex optimization and alternating optimization) are difficult to apply. Therefore, a distributed learning framework based on Multi-Agent Deep Deterministic Policy Gradient (MADDPG) is proposed to achieve model-free and communication-free collaborative optimization.

[0108] S31. Model each reflective surface as an intelligent agent and construct a local state based on the vehicle's position and its own phase.

[0109] Treating each IRS as an independent agent, a multi-agent decision-making system is constructed. This system has the following components:

[0110] 1. Intelligent Agent. Assume there are M IRSs in the system, each IRS equipped with... Each programming reflection unit has a reflection coefficient matrix as follows:

[0111] ;

[0112] In the formula, Let be the reflection coefficient matrix of the m-th IRS; Let n be the reflection phase of the nth unit, which is the variable to be optimized. The imaginary unit is used to construct phase adjustments in complex exponential form; The total number of reflective elements in the m-th IRS; This refers to the IRS number in the system;

[0113] 2. State space. Each agent... The local state input is defined as:

[0114] ;

[0115] In the formula, For the first Two-dimensional location of each vehicle user; This is the current phase vector of the m-th IRS.

[0116] 3. Action Space. Each agent's action is defined as the adjustment amount of the phase of its own reflector unit. At each decision-making moment, the agent outputs a continuous phase adjustment vector based on the current environmental state. The dimension of this vector is consistent with the number of reflectors controlled by the agent. Each component represents the phase change amplitude of the corresponding reflector unit. This design allows the agent to flexibly and precisely control the direction and intensity of the reflected beam, thereby progressively optimizing communication performance.

[0117] 4. Reward Function. The system adopts a global reward mechanism, where the reward value is determined jointly by the communication quality of all vehicle users. Specifically, the global reward is the total communication rate reported by all user receivers, and its mathematical expression is as follows:

[0118] ;

[0119] In the formula, The global reward represents the total communication rate of the system, expressed in bits per second per hertz. This represents the total number of vehicle users (communication users) in the system. From the base station to the Direct connection path channel status between users; This is the base station transmission matrix, which controls the direction of signal transmission. For from the first The intelligent reflective surface to the first Reflection path channel for each user; For the first The reflection coefficient matrix of a smart reflective surface is determined by the phase of each of its elements; From the base station to the An incident channel for a smart reflective surface; The summative white Gaussian noise power at the receiver is denoted as . This reward function comprehensively reflects the combined performance of the base station transmission strategy and the effects of all intelligent reflector modulation, and can effectively guide the agent to learn strategies to improve overall communication quality.

[0120] S32. The MADDPG algorithm is used for centralized training, with the total communication rate of the system as the global reward, and the optimal reflection strategy is learned collaboratively.

[0121] This invention employs the Multi-Agent Deep Deterministic Policy Gradient Algorithm (MADDPG) to achieve efficient collaboration among multiple intelligent reflectors. The core idea of ​​this mechanism is to utilize global information for centralized learning during the training phase, while supporting each agent to make independent decisions based solely on local information during the actual deployment phase, thus balancing learning efficiency and deployment feasibility.

[0122] During training, the system constructs a shared learning platform where the states and actions of all agents are monitored by a central controller. A global critic network is responsible for evaluating the long-term benefits of the joint actions of all agents, providing precise value feedback by analyzing the global state and the actions of each agent. This centralized criticism mechanism effectively captures the mutual influence between agents, avoiding learning biases caused by reward sparsity or delay. Simultaneously, each agent has an independent actor network for generating its own control policy. This network outputs specific phase adjustment instructions based on local observations and updates its policy using global value signals.

[0123] During training, the system continuously interacts with the environment, collecting experiential data such as state transitions, action execution, and reward feedback, and storing it in a shared replay buffer. This data is randomly sampled for iterative optimization of network parameters, significantly improving sample utilization and learning stability. Furthermore, the algorithm introduces a target network mechanism, which effectively mitigates fluctuations and divergence during training by slowly updating the target network parameters, ensuring smooth convergence of the learning process.

[0124] After training is complete, each intelligent reflector only needs to deploy its agent network to operate independently in a real-world environment. At this point, each agent no longer relies on information or communication coordination from other agents; it makes autonomous decisions based solely on its own perceived local state (such as the positions of surrounding vehicles and its own current configuration), achieving true distributed control. This "centralized training, decentralized execution" paradigm ensures the collaborative performance of the multi-agent system while meeting the hardware constraints of low power consumption, passivity, and no communication required by the intelligent reflectors, making it particularly suitable for large-scale vehicle-to-infrastructure (V2I) deployments.

[0125] S4. Alternating optimization and system deployment, the specific methods are as follows:

[0126] S41. S2 and S3 are executed alternately to form a closed loop of "optimizing base station beam - updating IRS phase";

[0127] This invention employs an alternating optimization strategy, iteratively executing steps S2 and S3 to form a collaborative closed loop of communication-sensing-reflection. First, with the IRS phase fixed, the base station beam is optimized based on equivalent transformation to ensure radar sensing performance meets requirements. Then, based on the updated beam, a multi-agent reinforcement learning (MADDPG) framework is initiated to collaboratively optimize the reflection phase of each IRS, improving communication quality in non-line-of-sight areas. This process iterates repeatedly until system performance (such as total rate and beam error) converges. This method decomposes a complex joint problem into two solvable subproblems, balancing optimization accuracy and computational efficiency to achieve a gradual improvement in sensing performance.

[0128] S42. Ultimately, the collaborative deployment of "cloud / roadside periodic optimization and autonomous response of lightweight model at the reflector end" is achieved in the vehicle-road-cloud architecture;

[0129] The system adopts a layered architecture of "cloud-controlled optimization, roadside execution, and end-side response." Equivalent conversion and progressive optimization algorithms run periodically in the cloud or roadside units, updating base station beams in conjunction with global information to ensure coordination between sensing and communication. A lightweight neural network model is deployed at the IRS end, requiring only vehicle location and communication rate feedback to autonomously adjust the reflection phase based on local conditions, without channel estimation or communication with other IRSs. This deployment method significantly reduces hardware costs and signaling overhead, fully leveraging the passive and low-power advantages of IRSs, making it suitable for large-scale vehicle-road cooperative scenarios.

[0130] In summary, this invention addresses the issues of communication interruption and sensing failure in non-line-of-sight (NLoS) areas by using an intelligent reflector (IRS)-assisted integrated sensing and communication (ISAC) system in urban road scenarios with obstacles such as buildings and large vehicles. This method achieves a combined improvement in communication rate and radar detection performance without requiring precise channel state information (CSI) by alternately optimizing the active beam of the base station and the passive beam of the IRS.

[0131] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A vehicle-road-cloud integrated communication method based on communication-sensing-reflection coordination, characterized in that, Includes the following steps: S1. Construct an integrated sensing and communication system model assisted by intelligent reflective surfaces; S11. Deploy the system; S12. Optimize target definition to maximize the total vehicle communication rate while ensuring that the radar beam pattern error does not exceed the threshold ε. S2. Optimization of active beamforming for base stations based on equivalent transformation; S21. With the phase of the fixed reflector, the non-convex communication rate target is equivalently transformed through fractional programming. S22. Construct a convex upper bound surrogate function for beammap error; S23. Iteratively solve the convex sub-protrusion problem to optimize the active beam of the base station; S3. Employ multi-agent reinforcement learning to optimize passive beamforming of intelligent reflectors; S31. Model each reflective surface as an intelligent agent and construct a local state based on the vehicle's position and its own phase. S32. The MADDPG algorithm is used for centralized training, with the total communication rate of the system as the global reward, and the optimal reflection strategy is learned collaboratively. S4, Alternating Optimization and System Deployment; S41. S2 and S3 are executed alternately to form a closed loop of "optimizing base station beam - updating reflector phase"; S42. Ultimately, the collaborative deployment of "cloud / roadside periodic optimization and autonomous response of lightweight model at the reflector end" is achieved in the vehicle-road-cloud architecture.

2. The vehicle-road-cloud integrated communication method based on communication-sensing-reflection coordination according to claim 1, characterized in that, The specific method of S11 is as follows: A multi-antenna base station is deployed on the roadside unit; the antenna base station performs two tasks simultaneously: communication function: transmitting data to multiple networked vehicles; sensing function: detecting multiple moving targets; multiple IRSs are deployed on the infrastructure on both sides of the road, each IRS consisting of dozens to hundreds of passive reflective units that can independently adjust the phase; The specific method of S12 is as follows: The system optimization objective is to maximize the total communication rate of all vehicles while ensuring radar detection accuracy. Specifically, radar performance is measured by beam pattern error, meaning the mean square error between the actual transmitted radar beam and the ideal beam must be below a preset threshold ε. The base station transmit power cannot exceed the maximum value P. max The IRS phase must be within the range of [0, 2π]. The formula for calculating the beammap error constraint is as follows: ; In the formula, The beammap error cost function; For base station transmission matrix; The guide vector describes the array at an angle. The response below; The ideal beam pattern is the desired beam gain. These are the weighting coefficients; It is the conjugate transpose; This is the actual beam pattern response; This represents the total number of angle sampling points; For the first Each sampling direction angle; The maximum permissible beammap error threshold.

3. The vehicle-road-cloud integrated communication method based on communication-sensing-reflection coordination according to claim 2, characterized in that, The specific method of S21 is as follows: The original optimization problem aims to maximize the communication rate, and is in the form of: ; In the formula, This is the channel vector; Noise power; The objective function is transformed into an equivalent convex optimization form using fractional programming techniques. An auxiliary variable t is introduced, and we let: ; The specific method of S22 is as follows: Perform a local approximation within the neighborhood of the current solution and construct a convex upper bound surrogate function. It satisfies the following conditions:

1. It is tangent to the original function at the current iteration point; 2. For all , 3. yes The convex function; using a first-order Taylor expansion approximation, the result is: ; In the formula, This is the current optimal solution at the k-th iteration; for The conjugate transpose of ; for The conjugate transpose of; The specific method of S23 is as follows: In each iteration, solve the following convex optimization subproblem: ; ; In the formula, The objective function for communication rate is convexized; For the convex upper bound surrogate function constructed; For transmit power constraints; After obtaining a new solution, the beam matrix is ​​updated, and the approximate constraints are reconstructed based on the new points to enter the next iteration.

4. The vehicle-road-cloud integrated communication method based on communication-sensing-reflection coordination according to claim 1, characterized in that, The specific method of S31 is as follows: Each IRS is treated as an independent agent, and a multi-agent decision-making system is constructed; the multi-agent system has the following components:

1. Intelligent Agent; Assume there are M IRSs in the system, each IRS is equipped with There are 1 programming reflection unit, and the reflection coefficient matrix is: ; In the formula, Let be the reflection coefficient matrix of the m-th IRS; Let n be the reflection phase of the nth unit, which is the variable to be optimized. The imaginary unit is used to construct phase adjustments in complex exponential form; The total number of reflective elements in the m-th IRS; This refers to the IRS number in the system; 2. State space; each agent The local state input is defined as: ; In the formula, For the first Two-dimensional location of each vehicle user; This is the current phase vector of the m-th IRS; 3. Action Space: The action of each agent is defined as the adjustment amount of the phase of its own reflection unit; at each decision moment, the agent outputs a continuous phase adjustment vector according to the current environmental state. The dimension of the phase adjustment vector is consistent with the number of reflection units controlled by the agent; each component represents the phase change amplitude of the corresponding reflection unit.

4. Reward Function; The system adopts a global reward mechanism, where the reward value is jointly determined by the communication quality of all vehicle users; the global reward is the total communication rate reported by all user receivers, mathematically expressed as follows: ; In the formula, The global reward represents the total communication rate of the system. This represents the total number of vehicle users in the system. From the base station to the Direct connection path channel status between users; This is the base station transmission matrix, which controls the direction of signal transmission. For from the first The intelligent reflective surface to the first Reflection path channel for each user; For the first The reflection coefficient matrix of each intelligent reflective surface is determined by the phase of each unit; From the base station to the An incident channel for a smart reflective surface; The additive white Gaussian noise power at the receiving end; This represents the total number of IRSs deployed in the system. The specific method of S32 is as follows: A multi-agent deep deterministic strategy gradient algorithm is used to achieve efficient collaboration among multiple intelligent reflective surfaces; During training, the system constructs a shared learning platform where the state and actions of all agents are monitored by a central controller. A global commentator network is responsible for evaluating the long-term benefits of the joint behavior of all agents, providing precise value feedback by analyzing the global state and the actions of each agent. Meanwhile, each agent has an independent executor network for generating its own control policy. The network outputs specific phase adjustment instructions based on local observations and updates its policy using global value signals. During training, the system continuously interacts with the environment, collects state transition, action execution, and reward feedback data, and stores them in a shared replay buffer. The data is randomly sampled for iterative optimization of network parameters. In addition, the algorithm introduces a target network mechanism to update the target network parameters. After training is completed, each intelligent reflector only needs to deploy an executor network to operate independently in a real environment. At this point, each agent no longer depends on the information or communication coordination of other agents, but makes autonomous decisions based solely on its own perceived local state to achieve distributed control.

5. The vehicle-road-cloud integrated communication method based on communication-sensing-reflection coordination according to claim 3, characterized in that, The specific method of S41 is as follows: An alternating optimization strategy is adopted, and S2 and S3 are executed cyclically to form a collaborative closed loop of communication-sensing-reflection. First, under the premise of fixed IRS phase, the base station beam is optimized based on equivalent transformation to ensure that the radar sensing performance meets the requirements. Then, based on the updated beam, a multi-agent reinforcement learning framework is launched to collaboratively optimize the reflection phase of each IRS. The specific method of S42 is as follows: The system adopts a layered architecture of "cloud control optimization, roadside execution, and terminal response"; the equivalent conversion and stepwise optimization algorithms run periodically in the cloud or roadside unit, and update the base station beam in combination with global information to ensure the coordination of perception and communication; a lightweight neural network model is deployed at the IRS end to receive vehicle position and communication rate feedback, that is, to autonomously adjust the reflection phase according to the local state.

Citation Information

Patent Citations

  • Waveform generation method and system of intelligent reflecting surface auxiliary safety communication integrated system based on reinforcement learning

    CN119697663A

  • Method for estimating arrival angle of user assisted by intelligent reflecting surface in multi-carrier ISAC system

    CN120957221A