Multi-modal wave beam and frequency spectrum joint optimization method for sensing integration
By using a nested structure of main beam and sub-beams and Bayesian deep reinforcement learning, the beam conflict and resource scheduling problems in the ISAC system were solved, realizing the organic integration and adaptive optimization of communication and sensing functions, and improving system performance and adaptability.
Patent Information
- Application Number
- CN202511043187.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-28
- Publication Date
- 2025-12-12
AI Technical Summary
Existing ISAC systems lack overall collaborative design under co-location conditions. Radar signal echo sidelobes interfere with communication, and static resource scheduling strategies lead to a decline in system performance, especially in rapidly changing network environments.
By adopting a nested structure of main beam and sub-beams, combined with Bayesian deep reinforcement learning, a dynamic spectral ratio function and an alternating optimization algorithm are designed to achieve the organic integration of communication and sensing functions and adaptive resource optimization.
It improves communication speed, sensing accuracy and spectrum utilization, enhances the system's adaptability and robustness in complex environments, reduces computational complexity, and meets real-time requirements.
Smart Images

Figure CN121124875A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of communication, more particularly to a multi-modal beam and spectrum joint optimization method for integrated communication and sensing. BACKGROUND
[0002] The current mainstream ISAC solution has three main defects. First, the traditional solution usually designs and optimizes communication and sensing functions independently, lacking overall collaborative consideration. This separate design results in the system failing to fully utilize the potential synergistic effect between the two functions. Second, in the co-sited ISAC system, the sidelobes of the radar signal will interfere with the communication receiver, reducing the communication signal-to-noise ratio. The traditional solution often lacks effective interference suppression mechanisms, resulting in a decline in communication link quality. Finally, the existing resource scheduling strategy mostly adopts static allocation or simple rules, which is difficult to cope with high-dynamic network environments. In particular, in fast-changing scenarios such as UAV swarms and Internet of Vehicles, the limitations of static strategies are more obvious, and the system performance declines significantly.
[0003] In view of this, the present application proposes a multi-modal beam and spectrum joint optimization method for integrated communication and sensing, which adopts a hierarchical structure of main beam and sub-beam nesting, organically integrates sensing and communication functions in the spatial domain, and introduces Bayesian deep reinforcement learning as an intelligent decision engine to realize adaptive optimization of resources. SUMMARY
[0004] In order to overcome the above-mentioned defects of the prior art, the present application provides a multi-modal beam and spectrum joint optimization method for integrated communication and sensing to solve the problems existing in the background art.
[0005] The present application provides the following technical solution: a multi-modal beam and spectrum joint optimization method for integrated communication and sensing, comprising the following steps: Step one, construct a beam weight vector through main beam and sub-beam nesting design; Step two, design a dynamic spectrum ratio function to adaptively adjust spectrum allocation according to the sensing detection probability; Step three, take the communication rate and sensing mutual information as the objective function, and optimize the decision variable through Bayesian deep reinforcement learning; Step four, use an alternating optimization algorithm for optimization.
[0006] Preferably, the beam weight vector is represented as: ; wherein, represents the total beam weight vector, represents the sensing main beam weight, represents the kth communication sub-beam weight, and respectively represent the power allocation factors for sensing and communication; by precisely controlling these weight parameters, the balance between sensing coverage and communication service can be achieved; denotes the total number of communication sub-beam weights; denotes communication.
[0007] Preferably, the dynamic spectrum ratio function is represented as: ; wherein, denotes the sensing detection probability, denotes the dynamic spectrum ratio function at time t, denotes the dynamic spectrum ratio function at time t, denotes the difference of dynamic spectrum ratio functions, denotes the maximum value of the dynamic spectrum ratio function, denotes the minimum value of the dynamic spectrum ratio function.
[0008] Preferably, the objective function in the objective function with communication rate and sensing mutual information as the target function is marked as a joint optimization objective function, and the joint optimization objective function is represented as: ; wherein, denotes the total rate of all communication users; denotes the mutual information metric of the radar; denotes the beam weight matrix; denotes the RIS phase vector; denotes the received signal-to-interference-and-noise ratio of the kth communication sub-beam; and respectively represent the corresponding weight factors.
[0009] Preferably, the optimization of the decision variable through the Bayesian deep reinforcement learning is specifically: The Bayesian deep reinforcement learning is introduced as an intelligent decision engine, the state space of the intelligent decision engine contains channel state information, target position information and historical sensing error, the action space covers beam pointing angle, spectrum allocation ratio and power allocation factor, and the reward function is set.
[0010] Preferably, the reward function is represented as: ; wherein, denotes the instantaneous reward function value obtained after being in state s t at time t and taking action a t ; denotes the state vector; denotes the action vector; Indicates the perceived mean square error; This represents the trade-off parameter.
[0011] Preferably, the decision-making process of the intelligent decision engine can be represented as follows: Select an action based on the current state; Perform the action and observe the environmental feedback; Bayesian posterior update.
[0012] Preferably, the method of reducing computational complexity by using an alternating optimization algorithm is as follows: a step-by-step solution strategy is adopted, specifically: First, the phase configuration of the RIS (Intelligent Reflector) is fixed, and the beam weight matrix is optimized using a semi-definite relaxation method. Then, with the beam configuration fixed, the RIS phase vector is optimized using a least-squares combined with manifold optimization method.
[0013] The technical effects and advantages of this invention are as follows: This invention, through steps one, two, and three, facilitates the use of a multimodal beam nesting structure. Through the hierarchical design of the main beam and sub-beams, it achieves the symbiotic integration of communication and sensing at the physical layer, effectively resolving beam conflict issues. By introducing a Bayesian method, multi-dimensional features such as channel state information (CSI), target location, and historical errors are incorporated into the state space. The uncertainty of strategy selection is quantified through a Bayesian posterior update mechanism, improving decision robustness by 40% and enhancing adaptability in complex environments. The multimodal beam nesting design, Bayesian deep reinforcement learning decision-making mechanism, and dual-index collaborative optimization method effectively address key issues in integrated communication and sensing, such as beam conflict, low spectrum utilization, and lack of intelligent decision-making. Communication rate, sensing accuracy, and spectrum utilization are effectively improved. Attached Figure Description
[0014] Figure 1 This is a flowchart of the multimodal beam and spectrum joint optimization method for the integration of sensing and communication in this invention. Detailed Implementation
[0015] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings. In addition, the forms of the various structures described in the following embodiments are merely illustrative. The multimodal beam and spectrum joint optimization method for sensor-integrated communication involved in the present invention is not limited to the structures described in the following embodiments. All other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0016] like Figure 1 As shown, this invention provides a method for joint optimization of multimodal beam and spectrum for integrated sensing and communication, comprising the following steps: Step 1: Construct a beam weight vector by nesting the main beam and sub-beams, i.e., a multi-mode beam nesting structure. The purpose is to resolve the spatial conflict between communication and sensing beams and achieve symbiosis of physical layer resources between the two. Through the hierarchical beam structure, the performance bottleneck of the traditional separate design is broken through, providing a hardware foundation for multi-user communication and wide-area sensing in high-density scenarios. Step 2: Design a dynamic spectrum ratio function to adaptively adjust spectrum allocation based on the sensing and detection probability. The purpose is to break the limitations of static spectrum allocation, improve spectrum resource utilization, and drive dynamic spectrum adjustment through sensing feedback to adapt to diverse service needs, such as the instantaneous switching between high-reliability communication and high-precision sensing. Step 3: Using communication rate and sensing mutual information as objective functions, optimize decision variables through Bayesian deep reinforcement learning. The purpose is to quantify the performance trade-off between communication and sensing, guide the intelligent decision engine through mathematical modeling, and maximize the overall system efficiency. This enables real-time resource optimization decisions in dynamic environments, improves system robustness through uncertainty quantification, and solves the convergence speed problem in highly dynamic scenarios. Step 4: Use an alternating optimization algorithm to reduce computational complexity and improve real-time performance. The purpose is to reduce computational complexity and meet real-time requirements. Through a step-by-step optimization strategy, millisecond-level resource scheduling can be achieved.
[0017] In this embodiment, it should be specifically noted that the beam weight vector is represented as follows: ;in, This represents the total beam weight vector. This indicates the perceived main beam weight. This represents the weight of the k-th communication sub-beam. and These represent the power allocation factors for sensing and communication, respectively; by precisely controlling these weighting parameters, a balance can be achieved between sensing coverage and communication services. Indicates the total number of communication sub-beam weights; It represents communication and is used to distinguish between two types of resources: sensing and communication. By adopting a hierarchical structure with nested main beams and sub-beams, sensing and communication functions are organically integrated in the spatial domain. The sensing main beam undertakes the task of environmental scanning and uses a ULA antenna array to achieve a wide-area coverage of ±60° while ensuring a gain level of 8 dBi. Within the spatial range of the main beam, multiple communication sub-beams are nested, with the beamwidth of each sub-beam controlled within 5°, providing high-quality point-to-point communication services for different users.
[0018] In this embodiment, it should be specifically noted that the dynamic spectrum ratio function is expressed as: ; in, This represents the probability of perception detection. express The dynamic spectral ratio function at time t. express The dynamic spectral ratio function at time t. This represents the difference between the dynamic spectrum ratio function. This represents the maximum value of the dynamic spectrum ratio function. This represents the minimum value of the dynamic spectrum ratio function; When the probability of perception detection When the threshold is below 0.8, the spectrum allocation for sensing functions is automatically increased. When sensing performance is sufficient, more spectrum resources are allocated to communication functions. This closed-loop mechanism ensures that the system can dynamically optimize resources according to real-time environmental changes.
[0019] In this embodiment, it should be specifically noted that the objective function in the objective function of communication rate and perceived mutual information is labeled as the joint optimization objective function, which is expressed as follows: ;in, This represents the total rate of all communication users, and its logarithmic form reflects the diminishing marginal utility of the rate. The mutual information metric of radar reflects the ability of its sensing function to acquire environmental information; Represents the beam weighting matrix; Represents the RIS phase vector; This represents the received signal-to-interference-plus-noise ratio (SIR) of the k-th communication sub-beam. It reflects the instantaneous state of the user link quality and is calculated using the rate R. k Key parameters; and These represent the corresponding weighting factors, respectively. and The design allows the system to be flexibly adjusted according to the needs of different application scenarios. In communication-intensive scenarios, it can increase... The value is optimized to ensure communication performance, and in perception-critical scenarios, it can increase... The value is used to ensure perception accuracy.
[0020] In this embodiment, it should be specifically explained that the optimization of decision variables through Bayesian deep reinforcement learning specifically refers to: Bayesian deep reinforcement learning is introduced as an intelligent decision engine to achieve adaptive resource optimization. The state space of the intelligent decision engine includes multi-dimensional environmental features such as channel state information, target location information, and historical perception errors, while the action space covers key decision variables such as beam pointing angle, spectrum allocation ratio, and power allocation factor. A reward function is set, which is expressed as: ;in, This indicates that the agent is in state s at time t. t And take action a t The resulting immediate reward function value indicates that a larger reward represents a better current decision. The state vector contains environmental features such as current channel state information, target location information, and historical sensing errors. This represents the action vector, which corresponds to the decision variables that the agent adjusts in this time slot, such as beam pointing angle, spectrum allocation ratio, power allocation factor, etc. This represents the mean square error of perception, which is the mean square value of the error in the radar's estimation of parameters such as target range and velocity. The smaller the value, the higher the perception accuracy. Indicates the trade-off parameters; The reward function balances the communication rate gain and the perception error penalty, enabling the agent to learn the optimal policy under different environmental conditions. The introduction of the Bayesian method enables the system to quantify the uncertainty of policy selection and improves the robustness of decision-making. The decision-making process of an intelligent decision engine can be represented as follows: Select an action based on the current state; Perform the action and observe the environmental feedback; Bayesian posterior update.
[0021] In this embodiment, it should be specifically noted that the method of reducing computational complexity by using an alternating optimization algorithm is as follows: a step-by-step solution strategy is adopted, specifically: First, the phase configuration of the RIS (Intelligent Reflector) is fixed, and the beam weight matrix is optimized using a semi-definite relaxation method. Then, with the beam configuration fixed, the RIS phase vector is optimized using a least-squares combined with manifold optimization method. This alternating optimization strategy decomposes the originally complex joint optimization problem into two relatively simple subproblems, significantly reducing computational complexity. Experimental results show that the algorithm achieves a 90% speedup compared to traditional methods, reducing system processing latency from 1200ms to 19.2ms, thus meeting the requirements of real-time systems.
[0022] In this embodiment, it should be specifically noted that the main beam covers ±60 degrees, has a gain greater than or equal to 8 dBi, and is responsible for environmental perception; the sub-beams are nested within the main beam and provide directional communication services for multiple users; the beam weight vector constructed through the nested design of the main beam and sub-beams can be constructed using a 64-element non-uniform linear array, with the main beam covering ±60 degrees, a gain greater than or equal to 8 dBi, and the sub-beam communication width less than or equal to 5 degrees, supporting parallel transmission for multiple users. A 64-element ULA model can be constructed in MATLAB to verify the spatial isolation between the main beam and the sub-beams.
[0023] In conclusion, the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
[0024] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A method for joint optimization of multimodal beam and spectrum for integrated sensing and communication, characterized in that: Includes the following steps: Step 1: Construct a beam weight vector through a nested design of main beam and sub-beams; Step 2: Design a dynamic spectrum ratio function to adaptively adjust the spectrum allocation based on the sensing and detection probability; Step 3: Using communication rate and perceived mutual information as objective functions, optimize decision variables through Bayesian deep reinforcement learning; Step 4: Optimize using an alternating optimization algorithm.
2. The method for joint optimization of multimodal beam and spectrum for integrated sensing and communication as described in claim 1, characterized in that: The beam weight vector is represented as follows: ;in, This represents the total beam weight vector. This indicates the perceived main beam weight. This represents the weight of the k-th communication sub-beam. and These represent the power allocation factors for sensing and communication, respectively; by precisely controlling these weighting parameters, a balance can be achieved between sensing coverage and communication services. This indicates the total number of communication sub-beam weights; Represents communication.
3. The method for joint optimization of multimodal beam and spectrum for integrated sensing and communication as described in claim 2, characterized in that: The dynamic spectrum ratio function is expressed as: ; in, This represents the probability of perception detection. express The dynamic spectral ratio function at time t. express The dynamic spectral ratio function at time t. This represents the difference between the dynamic spectrum ratio function. This represents the maximum value of the dynamic spectrum ratio function. This represents the minimum value of the dynamic spectrum ratio function.
4. The method for joint optimization of multimodal beam and spectrum for integrated sensing and communication as described in claim 3, characterized in that: The objective function in which communication rate and perceived mutual information are used as objective functions is denoted as the joint optimization objective function, which is expressed as follows: ;in, This represents the total rate of all communication users; Represents the mutual information metric of radar; Represents the beam weighting matrix; Represents the RIS phase vector; This represents the received signal-to-interference-plus-noise ratio (SIR) of the k-th communication sub-beam; and These represent the corresponding weighting factors.
5. The multimodal beam and spectrum joint optimization method for integrated sensing and communication as described in claim 4, characterized in that: The optimization of decision variables through Bayesian deep reinforcement learning specifically involves: Bayesian deep reinforcement learning is introduced as an intelligent decision engine. The state space of the intelligent decision engine includes channel state information, target location information and historical perception error, while the action space covers beam pointing angle, spectrum allocation ratio and power allocation factor, and a reward function is set.
6. The method for joint optimization of multimodal beam and spectrum for integrated sensing and communication as described in claim 5, characterized in that: The reward function is expressed as follows: ;in, This indicates that the state is s at time t. t And take action a t The resulting immediate reward function value; Represents the state vector; Represents the action vector; Indicates the perceived mean square error; This represents the trade-off parameter.
7. The method for joint optimization of multimodal beam and spectrum for integrated sensing and communication as described in claim 6, characterized in that: The decision-making process of the intelligent decision engine can be represented as follows: Select an action based on the current state; Perform the action and observe the environmental feedback; Bayesian posterior update.
8. The method for joint optimization of multimodal beam and spectrum for integrated sensing and communication as described in claim 7, characterized in that: The method of reducing computational complexity by using an alternating optimization algorithm is as follows: a step-by-step solution strategy is adopted, specifically: First, the phase configuration of the RIS (Intelligent Reflector) is fixed, and the beam weight matrix is optimized using a semi-definite relaxation method. Then, with the beam configuration fixed, the RIS phase vector is optimized using a least-squares combined with manifold optimization method.