Maximum entropy reinforcement learning-based communication fusion network resource allocation method and device

Through maximum entropy reinforcement learning, the antenna position and power allocation are optimized, and the complexity and dynamic adaptability of resource allocation in flexible antenna-assisted synesthesia fusion network is solved, and efficient resource management is achieved in a dynamic environment.

CN120302299AActive Publication Date: 2025-07-11UNIV OF SCI & TECH BEIJING
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510313982.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-17
Publication Date
2025-07-11
Estimated Expiration
2045-03-17

AI Technical Summary

Technical Problem

Traditional fixed-position antenna designs are difficult to meet the requirements of resource sharing efficiency and dynamic environmental adaptability in synesthesia fusion networks. Flexible antenna deployment increases the communication channel changes caused by resource allocation complexity and user mobility.

Method used

Using a method based on maximum entropy reinforcement learning, the state space, action space and reward functions of the optimized target reinforcement learning network are designed to train the network to maximize cumulative discount rewards and maximize strategic entropy, and the network is updated in combination with the experience playback mechanism to achieve the optimal antenna position and power allocation.

Benefits of technology

Effectively adjust resource allocation in a dynamic time-varying environment, maintain communication quality and perceptual accuracy, adapt to user mobility and network topology changes, and improve resource utilization efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120302299A_ABST
    Figure CN120302299A_ABST
Patent Text Reader

Abstract

The invention provides a method and a device for allocating resources of a communication and sensing fusion network based on maximum entropy reinforcement learning, and the method comprises the steps: exploring a flexible antenna-assisted communication and sensing fusion network architecture which considers a scene of a plurality of pinching antennas on the same waveguide and comprehensively considers a communication data rate and a sensing demand, establishing a problem model which maximizes the communication data rate and meets the sensing requirement in the flexible antenna-assisted communication-inductance fusion network; reconstructing an objective function of the problem model, wherein the objective function comprises designing and optimizing a state space, an action space and a reward function of the objective reinforcement learning network; training an optimization target reinforcement learning network by taking the entropy of the maximum strategy while maximizing the accumulated discount rewards as a criterion, and updating an evaluation network and a strategy network of the optimization target reinforcement learning network based on an experience playback mechanism; and after the training is completed, obtaining a solution of the problem model, including an optimal antenna position and a power allocation set. According to the invention, resource allocation can be carried out on the communication sensing fusion network.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of wireless communication technologies, and particularly to a method and apparatus for resource allocation in an integrated sensing and communication network based on maximum entropy reinforcement learning. Background Art

[0002] Integrated Sensing and Communication (ISAC) technology integrates wireless communication and radar sensing functions onto the same platform, and utilizes shared resources such as hardware, spectrum, and energy to improve operational efficiency, reduce costs, and promote sustainable development. The integrated sensing and communication network realizes sensing functions such as identification, positioning, and imaging by means of wireless communication signals, and further enhances and explores potential communication capabilities by using sensing information, enabling wireless signals to not only transmit effective communication information, but also sense, detect, and characterize the physical world. Integrated sensing and communication endows wireless communication with stronger sensing capabilities and gives rise to richer application scenarios. The utilization efficiency of shared resources and the adaptability to dynamic environments pose more challenges to integrated sensing and communication systems, especially their antenna systems. On the one hand, the resource sharing between sensing and communication functions requires the use of antennas with high directivity and flexibility to reduce signal interference and improve resource utilization efficiency. On the other hand, in a complex and time-varying environment, antennas must have the ability to dynamically adapt to environmental changes to ensure communication quality and sensing accuracy. However, traditional fixed-position antenna designs often struggle to meet requirements such as the utilization efficiency of shared resources and the adaptability to dynamic environments. Therefore, it is necessary to construct an integrated sensing and communication network architecture for a flexible antenna system.

[0003] In recent years, flexible antenna systems (such as pinching antennas) have received extensive attention. The pinching antenna system creates new line-of-sight links and / or enhances existing transceiver channels by applying low-cost dielectric materials at arbitrary positions on a dielectric waveguide. Different from traditional antennas, pinching antennas can be flexibly deployed, and increasing their number hardly incurs additional costs. The pinching antenna circumvents the high costs of other flexible antennas and the difficulty of combating large-scale path loss. Its flexible radiation pattern and strong layout adaptability demonstrate broad application prospects. The flexible deployment of pinching antennas and the dynamic changes in user mobility pose more profound challenges to the resource scheduling of the integrated network. On the one hand, the flexible antenna deployment may change the network topology and increase the complexity of resource allocation; on the other hand, user mobility causes the communication channels and sensing environments to change continuously, and the resource requirements fluctuate dynamically. How to adjust resources such as spectrum, power, and time in real time to meet communication and sensing requirements, and how to maintain service continuity during user movement have become urgent problems to be solved in the resource allocation research of integrated sensing and communication networks. Summary of the Invention

[0004] To solve the above technical problems existing in the prior art, the present invention provides a method and device for allocating resources of a communication-sensing fusion network based on maximum entropy reinforcement learning. The technical solution is as follows:

[0005] On the one hand, a method for allocating resources of a communication-sensing fusion network based on maximum entropy reinforcement learning is provided. The method includes:

[0006] S1. Explore a communication-sensing fusion network architecture assisted by flexible antennas. The network architecture considers the scenario of multiple pinching antennas on the same waveguide, and comprehensively considers the communication data rate and sensing requirements, and establishes a problem model for maximizing the communication data rate while meeting the sensing requirements in the flexible antenna-assisted communication-sensing fusion network;

[0007] S2. Reconstruct the objective function of the problem model, including designing the state space, action space, and reward function of the optimization objective reinforcement learning network;

[0008] S3. Train the optimization objective reinforcement learning network with the criterion of maximizing the cumulative discounted reward while maximizing the entropy of the policy, and update the evaluation network and policy network of the optimization objective reinforcement learning network based on the experience replay mechanism;

[0009] S4. After the training is completed, obtain the solution of the problem model. The solution of the problem model includes the optimal antenna position and power allocation set.

[0010] Optionally, the S1 specifically includes:

[0011] Maximizing the communication data rate of user m is expressed as:

[0012]

[0013] where M represents the total number of users, N represents the total number of pinching antennas, represents the set of antenna positions, p = (p m ) m∈M represents the power allocation set, the power of user m is p m , represents the position of the antenna, d represents the height of the antenna, ψ m =(x m , y m , 0) represents the position of user m, θ n is the phase shift of the transmitted signal passing through antenna n, α represents the spherical wave parameter, e is the natural base, j represents the imaginary unit, π is an irrational number, λ represents the wavelength, σ 2 represents additive white Gaussian noise;

[0014] Considering the signal-to-interference-plus-noise ratio Γ of the sensing target kk , set the threshold Γ of the sensing accuracy sen , the sensing needs to satisfy Γ k -Γ sen ≥0;

[0015] Establish a problem model that maximizes the communication data rate while meeting the sensing requirements as follows:

[0016]

[0017] Optionally, the state space, action space, and reward function of the optimized target reinforcement learning network designed in S2 are specifically as follows:

[0018] State space s t ∈S, where ψ t , respectively represent the set of pinching antenna positions, the set of user positions, and the set of sensing target positions at time slot t, and E t represents the remaining energy of the system;

[0019] Action space a t ∈A, where Δψ t , respectively represent the change in pinching antenna position and the change in mobile user position, and p t represents the power allocation vector at time slot t, and ΔE t represents the energy consumed by the system at time slot t;

[0020] Reward function r t ∈R, where τ is the weight factor.

[0021] Optionally, the criterion for training the optimized target reinforcement learning network is to maximize the cumulative discounted reward while maximizing the entropy of the policy, and the criterion function is:

[0022]

[0023] where Ε(·) represents the expectation calculation, γ∈(0,1) represents the discount factor, T is the total number of training time slots, represents the policy network, and its network parameters are is the entropy of the policy, and ρ represents the temperature factor, which is used to balance the importance of the reward and the entropy.

[0024] Optionally, the training process of the optimized target reinforcement learning network is as follows:

[0025] The first stage: Initialize the policy network and the evaluation network Initialize the experience replay pool B;

[0026] The second stage:

[0027] During the environment iteration, perform action updates With probability P, perform state update s t+1 = P(s t+1 |s t , a t ), generate a new experience tuple and store it in the replay pool B ← B ∪ {(s t , a t , r t (s t , a t ), s t+1 )};

[0028] During the gradient iteration, draw a mini - batch of experience tuples B from the experience replay pool B, where the number of experience tuples can be denoted as |B|. Subsequently, update the policy network and the evaluation network using gradient descent. The update formula for the policy network is as follows:

[0029]

[0030] The evaluation network update formula is as follows:

[0031]

[0032] If the termination condition is met, the training terminates; otherwise, loop through the second stage.

[0033] On the other hand, a resource allocation device for a joint sensing and communication fusion network based on maximum - entropy reinforcement learning is provided. The device includes:

[0034] A building module for exploring a flexible antenna - assisted joint sensing and communication fusion network architecture. The network architecture considers the scenario of multiple pinching antennas on the same waveguide and comprehensively considers the communication data rate and sensing requirements, and establishes a problem model for maximizing the communication data rate while meeting the sensing requirements in a flexible antenna - assisted joint sensing and communication fusion network;

[0035] A reconstruction module for reconstructing the objective function of the problem model, including designing the state space, action space, and reward function of the optimization - objective reinforcement learning network;

[0036] A training module for training the optimization - objective reinforcement learning network with the criterion of maximizing the cumulative discounted reward while maximizing the entropy of the policy, and updating the evaluation network and policy network of the optimization - objective reinforcement learning network based on the experience replay mechanism;

[0037] An obtaining module, configured to obtain the solution of the problem model after training is completed, where the solution of the problem model includes an optimal antenna position and a power allocation set.

[0038] Optionally, the establishing module is specifically configured to:

[0039] Maximizing the communication data rate of user m is expressed as:

[0040]

[0041] where M represents the total number of users, N represents the total number of pinching antennas, represents the set of antenna positions, p = (p m ) m∈M represents the power allocation set, the power of user m is p m , represents the position of the antenna, d represents the height of the antenna, ψ m =(x m ,y m ,0) represents the position of user m, θ n is the phase shift of the transmission signal passing through antenna n, α represents the spherical wave parameter, e is the natural base, j represents the imaginary unit, π is an irrational number, λ represents the wavelength, σ 2 represents additive white Gaussian noise;

[0042] Considering the signal-to-interference-plus-noise ratio Γ k of the sensing target k, setting the threshold Γ sen of the sensing accuracy, the sensing needs to satisfy Γ k -Γ sen ≥0;

[0043] Establish the following problem model that maximizes the communication data rate and meets the sensing requirements:

[0044]

[0045] Optionally, the state space, action space, and reward function of the optimized target reinforcement learning network designed in the reconstructing module specifically include:

[0046] The state space s t ∈S, where ψ t , respectively represent the set of pinching antenna positions, the set of user positions, and the set of sensing target positions at time slot t, and E t represents the remaining energy of the system;

[0047] The action space a t ∈A, where Δψ t , respectively represent the change amount of the pinching antenna position and the change amount of the mobile user position, p t represents the power allocation vector at time slot t, ΔE t represents the energy consumed by the system at time slot t;

[0048] The reward function r t ∈R, where τ is the weight factor.

[0049] Optionally, in the training module, with the criterion of maximizing the cumulative discounted reward while maximizing the entropy of the policy, the criterion function for training the optimization objective reinforcement learning network is:

[0050]

[0051] where, Ε(·) represents the expectation calculation, γ∈(0,1) represents the discount factor, T is the total number of training time slots, represents the policy network, and its network parameters are is the entropy of the policy, and ρ represents the temperature factor, which is used to balance the importance of the reward and the entropy.

[0052] Optionally, the training process of the optimization objective reinforcement learning network is as follows:

[0053] The first stage: Initialize the policy network and the evaluation network Initialize the experience replay pool B;

[0054] The second stage:

[0055] In the environment iteration, perform action update Update the state s with probability P t+1 =P(s t+1 |s t ,a t ), generate a new experience tuple and store it in the replay pool B←B∪{(s t ,a t ,r t (s t ,a t ),s t+1 )};

[0056] In the gradient iteration, draw a small batch of experience tuples B from the experience replay pool B, where the number of experience tuples can be expressed as |B|, and then update the policy network and the evaluation network by gradient descent, where the policy network The update formula is:

[0057]

[0058] The evaluation network The update formula is as follows:

[0059]

[0060] If the termination condition is met, the training terminates; otherwise, the second stage is looped.

[0061] The beneficial effects brought by the technical solution provided by the present invention at least include:

[0062] The present invention solves the problem of resource allocation in a communication-sensing integrated network assisted by reconfigurable antennas, and is particularly applicable to resource allocation in a communication-sensing integrated network in a dynamic time-varying environment, such as a scenario with time-varying channel states brought about by mobile users and reconfigurable antenna systems. Description of the Drawings

[0063] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0064] Figure 1 It is a flowchart of a method for resource allocation in a communication-sensing integrated network based on maximum entropy reinforcement learning provided by an embodiment of the present invention;

[0065] Figure 2 It is a block diagram of a device for resource allocation in a communication-sensing integrated network based on maximum entropy reinforcement learning provided by an embodiment of the present invention. Detailed Embodiments

[0066] To make the technical problems, technical solutions, and advantages to be solved by the present invention clearer, the following will be described in detail with reference to the accompanying drawings and specific embodiments.

[0067] An embodiment of the present invention provides a method for resource allocation in a communication-sensing integrated network based on maximum entropy reinforcement learning. This method can be implemented by an electronic device, which can be a terminal or a server. Figure 1 As shown in the flowchart of this method, the processing flow can include the following steps:

[0068] S1. Explore the communication-sensing integrated network architecture assisted by reconfigurable antennas. The network architecture considers the scenario of multiple pinching antennas on the same waveguide, and comprehensively considers the communication data rate and sensing requirements, and establishes a problem model that maximizes the communication data rate in the communication-sensing integrated network assisted by reconfigurable antennas while meeting the sensing requirements;

[0069] Optionally, S1 specifically includes:

[0070] Maximizing the communication data rate of user m is expressed as:

[0071]

[0072] Among them, M represents the total number of users, N represents the total number of pinching antennas, represents the set of antenna positions, p = (p m ) m∈M represents the power allocation set, and the power of user m is p m , represents the position of the antenna, d represents the height of the antenna, ψ m = (x m , y m , 0) represents the position of user m, θ n is the phase shift of the transmitted signal passing through antenna n, α represents the spherical wave parameter, e is the natural base, j represents the imaginary unit, π is an irrational number, λ represents the wavelength, σ 2 represents additive white Gaussian noise;

[0073] Consider the signal-to-interference-plus-noise ratio Γ k of the sensing target k, set the threshold Γ sen of the sensing accuracy, and the sensing needs to satisfy Γ k - Γ sen ≥0;

[0074] Establish the following problem model that maximizes the communication data rate while meeting the sensing requirements:

[0075]

[0076] S2. Reconstruct the objective function of the problem model, including designing the state space, action space, and reward function of the optimization objective reinforcement learning network;

[0077] Reinforcement learning learns the optimal policy through the interaction between the agent and the environment, and is particularly suitable for solving dynamic optimization problems in communication and sensing integration. Its advantages include being able to adapt to environmental changes and uncertainties, being able to optimize multiple objectives simultaneously by designing appropriate reward functions, and being able to achieve real-time decision-making based on online learning or offline pre-training. The application of reinforcement learning algorithms in communication and sensing integration provides a powerful tool for solving complex problems such as resource allocation, beamforming, and environmental sensing. With the continuous development of 6G technology, reinforcement learning will play an increasingly important role in communication and sensing integration, promoting the deep integration of communication and sensing.

[0078] The state space, action space, and reward function of the optimized objective reinforcement learning network designed in the embodiments of the present invention. According to the optimization problem model, the state space includes the position of the pinching antenna, the position of the communication user, the position of the sensing target, and the total remaining energy of the system; the action space includes the position quantity of the pinching antenna, the change quantity of the mobile user position, the power allocation vector, and the energy consumption; the reward function involves the communication data rate and the sensing accuracy, specifically as follows:

[0079] The state space, action space, and reward function of the optimized objective reinforcement learning network designed in S2 specifically include:

[0080] State space s t ∈S, where ψ t , respectively represent the set of pinching antenna positions, the set of user positions, and the set of sensing target positions at time slot t, and E t represents the remaining energy of the system;

[0081] Action space a t ∈A, where Δψ t , respectively represent the change in the position of the pinching antenna and the change in the position of the mobile user, p t represents the power allocation vector at time slot t, and ΔE t represents the energy consumed by the system at time slot t;

[0082] Reward function r t ∈R, where τ is the weight factor.

[0083] S3. With the criterion of maximizing the cumulative discounted reward while maximizing the entropy of the policy (this design idea aims to improve the exploration ability and robustness of the reinforcement learning algorithm), train the optimized objective reinforcement learning network, and update the evaluation network and policy network of the optimized objective reinforcement learning network based on the experience replay mechanism;

[0084] Optionally, the criterion function for training the optimized objective reinforcement learning network with the criterion of maximizing the cumulative discounted reward while maximizing the entropy of the policy in S3 is:

[0085]

[0086] where Ε(·) represents the expectation calculation, γ ∈ (0, 1) represents the discount factor, T is the total number of training time slots, represents the policy network, and its network parameters are is the entropy of the policy, and ρ represents the temperature factor, which is used to balance the importance of the reward and the entropy.

[0087] The experience replay mechanism updates network parameters by storing and reusing past experiences. During the algorithm training process, the data (s t , a t , r t , s t+1 ) generated by the interaction between the agent and the environment will be stored in a fixed-size experience pool B. The algorithm randomly samples a batch of data B from the buffer B to update the network parameters, rather than directly using the latest data. The main advantages of the experience replay mechanism are reusing historical data, reducing the need for environmental interaction; random sampling reduces data correlation and training variance; it supports offline learning and can be combined with data generated by other strategies.

[0088] Optionally, the training process of the optimization target reinforcement learning network is as follows:

[0089] The first stage: Initialize the policy network and the evaluation network Initialize the experience replay pool B;

[0090] The second stage:

[0091] In the environmental iteration, perform action updates Update the state s with probability P, where P(s t+1 = P(s t+1 | s t , a t ), generate new experience tuples and store them in the replay pool B ← B ∪ {(s t , a t , r t (s t , a t ), s t+1 )};

[0092] In the gradient iteration, draw a small batch of experience tuples B from the experience replay pool B, where the number of experience tuples can be expressed as |B|. Subsequently, update the policy network and the evaluation network using gradient descent, where the policy network The update formula is:

[0093]

[0094] The evaluation network The update formula is:

[0095]

[0096] If the termination condition (such as relative convergence) is met, the training terminates; otherwise, loop back to the second stage.

[0097] S4. After the training is completed, obtain the solution of the problem model, where the solution of the problem model includes the optimal antenna positions and power allocation set.

[0098] After the training of the optimization objective reinforcement learning network according to the embodiment of the present invention is completed, problem solving is realized, and the policy network outputs the optimal antenna positions and power allocation set.

[0099] As Figure 2 shown, the embodiment of the present invention further provides a resource allocation device for a communication and sensing fusion network based on maximum entropy reinforcement learning. The device includes:

[0100] A building module 210, configured to explore a communication and sensing fusion network architecture assisted by flexible antennas. The network architecture considers the scenario of multiple pinching antennas on the same waveguide, and comprehensively considers the communication data rate and sensing requirements, and establishes a problem model for maximizing the communication data rate while meeting the sensing requirements in the flexible antenna-assisted communication and sensing fusion network;

[0101] A reconstruction module 220, configured to reconstruct the objective function of the problem model, including designing the state space, action space, and reward function of the optimization objective reinforcement learning network;

[0102] A training module 230, configured to train the optimization objective reinforcement learning network with the criterion of maximizing the cumulative discounted reward and maximizing the entropy of the policy at the same time, and update the evaluation network and policy network of the optimization objective reinforcement learning network based on the experience replay mechanism;

[0103] An obtaining module 240, configured to obtain the solution of the problem model after the training is completed, where the solution of the problem model includes the optimal antenna positions and power allocation set.

[0104] Optionally, the building module is specifically configured to:

[0105] Maximizing the communication data rate of user m is expressed as:

[0106]

[0107] where M represents the total number of users, N represents the total number of pinching antennas, represents the set of antenna positions, p = (p m ) m∈M represents the power allocation set, the power of user m is p m , represents the position of the antenna, d represents the height of the antenna, ψ m = (x m , y m , 0) represents the position of user m, θ nFor the phase shift of the transmission signal passing through antenna n, α represents the spherical wave parameter, e is the base of the natural logarithm, j represents the imaginary unit, π is an irrational number, λ represents the wavelength, and σ 2 represents additive white Gaussian noise;

[0108] Consider the signal-to-interference-plus-noise ratio Γ of the sensing target k k , set the threshold Γ of the sensing accuracy sen , and the sensing needs to satisfy Γ k -Γ sen ≥0;

[0109] Establish the following problem model that maximizes the communication data rate while meeting the sensing requirements:

[0110]

[0111] Optionally, the state space, action space, and reward function of the optimized target reinforcement learning network designed in the reconstruction module are specifically as follows:

[0112] State space s t ∈S, where ψ t , respectively represent the set of pinching antenna positions, the set of user positions, and the set of sensing target positions at time slot t, and E t represents the remaining energy of the system;

[0113] Action space a t ∈A, where Δψ t , respectively represent the change in pinching antenna position and the change in mobile user position, and p t represents the power allocation vector at time slot t, and ΔE t represents the energy consumed by the system at time slot t;

[0114] Reward function r t ∈R, where τ is the weight factor.

[0115] Optionally, in the training module, with the criterion of maximizing the cumulative discounted reward while maximizing the entropy of the policy, the criterion function for training the optimized target reinforcement learning network is:

[0116]

[0117] where Ε(·) represents the expectation calculation, γ∈(0,1) represents the discount factor, T is the total number of training time slots, represents the policy network, and its network parameters are Here, \(H\) is the entropy of the policy, and \(\rho\) represents the temperature factor, which is used to balance the importance of the reward and the entropy.

[0118] Optionally, the training process of the optimization objective reinforcement learning network is as follows:

[0119] The first stage: Initialize the policy network and the evaluation network Initialize the experience replay pool \(B\);

[0120] The second stage:

[0121] In the environment iteration, perform action updates Perform state update \(s'\) with probability \(P\), where \(P(s'|s,a)\) t+1 \(= P(s' t+1 |s t ,a t ), and generate a new experience tuple and store it in the replay pool \(B \leftarrow B\cup\{(s t ,a t ,r t (s t ,a t ),s t+1 )\};

[0122] In the gradient iteration, sample a mini-batch of experience tuples \(B\) from the experience replay pool \(B\), where the number of experience tuples can be denoted as \(|B|\). Subsequently, update the policy network and the evaluation network using gradient descent. The update formula for the policy network is as follows:

[0123]

[0124] The update formula for the evaluation network is as follows:

[0125]

[0126] If the termination condition is met, the training terminates; otherwise, loop back to the second stage.

[0127] A resource allocation device for a communication-sensing fusion network based on maximum entropy reinforcement learning provided by an embodiment of the present invention has a functional structure corresponding to a resource allocation method for a communication-sensing fusion network based on maximum entropy reinforcement learning provided by an embodiment of the present invention, and will not be elaborated here.

[0128] The above are only the preferred embodiments of the present invention, and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for allocating resources of a cross-sensory fusion network based on maximum entropy reinforcement learning, characterized in that The method includes: S1. Explore a flexible antenna - assisted communication - sensing fusion network architecture. The network architecture considers the scenario of multiple pinching antennas on the same waveguide and comprehensively considers the communication data rate and sensing requirements, and establishes a problem model for maximizing the communication data rate while meeting the sensing requirements in the flexible antenna - assisted communication - sensing fusion network; S2. Reconstruct the objective function of the problem model, including designing the state space, action space, and reward function of the optimization - objective reinforcement learning network; S3. With the criterion of maximizing the cumulative discounted reward and maximizing the entropy of the policy, train the optimization - objective reinforcement learning network, and update the evaluation network and policy network of the optimization - objective reinforcement learning network based on the experience replay mechanism; S4. After the training is completed, obtain the solution of the problem model. The solution of the problem model includes the optimal antenna positions and power allocation sets.

2. The method according to claim 1, characterized in that, The specific content of S1 includes: Maximizing the communication data rate of user m is expressed as: Where, M represents the total number of users, N represents the total number of pinching antennas, represents the set of antenna positions, p = (p m ) m∈M represents the power allocation set, and the power of user m is p m , represents the position of the antenna, d represents the height of the antenna, ψ m = (x m , y m , 0) represents the position of user m, θ n is the phase shift of the transmitted signal passing through antenna n, α represents the spherical wave parameter, e is the natural base, j represents the imaginary unit, π is an irrational number, λ represents the wavelength, σ 2 represents additive white Gaussian noise; Consider the signal-to-interference-plus-noise ratio Γ of the sensing target k k , set the threshold Γ of the sensing accuracy sen , the sensing needs to satisfy Γ k -Γ sen ≥0; Establish the following problem model for maximizing the communication data rate while meeting the sensing requirements:

3. The method according to claim 2, wherein The state space, action space, and reward function of the optimization - objective reinforcement learning network designed in S2 specifically include: State space s t ∈S, where represent the set of pinching antenna positions, the set of user positions, and the set of sensing target positions at time slot t, respectively, and E t represents the remaining energy of the system; Action space a t ∈A, where respectively represent the change in the position of the pinching antenna and the change in the position of the mobile user, p t represents the power allocation vector at time slot t, ΔE t represents the energy consumed by the system at time slot t; Reward function r t ∈R, where τ is the weight factor.

4. The method according to claim 3, characterized in that The criterion function for training the optimization - objective reinforcement learning network in S3 with the criterion of maximizing the cumulative discounted reward and maximizing the entropy of the policy is: where Ε(·) represents the calculation of expectation, γ∈(0,1) represents the discount factor, T is the total number of training time slots, denotes the policy network, and its network parameters are is the entropy of the policy, and ρ represents the temperature factor, which is used to balance the importance of rewards and entropy.

5. The method according to claim 4, characterized in that The training process of the optimization - objective reinforcement learning network is: The first stage: Initialize the policy network and the evaluation network Initialize the experience replay pool B; The second stage: During the environment iteration, perform action updates With probability P, perform state update s t+1 = P(s t+1 |s t , a t ), generate a new experience tuple and store it in the replay pool B ← B ∪ {(s t , a t , r t (s t , a t ), s t+1 )}; In gradient iteration, a small batch of experience tuples B is sampled from the experience replay pool B, where the number of experience tuples can be denoted as |B|. Subsequently, the policy network and the evaluation network are updated using gradient descent, where the policy network The update formula is as follows: Evaluation network The update formula is as follows: If the termination condition is met, the training terminates; otherwise, loop through the second stage.

6. A resource allocation device for a cross-sensory fusion network based on maximum entropy reinforcement learning, characterized in that The device includes: A building module, used to explore a flexible antenna - assisted communication - sensing fusion network architecture. The network architecture considers the scenario of multiple pinching antennas on the same waveguide and comprehensively considers the communication data rate and sensing requirements, and establishes a problem model for maximizing the communication data rate while meeting the sensing requirements in the flexible antenna - assisted communication - sensing fusion network; A reconstruction module, used to reconstruct the objective function of the problem model, including designing the state space, action space, and reward function of the optimization - objective reinforcement learning network; A training module, used to train the optimization - objective reinforcement learning network with the criterion of maximizing the cumulative discounted reward and maximizing the entropy of the policy, and update the evaluation network and policy network of the optimization - objective reinforcement learning network based on the experience replay mechanism; An obtaining module, used to obtain the solution of the problem model after the training is completed. The solution of the problem model includes the optimal antenna positions and power allocation sets.

7. The device according to claim 6, characterized in that, The building module is specifically used for: Maximizing the communication data rate of user m is expressed as: Among them, M represents the total number of users, and N represents the total number of pinching antennas. represents the set of antenna positions, p = (p m ) m∈M represents the power distribution set, and the power of user m is p m , represents the position of the antenna, d represents the height of the antenna, ψ m = (x m , y m , 0) represents the position of user m, θ n is the phase shift of the transmission signal passing through antenna n, α represents the spherical wave parameter, e is the natural base, j represents the imaginary unit, π is an irrational number, λ represents the wavelength, σ 2 represents additive white Gaussian noise; Consider the signal-to-interference-plus-noise ratio Γ of the sensing target k k , set the threshold Γ of the sensing accuracy sen , the sensing needs to satisfy Γ k -Γ sen ≥0; Establish the following problem model for maximizing the communication data rate while meeting the sensing requirements:

8. The device according to claim 7, characterized in that, The state space, action space, and reward function of the optimization - objective reinforcement learning network designed in the reconstruction module specifically include: State space s t ∈S, where, respectively represent the set of pinching antenna positions, the set of user positions, and the set of sensed target positions at time slot t, and E t represents the remaining energy of the system; Action space a t ∈A, where respectively represent the change in the position of the pinching antenna and the change in the position of the mobile user, p t represents the power allocation vector at time slot t, ΔE t represents the energy consumed by the system at time slot t; Reward function r t ∈R, where τ is a weighting factor.

9. The device according to claim 8, characterized in that, The criterion function for training the optimization - objective reinforcement learning network in the training module with the criterion of maximizing the cumulative discounted reward and maximizing the entropy of the policy is: Among them, Ε(·) represents the expectation calculation, γ∈(0,1) represents the discount factor, T is the total number of training time slots, represents the policy network, and its network parameters are is the entropy of the policy, ρ represents the temperature factor, which is used to balance the importance of rewards and entropy.

10. The device according to claim 9, wherein, The training process of the optimization - objective reinforcement learning network is: The first stage: Initialize the policy network and the evaluation network Initialize the experience replay pool B; The second stage: During the environment iteration, perform action update Perform state update s with probability P t+1 = P(s t+1 |s t ,a t ), generate a new experience tuple and store it in the replay pool B ← B ∪ {(s t ,a t ,r t (s t ,a t ),s t+1 )}; In gradient iteration, a small batch of experience tuples B is sampled from the experience replay pool B, where the number of experience tuples can be expressed as |B|. Subsequently, the policy network and the evaluation network are updated using gradient descent, where the policy network The update formula is as follows: Evaluation network The update formula is as follows: If the termination condition is met, the training terminates; otherwise, loop through the second stage.

Citation Information

Patent Citations

  • Knowledge migration reinforcement learning network slice general calculation resource collaborative optimization method

    CN114615744A

  • Unmanned aerial vehicle intelligent trajectory planning and communication resource allocation method based on reinforcement learning

    CN116704823A

  • Cellular-free sensing fusion site planning method and system based on flexible strategy evaluation network

    CN118555576A

  • Sensitivity integration method based on mobile antenna

    CN118801934A

  • Task scheduling and resource allocation method in unmanned aerial vehicle assisted communication calculation system

    CN119047697A