A Synesthetic User Clustering and Resource Allocation Method Based on Deep Reinforcement Learning

By combining deep reinforcement learning methods with Actor-Critic and flexible Actor-Critic frameworks, user clustering and resource allocation in flexible antenna systems are optimized, solving the optimization problem of discrete-continuous action space and realizing efficient resource scheduling of the integrated sensing network in dynamic environments.

CN120416875BActive Publication Date: 2026-03-10UNIV OF SCI & TECH BEIJING
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-14
Publication Date
2026-03-10

AI Technical Summary

Technical Problem

Existing deep reinforcement learning frameworks struggle to handle optimization problems in discrete-continuous action spaces, leading to increased resource scheduling complexity for flexible antenna systems in integrated sensing networks and an inability to effectively optimize communication channel changes caused by user mobility.

Method used

We employ a deep reinforcement learning-based approach, combining the Actor-Critic framework and the flexible Actor-Critic framework, to optimize the discrete-continuous mixed variable problem of user clustering, antenna location, and power allocation. By establishing a flexible antenna link channel model and transmission signal model, we construct an optimization objective function and design a reward function to optimize the action space.

Benefits of technology

It effectively solves the problems of user clustering and resource allocation in flexible antenna-assisted integrated communication and sensing networks, and is particularly suitable for dynamic and time-varying environments, improving the efficiency and adaptability of resource allocation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120416875B_ABST
    Figure CN120416875B_ABST
Patent Text Reader

Abstract

This invention provides a method for user clustering and resource allocation in a sensor-integrated network based on deep reinforcement learning, relating to the field of wireless communication technology. The method includes: establishing a flexible antenna link channel model and a transmission signal model; constructing an optimization problem model that maximizes the communication data rate and sensing detection power based on the flexible antenna link channel model and the transmission signal model; designing an optimization objective function based on the optimization problem model; the optimization objective function includes a state space, an action space, and a reward function; and, based on the optimization objective function, employing a method combining an Actor-Critic framework and a flexible Actor-Critic framework to optimize the discrete-continuous hybrid variable problem of user clustering, antenna position, intra-cluster power allocation parameters, and inter-cluster power allocation in the action space, thereby obtaining the optimal user clustering, antenna position, and power allocation strategies. This invention can solve the problem of user clustering and resource allocation in a sensor-integrated non-orthogonal access network with flexible antenna assistance.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of wireless communication, in particular to a sensing-integrated user clustering and resource allocation method and device based on deep reinforcement learning. BACKGROUND

[0002] Sensing-integrated technology deeply integrates wireless communication and sensing functions into the same platform, utilizes hardware, spectrum and energy sharing resources, improves operational efficiency, reduces costs and promotes sustainable development. Sensing-integrated networks use communication signals to achieve identification, positioning and imaging sensing functions, and use sensing information to further enhance and exploit potential communication capabilities, so that wireless signals not only achieve the transmission of effective communication information, but also can sense, detect and characterize the physical world. Sensing-integrated gives wireless communication stronger sensing capabilities and gives birth to more abundant application scenarios. The utilization efficiency of shared resources and the adaptability of dynamic environments bring more challenges to sensing-integrated systems, especially their antenna systems. However, traditional fixed-position antenna designs often fail to meet the requirements of utilization efficiency of shared resources and adaptability of dynamic environments, so it is necessary to build a sensing-integrated network architecture for flexible antenna systems.

[0003] In recent years, flexible antenna systems such as pinching antennas have received attention from academia. They create new line-of-sight links and / or enhance existing transceiver channels by applying low-cost dielectric materials at arbitrary positions on the waveguide. Unlike traditional antennas, flexible antennas can be deployed flexibly, and increasing their number will not incur additional costs. Flexible antennas avoid the high cost of other flexible antennas and the difficulty of combating large-scale path loss problems. Their flexible radiation patterns and strong adaptability of layout capabilities show broad application prospects. The flexible deployment of flexible antennas and the dynamic changes in user mobility bring deeper challenges to the resource scheduling of integrated networks. On the one hand, flexible antenna deployment can change the network topology, increasing the complexity of resource allocation; on the other hand, user mobility causes communication channels and sensing environments to change constantly, and resource demand fluctuates dynamically. How to adjust spectrum, power and time resources in real time to meet communication and sensing needs, and how to maintain service continuity during user movement, have become difficult problems to be solved in the research of resource scheduling of sensing-integrated networks.

[0004] Deep reinforcement learning learns optimal strategies through the interaction between agents and the environment, making it particularly suitable for solving dynamic optimization problems in sensory integration. Its advantages include adaptability to environmental changes and uncertainties, the ability to simultaneously optimize multiple objectives by designing appropriate reward functions, and the capacity for real-time decision-making based on online learning or offline pre-training. The application of deep reinforcement learning algorithms in sensory integration provides a powerful tool for solving complex problems such as resource allocation, beamforming, and environmental perception. However, existing deep reinforcement learning frameworks typically only handle completely discrete or completely continuous action spaces. For optimization problems in discrete-continuous action spaces, a redesign of the algorithm optimization framework is necessary. Summary of the Invention

[0005] To address the limitation that existing deep reinforcement learning frameworks can typically only handle completely discrete or completely continuous action spaces, and the need to redesign algorithmic optimization frameworks for discrete-continuous action space optimization, this invention provides a method and apparatus for synesthetic user clustering and resource allocation based on deep reinforcement learning. The technical solution is as follows:

[0006] On the one hand, a method for synesthetic user clustering and resource allocation based on deep reinforcement learning is provided. This method is implemented by a synesthetic user clustering and resource allocation device based on deep reinforcement learning, and includes:

[0007] S1. Establish a flexible antenna link channel model and a transmission signal model; based on the flexible antenna link channel model and the transmission signal model, construct an optimization problem model that maximizes the communication data rate and sensing detection power;

[0008] S2. Based on the optimization problem model, design an optimization objective function; the optimization objective function includes: state space, action space, and reward function;

[0009] S3. Based on the state space, action space, and reward function, a method combining the Actor-Critic framework and the flexible Actor-Critic framework is used to optimize the discrete-continuous mixed variable problem of user clustering, antenna position, intra-cluster power allocation parameters, and inter-cluster power allocation in the action space, so as to obtain the optimal user clustering, antenna position, and power allocation strategy.

[0010] Optionally, the state space includes the location of the flexible antenna, the location of the communication user, the location of the sensing target, and the total remaining energy of the system.

[0011] The action space includes: the position change of the flexible antenna, user clustering variables, intra-cluster power allocation parameters, inter-cluster power allocation, the position change of the sensed target, and energy consumption.

[0012] The reward function includes: a reward function for communication data rate and a reward function for sensing and detection power.

[0013] Optionally, S1 involves establishing a flexible antenna link channel model and a transmission signal model; based on the flexible antenna link channel model and the transmission signal model, constructing an optimization problem model that maximizes the communication data rate and sensing detection power, including:

[0014] S11. Based on the integrated sensing network model of multiple flexible antennas and multiple sensing targets, the flexible antennas adopt non-orthogonal multiple access technology to divide multiple users into... Clustering; By utilizing user clustering variables, the relationship between flexible antennas and users is defined, and users are divided into several clusters;

[0015] S12. Define binary variables and set the positions of the antenna and each user; construct the antenna-to-user channel model based on the spherical wave channel model.

[0016] S13. Based on the signal transmission methods of all users in non-orthogonal multiple access, obtain... Cluster user signals;

[0017] S14. Based on the serial interference cancellation rules for non-orthogonal multiple access, obtain the data rates of weak users and strong users.

[0018] S15. Based on the data rates of weak users and strong users, obtain the data rate of the user cluster; based on the data rate of the user cluster, construct an optimization problem model that maximizes the communication data rate and the sensing detection power.

[0019] Optionally, the The cluster user's signal is represented by the following formula (1):

[0020] (1)

[0021] Where x represents Cluster user signals; This represents the total number of user clusters; Indicates the power allocation parameters within the first cluster; Indicates the power allocation parameters within the second cluster; , , Indicates that it is sent to the user communication signal.

[0022] Optionally, the data rate of the weak user is represented by the following formula (2):

[0023] (2)

[0024] in, Indicates the data rate for weak users; Represents user clusters The allocated power; Indicates that by antenna service clusters Weak users in Power allocation parameters; Indicates the channel gain for weak users; This indicates that the superimposed signal passes through the antenna. Phase shift; This represents intra-cluster interference and inter-cluster interference for weak users; This indicates inter-cluster interference from strong users; This represents additive white Gaussian noise; Indicates antenna with cluster Medium and weak users The correlation between them, and ,when When, it indicates a cluster Weak users By antenna Provide services, otherwise ;

[0025] The data rate of the strong user is expressed by the following formula (3):

[0026] (3)

[0027] in, Indicates the data rate of strong users; Indicates that by antenna service clusters Power allocation parameters for strong users s in the data; Indicates the channel gain for strong users; This indicates inter-cluster interference from strong users; Indicates antenna with cluster Medium-strength users The correlation between them, and ,when When, it indicates a cluster strong users By antenna Provide services, otherwise .

[0028] Optionally, the optimization problem model for maximizing the communication data rate and sensing detection power is expressed by the following formula (4):

[0029] (4)

[0030] in, The power of the detection signal indicating the direction of the target being sensed; Regularization parameters representing communication; Represents the regularization parameter for perception; Represents user clusters The data rate; C represents the total number of user clusters; K represents the total number of sensory network models of the sensing target; α represents the user clustering variable; represents the intra-cluster power allocation parameter vector; p represents the inter-cluster power allocation vector.

[0031] Optionally, S3 employs a method combining the Actor-Critic framework and the flexible Actor-Critic framework to optimize the discrete-continuous hybrid variable problem of user clustering, antenna location, intra-cluster power allocation parameters, and inter-cluster power allocation in the action space, thereby obtaining the optimal user clustering, antenna location, and power allocation strategy, including:

[0032] S31. Initialize discrete user clustering variables, continuously update the policy network of the Actor-Critic framework, obtain the optimal user clustering variables, and store the optimal user clustering variables in the state space.

[0033] S32. Based on the current state, the policy network of the Actor-Critic framework inputs actions into the environment, generates rewards, and produces new states;

[0034] S33. Based on the generated rewards, the evaluation network of the flexible Actor-Critic framework generates rewards and new states to update the evaluation strategy, and uses the minimum value of the two main evaluation networks as the output of the evaluation network of the flexible Actor-Critic framework.

[0035] S34. Based on the actions output by the policy network of the flexible Actor-Critic framework, obtain new states, actions, and rewards; merge the new states, actions, and rewards into experience tuples and store them in the experience replay pool; extract small batches of experience tuples to update the policy network and the main evaluation network of the flexible Actor-Critic framework.

[0036] S35. Based on the soft update strategy, the target evaluation network of the flexible Actor-Critic framework is updated.

[0037] S36. Iterate through steps S31-S35 until the cumulative discount reward converges, completing the training and obtaining the trained flexible Actor-Critic framework policy network. Based on the trained policy network, obtain the optimal user clustering, antenna location, and power allocation policies.

[0038] On the other hand, a sensory integration user clustering and resource allocation device based on deep reinforcement learning is provided. This device is applied to a sensory integration user clustering and resource allocation method based on deep reinforcement learning. The device includes:

[0039] The construction unit is used to establish a flexible antenna link channel model and a transmission signal model; based on the flexible antenna link channel model and the transmission signal model, an optimization problem model is constructed to maximize the communication data rate and the sensing detection power.

[0040] The design unit is used to design an optimization objective function based on the optimization problem model; the optimization objective function includes: a state space, an action space, and a reward function;

[0041] The optimization unit is used to optimize the discrete-continuous mixed variable problem of user clustering, antenna position, intra-cluster power allocation parameters and inter-cluster power allocation in the action space based on the state space, action space and reward function, using a method combining the Actor-Critic framework and the flexible Actor-Critic framework, to obtain the optimal user clustering, antenna position and power allocation strategy.

[0042] Optionally, the construction unit is used for:

[0043] Based on a sensing-integrated network model with multiple flexible antennas and multiple sensing targets, the flexible antennas employ non-orthogonal multiple access technology to divide multiple users into... Clustering; By utilizing user clustering variables, the relationship between flexible antennas and users is defined, and users are divided into several clusters;

[0044] Define binary variables and set the positions of the antenna and each user; construct the antenna-to-user channel model based on the spherical wave channel model;

[0045] Based on the signal transmission methods of all users in non-orthogonal multiple access, obtain Cluster user signals;

[0046] Based on the serial interference cancellation rules of non-orthogonal multiple access, the data rates of weak users and strong users are obtained.

[0047] Based on the data rates of weak users and strong users, the data rate of the user cluster is obtained; based on the data rate of the user cluster, an optimization problem model is constructed to maximize the communication data rate and the sensing detection power.

[0048] Optionally, the The cluster user's signal is represented by the following formula (1):

[0049] (1)

[0050] Where x represents Cluster user signals; This represents the total number of user clusters; Indicates the power allocation parameters within the first cluster; Indicates the power allocation parameters within the second cluster; , , Indicates that it is sent to the user communication signal.

[0051] Optionally, the data rate of the weak user is represented by the following formula (2):

[0052] (2)

[0053] in, Indicates the data rate for weak users; Represents user clusters The allocated power; Indicates that by antenna service clusters Weak users in Power allocation parameters; Indicates the channel gain for weak users; This indicates that the superimposed signal passes through the antenna. Phase shift; This represents intra-cluster interference and inter-cluster interference for weak users; This indicates inter-cluster interference from strong users; This represents additive white Gaussian noise; Indicates antenna with cluster Medium and weak users The correlation between them, and ,when When, it indicates a cluster Weak users By antenna Provide services, otherwise ;

[0054] The data rate of the strong user is expressed by the following formula (3):

[0055] (3)

[0056] in, Indicates the data rate of strong users; Indicates that by antenna service clusters Power allocation parameters for strong users s in the data; Indicates the channel gain for strong users; This indicates inter-cluster interference from strong users; Indicates antenna with cluster Medium-strength users The correlation between them, and ,when When, it indicates a cluster strong users By antenna Provide services, otherwise .

[0057] Optionally, the optimization problem model for maximizing the communication data rate and sensing detection power is expressed by the following formula (4):

[0058] (4)

[0059] in, The power of the detection signal indicating the direction of the target being sensed; Regularization parameters representing communication; Represents the regularization parameter for perception; Represents user clusters The data rate; C represents the total number of user clusters; K represents the total number of sensory network models of the sensing target; α represents the user clustering variable; represents the intra-cluster power allocation parameter vector; p represents the inter-cluster power allocation vector.

[0060] Optionally, the optimization unit is used for:

[0061] (1) Initialize discrete user clustering variables, continuously update the policy network of the Actor-Critic framework, obtain the optimal user clustering variables, and store the optimal user clustering variables in the state space;

[0062] (2) Based on the current state, the policy network of the Actor-Critic framework inputs actions into the environment, generates rewards, and produces new states;

[0063] (3) Based on the generated reward, the evaluation network of the flexible Actor-Critic framework generates a reward and a new state to update the evaluation strategy, and takes the minimum value of the two main evaluation networks as the output of the evaluation network of the flexible Actor-Critic framework.

[0064] (4) Based on the actions output by the policy network of the flexible Actor-Critic framework, obtain new states, actions and rewards; merge the new states, actions and rewards into experience tuples and store them in the experience replay pool; extract small batches of experience tuples to update the policy network and the main evaluation network of the flexible Actor-Critic framework.

[0065] (5) Based on the soft update strategy, the target evaluation network of the flexible Actor-Critic framework is updated;

[0066] (6) Iterate through steps (1)-(5) until the cumulative discount reward converges, complete the training, and obtain the policy network of the trained flexible Actor-Critic framework; based on the trained policy network, obtain the optimal user clustering, antenna location and power allocation policy.

[0067] On the other hand, a sensory integration user clustering and resource allocation device based on deep reinforcement learning is provided. The sensory integration user clustering and resource allocation device based on deep reinforcement learning includes: a processor; a memory, wherein the memory stores computer-readable instructions, and when the computer-readable instructions are executed by the processor, they implement any of the methods described above for sensory integration user clustering and resource allocation based on deep reinforcement learning.

[0068] On the other hand, a computer-readable storage medium is provided, wherein at least one instruction is stored in the storage medium, the at least one instruction being loaded and executed by a processor to implement any of the above-described methods of the synesthetic user clustering and resource allocation method based on deep reinforcement learning.

[0069] The beneficial effects of the technical solutions provided in the embodiments of the present invention include at least the following:

[0070] This invention first establishes a flexible antenna link channel model and a transmission signal model. Second, based on these models, it constructs an optimization problem model to maximize communication data rate and sensing detection power. Based on this model, an optimization objective function is designed, comprising a state space, an action space, and a reward function. Finally, based on the state space, action space, and reward function, a method combining an Actor-Critic framework and a flexible Actor-Critic framework is used to optimize the discrete-continuous hybrid variable problem of user clustering, antenna location, intra-cluster power allocation parameters, and inter-cluster power allocation in the action space, obtaining the optimal user clustering, antenna location, and power allocation strategies. This invention can solve the user clustering and resource allocation problem in flexible antenna-assisted inductive-sensing integrated non-orthogonal access networks, and is particularly suitable for inductive-sensing integrated network optimization problems in discrete-continuous variable spaces and dynamic time-varying environments. Attached Figure Description

[0071] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0072] Figure 1 This is a flowchart of a synesthetic user clustering and resource allocation method based on deep reinforcement learning provided by an embodiment of the present invention;

[0073] Figure 2 This is a block diagram of a synesthetic user clustering and resource allocation algorithm based on deep reinforcement learning provided in an embodiment of the present invention;

[0074] Figure 3 This is a block diagram of a synesthetic user clustering and resource allocation device based on deep reinforcement learning provided in an embodiment of the present invention;

[0075] Figure 4 This is a schematic diagram of the structure of a synesthetic user clustering and resource allocation device based on deep reinforcement learning, provided in an embodiment of the present invention. Detailed Implementation

[0076] The technical solution of the present invention will now be described with reference to the accompanying drawings.

[0077] In embodiments of the present invention, words such as "exemplarily," "for example," etc., are used to indicate that something is an example, illustration, or description. Any embodiment or design described as "exemplary" in the present invention should not be construed as being more preferred or advantageous than other embodiments or designs. Specifically, the use of the word "exemplary" is intended to present the concept in a concrete manner. Furthermore, in embodiments of the present invention, the meaning expressed by "and / or" can be both, or either one.

[0078] In the embodiments of this invention, the terms "image" and "picture" may sometimes be used interchangeably. It should be noted that, without emphasizing the distinction between them, they convey the same meaning. Similarly, the terms "of," "corresponding (relevant)," and "corresponding" may sometimes be used interchangeably. It should be noted that, without emphasizing the distinction between them, they convey the same meaning.

[0079] In this embodiment of the invention, sometimes a subscript such as W1 may be written in a non-subscript form such as W1. When the difference is not emphasized, the meaning they express is the same.

[0080] To make the technical problems, technical solutions and advantages of the present invention clearer, a detailed description will be given below in conjunction with the accompanying drawings and specific embodiments.

[0081] This invention provides a method for synesthetic user clustering and resource allocation based on deep reinforcement learning. This method can be implemented by a synesthetic user clustering and resource allocation device based on deep reinforcement learning, which can be a terminal or a server. Figure 1 The flowchart shown is for a synesthetic user clustering and resource allocation method based on deep reinforcement learning. The processing flow of this method may include the following steps:

[0082] S1. Establish a flexible antenna link channel model and a transmission signal model; based on the flexible antenna link channel model and the transmission signal model, construct an optimization problem model that maximizes the communication data rate and sensing detection power.

[0083] Specifically, for scenarios with multiple flexible antennas, flexible antenna link channel models and transmission signal models are established.

[0084] Optionally, the specific implementation process of S1 includes S11-S15:

[0085] S11. Based on the integrated sensing network model of multiple flexible antennas and multiple sensing targets, the flexible antennas adopt non-orthogonal multiple access technology to divide multiple users into... Clustering; By utilizing user clustering variables, the relationship between flexible antennas and users is defined, and users are divided into several clusters;

[0086] Each cluster is associated with a flexible antenna.

[0087] S12. Define binary variables and set the positions of the antenna and each user; construct the antenna-to-user channel model based on the spherical wave channel model.

[0088] Wherein, the binary variable is represented as , among which, when , indicating the first No. 1 in the cluster Each user is equipped with a flexible antenna. Serve, This indicates that there is no service association.

[0089] Each user cluster contains both strong and weak users.

[0090] In one feasible implementation, the antenna To users The channel model is expressed by the following formula (1):

[0091] (1)

[0092] in, s represents a strong user index; w represents a weak user index; Represents the parameters of a spherical wave; Represents the natural base; Represents the imaginary unit; Represent irrational numbers; Indicates wavelength. Represents norm operations, Indicates antenna Transmission power; Indicates antenna Location; Indicates the first Users in the cluster The location.

[0093] S13. Based on the signal transmission methods of all users in non-orthogonal multiple access, obtain... Cluster user signals;

[0094] Optionally, The signal of a cluster user is represented by the following formula (2):

[0095] (2)

[0096] Where x represents Cluster user signals; This represents the total number of user clusters; Indicates the power allocation parameters within the first cluster; Indicates the power allocation parameters within the second cluster; , , Indicates that it is sent to the user communication signal.

[0097] S14. Based on the serial interference cancellation rules for non-orthogonal multiple access, obtain the data rates of weak users and strong users.

[0098] Alternatively, the data rate for weak users can be expressed by the following formula (3):

[0099] (3)

[0100] in, Indicates the data rate for weak users; Represents user clusters The allocated power; Indicates that by antenna service clusters Weak users in Power allocation parameters; Indicates the channel gain for weak users; This indicates that the superimposed signal passes through the antenna. Phase shift; This represents intra-cluster interference and inter-cluster interference for weak users; This indicates inter-cluster interference from strong users; This represents additive white Gaussian noise; Indicates antenna with cluster Medium and weak users The correlation between them, and ,when When, it indicates a cluster Weak users By antenna Provide services, otherwise ;

[0101] The data rate of the strong user is expressed by the following formula (4):

[0102] (4)

[0103] in, Indicates the data rate of strong users; Indicates that by antenna service clusters Power allocation parameters for strong users s in the data; Indicates the channel gain for strong users; This indicates inter-cluster interference from strong users; Indicates antenna with cluster Medium-strength users The correlation between them, and ,when When, it indicates a cluster strong users By antenna Provide services, otherwise .

[0104] S15. Based on the data rates of weak users and strong users, obtain the data rate of the user cluster; based on the data rate of the user cluster, construct an optimization problem model that maximizes the communication data rate and the sensing detection power.

[0105] The data rate of the user cluster is expressed by the following formula (5):

[0106] (5)

[0107] in, This indicates the data rate of user cluster c.

[0108] Alternatively, the optimization problem model that maximizes the communication data rate and sensing detection power is expressed by the following formula (6):

[0109] (6)

[0110] in, The power of the detection signal indicating the direction of the target being sensed; Regularization parameters representing communication; Represents the regularization parameter for perception; Represents user clusters The data rate; C represents the total number of user clusters; K represents the total number of sensory network models of the sensing target; α represents the user clustering variable; represents the intra-cluster power allocation parameter vector; p represents the inter-cluster power allocation vector.

[0111] In one feasible implementation method, The power of the detection signal representing the direction of the sensed target can be expressed by the covariance matrix of the transmitted signal vector. With perceived target Channel vector to flexible antenna A decision can be expressed as .

[0112] S2. Based on the optimization problem model, design the optimization objective function; the optimization objective function includes: state space, action space and reward function.

[0113] Optionally, the state space includes the location of the flexible antenna, the location of the communication user, the location of the sensing target, and the total remaining energy of the system;

[0114] The state space is represented as , ,in, Indicates in Set of time-slot pin antenna locations; Represents a set of user locations; Represents the set of locations of the perceived target. This indicates the remaining energy of the system.

[0115] The action space includes: the positional change of the flexible antenna, user clustering variables, intra-cluster power allocation parameters, inter-cluster power allocation, the positional change of the sensed target, and energy consumption.

[0116] The action space is represented as It includes continuous actions and discrete actions; where continuous actions are represented as... ,in, This indicates the amount of change in the flexible antenna position. Indicates in Inter-cluster power allocation vector of time slot, Indicates in The intra-cluster power allocation parameter vector of the time slot, Indicates that the system is in Energy consumed in a time slot; where discrete actions are represented as , Indicates in User clustering vectors in time slots.

[0117] The reward functions include: the reward function for communication data rate and the reward function for sensing and detection power.

[0118] The reward function is expressed as follows: ,Right now ;in, Regularization parameters representing communication; Represents the regularization parameter for perception; express Time slot user cluster Communication data rate; express Time-slot sensing target The detection power; R represents the reward space.

[0119] S3. Based on the state space, action space, and reward function, a method combining the Actor-Critic framework and the flexible Actor-Critic framework is adopted to optimize the discrete-continuous mixed variable problem of user clustering, antenna position, intra-cluster power allocation parameters, and inter-cluster power allocation in the action space, so as to obtain the optimal user clustering, antenna position, and power allocation strategy.

[0120] In one feasible implementation, the present invention relates to a method based on a combination of the Actor-Critic framework and the flexible Actor-Critic framework, which designs a dual-module scheme to optimize the discrete-continuous mixed variable problem of user clustering, antenna position, intra-cluster power allocation parameters and inter-cluster power allocation in the action space.

[0121] Among them, such as Figure 2 The diagram illustrates a block diagram of a sensor-integrated user clustering and resource allocation algorithm based on deep reinforcement learning, provided by an embodiment of the present invention. In one feasible implementation, the dual modules include a clustering module and a resource allocation module. The clustering module processes discrete user clustering variables and uses an Actor-Critic framework to learn user clustering policies and determine the correlation between antennas and user clusters. The user clustering decision directly affects the resource allocation policy. Specifically, randomly initialized discrete user clustering variables are input into the Actor-Critic framework. Based on gradient updates of the policy network, the optimal user clustering variables are obtained, and these optimal user clustering variables become part of the state space in the resource allocation module.

[0122] The resource allocation module, designed for adaptive power resource allocation, transforms the problem into a Markov decision process and employs a flexible Actor-Critic framework to handle continuous power allocation decisions. This flexible Actor-Critic framework comprises one policy network, two primary evaluation networks, and two target evaluation networks. The policy network outputs power allocation decisions based on the system state, while the two primary evaluation networks and two target evaluation networks output their minimum values, avoiding over-estimation issues. This framework is used to evaluate communication and perception performance under given decisions and guides the updating of the flexible Actor-Critic framework's networks.

[0123] Among them, the clustering module and the resource allocation module work together to realize user clustering and resource allocation in the integrated sensory scenario.

[0124] Optionally, the specific implementation process of S3 includes S31-S36:

[0125] S31. Initialize discrete user clustering variables, continuously update the policy network of the Actor-Critic framework, obtain the optimal user clustering variables, and store the optimal user clustering variables in the state space.

[0126] S32. Based on the current state, the policy network of the Actor-Critic framework inputs actions into the environment, generates rewards, and produces new states;

[0127] S33. Based on the generated rewards, the evaluation network of the flexible Actor-Critic framework generates rewards and new states to update the evaluation strategy, and uses the minimum value of the two main evaluation networks as the output of the evaluation network of the flexible Actor-Critic framework.

[0128] To prevent overestimation, the minimum value of the two main evaluation networks is used as the output of the evaluation network.

[0129] S34. Based on the actions output by the policy network of the flexible Actor-Critic framework, obtain new states, actions, and rewards; merge the new states, actions, and rewards into experience tuples and store them in the experience replay pool; extract small batches of experience tuples to update the policy network and the main evaluation network of the flexible Actor-Critic framework.

[0130] Wherein, the empirical tuple is represented as ;in, Indicates spatial state; Indicates an action; Indicates a reward; It indicates a new state.

[0131] Among them, during the training process, the experience replay pool The number of tuples stored in the database gradually increases until it reaches a scale sufficient for sufficient sampling, allowing for the extraction of small batches of empirical tuples. Policy network and the main evaluation network Update.

[0132] S35. Based on the soft update strategy, the target evaluation network of the flexible Actor-Critic framework is updated.

[0133] S36. Iterate through steps S31-S35 until the cumulative discount reward converges, completing the training and obtaining the trained flexible Actor-Critic framework policy network. Based on the trained policy network, obtain the optimal user clustering, antenna location, and power allocation policies.

[0134] This invention first establishes a flexible antenna link channel model and a transmission signal model. Second, based on these models, it constructs an optimization problem model to maximize communication data rate and sensing detection power. Based on this model, an optimization objective function is designed, comprising a state space, an action space, and a reward function. Finally, based on the state space, action space, and reward function, a method combining an Actor-Critic framework and a flexible Actor-Critic framework is used to optimize the discrete-continuous hybrid variable problem of user clustering, antenna location, intra-cluster power allocation parameters, and inter-cluster power allocation in the action space, obtaining the optimal user clustering, antenna location, and power allocation strategies. This invention can solve the user clustering and resource allocation problem in flexible antenna-assisted inductive-sensing integrated non-orthogonal access networks, and is particularly suitable for inductive-sensing integrated network optimization problems in discrete-continuous variable spaces and dynamic time-varying environments.

[0135] Figure 3 This is a block diagram illustrating a sensory-integrated user clustering and resource allocation apparatus based on deep reinforcement learning, according to an exemplary embodiment. The apparatus is used in a sensory-integrated user clustering and resource allocation method based on deep reinforcement learning. (Refer to...) Figure 3 The device includes a construction unit 310, a design unit 320, and an optimization unit 330. Wherein:

[0136] Construction unit 310 is used to establish a flexible antenna link channel model and a transmission signal model; based on the flexible antenna link channel model and the transmission signal model, an optimization problem model is constructed to maximize the communication data rate and the sensing detection power.

[0137] Design unit 320 is used to design an optimization objective function based on the optimization problem model; the optimization objective function includes: a state space, an action space, and a reward function;

[0138] The optimization unit 330 is used to optimize the discrete-continuous mixed variable problem of user clustering, antenna position, intra-cluster power allocation parameters and inter-cluster power allocation in the action space based on the state space, action space and reward function, using a method combining the Actor-Critic framework and the flexible Actor-Critic framework, to obtain the optimal user clustering, antenna position and power allocation strategy.

[0139] Optionally, the construction unit 310 is used for:

[0140] Based on a sensing-integrated network model with multiple flexible antennas and multiple sensing targets, the flexible antennas employ non-orthogonal multiple access technology to divide multiple users into... Clustering; By utilizing user clustering variables, the relationship between flexible antennas and users is defined, and users are divided into several clusters;

[0141] Define binary variables and set the positions of the antenna and each user; construct the antenna-to-user channel model based on the spherical wave channel model;

[0142] Based on the signal transmission methods of all users in non-orthogonal multiple access, obtain Cluster user signals;

[0143] Based on the serial interference cancellation rules of non-orthogonal multiple access, the data rates of weak users and strong users are obtained.

[0144] Based on the data rates of weak users and strong users, the data rate of the user cluster is obtained; based on the data rate of the user cluster, an optimization problem model is constructed to maximize the communication data rate and the sensing detection power.

[0145] Optionally, the The cluster user's signal is represented by the following formula (1):

[0146] (1)

[0147] Where x represents Cluster user signals; This represents the total number of user clusters; Indicates the power allocation parameters within the first cluster; Indicates the power allocation parameters within the second cluster; , Indicates that it is sent to the user communication signal.

[0148] Optionally, the data rate of the weak user is represented by the following formula (2):

[0149] (2)

[0150] in, Indicates the data rate for weak users; Represents user clusters The allocated power; Indicates that by antenna service clusters Weak users in Power allocation parameters; Indicates the channel gain for weak users; This indicates that the superimposed signal passes through the antenna. Phase shift; This represents intra-cluster interference and inter-cluster interference for weak users; This indicates inter-cluster interference from strong users; This represents additive white Gaussian noise; Indicates antenna with cluster Medium and weak users The correlation between them, and ,when When, it indicates a cluster Weak users By antenna Provide services, otherwise ;

[0151] The data rate of the strong user is expressed by the following formula (3):

[0152] (3)

[0153] in, Indicates the data rate of strong users; Indicates that by antenna service clusters Power allocation parameters for strong users s in the data; Indicates the channel gain for strong users; This indicates inter-cluster interference from strong users; Indicates antenna with cluster Medium-strength users The correlation between them, and ,when When, it indicates a cluster strong users By antenna Provide services, otherwise .

[0154] Optionally, the optimization problem model for maximizing the communication data rate and sensing detection power is expressed by the following formula (4):

[0155] (4)

[0156] in, The power of the detection signal indicating the direction of the target being sensed; Regularization parameters representing communication; Represents the regularization parameter for perception; Represents user clusters The data rate; C represents the total number of user clusters; K represents the total number of sensory network models of the sensing target; α represents the user clustering variable; represents the intra-cluster power allocation parameter vector; p represents the inter-cluster power allocation vector.

[0157] Optionally, the optimization unit 330 is used to:

[0158] (1) Initialize discrete user clustering variables, continuously update the policy network of the Actor-Critic framework, obtain the optimal user clustering variables, and store the optimal user clustering variables in the state space;

[0159] (2) Based on the current state, the policy network of the Actor-Critic framework inputs actions into the environment, generates rewards, and produces new states;

[0160] (3) Based on the generated reward, the evaluation network of the flexible Actor-Critic framework generates a reward and a new state to update the evaluation strategy, and takes the minimum value of the two main evaluation networks as the output of the evaluation network of the flexible Actor-Critic framework.

[0161] (4) Based on the actions output by the policy network of the flexible Actor-Critic framework, obtain new states, actions and rewards; merge the new states, actions and rewards into experience tuples and store them in the experience replay pool; extract small batches of experience tuples to update the policy network and the main evaluation network of the flexible Actor-Critic framework.

[0162] (5) Based on the soft update strategy, the target evaluation network of the flexible Actor-Critic framework is updated;

[0163] (6) Iterate through steps (1)-(5) until the cumulative discount reward converges, complete the training, and obtain the policy network of the trained flexible Actor-Critic framework; based on the trained policy network, obtain the optimal user clustering, antenna location and power allocation policy.

[0164] This invention first establishes a flexible antenna link channel model and a transmission signal model. Second, based on these models, it constructs an optimization problem model to maximize communication data rate and sensing detection power. Based on this model, an optimization objective function is designed, comprising a state space, an action space, and a reward function. Finally, based on the state space, action space, and reward function, a method combining an Actor-Critic framework and a flexible Actor-Critic framework is used to optimize the discrete-continuous hybrid variable problem of user clustering, antenna location, intra-cluster power allocation parameters, and inter-cluster power allocation in the action space, obtaining the optimal user clustering, antenna location, and power allocation strategies. This invention can solve the user clustering and resource allocation problem in flexible antenna-assisted inductive-sensing integrated non-orthogonal access networks, and is particularly suitable for inductive-sensing integrated network optimization problems in discrete-continuous variable spaces and dynamic time-varying environments.

[0165] Figure 4 This is a schematic diagram of a synesthetic user clustering and resource allocation device based on deep reinforcement learning, provided by an embodiment of the present invention. Figure 4 As shown, the synesthetic user clustering and resource allocation device based on deep reinforcement learning can include the above-mentioned... Figure 3 The illustrated synesthetic user clustering and resource allocation device is based on deep reinforcement learning. Optionally, the synesthetic user clustering and resource allocation device 410 based on deep reinforcement learning may include a first processor 2001.

[0166] Optionally, the deep reinforcement learning-based synesthetic user clustering and resource allocation device 410 may also include a memory 2002 and a transceiver 2003.

[0167] The first processor 2001, memory 2002, and transceiver 2003 can be connected via a communication bus.

[0168] The following is combined Figure 4 The components of the deep reinforcement learning-based synesthetic user clustering and resource allocation device 410 are described in detail below:

[0169] The first processor 2001 is the control center of the deep reinforcement learning-based sensory integrated user clustering and resource allocation device 410. It can be a single processor or a collective term for multiple processing elements. For example, the first processor 2001 can be one or more central processing units (CPUs), application-specific integrated circuits (ASICs), or one or more integrated circuits configured to implement embodiments of the present invention, such as one or more digital signal processors (DSPs), or one or more field-programmable gate arrays (FPGAs).

[0170] Optionally, the first processor 2001 can execute various functions of the deep reinforcement learning-based synesthetic user clustering and resource allocation device 410 by running or executing software programs stored in the memory 2002 and calling data stored in the memory 2002.

[0171] In a specific implementation, as one example, the first processor 2001 may include one or more CPUs, for example... Figure 4 CPU0 and CPU1 are shown in the diagram.

[0172] In a specific implementation, as one example, the deep reinforcement learning-based synesthetic user clustering and resource allocation device 410 may also include multiple processors, for example... Figure 4 The first processor 2001 and the second processor 2004 are shown in the diagram. Each of these processors can be a single-core processor or a multi-core processor. Here, a processor can refer to one or more devices, circuits, and / or processing cores used to process data (such as computer program instructions).

[0173] The memory 2002 is used to store the software program that executes the present invention, and is controlled by the first processor 2001 to execute it. The specific implementation method can be referred to the above method embodiment, and will not be repeated here.

[0174] Optionally, the memory 2002 may be a read-only memory (ROM) or other type of static storage device capable of storing static information and instructions, random access memory (RAM) or other type of dynamic storage device capable of storing information and instructions, or electrically erasable programmable read-only memory (EEPROM), compact disc read-only memory (CD-ROM) or other optical disc storage, optical disc storage (including compressed optical discs, laser discs, optical discs, digital universal optical discs, Blu-ray discs, etc.), magnetic disk storage media or other magnetic storage devices, or any other medium capable of carrying or storing desired program code in the form of instructions or data structures and accessible by a computer, but not limited thereto. The memory 2002 may be integrated with the first processor 2001 or may exist independently, and may be connected via the interface circuit of the deep reinforcement learning-based sensory integrated user clustering and resource allocation device 410. Figure 4 (Not shown in the image) is coupled to the first processor 2001, and this embodiment of the invention does not specifically limit this.

[0175] The transceiver 2003 is used to communicate with network devices or with terminal devices.

[0176] Alternatively, transceiver 2003 may include a receiver and a transmitter. Figure 4 (Not shown separately). The receiver is used to implement the receiving function, and the transmitter is used to implement the transmitting function.

[0177] Optionally, the transceiver 2003 can be integrated with the first processor 2001 or exist independently, and can be connected to the interface circuit of the deep reinforcement learning-based synesthetic user clustering and resource allocation device 410. Figure 4 (Not shown in the image) is coupled to the first processor 2001, and this embodiment of the invention does not specifically limit this.

[0178] It should be noted that, ​ The structure of the deep reinforcement learning-based synesthetic user clustering and resource allocation device 410 shown in the figure does not constitute a limitation on the router. The actual knowledge structure recognition device may include more or fewer components than shown, or combine some components, or have different component arrangements.

[0179] Furthermore, the technical effects of the synesthetic user clustering and resource allocation device 410 based on deep reinforcement learning can be referred to the technical effects of the synesthetic user clustering and resource allocation method based on deep reinforcement learning described in the above method embodiments, and will not be repeated here.

[0180] It should be understood that the first processor 2001 in the embodiments of the present invention may be a central processing unit (CPU), or it may be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc.

[0181] It should also be understood that the memory in the embodiments of the present invention can be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. The volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of random access memory (RAM) are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate synchronous DRAM (DDR SDRAM), enhanced synchronous DRAM (ESDRAM), synchronous linked DRAM (SLDRAM), and direct rambus RAM (DR RAM).

[0182] The above embodiments can be implemented, in whole or in part, by software, hardware (such as circuits), firmware, or any other combination thereof. When implemented using software, the above embodiments can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, all or part of the processes or functions described in the embodiments of the present invention are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more sets of available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium. A semiconductor medium can be a solid-state drive.

[0183] It should be understood that the term "and / or" in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. A and B can be singular or plural. Additionally, the character " / " in this article generally indicates an "or" relationship between the preceding and following related objects, but it can also represent an "and / or" relationship. Please refer to the context for a more accurate understanding.

[0184] In this invention, "at least one" means one or more, and "more than one" means two or more. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of a single item or a plurality of items. For example, at least one of a, b, or c can represent: a, b, c, ab, ac, bc, or abc, where a, b, and c can be a single item or multiple items.

[0185] It should be understood that, in various embodiments of the present invention, the order of the above-mentioned process numbers does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0186] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.

[0187] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the devices, apparatuses, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0188] In the several embodiments provided by this invention, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another device, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.

[0189] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0190] In addition, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0191] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0192] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A method for user clustering and resource allocation based on deep reinforcement learning for integrated sensing and communication, characterized in that, The method comprises: S1, establishing a flexible antenna link channel model and a transmission signal model; and constructing an optimization problem model of maximizing a communication data rate and a sensing detection power according to the flexible antenna link channel model and the transmission signal model; S2, designing an optimization objective function based on the optimization problem model; the optimization objective function comprises a state space, an action space and a reward function; S3, based on the state space, the action space and the reward function, using a method combining an Actor-Critic framework and a flexible Actor-Critic framework to optimize discrete-continuous mixed variable problems of user clustering, antenna position, intra-cluster power allocation parameters and inter-cluster power allocation in the action space, and obtain optimal user clustering, antenna position and power allocation strategies.

2. The deep reinforcement learning based user clustering and resource allocation method for sensory integration according to claim 1, wherein, The state space comprises positions of flexible antennas, positions of communication users, positions of sensing targets and a total residual energy of the system; The action space comprises a position change amount of the flexible antenna, a user clustering variable, an intra-cluster power allocation parameter, an inter-cluster power allocation, a sensing target position change amount and an energy consumption. The reward function comprises a reward function of the communication data rate and a reward function of the sensing detection power.

3. The deep reinforcement learning based user clustering and resource allocation method for sensory integration according to claim 1, wherein, The S1 of establishing the flexible antenna link channel model and the transmission signal model; and the constructing the optimization problem model of maximizing the communication data rate and the sensing detection power according to the flexible antenna link channel model and the transmission signal model comprise: S11、According to the multiple flexible antenna and multiple sensing target integrated network model, the flexible antenna adopts the non-orthogonal multiple access technology, and multiple users are divided into Cluster; using the user clustering variable, the relationship between the flexible antenna and the user is defined, and the user is divided into several clusters; S12, defining a binary variable, setting positions of the antenna and each user; and constructing an antenna-to-user channel model based on a spherical wave channel model; S13, obtaining the signals of the cluster users according to the signal transmission mode of all users in the non-orthogonal multiple access cluster users; S14, obtaining a data rate of a weak user and a data rate of a strong user according to a serial interference cancellation rule of a non-orthogonal multiple access; S15, obtaining a data rate of a user cluster according to the data rate of the weak user and the data rate of the strong user; and constructing the optimization problem model of maximizing the communication data rate and the sensing detection power according to the data rate of the user cluster.

4. The deep reinforcement learning based user clustering and resource allocation method for sensory integration according to claim 3, wherein, The The signals of the cluster users are represented by the following equation (1): (1) where x represents signals of the cluster users; represents the total number of user clusters; represents a power allocation parameter within the first cluster; represents a power allocation parameter within the second cluster; , , represents a communication signal transmitted to the user .

5. The deep reinforcement learning based user clustering and resource allocation method for sensory integration according to claim 3, wherein, The data rate of the weak user is represented by the following formula (2): (2) in, Indicates the data rate for weak users; Represents user clusters The allocated power; Indicates that by antenna service clusters Weak users in Power allocation parameters; Indicates the channel gain for weak users; This indicates that the superimposed signal passes through the antenna. Phase shift; This represents intra-cluster interference and inter-cluster interference for weak users; This indicates inter-cluster interference from strong users; This represents additive white Gaussian noise; Indicates antenna with cluster Medium and weak users The correlation between them, and ,when When, it indicates a cluster Weak users By antenna Provide services, otherwise ; The data rate of the strong user is represented by the following formula (3): (3) in, Indicates the data rate of strong users; Indicates that by antenna service clusters Power allocation parameters for strong users s in the data; Indicates the channel gain for strong users; This indicates inter-cluster interference from strong users; Indicates antenna with cluster Medium-strength users The correlation between them, and ,when When, it indicates a cluster strong users By antenna Provide services, otherwise .

6. The deep reinforcement learning based user clustering and resource allocation method for sensory integration according to claim 3, wherein, The optimization problem model of maximizing the communication data rate and the sensing detection power is represented by the following formula (4): (4) wherein represents the power of the probe signal perceiving the target direction; represents the regularization parameter of the communication; represents the regularization parameter of the perception; represents the data rate of the user cluster C represents the total number of user clusters; K represents the total number of the perception-target integrated network models; and a represents the user clustering variable; represents the power allocation parameter vector within the cluster; and p represents the power allocation vector between the clusters.

7. The deep reinforcement learning based user clustering and resource allocation method for sensory integration according to claim 1, wherein, The S3 of using the method combining the Actor-Critic framework and the flexible Actor-Critic framework to optimize the discrete-continuous mixed variable problems of the user clustering, the antenna position, the intra-cluster power allocation parameter and the inter-cluster power allocation in the action space, and obtain the optimal user clustering, the antenna position and the power allocation strategy comprises: S31, initializing a discrete user clustering variable, constantly updating a policy network of the Actor-Critic framework, obtaining an optimal user clustering variable, and storing the optimal user clustering variable in the state space; S32, based on a current state, the policy network of the Actor-Critic framework inputs an action into an environment, generates a reward, and produces a new state. S33, according to the generated reward, the evaluation network of the flexible Actor-Critic framework generates the reward and the new state to update the evaluation policy, and the minimum value of the two main evaluation networks is taken as the output of the evaluation network of the flexible Actor-Critic framework; S34, according to the action output by the policy network of the flexible Actor-Critic framework, a new state, an action and a reward are obtained; the new state, the action and the reward are combined into an experience tuple and stored in an experience replay pool; a small batch of experience tuples are extracted to update the policy network of the flexible Actor-Critic framework and the main evaluation network of the flexible Actor-Critic framework; S35, based on a soft update policy, the target evaluation network of the flexible Actor-Critic framework completes the update; S36, the steps S31-S35 are iterated and repeated until the cumulative discounted reward converges, the training is completed, and the trained policy network of the flexible Actor-Critic framework is obtained; and the optimal user clustering, antenna position and power allocation strategy are obtained according to the trained policy network.

8. A deep reinforcement learning based sensor integration user clustering and resource allocation apparatus for implementing the deep reinforcement learning based sensor integration user clustering and resource allocation method according to any one of claims 1-7, wherein, The device comprises: a construction unit configured to establish a flexible antenna link channel model and a transmission signal model, and to construct an optimization problem model for maximizing a communication data rate and a sensing detection power based on the flexible antenna link channel model and the transmission signal model; a design unit configured to design an optimization objective function based on the optimization problem model, wherein the optimization objective function comprises a state space, an action space and a reward function; an optimization unit configured to optimize discrete-continuous mixed variable problems of user clustering, antenna position, intra-cluster power allocation parameters and inter-cluster power allocation in the action space based on the state space, the action space and the reward function by using a method combining an Actor-Critic framework and a flexible Actor-Critic framework, and to obtain an optimal user clustering, antenna position and power allocation strategy.

9. A device for user clustering and resource allocation based on deep reinforcement learning for integrated sensing, characterized in that, The device comprises: a processor; a memory having computer readable instructions stored thereon, wherein the computer readable instructions are executed by the processor to implement the method according to any one of claims 1 to 7.

10. A computer readable storage medium, characterized in that, The computer readable storage medium stores program codes, and the program codes are called and executed by the processor to implement the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Wireless resource allocation joint optimization method and device

    CN112566253A

  • Space-time domain resource allocation method based on deep reinforcement learning

    CN117715219A