Power terminal intelligent resource allocation method and system under ISAC heterogeneous service scene
By adopting an intelligent resource allocation method in the smart grid under ISAC technology, combining the characteristics of CT and ST services, and using proportional fairness algorithms and deep reinforcement learning algorithms, the fairness and efficiency of resource allocation in heterogeneous business scenarios are solved, and resource allocation with high throughput, low latency and high reliability is achieved.
Patent Information
- Application Number
- CN202510020034.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-07
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2045-01-07
AI Technical Summary
In 6G networks, in smart grids under ISAC technology, the existing technology is difficult to take into account the high throughput, high bandwidth requirements of CT services and the low latency and high reliability requirements of ST services in heterogeneous business scenarios, and cannot effectively ensure user fairness.
Using an intelligent resource allocation method for power terminals in ISAC heterogeneous business scenarios, by calculating the transmission rate of CT users, using a finite block long channel coding scheme to solve ST services, analyzing ST delay requirements, designing interrupt probability constraints, introducing a proportional fair algorithm to build a collaborative resource allocation model for CT and ST services, and using a dual-delay depth deterministic strategy gradient algorithm based on proportional fairness to solve the optimal resource allocation scheme.
On the premise of taking into account the ST delay requirements and CT user reliability, the average data rate of CT services is maximized, the fairness and efficiency of resource allocation are achieved, and the significant decline in user service quality is avoided.
Smart Images

Figure CN120034892A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a resource allocation method and system, and in particular to an intelligent resource allocation method and system for power terminals in an ISAC heterogeneous business scenario. Background Art
[0002] With the advent of 6G technology, smart grids will usher in a new era of deep integration of perception, artificial intelligence and communication, including the combination of artificial intelligence and communication, synaesthesia integration and pan-connectivity. In this landscape of Internet of Things, People and Intelligence, ISAC uses a single waveform and a single device to unify sensing and communication functions in a more efficient way. With synaesthesia integration, it greatly enhances the capabilities of 6G networks and provides strong support for scenarios such as optimized operation, fault detection and energy scheduling of smart grids.
[0003] In smart grids, ISAC technology can improve the grid's real-time perception of the environment and equipment status and enhance the reliability of communication networks. There are different types of business traffic in the ISAC network. CT tasks include data interaction in distributed energy management and household power load information transmission. Such services require high-throughput and high-bandwidth network capabilities; ST tasks include grid fault location and power equipment status monitoring. Such services have low latency and high reliability requirements.
[0004] Existing research focuses on the optimization of a single task, lacks a mechanism for the fair allocation of multiple types of communication tasks, and cannot take into account the fair coordination of heterogeneous tasks. Especially in the case of limited 6G network resources, how to balance the high throughput and high bandwidth requirements of communication tasks and the low latency and high reliability requirements of perception tasks has become a key challenge. In addition, user fairness between different communication tasks must also be fully considered to avoid a significant decline in the quality of service for some users. Summary of the invention
[0005] Purpose of the invention: The purpose of the present invention is to provide an intelligent resource allocation method for power terminals in ISAC heterogeneous business scenarios to take into account ST low latency and CT reliability while maximizing the average data rate of CT services. On the other hand, to provide an intelligent resource allocation system for power terminals in ISAC heterogeneous business scenarios.
[0006] Technical solution: The present invention provides a method for intelligent resource allocation of power terminals in an ISAC heterogeneous service scenario, comprising the following steps:
[0007] (1) Calculate the CT user transmission rate based on the CT link channel model;
[0008] (2) Using a finite block length channel coding scheme to solve the achievable transmission rate of ST services;
[0009] (3) Analyze ST delay requirements and design ST outage probability constraints;
[0010] (4) Introduce the proportional fairness algorithm to build a collaborative resource allocation model for CT and ST services;
[0011] (5) setting corresponding Markov model parameters, wherein the Markov model parameters include a state space, an action space, and a reward function;
[0012] (6) Use the double-delay deep deterministic policy gradient algorithm based on proportional fairness to solve the optimal resource allocation solution.
[0013] Preferably, the CT link channel model formula in step 1 is as follows:
[0014]
[0015] in, represents the transmission rate of the i-th CT user in resource block b in time slot t, ω i,b (t) represents the total number of punctured mini-slots, ω i,b (t)∈{0,1,…,K}, K represents the total number of mini-slots, η i,b,k (t) represents the puncturing of the kth mini-time slot resource block b of the CT user in time slot t, η i,b,k (t) = 1 means perforated, otherwise 0;
[0016]
[0017] in, represents the signal-to-noise ratio of CT user i in resource block b in time slot t, σ 2 represents the noise power, represents the downlink transmission power of the i-th CT user in resource block b in time slot t, represents the transmission channel gain, and Indicates the interference caused by other CT users and ST users;
[0018]
[0019] in, represents the downlink transmission rate of the i-th CT user on all resource blocks after puncturing, x i,b (t)=1 indicates that resource block b is allocated to the i-th CT user, otherwise it is 0.
[0020] Preferably, the formula of the finite block length channel coding scheme in step 2 is as follows:
[0021]
[0022] in, represents the achievable transmission rate of the jth ST user at time slot t, represents the signal-to-noise ratio of the ST user, Indicates the number of symbols in a mini-slot, Q -1 (θ) represents the inverse Gaussian function, θ represents the transmission error probability, represents channel dispersion, which represents the randomness of the user channel;
[0023]
[0024] Preferably, the ST interruption probability constraint in step 3 is as follows:
[0025]
[0026] Where V(t) represents the total number of ST packets arriving in time slot t, expressed as V(t)=∑ k∈K V k (t) calculation, V k (t) represents the number of ST packets arriving in mini-slot k, and s represents the size of the ST packet.
[0027] Preferably, the proportional fairness algorithm formula in step 4 is:
[0028]
[0029] in, represents the instantaneous data rate of the CT user in time slot t, represents the historical average throughput at time slot t and is updated according to the following formula:
[0030]
[0031] Where T represents the length of the time window for calculating the average throughput.
[0032] Preferably, the resource allocation model formula in step 4 is as follows:
[0033]
[0034] Among them, P max represents the maximum transmission power of the base station, x represents the RB allocation matrix, p represents the power allocation vector, and η represents the mini-time slot puncturing number matrix. The goal of the optimization problem is to find the optimal value x of the above three parameters. * 、p * and η * Maximize the average CT transmission rate considering the fairness of CT users;
[0035] Indicates resource block allocation constraints, a resource block is only associated with a single user;
[0036] It means to ensure that the sum of transmission power does not exceed the maximum transmission power of the base station;
[0037] Indicates the reliability of ST business, ω i,b (t)∈{0,1,…,K} represents the value range of the number of punctured time slots of resource block b.
[0038] Preferably, the state space and action space in step 5 are as follows:
[0039] State space: s(t) = {g e (t),g u (t),λ,I,J}, where g e (t) and g u (t) represents the channel conditions of ST and CT users, respectively, λ represents the Poisson arrival rate of ST traffic, and I and J represent the number of CT and ST users, respectively;
[0040] Action space: a(r) = {x, p, η}, where x represents resource block allocation, p represents transmission power allocation, and η represents the ST scheduling strategy.
[0041] Preferably, the reward function in step 5 is as follows:
[0042] R(t) represents R e (t) and R u The sum of the weights of (t);
[0043]
[0044] Among them, R e (t) represents the demand of CT users, R u (t) represents the ST reliability requirement.
[0045] Preferably, the specific steps of solving the optimal resource allocation solution in step 6 are as follows:
[0046] (61) Initialize the state space s t , training space θ 1 ,θ 2 , φ, target network θ 1 ,θ 2 , φ, experience replay pool R and delay update parameter d;
[0047] (62) According to the policy network plus noise ε~N(0.6), action a is selected at each time step t ;
[0048] (63) According to action a t Get rewards t and the next state s t+1 , will (s t ,a t ,r t ,s t+1 ) is stored in the experience replay pool R;
[0049] (64) Randomly sample N data (s) from the experience replay pool R t ,a t ,r t ,s t+1 );
[0050] (65) Use the Bellman equation to calculate the target value y = r + γmin i=1,2 Q θi′ (s′, a′), where γ represents the discount factor;
[0051] (66) Using loss function loss = E[(yQ θi (s,a)) 2 ]Update critic network;
[0052] (67) After every d time steps, calculate the gradient Update the actor network;
[0053] (68) Use soft update to update the target network parameters: θ′ i =τθ i +(1-τ)θ′ i , i=1,2 and φ'=τφ+(1-τ)φ', τ<<1;
[0054] (69) Repeat steps 62 to 69 until the maximum number of training steps T is reached.
[0055] The present invention provides a power terminal intelligent resource allocation system in an ISAC heterogeneous service scenario, comprising:
[0056] Channel monitoring and data acquisition module: It consists of multiple types of sensors deployed at the power terminal, responsible for real-time collection of various physical parameters of the CT link channel, and transmits the data to the CT link channel model calculation unit. At the same time, it can monitor and collect network status indicators related to ST business, providing data support for subsequent ST business analysis;
[0057] CT link channel model calculation unit: It has a built-in CT link channel model, receives data from the channel monitoring and data acquisition module, calculates the CT user transmission rate using an algorithm, and outputs the result to the resource allocation collaborative control module;
[0058] Finite block length channel coding solution unit: configures the finite block length channel coding scheme optimized for ST service characteristics, receives part of the ST service related data from the channel monitoring and data acquisition module, uses the coding algorithm to solve the achievable transmission rate of the ST service, and feeds the result back to the resource allocation collaborative control module;
[0059] ST service analysis and constraint design unit: Use big data analysis and real-time monitoring technology to analyze ST latency requirements, design ST interruption probability constraints, and pass constraint conditions and policy information to the resource allocation collaborative control module;
[0060] Proportional fairness algorithm and collaborative model building module: Built-in proportional fairness algorithm, receiving business data and parameters from the above units, building CT and ST business collaborative resource allocation model, determining the objective function, and outputting the model and function information to the dual-delay deep deterministic policy gradient algorithm solving unit;
[0061] Markov model parameter configuration unit: responsible for setting the state space, action space and reward function of the Markov model according to the real-time business status of the power terminal, and passing the configured parameters to the double-delay deep deterministic policy gradient algorithm solving unit;
[0062] Double-delayed deep deterministic policy gradient algorithm solving unit: Based on the objective function constructed by Markov model feedback and proportional fairness algorithm, the double-delayed deep deterministic policy gradient algorithm is used to solve the optimal resource allocation plan, and the plan is sent to the resource scheduling execution module of the power terminal in real time;
[0063] Resource scheduling execution module: Receives the optimal resource allocation plan and is responsible for accurately allocating resources to CT users and ST services according to the plan. It has the ability to monitor the execution of resource allocation in real time and dynamically adjust the allocation strategy to deal with emergencies.
[0064] Beneficial effects: Compared with the prior art, the present invention has the following advantages: by constructing a resource allocation model, introducing a proportional fairness algorithm, and combining deep reinforcement learning to dynamically allocate resources for the two services through interaction with the environment, the CT average data rate is maximized while taking into account the ST delay requirements and CT user reliability. BRIEF DESCRIPTION OF THE DRAWINGS
[0065] Figure 1 This is a schematic diagram of the ISAC heterogeneous service multiplexing scenario model of the present invention;
[0066] Figure 2 Schematic diagram of the policy gradient algorithm of the dual-delay deep deterministic strategy of the present invention. DETAILED DESCRIPTION
[0067] The technical solution of the present invention is described in detail below in conjunction with the accompanying drawings.
[0068] This embodiment considers a base station serving I CT users and J ST users, i and j are used to represent the i-th CT user and the j-th ST user respectively, i∈{1,2,…,I}, and the downlink CT request and ST link request corresponding to these users are considered. The two service requests adopt OFDMA multiplexing mode, and the network resources are divided into B resource blocks RB, one resource block b contains 12 subcarriers, b∈{1,2,…,B}.
[0069] This embodiment uses a perforation method to schedule ST traffic to meet the low latency requirements of ST services. By using a short TTI to schedule ST data packets, when ST data packets arrive, they will seize the frequency resources of the CT service being sent, thereby achieving the low latency requirements of ST services. However, this approach will inevitably affect the transmission quality of CT, so while considering the low latency requirements of ST services, the transmission quality of CT services should also be considered.
[0070] (1) Calculate the CT user transmission rate based on the CT link channel model.
[0071] The data rate of CT users can be calculated using the Shannon formula, and the signal-to-noise ratio of CT user i in time slot t resource block b is as follows:
[0072]
[0073] in, represents the downlink transmission power of the i-th CT user in resource block b in time slot t, in Hz, represents the transmission channel gain, and denote the interference caused by other CT users and ST users, σ 2 represents the noise power;
[0074] Considering the impact of ST traffic on CT traffic, we use η i,b,k (t) represents the puncturing of the kth mini-time slot resource block b of the CT user in time slot t, η i,b,k (t) = 1 means perforated, otherwise 0;
[0075] The transmission rate of the i-th CT user in resource block b in time slot t can be expressed as:
[0076]
[0077] in, represents the total number of punctured mini-slots, ω i,b(t) ∈ {0, 1, 2, …, K}, where K represents the total number of mini - slots. The downlink transmission rate of the i - th user after puncturing on all resource blocks can be expressed as:
[0078]
[0079] Among them, x i,b (t) = 1 indicates that resource block b is allocated to the i - th CT user, otherwise it is 0.
[0080] (2) Calculate the transmission rate of CT users according to the CT link channel model.
[0081] Since the size of ST packets is usually much smaller than that of CT packets, the Shannon formula cannot be used to solve it as in the case of CT services. A finite - block - length channel coding scheme can be adopted for solving. The achievable transmission rate of the j - th ST user at time slot t can be expressed as:
[0082]
[0083] Among them, represents the signal - to - noise ratio of the ST user, represents the number of symbols in the mini - slot, Q -1 (θ) represents the Gaussian inverse function, and θ represents the transmission error probability. represents the channel dispersion, which can reflect the randomness of the user channel and can be calculated by the following formula:
[0084]
[0085] (3) Analyze the ST delay requirements and design the ST outage probability constraint.
[0086] The reliability of ST is achieved by ensuring that its outage probability is less than a specified threshold. When the size of the ST packet arriving within time slot t is greater than the transmission rate that the ST user can achieve, it is regarded as a transmission outage. Using ε to represent the maximum tolerable outage threshold, the outage probability constraint of ST can be expressed by the following formula:
[0087]
[0088] Among them, V(t) represents the total number of ST packets arriving within time slot t, and is calculated by V(t) = ∑ k∈K V k (t), where V k (t) represents the number of ST packets arriving within mini - slot k, and s represents the size of the ST packet.
[0089] (4) Introduce the proportional fairness algorithm to construct a collaborative resource allocation model for CT and ST services.
[0090] Considering that ST service puncturing will affect the service quality of CT service and the ST puncturing location is uncertain, this may cause CT users with lower transmission rates, such as users at the edge of the cell with poor channel conditions, to suffer losses caused by more puncturing. Therefore, in order to avoid a significant impact on fairness between users, a priority strategy needs to be formulated to balance the fairness of total throughput and user satisfaction. This solution adopts a proportional fairness algorithm and calculates the priority by introducing the historical average throughput of the user. The objective function of the PF scheduler is:
[0091]
[0092] in, represents the instantaneous data rate of the CT user in time slot t, represents the historical average throughput at time slot t and is updated according to the following formula:
[0093]
[0094] Where T represents the length of the time window for calculating the average throughput, and the PF scheduler is able to establish proportional fairness among users over time.
[0095] This embodiment expresses the resource allocation problem of CT and ST as an optimization problem. The goal is to maximize the throughput of CT users under the premise of satisfying the QoS constraints of ST services. Different from the traditional strategy of maximizing the average transmission rate of all CT users, the PF algorithm is introduced to make a trade-off between the throughput of CT users and the fairness of satisfaction, so as to avoid large differences in the service quality of CT users. The following is the formula of the collaborative resource allocation model of CT and ST services:
[0096]
[0097]
[0098] Among them, P max represents the maximum transmission power of the base station, x represents the RB allocation matrix, p represents the power allocation vector, and η represents the mini-time slot puncturing number matrix. The goal of the optimization problem is to find the optimal value x of the above three parameters. * , p * and η * Maximize the average CT transmission rate considering the fairness of CT users; Indicates resource block allocation constraints, meaning that a resource block is only associated with a single user; It means to ensure that the sum of transmission power does not exceed the maximum transmission power of the base station; Indicates that the reliability of ST business is guaranteed, and constraints ω i,b(t)∈{0,1,…,K}, represents the value range of the number of punctured time slots of resource block b.
[0099] (5) Set the corresponding Markov model parameters, including state space, action space and reward function.
[0100] In an embodiment, a Markov decision process is used to represent a wireless resource scheduling problem, wherein the state s(t) represents the current situation of the system, the action a(t) represents the resource allocation strategy, and the reward R(t) represents the optimization target representing the system performance;
[0101] In the state space, s(t) = {g e (t),g u (t),λ,I,J},g e (t) and g u (t) represents the channel conditions of ST and CT users, respectively, λ represents the Poisson arrival rate of ST traffic, and I and J represent the number of CT and ST users, respectively;
[0102] The action set of the agent is defined as a(t) = {x, p, η}, where x, p, η represent resource block allocation, transmission power allocation, and ST scheduling strategy, respectively;
[0103] The reward at time t is represented by R(t). The requirements of CT and ST users are defined in the reward function, and the agent selects the action that can obtain higher rewards. In this scheme, the reward function is represented by the following two equations:
[0104]
[0105] Among them, R e (t) represents the CT user requirements, including CT data rate and reliability, R u (t) represents the ST reliability requirement. Considering the business requirements of CT and ST, it is expressed as the sum of the weights of and:
[0106]
[0107] Among them, the parameters The value of is adjusted during the training process. θ(t) represents the estimated outage probability, and ε represents the maximum tolerable outage probability.
[0108] Evaluation Status t Next action a t The action value function of the value is Q π (s t ,a t ), the goal of reinforcement learning is to find the optimal strategy π * Maximizing Qπ (s t ,a t ), where Q π (s t ,a t ) is calculated according to the Bellman equation, and the calculation formula is:
[0109] Q π (s t ,a t )=E[R(s t ,a t )+Q π (s t+1 ,a t+1 )].
[0110] (6) Use the double-delay deep deterministic policy gradient algorithm based on proportional fairness to solve the optimal resource allocation plan.
[0111] The dual-delayed deep deterministic policy gradient algorithm adopts a dual network architecture, target policy smoothing regularization and delayed update technology to solve the overestimation problem in wireless resource scheduling. TD3 involves a total of six neural networks. The training network includes the actor network μ φ , critic1 network Q θ1 and critic2 network Q θ2 , the target network includes the target network μ′ φ , target critic1 network Q′ θ1 and target critic2 network Q′ θ2 ;
[0112] The critic network is updated by minimizing the error between the evaluation value and the target value. The target actor network calculates the action a′ under the state s′ and selects the smaller of the two Q values as the target Q value in the current state:
[0113]
[0114] Among them, the resource allocation strategy is adjusted according to the Q value update rule so that the scheduling decision in each state can maximize the overall performance of the system. The loss function of the critic network is:
[0115] loss=E[(yQ θi (s,a)) 2 ];
[0116] Among them, i=1,2 corresponds to two critic networks, and the parameters of the critic network are updated by minimizing these two loss functions;
[0117] Use the actor network to calculate the action a=μ in state sφ (s), using critic1 or critic2 to calculate the evaluation value of the state action pair (s, a) is Q θ1 (s,a). Use the gradient ascent algorithm to maximize Q θ1 (s,a), complete the update of the actor network:
[0118]
[0119] TD3 uses a soft update method to update the target network parameters, that is, introducing a learning rate τ, taking the weighted average of the old target network parameters and the new corresponding network parameters and assigning them to the target network, which can be expressed as follows:
[0120] θ′ i =τθ i +(1-τ)θ′ i , i=1,2;
[0121] φ′=τφ+(1-σ)φ′, τ<<1;
[0122] Based on the DRL process and TD3 architecture, a double-delay deep deterministic policy gradient algorithm based on proportional fairness is proposed. The algorithm is based on the state space s t To select action a(t) = {x, p, η}, and according to the reward r t Update strategy, t ,a t ,r t ,s t+1 ) is stored in the experience replay pool. At each time step, the system samples data from the experience replay pool and iteratively updates the actor network, critic network, and target network based on this data.
[0123] Among them, the actor network represents the scheduling strategy, and the critic network evaluates the performance of the current strategy. In order to ensure network stability, the actor and target networks are updated every d iterations. Through this process, the PF-TD3A algorithm can find the optimal solution for RB resource blocks, power allocation, and ST scheduling. * , p * and η * , thereby improving the overall performance of the system.
Claims
1. A method for intelligent resource allocation of power terminals in ISAC heterogeneous business scenarios, characterized in that: The following steps are involved: (1) Calculate the CT user transmission rate based on the CT link channel model; (2) Using a finite block length channel coding scheme to solve the achievable transmission rate of ST services; (3) Analyze ST delay requirements and design ST outage probability constraints; (4) Introduce the proportional fairness algorithm to build a collaborative resource allocation model for CT and ST services; (5) setting corresponding Markov model parameters, wherein the Markov model parameters include a state space, an action space, and a reward function; (6) Use the double-delay deep deterministic policy gradient algorithm based on proportional fairness to solve the optimal resource allocation solution.
2. The method for intelligent resource allocation of power terminals according to claim 1, characterized in that: The CT link channel model formula in step 1 is as follows: in, represents the transmission rate of the i-th CT user in resource block b in time slot t, ω i,b (t) represents the total number of punctured mini-slots, K represents the total number of mini-slots, η i,b,k (t) represents the puncturing of the kth mini-time slot resource block b of the CT user in time slot t, η i,b,k (t) = 1 means perforated, otherwise 0; in, represents the signal-to-noise ratio of CT user i in resource block b in time slot t, σ 2 represents the noise power, represents the downlink transmission power of the i-th CT user in resource block b in time slot t, represents the transmission channel gain, and Indicates the interference caused by other CT users and ST users; in, represents the downlink transmission rate of the i-th CT user on all resource blocks after puncturing, x i,b (t)=1 indicates that resource block b is allocated to the i-th CT user, otherwise it is 0.
3. The method for intelligent resource allocation of power terminals according to claim 1, characterized in that: The formula of the finite block length channel coding scheme in step 2 is as follows: in, represents the achievable transmission rate of the jth ST user at time slot t, represents the signal-to-noise ratio of the ST user, represents the number of symbols in a mini-slot, represents the inverse Gaussian function, represents the transmission error probability, represents channel dispersion, which represents the randomness of the user channel; 4. The method for intelligent resource allocation of power terminals according to claim 1, characterized in that: The ST interruption probability constraint in step 3 is as follows: Where V(t) represents the total number of ST packets arriving in time slot t, expressed as V(t)=∑ k∈K V k (t) calculation, V k (t) represents the number of ST packets arriving in mini-slot k, and s represents the size of the ST packet.
5. The method for intelligent resource allocation of power terminals according to claim 1, characterized in that: The proportional fairness algorithm formula in step 4 is: in, represents the instantaneous data rate of the CT user in time slot t, represents the historical average throughput at time slot t and is updated according to the following formula: Where T represents the length of the time window for calculating the average throughput.
6. The method for allocating intelligent resources of power terminals according to claim 1, characterized in that: The resource allocation model formula in step 4 is as follows: Among them, P max represents the maximum transmission power of the base station, x represents the RB allocation matrix, p represents the power allocation vector, and η represents the mini-time slot puncturing number matrix. The goal of the optimization problem is to find the optimal value x of the above three parameters. * 、p * and η * Maximize the average CT transmission rate considering the fairness of CT users; Indicates resource block allocation constraints. A resource block is only associated with a single user. It means to ensure that the sum of transmission power does not exceed the maximum transmission power of the base station; Indicates the reliability of ST business, ω i,b (t)∈{0,1,…,K} represents the value range of the number of punctured time slots of resource block b.
7. The method for intelligent resource allocation of power terminals according to claim 1, characterized in that: The state space and action space described in step 5 are as follows: State space: s(t) = {g e (t),g u (t),λ,I,J}, where g e (t) and g u (t) represents the channel conditions of ST and CT users, respectively, λ represents the Poisson arrival rate of ST traffic, and I and J represent the number of CT and ST users, respectively; Action space: a(t) = {x, p, η}, where x represents resource block allocation, p represents transmission power allocation, and η represents the ST scheduling strategy.
8. The method for allocating intelligent resources of power terminals according to claim 1, characterized in that: The reward function described in step 5 is as follows: R(t) represents R e (t) and R u (t) is the sum of the weights, parameters The value of is adjusted during the training process; Among them, R e (t) represents the demand of CT users, R u (t) represents the ST reliability requirement.
9. The method for intelligent resource allocation of power terminals according to claim 1, characterized in that: The specific steps for solving the optimal resource allocation solution described in step 6 are as follows: (61) Initialize the state space s t , training space θ1, θ2, φ, target network θ1, θ2, φ, experience replay pool R and delay update parameter d; (62) According to the policy network plus noise ε~N(0.6), action a is selected at each time step t ; (63) According to action a t Get rewards t and the next state s t+1 , will (s t ,a t ,r t ,s t+1 ) is stored in the experience replay pool R; (64) Randomly sample N data (s) from the experience replay pool R t ,a t ,r t ,s t+1 ); (65) Use the Bellman equation to calculate the target value y = r + γmin i=1,2 Q θi′ (s′, a′), where γ represents the discount factor; (66) Using the loss function loss = E[(yQ θi (s,a)) 2 ]Update critic network; (67) After every d time steps, calculate the gradient Update the actor network; (68) Use soft update to update the target network parameters: θ′ i =τθ i +(1-τ)θ′ i , i=1,2 and φ'=τφ+(1-τ)φ', τ<<1; (69) Repeat steps 62 to 69 until the maximum number of training steps T is reached.
10. An intelligent resource allocation system for power terminals in ISAC heterogeneous business scenarios, characterized in that: include: Channel monitoring and data acquisition module: It consists of multiple types of sensors deployed at the power terminal, responsible for real-time collection of various physical parameters of the CT link channel, and transmits the data to the CT link channel model calculation unit. At the same time, it can monitor and collect network status indicators related to ST business, providing data support for subsequent ST business analysis; CT link channel model calculation unit: It has a built-in CT link channel model, receives data from the channel monitoring and data acquisition module, calculates the CT user transmission rate using an algorithm, and outputs the result to the resource allocation collaborative control module; Finite block length channel coding solution unit: configures the finite block length channel coding scheme optimized for ST service characteristics, receives part of the ST service related data from the channel monitoring and data acquisition module, uses the coding algorithm to solve the achievable transmission rate of the ST service, and feeds the result back to the resource allocation collaborative control module; ST service analysis and constraint design unit: Use big data analysis and real-time monitoring technology to analyze ST latency requirements, design ST interruption probability constraints, and pass constraint conditions and policy information to the resource allocation collaborative control module; Proportional fairness algorithm and collaborative model building module: Built-in proportional fairness algorithm, receiving business data and parameters from the above units, building CT and ST business collaborative resource allocation model, determining the objective function, and outputting the model and function information to the dual-delay deep deterministic policy gradient algorithm solving unit; Markov model parameter configuration unit: responsible for setting the state space, action space and reward function of the Markov model according to the real-time business status of the power terminal, and passing the configured parameters to the double-delay deep deterministic policy gradient algorithm solving unit; Double-delayed deep deterministic policy gradient algorithm solving unit: Based on the objective function constructed by Markov model feedback and proportional fairness algorithm, the double-delayed deep deterministic policy gradient algorithm is used to solve the optimal resource allocation plan, and the plan is sent to the resource scheduling execution module of the power terminal in real time; Resource scheduling execution module: Receives the optimal resource allocation plan and is responsible for accurately allocating resources to CT users and ST services according to the plan. It has the ability to monitor the execution of resource allocation in real time and dynamically adjust the allocation strategy to deal with emergencies.
Citation Information
Patent Citations
Dynamic multi-access business distributing method in isomerism cooperative network
CN103002465A
Edge resource allocation method and device
CN113760541A
Multi-service joint downlink resource allocation method
CN116489774A
Resource scheduling optimization method based on hybrid digital theory in 5G scene
CN117812727A
Cellular network user association and resource allocation method based on deep reinforcement learning
CN118828603A