Wireless network control system for joint communication control based on deep reinforcement learning
By adding a communication controller and a downlink estimator to the wireless network control system, and using deep reinforcement learning to optimize control input and communication resource allocation, the performance bottleneck of the wireless network control system in scenarios with fading channels and unknown rules of the controlled system is solved, and the system performance is significantly improved.
Patent Information
- Application Number
- CN202511690068.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-18
- Publication Date
- 2026-01-13
AI Technical Summary
Existing wireless network control systems face difficulties in joint communication control in scenarios with fading channels and unknown rules of the controlled system, leading to performance bottlenecks.
A wireless network control system based on deep reinforcement learning is adopted, with the addition of a communication controller and a downlink estimator. Joint communication control is performed using an AC neural network, and control input and communication resource allocation are optimized through a deep deterministic policy gradient algorithm and heterogeneous policy training method.
It significantly improves the performance of the wireless control system, optimizes the quality of uplink and downlink transmission information, and realizes joint optimization of control input, transmission power and bandwidth allocation under the constraint of limited communication resources.
Smart Images

Figure CN121334858A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of wireless control technology, and more specifically, to a wireless network control system based on deep reinforcement learning for joint communication control. Background Technology
[0002] Wireless communication refers to the technology of exchanging information through electromagnetic waves without physical conductors, enabling long-distance communication. Common applications include mobile phone communication, wireless networks, and satellite communication. Today, wireless communication has evolved to the sixth generation (6G) wireless network—which boasts superior characteristics such as ultra-low latency communication, extremely high throughput, large-scale self-organizing networks, and artificial intelligence capabilities, promising to bring higher service quality and entirely new experiences.
[0003] However, the separation of network communication and control design in traditional wireless networks prevents the corresponding wireless network control system from dynamically adjusting to the current channel environment. This results in a performance bottleneck for the wireless network control system, which has extremely high requirements for the wireless channel. To solve this problem, a joint communication control approach has been considered to improve the performance of the wireless control system. However, existing wireless network control systems are prone to fading channels and scenarios where the rules of the controlled system are unknown, making joint communication control difficult and almost impossible to complete the wireless control task. Summary of the Invention
[0004] Therefore, it is necessary to provide a wireless network control system based on deep reinforcement learning to address the difficulty of joint communication control in scenarios with fading channels and unknown rules of the controlled system, which is a problem of existing wireless network control systems.
[0005] This invention is achieved using the following technical solution: This invention discloses a wireless network control system based on deep reinforcement learning for joint communication control, comprising: a base station on the control side and a controller on the controlled side. M Each subsystem includes: a communication controller installed at the base station, and a controllable system located on the controlled side. M A downlink estimator.
[0006] The communication controller includes: an AC neural network module, a control experience pool, and an estimation experience pool. M An uplink estimator.
[0007] The wireless network control system operates periodically; among them, the first n One operating cycle T [ n Including: the first n An upward phase , No. n Strategy generation phase , No.n The downward phase .
[0008] Used for: M Each subsystem uploads its post-execution observation state to the corresponding data. M An uplink estimator, M Each uplink estimator combines uplink historical data provided by the estimation experience pool to improve the quality of uplink transmission information. For: AC neural network module based M The output of each uplink estimator and the control history data provided by the control experience pool are used to generate augmented actions that include control inputs and communication resource allocation using a deep deterministic policy gradient algorithm. Used for: downloading augmentation actions to M A downlink estimator, M Each downlink estimator processes historical downlink data provided by an estimation experience pool to improve the quality of downlink transmission information. M The output of each downlink estimator is transmitted to M Each subsystem is used to implement execution.
[0009] This wireless network control system based on deep reinforcement learning for joint communication control implements the methods or processes according to embodiments of this disclosure.
[0010] Compared with the prior art, the present invention has the following beneficial effects: 1. Based on the existing wireless network control system, this invention adds a communication controller with an uplink estimator to the base station and a downlink estimator to the controlled side to improve the quality of uplink and downlink transmission information. It also utilizes an AC neural network to jointly generate control input and communication resource allocation schemes based on historical data and uplink transmission information, thereby achieving joint optimization of control input, transmission power and bandwidth allocation under the constraint of limited communication resources, which significantly improves the overall performance of the wireless control system.
[0011] 2. The upper and lower estimators of this invention are built based on a long short-term memory neural network, which can predict lost observation data and control input based on corresponding historical data, thereby optimizing the quality of transmitted data.
[0012] 3. This invention adopts a dual-branch structure design for AC neural networks. By combining current input and historical data, it obtains augmented actions that include control input and communication resource allocation with the goal of maximizing system rewards, thereby achieving joint optimization of control input, transmit power and bandwidth allocation.
[0013] 4. This invention also designs a heterogeneous strategy training method for AC neural networks. By first calculating the importance of the experience samples stored in the control experience pool based on time difference error and time slot index, and then reordering them, and then performing random sampling, it fits the strong temporal correlation of the entire control system and helps to improve the performance of AC neural networks in solving augmented actions. Attached Figure Description
[0014] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0015] Figure 1 This is a structural diagram of a wireless network control system based on joint communication control using deep reinforcement learning. Figure 2 for Figure 1 Architecture diagram of the communication controller; Figure 3 for T [ n The composition diagram of ]; Figure 4 For the first m A diagram of the neural network structure of an uplink estimator; Figure 5 This is a diagram of the mesh structure of the online policy network in the AC neural network module. Figure 6 This is a diagram of the mesh structure of the online value network in the AC neural network module. Figure 7 For the first m A diagram of the neural network structure of a downlink estimator. Detailed Implementation
[0016] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0017] It should be noted that when a component is said to be "installed on" another component, it can be directly on the other component or it may be in a component that is centered on it. When a component is said to be "set on" another component, it can be directly set on the other component or it may also be in a component that is centered on it. When a component is said to be "fixed to" another component, it can be directly fixed to the other component or it may also be in a component that is centered on it.
[0018] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. The terminology used herein in the specification of this invention is for the purpose of describing particular embodiments only and is not intended to be limiting of the invention. The term "or / and" as used herein includes any and all combinations of one or more of the associated listed items.
[0019] First, it should be noted that the scenario addressed by this invention is: a remote base station responsible for management. M The control and communication tasks of each subsystem.
[0020] The base station is located on the control side. M If each subsystem is located on the controlled side, then communication between the controlled side and the control side is achieved through the uplink channel, and communication between the control side and the controlled side is achieved through the downlink channel.
[0021] Each subsystem contains a controlled object, an associated sensor, and an associated actuator. That is, the... m The subsystem includes: the first m The actuator, the first m The first controlled object, the first m One sensor; m ∈[1, M ].
[0022] See Figure 1 This demonstrates the wireless network control system based on deep reinforcement learning for joint communication control provided in this embodiment, which adds the following to the above scenario: a component located on the controlled side. M A downlink estimator is installed in the communication controller of the base station.
[0023] First look M A downlink estimator, which is... M Each subsystem has a one-to-one correspondence. Generally, the first subsystem can be considered as... m The downlink estimator is integrated in the first... m In one actuator, and the two are communicatively connected; it can also be designed as follows: the first m The downlink estimator is independent in the first... m The actuators are external and the two are connected in communication. Figure 1 This demonstrates the integrated design of the former.
[0024] Next, let's look at the communication controller, which replaces the base station to take over management. M The control and communication tasks of each subsystem. See also Figure 2 The communication controller includes: an AC neural network module, a control experience pool, and an estimation experience pool. M An uplink estimator.
[0025] In summary, the entire wireless network control system employs the frequency division multiple access protocol. M Each subsystem shares bandwidth for uploading and downloading. Therefore, the entire wireless network control system operates periodically, with each operating cycle divided into three time slots of fixed duration.
[0026] See Figure 3 , with the first n One operating cycle T [ n For example, it includes: the first n An upward phase , No. n Strategy generation phase , No. n The downward phase .
[0027] 1. Used for: M Each subsystem uploads its post-execution observation state to the corresponding data. M An uplink estimator, M An uplink estimator processes uplink historical data provided by an estimation experience pool to improve the quality of uplink transmission information.
[0028] 2. For: AC neural network module based M The output of an uplink estimator and control history data provided by a control experience pool are used to generate augmented actions that include control inputs and communication resource allocations using a deep deterministic policy gradient algorithm.
[0029] 3. Used for: downloading the corresponding augmentation action group to M A downlink estimator, M Each downlink estimator processes historical downlink data provided by an estimation experience pool to improve the quality of downlink transmission information. M The output of each downlink estimator is transmitted to M Each subsystem is used to implement execution.
[0030] The following is a detailed introduction. T [ n The working mode of each component in the system and the corresponding data flow: ①、No.m One sensor is used for: data collection and superimposed To obtain Then With uplink transmission status The m Uplink channel transmission to the first m An uplink estimator.
[0031] in, express T [ n [Middle] m The controlled state of a controlled object; express T [ n Measurement noise (satisfying a Gaussian distribution); express T [ n The observed state (e.g., position, velocity, temperature, and other state parameters).
[0032] In other words, The expression is: .
[0033] It is important to note that the first m Each subsystem is based on , Will Uploaded to m One uplink estimator; among which... express T [ n [Middle] m The transmission power of each uplink channel; express T [ n Assigned to the first m The bandwidth of each subsystem.
[0034] Because it uses a frequency division multiple access method, the first m achievable rate of uplink channel Represented as: ; ; ; In the formula, express T [ n [Middle] m The transmission signal-to-noise ratio of the uplink channel; Represented as the first m Channel gain of each uplink channel;N 0 represents the noise power spectral density; β 0 indicates the channel gain at a reference distance of 1m; d m Indicates the first m The distance of each controlled object from the base station; κ This represents the path loss attenuation index. Let $\mathbf{a}$ be a random variable that follows an exponential distribution and has a mean $\mathbf{a}$. .
[0035] So, let This indicates that the uplink transmission was successful; This indicates that the upstream transmission has been interrupted. The calculation formula is: ; In the formula, Represented as The amount of data transmitted upstream.
[0036] In conclusion, based on , The data is constructed and stored in the control experience pool.
[0037] ② The estimated experience pool is used to: provide information to the first... m An uplink estimator provides To the first m Each downlink estimator provides .in, express T [ n The corresponding historical data; express T [ n The corresponding historical data for the downlink.
[0038] because , Data storage involving other components will not be discussed in detail here.
[0039] ③、No. m An uplink estimator is used for: When it is 1, Direct output is ,exist When it is 0, it is for Perform neural network processing to obtain and will The data is transmitted to the AC neural network module, the estimation experience pool, and the control experience pool. express T [ n The estimated observations.
[0040] because The uplink transmission was successful, meaning there was no data loss, therefore... Direct output is .and The uplink transmission is interrupted, meaning there is data loss, which significantly degrades the performance of the control system. Therefore, neural networks are needed to predict information that cannot be received due to the interruption—that is, information that cannot be received due to the interruption. Perform neural network processing to obtain This ensures the integrity of the uplink data stream and effectively improves the performance of the control system.
[0041] so, The expression is: ; In the formula, Indicates the first m The operation of the neural network within the uplink estimator; Indicates the first m The network parameters of the neural network within the uplink estimator.
[0042] In addition, see Figure 4 , No. m The neural network within the uplink estimator is designed based on a long short-term memory network, specifically including: 2 linear layers, 1 linear correction layer, and 1 long short-term memory layer.
[0043] In the m In the uplink estimator: the first linear layer is used for the input. The linear correction layer processes the output of the first linear layer using the ReLU activation function; the long short-term memory layer processes the output of the linear correction layer using a long short-term memory network; the second linear layer performs a linear transformation on the output of the long short-term memory network layer to obtain... .
[0044] No. m The operation of the neural network within each uplink estimator can also be further represented as: ; In the formula, Linear(.) represents the linear layer processing procedure; LSTM(.) represents the long short-term memory layer processing procedure; and Relu(.) represents the linear correction layer processing procedure.
[0045] Of course, since neural networks are used for prediction, then in T [ n In the middle, the first m An uplink estimator is based on Label data Constructing the minimum mean square error loss function And train the neural network within it using gradient descent, then... When the value is 0, the trained neural network will be used. Processed into .
[0046] in, The expression is: ; In the formula, This represents the L2 norm.
[0047] ④ The control experience pool is used to: provide the AC neural network module with... .in, express T [ n [Control historical data in the middle.]
[0048] because Data storage involving other components will not be discussed in detail here.
[0049] ⑤ The AC neural network module is used to: employ the Deep Deterministic Policy Gradient (DDPG) algorithm based on , Solve the objective optimization problem P1 to obtain ;Will via downlink transmission status The m downlink channel transmission to the first m A downlink estimator; will , , Transmitted to the control experience pool.
[0050] in, express T [ n The first m An augmentation action; express T [ n [The text appears to be incomplete and contains several grammatical errors. A more accurate translation would require the full context.] m Control inputs for each subsystem; express T [ n +1] allocated to the first m The bandwidth of each subsystem; express T [ n +1] in the middle m The transmission power of each uplink channel; express T [ n +1] in the middle mThe transmission power of the downlink channel.
[0051] For the entire control system, the objective optimization problem P1 is: to obtain the maximum long-term cumulative reward while satisfying the constraints of limited communication resources.
[0052] So, let T [ n [Middle] m The augmented instantaneous reward of each subsystem is ,and The expression is: ; in, The first two items constitute the control reward, and the last item is the communication cost.
[0053] In the formula, Represents the value function; w m This represents the penalty coefficient (the value is determined based on the actual situation, but is generally 10), which is used to balance control rewards and communication costs.
[0054] but, M Total rewards of each subsystem ; Therefore, the expression for the objective optimization problem P1 is: ; In the formula, J This indicates a long-term cumulative reward. γ A discount factor representing the impact of future rewards on the present; This indicates the maximum uplink transmission power (the value is determined based on the actual situation, but is generally taken as 10W). This indicates the maximum downlink transmission power (the value is determined based on the actual situation, but is generally taken as 20W). B This indicates the maximum bandwidth (the value is determined based on the actual situation, but is generally 20MHz).
[0055] Therefore, the AC neural network module actually solves P1 through deep reinforcement learning based on the Actor-Critic algorithm. In this way, regardless of the specific rules of the subsystem, it can adapt through continuous iteration of deep reinforcement learning.
[0056] Specifically, the AC neural network module includes: an online policy network, a target policy network, an online value network, and a target value network. The online policy network and the target policy network have the same structure; the online value network and the target value network also have the same structure.
[0057] AC neural network module in Construct augmented state group Augmentation motion group The online policy network and online value network are trained using a heterogeneous policy training method, and the target policy network and target value network are updated using a soft update method. in, ; .
[0058] The AC neural network module can use existing commonly used neural networks, or it can design its own structure to improve network performance.
[0059] See Figure 5 This demonstrates the self-designed online strategy network structure diagram, which includes: 4 linear layers, 3 linear correction layers, 1 long short-term memory layer, 1 aggregation layer, 1 splitting layer, 1 hyperbolic tangent function layer, 1 sigmoid function layer, 2 flexible maximum function layers, and 4 product layers.
[0060] like Figure 5 As shown, the inputs to the online policy network include INPUT1~INPUT2, and the output includes OUTPUT.
[0061] in, ; ; .
[0062] Specifically, in the online policy network: the first linear layer is used as input to INPUT1; the first linear correction layer processes the output of the first linear layer using the ReLU activation function; the long short-term memory layer processes the output of the first linear correction layer using the long short-term memory network; the second linear layer is used as input to INPUT2; the second linear correction layer processes the output of the second linear layer using the ReLU activation function; the aggregation layer aggregates the outputs of the long short-term memory layer and the second linear correction layer to obtain aggregated data; the third linear layer is used as input to the aggregated data; the third linear correction layer processes the output of the third linear layer using the ReLU activation function. The output is processed; the fourth linear layer performs a linear transformation on the output of the third linear correction layer; the splitting layer is used to: split the output of the fourth linear layer into a control input part, an uplink power allocation part, a downlink power allocation part, and a bandwidth allocation part; the hyperbolic tangent function layer is used to process the control input part using the hyperbolic tangent function; the sigmoid function layer is used to process the uplink power allocation part using the sigmoid function; the first flexible maximum function layer is used to process the downlink power allocation part using the flexible maximum function; the second flexible maximum function layer is used to process the bandwidth allocation part using the flexible maximum function; the first product layer is used to combine the output of the hyperbolic tangent function layer with the action maximum value. Multiply to obtain The second product layer is used to combine the output of the sigmoid function with... Multiply to obtain The third product layer is used to combine the output of the first flexible maximum function with... Multiply to obtain The fourth product layer is used to combine the output of the second flexible maximum function with... B Multiply to obtain ; , , , That is, to form OUTPUT.
[0063] See Figure 6 This demonstrates a self-designed online value network structure diagram, which includes: 4 linear layers, 3 linear correction layers, 1 long short-term memory layer, and 1 aggregation layer.
[0064] like Figure 6 As shown, the inputs of the online value network include IN1~IN2, and the output includes OUT.
[0065] in, ; ; ; Value [ n ] indicates the generated Q The values are used to construct the loss function for updating the online policy network.
[0066] In the online value network: the first linear layer is used as input IN1; the first linear correction layer processes the output of the first linear layer through the ReLU activation function; the long short-term memory layer processes the output of the first linear correction layer through the long short-term memory network; the second linear layer is used as input IN2; the second linear correction layer processes the output of the second linear layer through the ReLU activation function; the aggregation layer aggregates the output of the long short-term memory layer and the output of the second linear correction layer to obtain aggregated data; the third linear layer is used as input to the aggregated data; the third linear correction layer processes the output of the third linear layer through the ReLU activation function; the fourth linear layer performs a linear transformation on the output of the third linear correction layer to obtain OUT.
[0067] Furthermore, since a heterogeneous policy training method is used, it can be designed such that during training, Δ is stored in the control experience pool for each direction (a region can be partitioned off from the control experience pool for storage). N After a certain number of experience samples, the online policy network is updated by sampling from the control experience pool according to a priority sampling strategy; Δ N This indicates the quantity threshold.
[0068] Among them, Time-directed control experience pool stores experience samples .
[0069] In this embodiment, the priority sampling strategy is designed based on time difference error and time slot index. Specifically, it includes: reordering all experience samples in the control experience pool according to their importance, and then performing random sampling.
[0070] The formula for calculating importance is as follows: ; In the formula, This indicates the first control experience pool before sorting. i The importance of each empirical sample; sigmoid(.) represents the sigmoid function; This indicates the first control experience pool before sorting. i The time difference error of each empirical sample; index( i ) indicates the first control experience pool before sorting. i A time-series index of an empirical sample; N C This indicates the length of the region used to store experience samples in the control experience pool.
[0071] The sampling probability of random sampling is: ; In the formula, p (j) This indicates the sorted control experience pool. j The sampling probability of an empirical sample; α This indicates the degree of priority given to its use.
[0072] This priority sampling strategy, which aligns with the strong temporal correlation of the entire control system, helps improve the performance of AC neural networks in solving augmented actions.
[0073] It should be noted that the AC neural network module is based on , Will Download to the m One downlink estimator; among which... express T [ n [Middle] m The transmission power of the downlink channel.
[0074] Because it uses a frequency division multiple access method, the first m achievable rate of downlink channel Represented as: ; ; ; In the formula, express T [ n [Middle] m The transmission signal-to-noise ratio of each downlink channel; Represented as the first m Channel gain of each downlink channel.
[0075] So, let This indicates that the downlink transmission was successful; This indicates that the downlink transmission has been interrupted. The calculation formula is: ; In the formula, Represented as The amount of data transmitted downlink.
[0076] In conclusion, based on , The data is constructed and stored in the control experience pool.
[0077] ⑥、No. m A downlink estimator is used for: When it is 1, Direct output is ,exist When it is 0, it is for Perform neural network processing to obtain ;Will Continue transmitting to the first m One actuator; will The data is fed back to the estimation experience pool and the control experience pool; among them, express T [ n [The text appears to be incomplete and contains several grammatical errors. A more accurate translation would require the full context.] m The predicted inputs for each subsystem.
[0078] because The current downlink transmission is successful, meaning there was no data loss, therefore... Direct output is .and The downlink transmission is interrupted, meaning there is data loss, which significantly degrades the performance of the control system. Therefore, neural networks are needed to predict information that cannot be received due to the interruption—that is, information that cannot be received due to the interruption. Perform neural network processing to obtain This ensures the integrity of the downlink data stream and effectively improves the performance of the control system.
[0079] so, The expression is: ; In the formula, Indicates the first m The operation of the neural network within the downlink estimator; Indicates the first m The network parameters of the neural network within the downlink estimator.
[0080] In addition, see Figure 7 , No. m The neural network within the downlink estimator is designed based on a long short-term memory network, specifically including: 2 linear layers, 1 linear correction layer, and 1 long short-term memory layer.
[0081] In the m In the downlink estimator: the first linear layer is used for the input. The linear correction layer processes the output of the first linear layer using the ReLU activation function; the long short-term memory layer processes the output of the linear correction layer using a long short-term memory network; the second linear layer performs a linear transformation on the output of the long short-term memory network layer to obtain... .
[0082] No. m The operation of the neural network within each downlink estimator can also be further represented as: ; In the formula, Linear(.) represents the linear layer processing procedure; LSTM(.) represents the long short-term memory layer processing procedure; and Relu(.) represents the linear correction layer processing procedure.
[0083] Of course, since neural networks are used for prediction, then in T [ n In the middle, the first m A downlink estimator is based on Label data Constructing the minimum mean square error loss function And train the neural network within it using gradient descent, then... When the value is 0, the trained neural network will be used. Processed into .
[0084] in, The expression is: .
[0085] ⑦、No. m One actuator is used for: according to For the first m Each controlled object is executed to generate ;in, express T [ n +1] in the middle m The controlled state of a controlled object.
[0086] Satisfy the first m The discrete state transition equation of a controlled object is expressed as follows: ; In the formula, A nonlinear function representing a state transition; express T [ n The system perturbation (satisfies a Gaussian distribution).
[0087] In addition, for , , Its specific composition is as follows: I. This includes: estimated observations and predictive inputs for historical operating cycles. That is, The expression is: ; In the formula, L This indicates the threshold number of cycles.
[0088] because n Is it from 1 or counting, then in n Not achieved L When needed Use 0 as a placeholder. It's important to note that... In the other case, L Group (0,0).
[0089] II. This includes: predictive inputs from historical operating cycles. That is, The expression is: ; and Similarly, due to n Is it from 1 or counting, then in n Not achieved L hour We need to use 0 as a placeholder. It's important to note that... In the other case, L Group (0,0).
[0090] III. This includes: estimated observations of historical operating cycles, uplink transmission status, downlink transmission status, predicted input, uplink transmission power, downlink transmission power, and bandwidth. In other words, The expression is: ; In the formula, Indicates inclusion , The power set; .
[0091] and , Similarly, due to n Is it from 1 or counting, then in n Not achieved L hour We need to use 0 as a placeholder. It's important to note that... In the other case, L Group (0,0).
[0092] In summary, as the operating cycle increases, deep reinforcement learning converges, and the entire control system gradually acquires the ability to make decisions in the faceted wireless channel. This enables joint optimization of control input, transmit power, and bandwidth allocation under the constraint of limited communication resources, thereby significantly improving the overall performance of the wireless control system.
[0093] Furthermore, simulation experiments were conducted on the aforementioned wireless network control system based on deep reinforcement learning for joint communication control, and the results were compared with existing methods. The results show that the wireless network control system based on deep reinforcement learning for joint communication control proposed in this invention achieves higher control rewards, lower communication costs, and higher system rewards. This demonstrates that it achieves joint optimization of control input, transmit power, and bandwidth allocation under limited communication resource constraints, thereby significantly improving the overall performance of the wireless control system.
[0094] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0095] The embodiments described above are merely illustrative of several implementations of the present invention, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these all fall within the protection scope of the present invention. Therefore, the protection scope of this invention patent should be determined by the appended claims.
Claims
1. A wireless network control system based on deep reinforcement learning for joint communication control, comprising: Base stations located on the control side, and base stations located on the controlled side M A subsystem; characterized in that it further includes: Located on the controlled side M A downlink estimator; and The communication controller installed at the base station includes: an AC neural network module, a control experience pool, and an estimation experience pool. M One uplink estimator; The wireless network control system operates periodically; wherein, the first n One operating cycle T [ n Including: the first n An upward phase , No. n Strategy generation phase , No. n The downward phase ; Used for: M Each subsystem uploads its post-execution observation state to the corresponding data. M An uplink estimator, M Each uplink estimator combines uplink historical data provided by the estimation experience pool to improve the quality of uplink transmission information. For: AC neural network module based M The output of each uplink estimator and the control history data provided by the control experience pool are used to generate augmented action groups containing control inputs and communication resource allocations using a deep deterministic policy gradient algorithm. Used for: downloading the corresponding augmentation action group to M A downlink estimator, M Each downlink estimator processes historical downlink data provided by an estimation experience pool to improve the quality of downlink transmission information. M The output of each downlink estimator is transmitted to M Each subsystem is used to implement execution.
2. A wireless network control system based on deep reinforcement learning for joint communication control according to claim 1, characterized in that: No. m The subsystem includes: the first m The actuator, the first m The first controlled object, the first m One sensor; m ∈[1, M ]; No. m The downlink estimator is integrated in the first... m There are multiple actuators, and the two are communicatively connected; Or, the first m The downlink estimator is independent in the first... m The actuators are external and the two are connected in communication.
3. A wireless network control system based on deep reinforcement learning for joint communication control according to claim 2, characterized in that: exist T [ n ]middle: No. m One sensor is used for: data collection and superimposed To obtain Then With uplink transmission status The m Uplink channel transmission to the first m One uplink estimator; among which... express T [ n [Middle] m The controlled state of a controlled object; express T [ n Measurement noise; express T [ n The observation status; based on , The system is constructed and stored in the control experience pool; express T [ n Assigned to the first m The bandwidth of each subsystem; express T [ n [Middle] m The transmission power of each uplink channel; The estimated experience pool is used to: provide information to the first... m An uplink estimator provides To the first m Each downlink estimator provides ;in, express T [ n The corresponding historical data; express T [ n The corresponding historical downlink data; No. m An uplink estimator is used for: When it is 1, Direct output is ,exist When it is 0, it is for Perform neural network processing to obtain and will The data is transmitted to the AC neural network module, the estimation experience pool, and the control experience pool; among which, express T [ n Estimated observations; The control experience pool is used to: provide the AC neural network module ;in, express T [ n [Control historical data in the middle;] The neural network module is used to: employ a deep deterministic gradient algorithm based on... , Solve the objective optimization problem P1 to obtain ;Will via downlink transmission status The m downlink channel transmission to the first m A downlink estimator; will , , Transmitted to the control experience pool; among which, express T [ n The first m An augmentation action; express T [ n [The text appears to be incomplete and contains several grammatical errors. A more accurate translation would require the full context.] m Control inputs for each subsystem; express T [ n +1] allocated to the first m The bandwidth of each subsystem; express T [ n +1] in the middle m The transmission power of each uplink channel; express T [ n +1] in the middle m The transmission power of each downlink channel; based on , The system is constructed and stored in the control experience pool; express T [ n [Middle] m The transmission power of each downlink channel; No. m A downlink estimator is used for: When it is 1, Direct output is ,exist When it is 0, it is for Perform neural network processing to obtain ;Will Continue transmitting to the first m One actuator; will The data is fed back to the estimation experience pool and the control experience pool; among them, express T [ n [The text appears to be incomplete and contains several grammatical errors. A more accurate translation would require the full context.] m Predicted inputs for each subsystem; No. m One actuator is used for: according to For the first m Each controlled object is executed to generate ;in, express T [ n +1] in the middle m The controlled state of a controlled object.
4. A wireless network control system based on deep reinforcement learning for joint communication control according to claim 3, characterized in that: The expression is: ; The calculation formula is: ; In the formula, Represented as Uplink data volume; Represented as the first m Channel gain of each uplink channel; N 0 represents the noise power spectral density; The expression is: ; In the formula, L Indicates the threshold number of cycles; The expression is: ; The expression is: ; In the formula, Indicates the first m The operation of the neural network within the uplink estimator; Indicates the first m The network parameters of the neural network within the uplink estimator; The expression is: ; In the formula, Indicates inclusion , The power set; ; The expression for P1 is: ; In the formula, J This indicates a long-term cumulative reward. γ A discount factor representing the impact of future rewards on the present; express M The sum of rewards for each subsystem; Indicates the maximum uplink transmission power; Indicates the maximum downlink transmission power; B Indicates the maximum bandwidth; The calculation formula is: ; In the formula, Represented as Downlink data volume; Represented as the first m Channel gain of each downlink channel; The expression is: ; In the formula, Indicates the first m The operation of the neural network within the downlink estimator; Indicates the first m Network parameters of the neural network within the downlink estimator; The expression is: ; In the formula, A nonlinear function representing a state transition; express T [ n System disturbances.
5. A wireless network control system based on deep reinforcement learning for joint communication control according to claim 4, characterized in that: , The expression is: ; In the formula, β 0 indicates the channel gain at a reference distance of 1m; d m Indicates the first m The distance of each controlled object from the base station; κ This represents the path loss attenuation index. Let $\mathbf{a}$ be a random variable that follows an exponential distribution and has a mean $\mathbf{a}$. ; The expression is: ; In the formula, express T [ n [Middle] m Augmented instantaneous rewards for each subsystem; Represents the value function; w m This represents the coefficient of the penalty term.
6. A wireless network control system based on deep reinforcement learning for joint communication control according to claim 4, characterized in that: No. m The neural network within the uplink estimator consists of: 2 linear layers, 1 linear correction layer, and 1 long short-term memory layer; In the m In the uplink estimator: the first linear layer is used for the input. The linear correction layer processes the output of the first linear layer using the ReLU activation function; the long short-term memory layer processes the output of the linear correction layer using a long short-term memory network; the second linear layer performs a linear transformation on the output of the long short-term memory network layer to obtain... ; No. m The neural network within the downlink estimator consists of: two linear layers, one linear correction layer, and one long short-term memory layer; In the m In the downlink estimator: the first linear layer is used for the input. The linear correction layer processes the output of the first linear layer using the ReLU activation function; the long short-term memory layer processes the output of the linear correction layer using a long short-term memory network; the second linear layer performs a linear transformation on the output of the long short-term memory network layer to obtain... .
7. A wireless network control system based on deep reinforcement learning for joint communication control according to claim 4, characterized in that: exist T [ n ]middle, No. m An uplink estimator is based on Label data Constructing the minimum mean square error loss function And train the neural network within it using gradient descent, then... When the value is 0, the trained neural network will be used. Processed into ; No. m A downlink estimator is based on Label data Constructing the minimum mean square error loss function And train the neural network within it using gradient descent, then... When the value is 0, the trained neural network will be used. Processed into .
8. A wireless network control system based on deep reinforcement learning for joint communication control according to claim 4, characterized in that: The AC neural network module includes: online policy network, target policy network, online value network, and target value network; AC neural network module in Construct augmented state group Augmentation motion group The online policy network and online value network are trained using a heterogeneous policy training method, and the target policy network and target value network are updated using a soft update method. in, ; .
9. A wireless network control system based on deep reinforcement learning for joint communication control according to claim 8, characterized in that: The online policy network and the target policy network have the same structure, both including: 4 linear layers, 3 linear correction layers, 1 long short-term memory layer, 1 aggregation layer, 1 splitting layer, 1 hyperbolic tangent function layer, 1 sigmoid function layer, 2 flexible maximum function layers, and 4 product layers; The inputs to the online policy network include INPUT1~INPUT2, and the output includes OUTPUT. in, ; ; ; In the online policy network: the first linear layer is used as input to INPUT1; the first linear correction layer processes the output of the first linear layer using the ReLU activation function; the long short-term memory layer processes the output of the first linear correction layer using the long short-term memory network; the second linear layer is used as input to INPUT2; the second linear correction layer processes the output of the second linear layer using the ReLU activation function; the aggregation layer aggregates the outputs of the long short-term memory layer and the second linear correction layer to obtain aggregated data; the third linear layer is used as input to the aggregated data; the third linear correction layer processes the output of the third linear layer using the ReLU activation function. The process involves: The fourth linear layer performs a linear transformation on the output of the third linear correction layer; the splitting layer divides the output of the fourth linear layer into a control input, uplink power allocation, downlink power allocation, and bandwidth allocation; the hyperbolic tangent function layer processes the control input using the hyperbolic tangent function; the sigmoid function layer processes the uplink power allocation using the sigmoid function; the first flexible maximum function layer processes the downlink power allocation using the flexible maximum function; the second flexible maximum function layer processes the bandwidth allocation using the flexible maximum function; and the first product layer combines the output of the hyperbolic tangent function layer with the action maximum value. Multiply to obtain The second product layer is used to combine the output of the sigmoid function with... Multiply to obtain The third product layer is used to combine the output of the first flexible maximum function with... Multiply to obtain The fourth product layer is used to combine the output of the second flexible maximum function with... B Multiply to obtain ; , , , That is, to form an OUTPUT; The online value network and the target value network have the same structure, both including: 4 linear layers, 3 linear correction layers, 1 long short-term memory layer, and 1 aggregation layer; The inputs to the online value network include IN1~IN2, and the output includes OUT. in, ; ; ; Value [ n ] indicates the generated Q value; In the online value network: the first linear layer is used as input IN1; the first linear correction layer processes the output of the first linear layer through the ReLU activation function; the long short-term memory layer processes the output of the first linear correction layer through the long short-term memory network; the second linear layer is used as input IN2; the second linear correction layer processes the output of the second linear layer through the ReLU activation function; the aggregation layer aggregates the output of the long short-term memory layer and the output of the second linear correction layer to obtain aggregated data; the third linear layer is used as input to the aggregated data; the third linear correction layer processes the output of the third linear layer through the ReLU activation function; the fourth linear layer performs a linear transformation on the output of the third linear correction layer to obtain OUT.
10. A wireless network control system based on deep reinforcement learning for joint communication control according to claim 9, characterized in that: During training, Δ is stored in the control experience pool for each direction. N After a certain number of experience samples, the online policy network is updated by sampling from the control experience pool according to a priority sampling strategy; Δ N Indicates a quantity threshold; Among them, Time-directed control experience pool stores experience samples ; The priority sampling strategies include: All experience samples in the control experience pool are reordered according to their importance, and then random sampling is performed. The formula for calculating importance is as follows: ; In the formula, This indicates the first control experience pool before sorting. i The importance of each empirical sample; sigmoid(.) represents the sigmoid function; This indicates the first control experience pool before sorting. i The time difference error of each empirical sample; index( i ) indicates the first control experience pool before sorting. i A time-series index of an empirical sample; N C This indicates the length of the region in the control experience pool used to store experience samples. The sampling probability of random sampling is: ; In the formula, p (j) This indicates the sorted control experience pool. j The sampling probability of an empirical sample; α This indicates the degree of priority given to its use.