Control performance-oriented reinforcement learning network training and transmission bandwidth allocation method
By using a reinforcement learning network training method oriented towards control performance, the communication bandwidth in the unmanned operation system is rationally allocated, solving the problem of low resource utilization in the existing technology and realizing efficient resource utilization and improved machine operation performance in the unmanned operation system.
Patent Information
- Application Number
- CN202511108057.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-07
- Publication Date
- 2025-10-31
AI Technical Summary
In unmanned operation scenarios, existing communication bandwidth allocation strategies fail to maximize machine operation performance. They only optimize network communication performance while ignoring the closed-loop process of perception, communication, computing and control, resulting in low resource utilization.
A reinforcement learning network training method oriented towards control performance is adopted. By determining the channel transmission bandwidth between the communication device and multiple sensors, the weights of the reinforcement learning network are optimized. With the control performance of the communication device on the execution unit as the objective, bandwidth resources are rationally allocated by combining the perception-communication-computation-control closed-loop process.
It improves the resource utilization and machine operation performance of unmanned operation systems, and enhances the control performance of communication equipment over execution units by optimizing bandwidth allocation strategies.
Smart Images

Figure CN120880915A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of bandwidth allocation technology, and in particular to a method, apparatus, device, and computer program product for training and allocating transmission bandwidth for reinforcement learning networks oriented towards control performance. Background Technology
[0002] With the continuous development of automation and robotics technologies, the demand for unmanned operations in remote areas is growing. Unmanned operations require the deployment of multiple sensors to obtain global environmental information, which guides the machine to operate in complex environments. This necessitates the construction of a network of aerial infrastructure such as satellites and drones to control the machine and enable it to perform its tasks.
[0003] However, for satellite and drone networks, communication bandwidth is severely limited, and how to effectively allocate bandwidth to meet the needs of multiple sensors has become an urgent problem to be solved.
[0004] In related technologies, the communication bandwidth is usually evenly distributed among multiple sensors, or the bandwidth allocation strategy is determined with the rate of the bandwidth communication channel as the optimization target, so as to allocate the bandwidth.
[0005] However, in the closed-loop control process of machine operation, sensing, communication, computing and control are closely related processes. Optimizing only the network communication performance is not necessarily the optimal allocation scheme for communication resources, nor can it maximize the machine's operating performance. Summary of the Invention
[0006] To overcome the problems existing in related technologies, this disclosure provides a method, apparatus, device, and computer program product for training and allocating transmission bandwidth for reinforcement learning networks oriented towards control performance, which can solve the above-mentioned problems.
[0007] According to a first aspect of the present disclosure, a reinforcement learning network training method for control performance is provided. The reinforcement learning network is used to determine the channel transmission bandwidth between a communication device and multiple sensors, wherein the communication device controls the operation of an execution unit based on observation data from the multiple sensors. The method includes: determining an environmental state value for a first control period based on environmental parameters of the first control period; determining a corresponding bandwidth allocation strategy based on the environmental state value of the first control period; evolving the environment according to the bandwidth allocation strategy; determining a reward for the environmental state value of the first control period based on the control performance of the communication device on the execution unit; and determining an environmental state value for a second control period, wherein the second control period is the next control period after the first control period; determining interaction information between the execution unit and the environment, the interaction information including the environmental state value of the first control period, the reward of the first control period, the bandwidth allocation strategy of the first control period, and the environmental state value of the second control period; and optimizing the weights of the reinforcement learning network according to the interaction information.
[0008] According to a second aspect of the present disclosure, a transmission bandwidth allocation method is provided, executed by a communication device, wherein the communication device controls the operation of an execution unit based on observation data from multiple sensors. The method includes: determining environmental parameters and determining environmental state values; inputting the environmental state values into a reinforcement learning network trained through a first aspect embodiment, determining a bandwidth allocation strategy output by the reinforcement learning network; determining a transmission bandwidth allocation result for the multiple sensors based on the bandwidth allocation strategy, and allocating bandwidth according to the transmission bandwidth allocation result.
[0009] According to a third aspect of the present disclosure, a reinforcement learning network training apparatus for control performance is provided. The reinforcement learning network is used to determine the channel transmission bandwidth between a communication device and multiple sensors, wherein the communication device controls the operation of an execution unit based on observation data from the multiple sensors. The apparatus includes: a policy estimation module configured to determine an environmental state value for a first control period based on environmental parameters of the first control period, and to determine a corresponding bandwidth allocation policy based on the environmental state value of the first control period; a reward estimation module configured to evolve the environment according to the bandwidth allocation policy, determine a reward for the environmental state value of the first control period based on the control performance of the communication device on the execution unit, and determine an environmental state value for a second control period, wherein the second control period is the next control period after the first control period; an information determination module configured to determine interaction information between the execution unit and the environment, the interaction information including the environmental state value of the first control period, the reward of the first control period, the bandwidth allocation policy of the first control period, and the environmental state value of the second control period; and a parameter optimization module configured to optimize the weights of the reinforcement learning network based on the interaction information.
[0010] According to a fourth aspect of the present disclosure, an electronic device is provided, comprising: a processor and a memory; the memory being used to store a computer program; and the processor being used to execute, by invoking the computer program, a reinforcement learning network training method for control performance as described in the first aspect.
[0011] According to a fifth aspect of the present disclosure, a computer-readable storage medium is provided having a computer program stored thereon that, when executed by a processor, implements the reinforcement learning network training method for control performance as described in the first aspect.
[0012] According to a sixth aspect of the present disclosure, a computer program product is provided, including a computer program that, when executed by a processor, implements the method as described in the first aspect.
[0013] The technical solutions provided by the embodiments of this disclosure may include the following beneficial effects:
[0014] The reinforcement learning network training method proposed in this disclosure, which focuses on control performance, is a reinforcement learning network designed for unmanned operation scenarios to determine the allocation strategy of channel transmission bandwidth between communication devices and multiple sensors. In the training phase of the reinforcement learning network, this disclosure optimizes the reinforcement learning network by using the control performance of the communication device over the execution unit as the optimization objective.
[0015] To optimize reinforcement learning networks, sample data is required. In this disclosure, the sample data is interactive information, which contains four values: the bandwidth allocation strategy adopted in the current control cycle, the environmental state value before adopting the bandwidth allocation strategy in the current control cycle, the environmental state value after adopting the bandwidth allocation strategy in the current control cycle, and the reward corresponding to the current bandwidth allocation strategy. Based on this interactive information containing these four values, optimization can be achieved with the goal of improving the control performance of the communication device on the execution unit.
[0016] In unmanned operation scenarios, the process typically involves four stages: perception, communication, computation, and control. Related technologies, when considering bandwidth allocation strategies, focus solely on the communication process, optimizing communication efficiency. However, since the entire operation scenario includes computation and control after communication, optimal communication efficiency does not necessarily equate to optimal control results. In contrast to related technologies, this disclosure optimizes the reinforcement learning network by taking the control performance of the communication device over the execution unit as the optimization objective and the final control result of the entire operation scenario as the optimization object. This reinforcement learning network training method enables the bandwidth allocation strategy output by the trained reinforcement learning network to consider each stage of the machine's closed-loop operation process, thereby maximizing machine performance and improving resource utilization.
[0017] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description
[0018] The accompanying drawings, which are incorporated in and form part of this disclosure, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure.
[0019] Figure 1 This is a schematic diagram illustrating a scenario of operation of a communication device control execution unit according to an exemplary embodiment of the present disclosure.
[0020] Figure 2 This is a schematic flowchart illustrating a reinforcement learning network training method for control performance according to an exemplary embodiment of the present disclosure.
[0021] Figure 3 This is a logical illustration of a reinforcement learning network training method for control performance, as shown in an exemplary embodiment of this disclosure.
[0022] Figure 4 This disclosure is a schematic flowchart illustrating a transmission bandwidth allocation according to an exemplary embodiment.
[0023] Figure 5 This is a schematic diagram illustrating the result of a bandwidth allocation scheme according to an exemplary embodiment of the present disclosure.
[0024] Figure 6 This disclosure is a block diagram illustrating a reinforcement learning network training apparatus for control performance according to an exemplary embodiment.
[0025] Figure 7 This disclosure is a schematic block diagram illustrating a reinforcement learning network training apparatus for control performance according to an exemplary embodiment. Detailed Implementation
[0026] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numerals in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this disclosure. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this disclosure as detailed in the appended claims.
[0027] The terminology used in this disclosure is for the purpose of describing particular embodiments only and is not intended to be limiting of the disclosure. The singular forms “a,” “the,” and “the” as used in this disclosure and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to and includes any and all possible combinations of one or more of the associated listed items.
[0028] It should be understood that although the terms first, second, third, etc., may be used in this disclosure to describe various information, such information should not be limited to these terms. These terms are used only to distinguish information of the same type from one another. For example, without departing from the scope of this disclosure, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to determination."
[0029] With the continuous development of automation and robotics technologies, the demand for unmanned operations in remote areas is increasing, such as environmental monitoring, disaster relief, and mining automation. Because the capabilities of a single machine are limited, making it difficult to simultaneously handle both tasks and perception, unmanned operations often require the deployment of multiple sensors. These sensors obtain comprehensive environmental information, which guides the machine to cope with complex environments and complete the task.
[0030] However, in remote and disaster-stricken areas, ground-based communication infrastructure is often difficult to deploy or is damaged. Therefore, aerial infrastructure such as satellites and drones is needed to build communication networks to serve machines and instruct them to complete tasks. For satellite-drone networks, communication bandwidth is severely limited. Therefore, how to effectively allocate bandwidth in communication channels to meet the communication needs of multiple sensors has become a pressing problem. Through reasonable bandwidth allocation, sensor data can be effectively utilized, improving machine performance and reducing communication pressure. Therefore, researching the bandwidth allocation problem for multiple sensors to meet the requirements of machine operation in environments with limited wireless communication resources is a key issue currently under extensive research.
[0031] Figure 1 This is a schematic diagram illustrating a scenario of operation of a communication device control execution unit according to an embodiment of the present disclosure.
[0032] like Figure 1 As shown, the unmanned operation system includes:
[0033] Communication equipment, which can be a drone, can carry a processor (such as a mobile edge computing server) and a communication module. During operation, the communication equipment is responsible for receiving transmission data sent by multiple sensors, analyzing the received transmission data, and generating control commands. The generated control commands are then transmitted to the execution unit (such as a robot) at the work site through the communication module, thereby controlling the execution unit to perform the operation.
[0034] Satellites can monitor and control communication devices (such as drones) via telemetry and control links;
[0035] Multiple sensors are used to send the collected sensing data to the communication device during the operation of the execution unit, so that the processor on the communication device can perform sensing data fusion and enable the communication device to determine appropriate control commands.
[0036] An execution unit, such as a robot, is used to receive control commands sent by a communication device and perform tasks on the controlled object.
[0037] The various components of the unmanned operation system form a closed-loop control system of "perception-communication-computation-control," which operates cyclically to instruct the execution units to complete the task. However, during this process, the communication bandwidth between the communication equipment and multiple sensors is limited. How to rationally allocate these limited bandwidth resources to improve the operational performance of the execution units is the main challenge in this scenario.
[0038] Related technologies for multi-sensor bandwidth allocation often focus on network-side communication capabilities, such as communication capacity or latency. However, in the closed-loop control process of machine operation, sensing, communication, computing, and control are all closely related. Optimizing only the communication process may not maximize the overall efficiency of machine operation, making it difficult to achieve optimal resource utilization and maximize machine performance.
[0039] To address the aforementioned technical issues, this disclosure proposes a transmission bandwidth allocation method, which requires implementation based on a trained reinforcement learning network. The training of this reinforcement learning network will be described below.
[0040] Figure 2 This is a schematic flowchart illustrating a reinforcement learning network training method for control performance according to embodiments of the present disclosure. The reinforcement learning network is used to determine the channel transmission bandwidth between a communication device and multiple sensors, and the communication device controls the operation of an execution unit based on observation data from the multiple sensors. The reinforcement learning network training method can be as follows: Figure 1 The processor (e.g., a mobile edge computing server) on the communication device (e.g., a drone) shown in the diagram performs the operation.
[0041] like Figure 2 As shown, reinforcement learning network training methods include:
[0042] In step S201, the environmental state value of the first control cycle is determined according to the environmental parameters of the first control cycle, and the corresponding bandwidth allocation strategy is determined based on the environmental state value of the first control cycle.
[0043] In step S202, the environment is evolved according to the bandwidth allocation strategy, the reward for the environment state value of the first control cycle is determined based on the control performance of the communication device on the execution unit, and the environment state value of the second control cycle is determined, wherein the second control cycle is the next control cycle after the first control cycle.
[0044] In step S203, the interaction information between the execution unit and the environment is determined. The interaction information includes the environment state value of the first control cycle, the reward of the first control cycle, the bandwidth allocation strategy of the first control cycle, and the environment state value of the second control cycle.
[0045] In step S204, the weights of the reinforcement learning network are optimized based on the interaction information.
[0046] It should be noted that, in the application scenario of this disclosure, the communication device can determine the control instructions to be sent to the execution unit based at least on observation data acquired from multiple sensors. These control instructions are used to control the execution unit to perform the target operation. However, the method by which the communication device determines the control instructions, and the method by which the execution unit is controlled based on the control instructions, are not the focus of this disclosure. This disclosure focuses on protecting a suitable bandwidth allocation strategy employed by the communication device to allocate bandwidth resources to multiple sensors under limited bandwidth resources.
[0047] In some embodiments, the reinforcement learning network training method of this disclosure is executed by the communication device after setting up multiple sensors, communication devices, and execution units.
[0048] Many parameters in different unmanned operation systems will change. Therefore, for the target scenario, the communication equipment needs to train a suitable reinforcement learning network based on the parameters obtained on site, and then determine the bandwidth allocation strategy based on the trained reinforcement learning network.
[0049] In some embodiments, the weights of the reinforcement learning network are optimized based on the interaction information.
[0050] Figure 3 This is a logical schematic diagram of a reinforcement learning network training method for control performance, as shown in an embodiment of the present disclosure.
[0051] like Figure 3 As shown, the core purpose of this disclosure is to obtain interactive information (sample data) that can be used for training reinforcement learning networks. Once suitable interactive information is obtained, it can be used to train reinforcement learning networks.
[0052] In some embodiments, interaction information between the execution unit and the environment is determined, the interaction information including the environment state value of the first control cycle, the reward of the first control cycle, the bandwidth allocation strategy of the first control cycle, and the environment state value of the second control cycle.
[0053] It should be noted that a complete control cycle of an unmanned operation system can include the process of "perception-communication-computation-control". When the communication device instructs the execution unit to complete the operation, which requires multiple control cycles, the first control cycle can be any of these multiple control cycles, and the second control cycle is the control cycle following the first control cycle.
[0054] In order for the trained reinforcement learning network to achieve the expected technical effect of this disclosure, the reinforcement learning network is optimized based on the control performance of the execution unit, and the most suitable bandwidth allocation strategy is determined accordingly. The interaction information needs to include four key values: the environment state value of the first control period (time t), the reward of the first control period (time t), the bandwidth allocation strategy of the first control period (time t), and the environment state value of the second control period (time t+1).
[0055] The following is an introduction to these four key values.
[0056] In some embodiments, the environmental state value of the first control cycle is determined based on the environmental parameters of the first control cycle.
[0057] Environmental state values can reflect the state of the entire unmanned operation system. Communication devices can acquire environmental parameters through multiple sensors and determine environmental state values based on these parameters. Furthermore, communication devices can determine control commands for the execution units based on these environmental state values. Since the operation process of the execution units affects the environmental state values, these values can also be used to characterize the operation results after the execution units perform their tasks according to the control commands.
[0058] In some embodiments, the corresponding bandwidth allocation strategy is determined based on the environmental state value of the first control cycle.
[0059] The initial reinforcement learning network can determine the bandwidth allocation strategy corresponding to the first control cycle based on the environmental state values. Of course, this initial reinforcement learning network has not yet been optimized, so its determined bandwidth allocation strategy is far from the target's optimal bandwidth allocation strategy. The initial reinforcement learning network needs to be trained and adjusted before it can output a bandwidth allocation strategy that is closer to the target.
[0060] In some embodiments, the environment evolves according to a bandwidth allocation strategy.
[0061] In the unmanned operation system disclosed herein, determining the bandwidth allocation strategy and the allocation of bandwidth resources for multiple sensors will affect the reception of observation data sent by the communication equipment from the multiple sensors, thereby affecting the decision of the communication equipment to determine control commands. The affected control commands will in turn affect the operation results in the control cycle, and thus lead to the environmental state value in the next control cycle determined by the observation data sent by the sensors.
[0062] It should be noted that since the environment can evolve according to the bandwidth allocation strategy, thereby determining the environment state value of the second control cycle, the bandwidth allocation strategy for the second control cycle can be directly determined based on the environment state value of the second control cycle obtained from this evolution. Then, the second control cycle can be evolved again to determine the reward for the second control cycle and the environment state value for the next control cycle, thus determining the interaction information corresponding to the second control cycle. This process can be repeated multiple times to obtain multiple consecutive interaction information corresponding to each control cycle.
[0063] In some embodiments, a reward for the environmental state value of the first control cycle is determined based on the control performance of the execution unit by the communication device, and an environmental state value for the second control cycle is determined, wherein the second control cycle is the next control cycle after the first control cycle.
[0064] After evolving the environment based on the bandwidth allocation strategy, the rewards and the environmental state values for the next control cycle can be determined.
[0065] Since the reward is determined based on the control performance of the execution unit by the communication device, the reward can reflect the impact of the bandwidth allocation strategy on the overall control performance. Thus, the reinforcement learning network can be trained through the reward, so that the bandwidth allocation strategy output by the trained reinforcement learning network can enable the unmanned operation system to have better control performance.
[0066] In summary, the reward is determined based on the control performance of the communication device over the execution unit, and is used to reflect the optimization objective during the training process of the reinforcement learning network. The bandwidth allocation strategy determined by the reinforcement learning network will affect the state of the entire unmanned operation system. The environmental state values before and after the implementation of the bandwidth allocation strategy are the environmental state values of the first control cycle and the environmental state values of the second control cycle after the environmental evolution.
[0067] This disclosure defines environmental state values and, based on the evolution of the environment, allows for the evolution and derivation of environmental state values for the next control cycle, simulating the operation of an unmanned operating system, given a determined bandwidth allocation strategy. Furthermore, the reward in this disclosure is determined with the control performance of the communication device over the execution unit as the optimization objective. Therefore, optimizing the reinforcement learning network based on this reward allows the bandwidth allocation strategy determined by the reinforcement learning network to move closer to the optimal target bandwidth allocation strategy, which enables the communication device to achieve the best control performance over the execution unit. Thus, the reinforcement learning network trained using the method proposed in this disclosure can more rationally allocate bandwidth resources from multiple sensors, improve resource utilization, and enable the unmanned operating system to achieve maximized operational performance.
[0068] In some embodiments, optimizing the weights of the reinforcement learning network based on the interaction information includes: optimizing the weights of the reinforcement learning network based on the interaction information when the number of interaction information is greater than a first quantity threshold.
[0069] When training a reinforcement learning network, the more interaction information there is, the more accurate the training result will be. Therefore, a first quantity threshold can be set, and interaction information can be repeatedly acquired until the number of interaction information exceeds the first quantity threshold. Then, the parameters of the reinforcement learning network can be optimized based on the number of interaction information that exceeds the first quantity threshold.
[0070] In some embodiments, the method further includes: if the number of control cycles is greater than a second quantity threshold, or if the cost for determining the reward is greater than a cost threshold, retaining the determined interaction information and re-determining the environmental parameters of the first control cycle to generate interaction information for the next round.
[0071] For a stable control system, the cost will not exceed a preset cost threshold. Therefore, when the cost used to determine the reward exceeds the cost threshold, it can be determined that the current bandwidth allocation strategy makes the control system unstable. In this case, the absolute value of the reward obtained by continuing to evolve is very large, making the training process unstable. Therefore, after retaining the determined interaction information, the environment can be re-initialized, that is, the environmental parameters can be re-determined through sensors, so as to determine the environmental state value of the first control cycle of the new round based on the environmental parameters. Here, the first control cycle of the new round can be the initial control cycle of the new round.
[0072] Each time the environment is initialized, the interaction information can be retained to accumulate the number of interaction information until the number of interaction information exceeds a first threshold.
[0073] It should be noted that the first threshold is greater than the second threshold; as an example, the first threshold can be much larger than the second threshold. Thus, when the required amount of interaction information is determined for training the reinforcement learning network, this interaction information comes from multiple control cycles within a multi-round interaction process.
[0074] The following specific example will further illustrate the solution disclosed herein.
[0075] exist Figure 1 The unmanned operation system shown contains K sensors and one execution unit controlling the operation. The execution unit's operation process is modeled as a discrete linear time-invariant system. In the t-th period, the state evolution of this system can be represented by the equation:
[0076] x t+1 =Ax t+Bu t +v t
[0077] Where A is the state transition matrix; B is the control matrix; x t The system state at the t-th period is represented by an n-dimensional vector; u t Let m represent the system's control input during the t-th cycle; Let A represent the noise in the control process, which follows a Gaussian distribution with a mean vector of zero and a covariance matrix of V. Here, A, B, and V are all known parameters.
[0078] There are K sensors simultaneously sensing the state of the system, and the sensing process is linear, which can be represented by the following equation:
[0079]
[0080] Among them, y k,t It is a scalar used to represent the observation results, that is, the observation data collected by the sensor; w represents the sensing coefficient of the k-th sensor. k,t To perceive noise, it follows a vector with a mean of zero and a variance matrix of... The Gaussian distribution. The perception matrix C = [c1 c2 …c K ] T .
[0081] The sensor transmits its observations to a communication device (e.g., a drone) via a wireless channel. The channel coefficient from the k-th sensor to the communication device is h. k,t The channel is modeled as Rayleigh fading, and the transmission power is p. k The noise power spectral density is N0, and the transmission bandwidth constraint is B. max The transmission time for each control cycle is T, and the packet error probability is ∈.
[0082] Due to the limited communication bandwidth between the sensor and the communication device, the observation results transmitted by the sensor are distorted. We can define the observation result received by the communication device from the k-th sensor as... Its formula can be expressed as:
[0083]
[0084] Where, n k,t This represents distortion caused by limited transmission bandwidth. n can be... k,t The model is based on zero-mean Gaussian noise, and it is assumed that the sensor's observation y k,t It also follows a Gaussian distribution with a variance of . Assume the effective information content during transmission is I. k,Based on the relevant formulas for interactive information, n can be... k,t The variance is calculated as follows The information extraction rate ρ is defined as the rate at which effective information is extracted from perceived data.
[0085] In some embodiments, the LQR cost (Linear Quadratic Regulator) can be used to measure overall control performance. The weight matrix of the LQR cost is Q and R.
[0086] In unmanned operation systems, communication devices control execution units to perform operations through control commands. The control performance of the execution units can be characterized by LQR cost, thereby digitizing the control performance for easier calculation.
[0087] In some embodiments, the Proximal Policy Optimization (PPO) algorithm in deep reinforcement learning can be used to design bandwidth allocation strategies.
[0088] Based on this PPO algorithm, the parameters of the reinforcement learning network during the training process include the maximum step size t of the environment. max Training step size n train Clipping range ∈, discount factor γ, trace decay factor λ, value function loss weight coefficient c1, entropy loss weight coefficient c2, reward scaling factor α, LQR threshold LQR max .
[0089] The training process of a reinforcement learning network involves training the weights within the network to adjust them so that the network can produce a more ideal output based on the provided input (bandwidth allocation strategy). To achieve more accurate training, sample data is needed; in this disclosure, this sample data is referred to as interaction information. Therefore, the following section will detail the process of obtaining interaction information and training the reinforcement learning network.
[0090] In some embodiments, the path loss of the sensor-to-communication device channel can be determined based on the locations of multiple sensors and communication devices.
[0091] Calculate path loss PL k The formula is as follows:
[0092]
[0093] After determining the path loss, the large-scale fading coefficient of the channel can then be determined. It should be noted that η in the formula LOS η NLOSa and b are parameters related to the channel environment, f is the carrier frequency, c is the speed of light, and d is the speed of light. k and θ k These represent the distance and elevation angle from the k-th sensor to the communication device, respectively.
[0094] In some embodiments, the environmental state value includes: the channel coefficients of the plurality of sensors, the prior estimate of the operation state of the execution unit by the communication device, and the covariance matrix corresponding to the prior estimate.
[0095] The environmental state value is used to describe the state of the entire unmanned operation system, and can also be understood as the state of the intelligent agent in the reinforcement learning network.
[0096] For example, after determining the path loss and the large-scale fading coefficient of the channel, a reinforcement learning agent can be constructed, and the agent's state s in the t-th control cycle can be defined. t Defined as:
[0097]
[0098] Among them, h t =[h 1,t h 2,t … h K,t ], h k,t This represents the channel coefficient of the k-th sensor during the t-th control cycle;
[0099] This represents the prior estimate of the state of the control module of the unmanned operation system in the t-th control cycle, and is used to characterize the prior estimate of the control result of the communication device on the execution unit.
[0100] P t|t-1 This represents the covariance matrix of the prior estimate.
[0101] The bandwidth allocation strategy (also known as action) determined by the agent. t The dimension is K, and the value range of each component is [-1,1], representing the bandwidth allocation strategy at the current time (the t-th control cycle).
[0102] In some embodiments, the reinforcement learning network includes a neural network.
[0103] In some embodiments, a deep reinforcement learning neural network can be constructed and the network parameters initialized.
[0104] The constructed deep reinforcement learning neural network can include a feature extraction module, an action estimation module, and a value estimation module. The latter two modules can output an action with dimension K and a value estimate with dimension 1, respectively.
[0105] After constructing a learning neural network, it can interact with the environment to collect sample data (interaction information) needed to train the neural network.
[0106] The process of collecting interaction information is as follows:
[0107] 1. Initialize the environment by setting the control cycle number t = 0, initializing the environment coefficients, and randomly generating the initial small-scale channel fading coefficients s according to the standard complex Gaussian distribution. k,0 Calculate the initial channel coefficient h k,0 =|l k s k,0 Randomly generate a positive definite initial state covariance matrix P0, and randomly generate the initial system state. in This represents a Gaussian distribution.
[0108] Initial observation noise can be generated based on the Gaussian distribution. And calculate the initial observation values. Record the initial environment state value as follows:
[0109] s0 = h0, 0, P0,
[0110] The initialized environmental state value can be used as the environmental state value for the first control cycle.
[0111] 2. Neural networks can adjust based on the environmental state value s. t By utilizing the action estimation module in a neural network, a suitable action 'a' is determined. t (Bandwidth allocation strategy).
[0112] 3. Based on a defined action a t It can evolve the environment to calculate rewards and environmental state values for the next control cycle.
[0113] Specifically, we can first set action a t Scaling to the [0,1] interval yields... Then the bandwidth allocation result is calculated. Where, sum a′ t This is a summation formula used to calculate the sum of each term vector.
[0114] Bandwidth allocation result B for multiple sensors t =[B 1,t B 2,t … B K,t ], B k,t This represents the bandwidth allocated to the k-th sensor.
[0115] Based on the bandwidth allocated to each sensor, the amount of data transmitted by each sensor can be calculated to be approximately:
[0116]
[0117] Among them, Q -1 • This represents the inverse function of the Q-function, which is defined as:
[0118]
[0119] V k,t The channel dispersion can be calculated using the following formula:
[0120]
[0121] Using the above formula, the effective information I in the observation data sent by the sensor can be calculated. k,t =ρD k,t Then the covariance S of the prior perceived data can be calculated. t =CP t|t-1 C T +W, the variance of the sensor's observations. That is, S t The kth element on the diagonal.
[0122] Since the evolution of the environment is based on virtual simulations of real-world scenarios, noise is present in all these processes: sensor observations, communication equipment receiving observations, and control execution units performing operations. Therefore, if noise is not included in the virtual simulation, the results obtained will differ from those in the real-world scenario, leading to a training reinforcement learning network that cannot adapt to the actual situation. Thus, it is necessary to introduce the influence of noise into the virtual simulation process.
[0123] Next, Gaussian-distributed distortion noise can be randomly generated. Calculate the actual observation data received by the receiving end (communication equipment) The actual received observation data is not the data received in the actual scenario, but rather the observation data that the communication equipment can receive in the virtual simulation based on noise, which is closer to the actual scenario.
[0124] Using the Kalman filter principle, the Kalman filter gain is calculated as follows:
[0125] K t =P t|t-1 C T CP t|t-1 C T +W+N t -1
[0126] in:
[0127] C = [c1, c2, ..., c K ] T diag· represents generating a diagonal matrix. After adding noise interference, the posterior estimate can be determined based on the observation data acquired by the communication equipment under the influence of noise, thus obtaining the optimal estimate of the system state.
[0128]
[0129] And calculate the covariance of the optimal estimate:
[0130] P t =IK t CP t|t-1
[0131] Calculate control input Control input is the control action. For example, if the task of the execution unit is to lift an object, then the control input can be regarded as the force applied to the object. Here, G is the optimal LQR control gain matrix, which can be calculated by solving the Ricatti equation.
[0132] Calculate the LQR cost for the current control cycle:
[0133]
[0134] After determining the LQR cost, the reward for the current control period is the LQR cost divided by the scaling factor, i.e.:
[0135]
[0136] Once the reward is determined, the environmental state value can be updated to determine the environmental state value for the next control cycle.
[0137] In some embodiments, determining the environmental state value of the second control period includes: generating random noise, and determining a prior estimate of the second control period and the covariance matrix of the second control period based on the random noise.
[0138] Specifically, random system state noise can be generated, and this noise follows the distribution as follows: Based on the generated random noise, the control system state vector x for the next control cycle can be updated. t+1 =Ax t +Bu t +v t .
[0139] Observation noise that generates a Gaussian distribution Calculate the observation results for the next control cycle. This observation result represents the observation data that the sensor can obtain after taking observation noise into account.
[0140] The noise-based covariance matrix can be used to update the prior estimate covariance P for the next control cycle. t+1|t =AP t A T +V, and the prior estimate of the system state value for the next control cycle.
[0141]
[0142] The initial small-scale channel fading coefficient s is randomly generated in the next control cycle based on the standard complex Gaussian distribution. k,t+1 And calculate the channel coefficient h for the next control cycle. k,t+1 =|l k s k,t+1 |
[0143] Set the environmental state value s at time t+1 of the next control cycle. t+1 Recorded as The information recording the interaction between the intelligent agent and the environment is s. t r t a t s t+1 The interactive information includes the environmental state value of the first control cycle, the reward corresponding to the first control cycle, the action (bandwidth allocation strategy) of the first control cycle, and the environmental state value of the second control cycle.
[0144] As can be seen from the above formulas, the determined bandwidth allocation strategy affects the communication bandwidth of each sensor, which in turn affects the amount of data transmitted by the sensors, and consequently, the observation data that the communication equipment can receive. When the observation data received by the communication equipment is affected, the control of the execution unit by the communication equipment will inevitably be affected as well. This allows the unmanned operation system to evolve based on the bandwidth allocation strategy, thereby determining the relationship between the bandwidth allocation strategy and the control results of the communication equipment on the execution unit. The control results of the communication equipment on the execution unit can be characterized based on environmental state values.
[0145] 4. The above steps describe one evolution of the environment within a control cycle, corresponding to the acquisition of a set of interaction information. The more interaction information there is, the more beneficial it is for training the reinforcement learning network and adjusting the parameters within it.
[0146] Therefore, in some embodiments, training of the reinforcement learning network (reinforcement learning network) can begin based on the interaction information if the amount of interaction information is greater than a first quantity threshold.
[0147] If the amount of interactive information is less than or equal to the first threshold, the process continues to evolve to obtain more interactive information.
[0148] Furthermore, in terms of LQR cost LQR t Continuously greater than the cost threshold LQR max If this happens three times, the environment can be re-initialized, and the agent can continue to interact with the environment to obtain interaction information.
[0149] Alternatively, if the number of control cycles t is greater than or equal to the second quantity threshold t max In such cases, the environment can be re-initialized, and the agent can continue to interact with the environment to obtain interaction information.
[0150] It should be noted that regardless of whether the environment is initialized and the interaction with the environment is restarted, the acquired interaction information is retained and gradually accumulated until the amount of interaction information exceeds the first threshold before the reinforcement learning network training begins.
[0151] 5. Based on the collected interaction information exceeding the first threshold, update the parameter weights of the reinforcement learning network using the PPO algorithm to achieve reinforcement learning network training.
[0152] In some embodiments, optimizing the weights of the reinforcement learning network based on the interaction information includes: determining a clip loss based on the interaction information; determining a value function loss based on the interaction information; determining an entropy loss based on the interaction information; and optimizing the parameters of the reinforcement learning network based on the clip loss, the value function loss, and the entropy loss.
[0153] Specifically, for each set of interactive information s t r t a t s t+1 First, we can calculate the clip loss of the algorithm:
[0154]
[0155] Among them, rewards π represents the ratio of the policies of the new and old reinforcement learning networks. θ (a t |s t ) and π old (a t |s t ) represent the probability values of the actions output by the current reinforcement learning network and the old reinforcement learning network, respectively. clip(r) t (θ), 1-∈, 1+∈) represents r t(θ) is clipped to the interval [1-∈, 1+∈], A t The estimated value of the advantage function can be calculated using the following formula:
[0156] A t =δ t +(γλ)δ t+1 +…+(γλ) T-t+1 δ T-1
[0157] Where, δ t =r t +γV(s t+1 )-V(s t V(s) represents the difference between the actual reward and the expected reward. t ) indicates that the reinforcement learning network adjusts its behavior based on the environmental state value s. t The estimated value, where T represents the total time step taken for the corresponding trajectory.
[0158] After determining the clip loss, the value function loss is determined based on the interaction information. For example, the value function loss L can be calculated. VF =Vs t -r t 2 The entropy loss is determined based on the interaction information. For example, the entropy loss can be calculated.
[0159] After determining the clip loss, value function loss, and entropy loss, the total loss function can be determined as follows:
[0160] L = -L CLIP +c1L VF -c2L entropy
[0161] The parameters in the reinforcement learning network can be updated using gradient descent based on this total loss function. Training ends when the maximum training step size is reached and the number of updates to the reinforcement learning network reaches a preset value.
[0162] In some embodiments, if the training step size is less than the preset maximum training step size, interaction information can be collected again to retrain the reinforcement learning network.
[0163] Based on the aforementioned reinforcement learning network training method, this disclosure introduces a method to characterize the perceptual distortion noise caused by limited communication bandwidth through interactive information, uses the Kalman filtering method for perceptual data fusion, and takes LQR cost as the direct optimization objective to determine the control performance of the communication device on the execution unit as the optimization objective. A new method for allocating the transmission bandwidth of multiple sensors is designed through reinforcement learning, thereby improving the control performance of the unmanned operation system in scenarios with limited bandwidth resources.
[0164] After training the reinforcement learning network, this disclosure also proposes a method for allocating transmission bandwidth.
[0165] Figure 4 This is a schematic flowchart illustrating a transmission bandwidth allocation method according to embodiments of the present disclosure. The method comprises, as... Figure 1 The communication device shown executes the operation of the execution unit based on observation data from multiple sensors.
[0166] like Figure 4 As shown, the method includes:
[0167] In step S401, environmental parameters are determined and environmental state values are determined;
[0168] In step S402, the environmental state value is input into the reinforcement learning network trained by the reinforcement learning network training method for control performance shown in any of the above embodiments of this disclosure, and the bandwidth allocation strategy output by the reinforcement learning network is determined.
[0169] In step S403, the transmission bandwidth allocation result for the multiple sensors is determined according to the bandwidth allocation strategy, and the bandwidth is adjusted according to the transmission bandwidth allocation result.
[0170] Bandwidth resources can be allocated using the reinforcement learning network trained by the above-described reinforcement learning network training method for control performance as disclosed in this disclosure.
[0171] In some embodiments, environmental parameters may include environmental channel information, prior estimates, and the covariance of the prior estimates.
[0172] Specifically, it can obtain the current environmental channel information h. t The Kalman filter method is used to obtain prior estimates of the control system state. And the estimated covariance P t|t-1 To obtain environmental state values
[0173] The acquired environmental state values can be input into the trained reinforcement learning network to obtain action (bandwidth allocation strategy) a. t , will a t Scaling to the [0,1] interval to obtain Then the bandwidth allocation result is calculated.
[0174] It should be noted that the specific calculation method of the data mentioned above can be found in the embodiments proposed in the reinforcement learning network training method of this disclosure.
[0175] In some embodiments, bandwidth is allocated according to the transmission bandwidth allocation result.
[0176] The technical effects of this disclosure are illustrated below with a specific example.
[0177] The solutions disclosed herein can be applied to, for example... Figure 1 The diagram illustrates a scenario where a machine communication network serves multiple sensors and a robot performs unmanned tasks.
[0178] exist Figure 1 In the scenario shown, there are four sensors, randomly and evenly distributed within a circle with a radius of 5000 meters. The drone is located at the center of the circle at a height of 100 meters. Channel-related parameters are set to a = 4.88, b = 0.43, and η... LOS =0.1, η NLOS =21, carrier frequency f = 2GHz, speed of light c = 3 × 10 8 m / s, transmission power p k =0.1W, channel power spectral density is N0 = -174dBm / Hz, transmission time T = 1ms, packet error probability ∈ = 10 -5 The information extraction rate ρ = 0.005.
[0179] The relevant control parameters are set as follows, and the control parameter matrix is set as follows:
[0180]
[0181] The noise covariance is controlled to be set as follows:
[0182]
[0183] In LQR, both the weight matrices Q and R are set to identity matrices. The perception matrix C is set to an identity matrix, and the perception noise variance is...
[0184] The training parameters for the reinforcement learning network are set to control the maximum step size t of the cycle. max =40, Training step size n train =500000, Clipping range ∈ =0.2, Discount factor γ =0.99, Trace decay factor λ =0.95, Value function loss weight coefficient c1 =0.5, Entropy loss weight coefficient c2 =0, Reward scaling factor α =100, LQR threshold LQR max =100000.
[0185] The reinforcement learning network is configured with three fully connected layers of sizes 4×16, 4×16, and 10×32, respectively, for inputting... h t and matrix P t|t-1The upper part of the algorithm is then used to concatenate the outputs of the three networks, pass them through a 64×64 hidden layer, and then cascade a residual network. The residual network module includes a 64×64 fully connected layer. The output of the fully connected layer is added to the input of the residual network to obtain the final output. The activation function of all these networks is ReLU. This output serves as the input to the action estimation network and the value estimation network. The action estimation network consists of cascaded 64×64, 64×64, and 64×64 fully connected layers, while the value estimation network consists of cascaded 64×64, 64×64, and 64×1 fully connected layers. The activation function of the 64×64 fully connected layer is tanh.
[0186] Under the above simulation conditions, this example simulates the LQR cost under the conditions of bandwidth constraints of [40,45,50,55,60,65]kHz.
[0187] Figure 5 This is a schematic diagram illustrating the result of a bandwidth allocation scheme according to an embodiment of the present disclosure.
[0188] like Figure 5 As shown, the performance of this scheme is compared with that of the scheme maximizing channel rate allocation and the scheme of average bandwidth allocation. The curves marked with crosses represent the simulation results of this scheme, the curves marked with circles represent the simulation results of the scheme maximizing channel rate allocation, and the curves marked with squares represent the simulation results of the scheme of average bandwidth allocation. It is evident that this scheme can effectively reduce the LQR cost of the unmanned operation system and improve the system's control performance.
[0189] Corresponding to the embodiments of the reinforcement learning network training method for control performance disclosed herein, this disclosure also provides embodiments of a corresponding reinforcement learning network training apparatus for control performance.
[0190] The reinforcement learning network is used to determine the channel transmission bandwidth between the communication device and multiple sensors, wherein the communication device controls the operation of the execution unit based on the observation data of the multiple sensors.
[0191] Please see Figure 6 , Figure 6 This is a block diagram of a reinforcement learning network training apparatus for control performance according to one embodiment of this disclosure. Figure 6 As shown, the reinforcement learning network training device includes:
[0192] The strategy estimation module 610 is configured to determine the environmental state value of the first control cycle based on the environmental parameters of the first control cycle, and to determine the corresponding bandwidth allocation strategy based on the environmental state value of the first control cycle.
[0193] The reward estimation module 620 is configured to evolve the environment according to the bandwidth allocation strategy, determine the reward for the environment state value for the first control period based on the control performance of the communication device on the execution unit, and determine the environment state value for the second control period, wherein the second control period is the next control period after the first control period.
[0194] The information determination module 630 is configured to determine the interaction information between the execution unit and the environment, the interaction information including the reward of the first control cycle, the bandwidth allocation strategy of the first control cycle, and the environment state value of the second control cycle.
[0195] The parameter optimization module 640 is configured to optimize the weights of the reinforcement learning network based on the interaction information.
[0196] In some embodiments, optimizing the weights of the reinforcement learning network based on the interaction information includes: optimizing the weights of the reinforcement learning network based on the interaction information when the number of interaction information is greater than a first quantity threshold.
[0197] In some embodiments, the apparatus is further configured to: retain the determined interaction information and re-determine the environmental parameters of the first control cycle to generate the interaction information for the next round when the number of control cycles is greater than a second quantity threshold, or when the cost for determining the reward is greater than a cost threshold.
[0198] In some embodiments, the environmental state value includes: the channel coefficient h of the plurality of sensors, the prior estimate x of the control result of the communication device on the execution unit, and the covariance matrix P corresponding to the prior estimate.
[0199] In some embodiments, determining the environmental state value of the second control period includes: generating random noise, and determining a prior estimate of the second control period and the covariance matrix of the second control period based on the random noise.
[0200] In some embodiments, optimizing the weights of the reinforcement learning network based on the interaction information includes: determining a clip loss based on the interaction information; determining a value function loss based on the interaction information; determining an entropy loss based on the interaction information; and optimizing the weights of the reinforcement learning network based on the clip loss, the value function loss, and the entropy loss.
[0201] The specific implementation process of the functions and roles of each unit in the above device can be found in the implementation process of the corresponding steps in the above method, and will not be repeated here.
[0202] Embodiments of this disclosure also provide an electronic device, including: a processor and a memory; the memory for storing a computer program; and the processor for executing a reinforcement learning network training method as described in any of the foregoing embodiments by invoking the computer program.
[0203] Embodiments of this disclosure also propose a computer-readable storage medium having a computer program stored thereon that, when executed by a processor, implements the reinforcement learning network training method as described in any of the above embodiments.
[0204] Embodiments of this disclosure also provide a computer program product, including a computer program that, when executed by a processor, implements the method as described in any of the foregoing embodiments.
[0205] Figure 7 This is a schematic block diagram illustrating a reinforcement learning network training device 700 according to embodiments of the present disclosure. For example, device 700 may be a mobile phone, computer, digital broadcasting terminal, messaging device, game console, tablet device, medical device, fitness equipment, personal digital assistant, etc.
[0206] Reference Figure 7 The device 700 may include one or more of the following components: processing component 702, memory 704, power supply component 706, multimedia component 708, audio component 710, input / output (I / O) interface 712, sensor component 714, and communication component 716.
[0207] Processing component 702 typically controls the overall operation of device 700, such as operations associated with display, telephone calls, data communication, camera operation, and recording. Processing component 702 may include one or more processors 720 to execute instructions to complete all or part of the steps of the reinforcement learning network training method described above. Furthermore, processing component 702 may include one or more modules to facilitate interaction between processing component 702 and other components. For example, processing component 702 may include a multimedia module to facilitate interaction between multimedia component 708 and processing component 702.
[0208] Memory 704 is configured to store various types of data to support the operation of device 700. Examples of such data include instructions for any application or method operating on device 700, contact data, phonebook data, messages, pictures, videos, etc. Memory 704 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.
[0209] Power supply assembly 706 provides power to various components of device 700. Power supply assembly 706 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to device 700.
[0210] Multimedia component 708 includes a screen that provides an output interface between the device 700 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touchscreen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors may sense not only the boundaries of the touch or swipe action but also the duration and pressure associated with the touch or swipe operation. In some embodiments, multimedia component 708 includes a front-facing camera and / or a rear-facing camera. When the device 700 is in an operating mode, such as a shooting mode or a video mode, the front-facing camera and / or the rear-facing camera may receive external multimedia data. Each front-facing camera and rear-facing camera may be a fixed optical lens system or have focal length and optical zoom capabilities.
[0211] Audio component 710 is configured to output and / or input audio signals. For example, audio component 710 includes a microphone (MIC) configured to receive external audio signals when device 700 is in an operating mode, such as call mode, recording mode, and voice recognition mode. The received audio signals may be further stored in memory 704 or transmitted via communication component 716. In some embodiments, audio component 710 also includes a speaker for outputting audio signals.
[0212] I / O interface 712 provides an interface between processing component 702 and peripheral interface modules, such as keyboards, click wheels, buttons, etc. These buttons may include, but are not limited to, home buttons, volume buttons, power buttons, and lock buttons.
[0213] Sensor assembly 714 includes one or more sensors for providing state assessments of various aspects of device 700. For example, sensor assembly 714 may detect the on / off state of device 700, the relative positioning of components such as the display and keypad of device 700, changes in the position of device 700 or a component of device 700, the presence or absence of user contact with device 700, the orientation or acceleration / deceleration of device 700, and temperature changes of device 700. Sensor assembly 714 may include a proximity sensor configured to detect the presence of nearby objects without any physical contact. Sensor assembly 714 may also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, sensor assembly 714 may also include an accelerometer, a gyroscope, a magnetometer, a pressure sensor, or a temperature sensor.
[0214] Communication component 716 is configured to facilitate wired or wireless communication between device 700 and other devices. Device 700 can access wireless networks based on communication standards, such as WiFi, 2G, 3G, 4G LTE, 5G NR, or combinations thereof. In one exemplary embodiment, communication component 716 receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In one exemplary embodiment, communication component 716 also includes a near-field communication (NFC) module to facilitate short-range communication. For example, the NFC module may be implemented based on radio frequency identification (RFID) technology, Infrared Data Association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.
[0215] In an exemplary embodiment, the apparatus 700 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the reinforcement learning network training method described above.
[0216] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 704 including instructions, which can be executed by the processor 720 of the device 700 to complete the reinforcement learning network training method described above. For example, the non-transitory computer-readable storage medium may be a ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc.
[0217] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the disclosure herein. This disclosure is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the following claims.
[0218] It should be understood that this disclosure is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this disclosure is limited only by the appended claims.
[0219] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. The terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0220] The methods and apparatus provided in the embodiments of this disclosure have been described in detail above. Specific examples have been used to illustrate the principles and implementation methods of this disclosure. The descriptions of the embodiments above are only for the purpose of helping to understand the methods and core ideas of this disclosure. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this disclosure. Therefore, the content of this specification should not be construed as a limitation of this disclosure.
Claims
1. A reinforcement learning network training method for control performance, characterized in that, The reinforcement learning network is used to determine the channel transmission bandwidth between the communication device and multiple sensors, wherein the communication device controls the operation of the execution unit based on the observation data of the multiple sensors, and the method includes: The environmental state value of the first control cycle is determined based on the environmental parameters of the first control cycle, and the corresponding bandwidth allocation strategy is determined based on the environmental state value of the first control cycle. The environment is evolved according to the bandwidth allocation strategy. The reward for the environment state value of the first control cycle is determined based on the control performance of the communication device on the execution unit, and the environment state value of the second control cycle is determined, wherein the second control cycle is the next control cycle after the first control cycle. The interaction information between the execution unit and the environment is determined. The interaction information includes the environment state value of the first control cycle, the reward of the first control cycle, the bandwidth allocation strategy of the first control cycle, and the environment state value of the second control cycle. The weights of the reinforcement learning network are optimized based on the interaction information.
2. The method according to claim 1, characterized in that, The step of optimizing the weights of the reinforcement learning network based on the interaction information includes: If the number of interaction information is greater than a first threshold, the weights of the reinforcement learning network are optimized based on the interaction information.
3. The method according to claim 2, characterized in that, The method further includes: If the number of control cycles exceeds a second quantity threshold, or if the cost for determining the reward exceeds a cost threshold, the determined interaction information is retained, and the environmental parameters for the first control cycle are redefined to generate the interaction information for the next round.
4. The method according to claim 1, characterized in that, The environmental status values include: The channel coefficients of the plurality of sensors, the prior estimate of the working state of the execution unit, and the covariance matrix corresponding to the prior estimate.
5. The method according to claim 4, characterized in that, Determining the environmental state value for the second control cycle includes: Random noise is generated, and the prior estimate of the second control period and the covariance matrix of the second control period are determined based on the random noise.
6. The method according to claim 1, characterized in that, The step of optimizing the weights of the reinforcement learning network based on the interaction information includes: Determine the clip loss based on the interaction information; Determine the value function loss based on the interaction information; Determine the entropy loss based on the interaction information; The weights of the reinforcement learning network are optimized based on the clip loss, the value function loss, and the entropy loss.
7. A method for allocating transmission bandwidth, characterized in that, Performed by a communication device, which controls the operation of the execution unit based on observation data from multiple sensors, the method includes: Determine environmental parameters and environmental state values; The environmental state value is input into the reinforcement learning network trained by any one of claims 1-6, and the bandwidth allocation strategy output by the reinforcement learning network is determined. The bandwidth allocation results for the multiple sensors are determined according to the bandwidth allocation strategy, and the bandwidth is adjusted according to the bandwidth allocation results.
8. A reinforcement learning network training device for control performance, characterized in that, The reinforcement learning network is used to determine the channel transmission bandwidth between the communication device and multiple sensors, wherein the communication device controls the operation of the execution unit based on the observation data of the multiple sensors, and the device includes: The strategy estimation module is configured to determine the environmental state value of the first control cycle based on the environmental parameters of the first control cycle, and to determine the corresponding bandwidth allocation strategy based on the environmental state value of the first control cycle. The reward estimation module is configured to evolve the environment according to the bandwidth allocation strategy, determine the reward for the environment state value for the first control period based on the control performance of the communication device on the execution unit, and determine the environment state value for the second control period, wherein the second control period is the next control period after the first control period. The information determination module is configured to determine the interaction information between the execution unit and the environment. The interaction information includes the environment state value of the first control cycle, the reward of the first control cycle, the bandwidth allocation strategy of the first control cycle, and the environment state value of the second control cycle. The parameter optimization module is configured to optimize the weights of the reinforcement learning network based on the interaction information.
9. An electronic device, characterized in that, include: Processor, memory; The memory is used to store computer programs; The processor is configured to execute, by invoking the computer program, the reinforcement learning network training method for control performance as described in any one of claims 1-6.
10. A computer program product, characterized in that, Includes a computer program that, when executed by a processor, implements the method described in any one of claims 1-6.