Network modal configuration and adjustment method and device based on double-time-scale reinforcement learning

By using the network mode configuration and adjustment method of dual-time scale reinforcement learning in the dual-time scale optimization technology, dynamically adjusting the window size of the network mode is solved, and the problem of difficulty in dynamically adjusting the long-time scale time window size in the dynamic environment in the existing technology is solved, and the system performance and efficiency improvement is achieved.

CN120075071APending Publication Date: 2025-05-30UNIV OF SCI & TECH BEIJING
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510191064.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-20
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

In the prior art, the dual-time-scale optimization technology is difficult to dynamically adjust the time window size of the long-time-scale in a dynamic environment, resulting in limited system performance optimization.

Method used

The network modal configuration and adjustment method based on dual-time scale reinforcement learning is adopted to decouple the network modal level resource allocation from the user-level resource allocation and drone trajectory planning problems, and optimize them on different time scales. Through the critic network V in reinforcement learning, the window size of the network mode is dynamically adjusted to achieve dynamic adjustment of the time window size of a long-term scale.

Benefits of technology

It realizes the reduction of the total system cost, simplifies the complexity of the problem, and more effectively optimizes the characteristics of each time scale, improving system performance and efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120075071A_ABST
    Figure CN120075071A_ABST
Patent Text Reader

Abstract

The invention provides a network mode configuration and adjustment method and device based on double-time-scale reinforcement learning, and relates to the technical field of communication. The method comprises the following steps: constructing a system model of an unmanned aerial vehicle auxiliary network mode; constructing an optimization problem model according to the system model; reconstructing the optimization problem model into a dual-time scale Markov decision process under a reinforcement learning framework; and solving is carried out through a scheme based on double-time-scale reinforcement learning. According to the method, an optimization problem is decoupled into optimization sub-problems on two different time scales of a network modal level and a user level, and based on the characteristic that a reviewer network in reinforcement learning is used for expected accumulated discount rewards in a certain specific state, reconfiguration and adjustment benefits of network modal parameters are predicted by using the reviewer network; and then the cost is compared with the cost of reconfiguration and adjustment of the network modal parameters, so that the self-adaptive dynamic configuration and adjustment of the network modal parameters are realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of communication technologies, and in particular, to a method and device for mode configuration and adjustment based on a dual-time-scale reinforcement learning network. Background Art

[0002] The background of the dual-time-scale optimization technology mainly stems from the dynamic characteristics and optimization requirements exhibited by complex systems in different time dimensions. This optimization method is particularly applicable to systems that require decision-making and scheduling on different fast and slow time scales, such as communication systems and control systems involving reinforcement learning. The following is a detailed elaboration of the background of the dual-time-scale optimization technology:

[0003] 1 Technical Origin and Theoretical Basis:

[0004] The dual-time-scale optimization technology is not an independent technical field, but a specific manifestation of optimization theory in the application of complex systems. Its theoretical basis mainly stems from optimization theory, dynamic system theory, and control theory, etc. With the continuous in-depth research in these fields, scholars have gradually realized that in complex systems, the change speeds of different variables and states often vary significantly, so it is necessary to make optimization decisions on different time scales.

[0005] 2 Application Background in Communication Systems:

[0006] In communication systems, the dual-time-scale optimization technology is mainly used to handle signal transmission and resource allocation problems on different time scales. The following are some specific application scenarios:

[0007] 1) In 6G communication systems, ISAC (Integrated Sensing And Communication) is an important research direction. The ISAC system needs to share resources between the communication and sensing tasks, and these two tasks often have different time-scale characteristics. The communication task requires real-time data transmission and has high requirements for time delay; while the sensing task pays more attention to the long-term monitoring and analysis of the environmental state. Therefore, the dual-time-scale optimization technology can be used to reasonably allocate resources in the ISAC system to ensure that both the communication and sensing tasks can be effectively supported.

[0008] 2) In a heterogeneous cellular network powered by the smart grid, renewable energy and wireless channels evolve dynamically on different time scales. To minimize the long-term average energy trading cost, a two-time-scale dynamic resource management scheme is required. This scheme decomposes the real-time joint problem into two sub-problems, which are solved on fast and slow time scales respectively by different optimization methods. For example, power allocation and energy sharing strategies are adjusted in real time on the fast time scale to cope with the dynamic changes of wireless channels; while on the slow time scale, long-term two-way energy trading plans are formulated to optimize energy usage efficiency.

[0009] 3) In high-frequency communication (such as millimeter-wave communication), the Doppler frequency offset of the channel becomes more and more serious, and the performance of traditional communication waveforms, such as OFDM (Orthogonal Frequency Division Multiplexing), is limited. To solve this problem, new signal waveforms can be designed using the two-time-scale characteristics. For example, the communication process can be divided into multiple time scales according to the change speed of the channel state, and corresponding signal waveforms and transmission strategies are designed for each time scale. This can better adapt to the dynamic changes of the channel and improve communication performance.

[0010] 3 Application background in reinforcement learning:

[0011] In the field of reinforcement learning, two-time-scale optimization techniques are also widely used. Reinforcement learning is a method of learning the optimal strategy through trial and error, and its core idea is to continuously adjust the strategy to maximize the cumulative reward during the interaction with the environment. In reinforcement learning algorithms such as Actor-Critic, Actor and Critic learn and update on different time scales respectively. Actor is responsible for generating action strategies, and its update speed is relatively slow; while Critic is responsible for evaluating the value of actions, and its update speed is relatively fast. This two-time-scale update mechanism helps the algorithm converge to the optimal strategy more stably. By separating the two processes of policy generation and value evaluation and updating them on different time scales, two-time-scale optimization techniques help the algorithm converge to the optimal strategy more stably. The fast update of Critic can provide immediate feedback to help Actor adjust its strategy more accurately; while the slow update of Actor helps to avoid excessive fluctuations in the strategy and maintain the stability of learning.

[0012] 4 Technical advantages and challenges:

[0013] The advantage of the dual-time-scale optimization technology lies in its ability to more accurately reflect the dynamic characteristics of complex systems, improving the accuracy and real-time performance of optimization decisions. However, this technology also faces some challenges, such as how to handle the coupling relationship between different time scales and how to determine reasonable optimization objectives and constraints. In addition, as the scale and complexity of the system increase, the difficulty of solving the dual-time-scale optimization problem will gradually increase.

[0014] In summary, the dual-time-scale optimization technology is a method with broad application prospects in the optimization of complex systems. With the continuous in-depth research and the continuous expansion of application fields, this technology is expected to play a more important role in the future.

[0015] To meet the diverse service requirements of vehicles in a dynamic vehicle environment, the literature [1] (Cui Y, Huang X, He P, et al. A two-timescale resource allocation scheme in vehicular network slicing [C] / / 2021 IEEE 93rd Vehicular Technology Conference (VTC2021-Spring). IEEE, 2021: 1-5.) proposed a two-time-scale radio resource allocation scheme LST-MDDPG to provide stable services for vehicles. Specifically, for the long-term dynamic characteristics of vehicle service requests, the literature [1] uses LSTM (Long Short-Term Memory) to track trajectories and performs dedicated resource allocation on a long time scale using historical data. On the other hand, for the impact of channel changes caused by high-speed movement in a short time, the literature [1] uses the DRL (Deep Reinforcement Learning) algorithm, namely DDPG (Deep Deterministic Policy Gradient), to adjust resource allocation.

[0016] Reference [2] (Ye F, Li J, Zhu P, et al. Intelligent hierarchical NOMA-based network slicing in cell-free RAN for 6G systems[J]. IEEE Transactions on Wireless Communications, 2023.) proposed a network slicing architecture based on double time-scale reinforcement learning in the scalable cell-free radio access network of the 6G new full-spectrum mobile edge computing network, and jointly allocated communication, computing, and caching resources with different resource granularities to meet the requirements of delay-critical applications with different delays. Using reinforcement learning, the network slice-level resource allocation problem is optimized on a long time scale, and the user association and user-level resource allocation problems are optimized on a short time scale.

[0017] In the above solutions, the scenarios where the computing tasks are heterogeneous and dynamic are not considered, and the network slice window size is set to be fixed and cannot change dynamically with service requests.

[0018] As an effective decision-making method, reinforcement learning has demonstrated its unique advantages in solving double time-scale optimization problems. Such problems usually involve decision-making on different time scales, where the long time scale focuses on the macro planning and long-term goals of the system, while the short time scale focuses on immediate adjustments and optimizations. Existing reinforcement learning solutions for double time-scale optimization problems can be mainly classified into three categories.

[0019] The first category of solutions:

[0020] In this category of solutions, the decision maker adopts a non-reinforcement learning method on the long time scale and uses reinforcement learning for decision-making on the short time scale. The advantage of this method lies in combining the advantages of non-reinforcement learning in macro planning and the flexibility of reinforcement learning in micro adjustment. However, since the non-reinforcement learning method only aims to maximize the system performance at the current moment and ignores the consideration of future impacts, this may lead to the sub-optimization of the system performance over the entire time period. In addition, the integration and docking efficiency between the non-reinforcement learning method and reinforcement learning is not high, limiting the improvement of the overall performance.

[0021] The second category of solutions:

[0022] Contrary to the first category, this category of solutions adopts a non-reinforcement learning method on the short time scale and uses reinforcement learning on the long time scale. Similarly, although this approach can quickly respond to changes in the short term, it lacks in-depth consideration of the global goal in the long term, resulting in limitations in system performance. The inherent limitations of the non-reinforcement learning method also restrict its effective combination with reinforcement learning, affecting the optimization of the overall performance.

[0023] The third type of solution:

[0024] The third type of solution introduces the concept of hierarchical reinforcement learning, aiming to use reinforcement learning for decision-making on both long and short time scales. In this architecture, the decision-making on the long time scale is responsible for the upper-level policy, which plans and makes decisions over a relatively long time span; while the decision-making on the short time scale is executed by the lower-level policy, which focuses on a relatively short time span. The upper and lower-level policies interact with each other through a feedback mechanism. The upper-level policy provides guidance to the lower-level policy, and at the same time, the lower-level policy feeds back the execution results to the upper-level policy. This hierarchical structure can effectively optimize the system performance over the entire time period. However, it has an obvious defect: the time window size on the long time scale is fixed and cannot be dynamically adjusted according to the real-time state of the system. In a dynamic environment, such as the wireless communication field, this static time window design may seriously hinder the performance optimization of the system. Summary of the Invention

[0025] In order to solve the technical problem that the time window size on the long time scale in the existing technology is fixed and cannot be dynamically adjusted according to the real-time state of the system. In a dynamic environment, such as the wireless communication field, this static time window design may seriously hinder the performance optimization of the system, the embodiments of the present invention provide a method and device for configuring and adjusting the network mode based on the dual-time-scale reinforcement learning network. The technical solution is as follows:

[0026] On the one hand, a method for configuring and adjusting the network mode based on the dual-time-scale reinforcement learning network is provided. This method is implemented by a network mode configuration and adjustment device, and the method includes:

[0027] S1. Construct a system model of the drone-assisted network mode.

[0028] S2. Construct an optimization problem model according to the system model.

[0029] S3. Reconstruct the optimization problem model into a dual-time-scale Markov decision process under the framework of reinforcement learning.

[0030] S4. Through the solution based on the dual-time-scale reinforcement learning, optimize the resource allocation problem at the network mode level on the long time scale, and optimize the resource allocation at the user level and the drone trajectory planning problem on the short time scale, and solve the dual-time-scale Markov decision process to obtain the network mode configuration and adjustment result based on the dual-time-scale reinforcement learning.

[0031] Optionally, the system model in S1 includes drones and a set of Internet of Things devices.

[0032] Among them, the unmanned aerial vehicle (UAV) includes multiple independent virtual network modes; the Internet of Things (IoT) devices have heterogeneous computing tasks, and the types of computing tasks change dynamically over time.

[0033] Building the system model of the UAV-assisted network mode in S1 includes:

[0034] S11. Define the channel gain between the UAV and the m-th IoT device at time slot t.

[0035] S12. Build the total delay of the computing tasks generated by the m-th IoT device at time slot t.

[0036] S13. Build the total cost of the UAV-assisted network mode system at time slot t.

[0037] Optionally, the channel gain is shown in the following formula (1):

[0038]

[0039] In the formula, represents the channel gain between the UAV and the m-th IoT device at time slot t, g T represents the transmitting antenna gain of the IoT device, g R represents the receiving antenna gain of the UAV, λ C represents the wavelength of the signal, represents the coordinates of the UAV at the position of time slot t, represents the position coordinates of the m-th IoT device, and p represents the path loss exponent.

[0040] The total delay is shown in the following formula (2):

[0041]

[0042] In the formula, T m,t represents the total delay, represents the delay caused by the m-th IoT device offloading the computing task to the UAV edge server at time slot t, represents the delay caused by the UAV edge server processing the computing task offloaded by the m-th IoT device at time slot t.

[0043] The total cost is shown in the following formula (3):

[0044]

[0045] In the formula, represents the total cost, represents the resource consumption cost, represents the network mode reconfiguration cost, represents the computing task processing failure cost.

[0046] Optionally, the optimization problem model in S2 is shown in the following formula (4):

[0047]

[0048] In the formula, represents the proportion of computing resources allocated to the network mode NS in time slot t, CI and represents the proportion of computing resources allocated to the network mode NS in time slot t, DI and represents the proportion of computing resources allocated to the network mode NS in time slot t, DS and represents the bandwidth resource allocated to the network mode NS in time slot t, CI and represents the bandwidth resource allocated to the network mode NS in time slot t, DI and represents the bandwidth resource allocated to the network mode NS in time slot t, DS and represents the set of time slots, represents the proportion of the computing resources allocated by the UAV to the m-th Internet of Things device in its corresponding network mode's total available computing resources, represents the proportion of the bandwidth allocated by the UAV to the m-th Internet of Things device in its corresponding network mode's total available bandwidth, represents the set of Internet of Things devices, T represents the decision time period, and

[0049] Optionally, the double-time-scale Markov decision process in S3 is shown in the following formula (5):

[0050]

[0051] Among them, is the state space, is the long-time-scale action space for the network mode-level resource allocation problem, is the short-time-scale action space for the user-level resource allocation and UAV trajectory planning problem, is the reward function, and P is the state transition probability.

[0052] At time slot t, the global state of the system is expressed as:

[0053]

[0054] In the formula, represents the coordinates of the UAV at time slot t, Indicates the proportion of computing resources in time slot t allocated to network mode NS CI of computing resources, Indicates the proportion of computing resources in time slot t allocated to network mode NS DI of computing resources, Indicates the proportion of computing resources in time slot t allocated to network mode NS DS of computing resources, Indicates the proportion of computing resources in time slot t allocated to network mode NS CI of bandwidth resources, Indicates the proportion of bandwidth resources in time slot t allocated to network mode NS DI of bandwidth resources, Indicates the proportion of bandwidth resources in time slot t allocated to network mode NS DS of bandwidth resources, Indicates the position coordinates of the m-th Internet of Things device, Indicates the data size of the computing task generated by the m-th Internet of Things device in time slot t, Indicates the number of CPU cycles required for the m-th Internet of Things device to process the computing task in time slot t, Indicates the latency constraint of the computing task generated by the m-th Internet of Things device in time slot t, Indicates the type of computing task generated by the m-th Internet of Things device in time slot t, Indicates the set of Internet of Things devices.

[0055] The long-term scale action of the system in time slot t Is expressed as:

[0056]

[0057] In the formula,

[0058] The long-term scale action of the system in time slot t Is expressed as:

[0059]

[0060] In the formula, Indicates the proportion of computing resources allocated by the drone to the m-th Internet of Things device in its corresponding network mode's total available computing resources, Indicates the proportion of bandwidth allocated by the drone to the m-th Internet of Things device in its corresponding network mode's total available bandwidth, Indicates the coordinates of the drone's position in time slot t.

[0061] The reward of the system in time slot t Is expressed as:

[0062]

[0063] In the formula, represents the total cost, represents the resource consumption cost, represents the network mode reconfiguration cost, represents the computational task processing failure cost.

[0064] The state transition probability is expressed as:

[0065] P := P(s(t + 1)|s(t)) (10)

[0066] In the formula, P represents the state transition probability, s(t) represents the Markov state, and s(t + 1) represents the successor Markov state.

[0067] Optionally, the scheme based on double - time - scale reinforcement learning in S4 includes a policy network π with parameter θ and a critic network V with parameter φ.

[0068] The policy network π includes a long - time - scale policy network π S with parameter θ S and a short - time - scale policy network π U with parameter θ U .

[0069] The optimization objective of the scheme based on double - time - scale reinforcement learning is to maximize the cumulative discounted reward, as shown in the following formula (11):

[0070]

[0071] In the formula, T represents the decision time period, γ t-1 represents the discount factor, and r(t) represents the reward of the system.

[0072] The policy network optimizes the parameters by maximizing the truncated objective function, as shown in the following formula (12):

[0073]

[0074] In the formula, J(θ) represents the truncated objective function, d θ represents the policy probability ratio, represents the estimated advantage function, and ∈ represents the hyperparameter.

[0075] The evaluation network updates the parameters by minimizing the mean - square error, as shown in the following formula (13):

[0076]

[0077] In the formula, L(φ) represents the mean - square error, represents the cumulative discounted reward starting from state s(t), and Vφ (s) represents the critic network V with parameter φ.

[0078] Optionally, the scheme based on double - time - scale reinforcement learning in S4 further includes: dynamically adjusting the window size of the network mode.

[0079] Among them, dynamically adjusting the window size of the network mode includes:

[0080] S41. Based on the long - time - scale policy network π S The new long - time - scale action output Calculate the new network mode resource configuration.

[0081] S42. Based on the new network mode resource configuration, obtain the new state

[0082] S43. Based on the critic network V, calculate the cumulative discounted reward V(s(t)) of the old state s(t) and the cumulative discounted reward of the new state And calculate the corresponding network mode re - configuration cost;

[0083] S44. Based on the cumulative discounted reward V(s(t)) of the old state s(t), the cumulative discounted reward of the new state And The network mode re - configuration cost, determine whether to re - configure the network mode and re - configure the resources of the network mode.

[0084] On the other hand, a device for network mode configuration and adjustment based on double - time - scale reinforcement learning is provided. This device is applied to the method for network mode configuration and adjustment based on double - time - scale reinforcement learning. The device includes:

[0085] A system construction module, used to construct a system model of the drone - assisted network mode.

[0086] An optimization problem construction module, used to construct an optimization problem model according to the system model.

[0087] A reconstruction module, used to reconstruct the optimization problem model into a double - time - scale Markov decision process under the framework of reinforcement learning.

[0088] An output module, used to solve the double - time - scale Markov decision process by optimizing the resource allocation problem at the network mode level on the long - time scale and optimizing the resource allocation at the user level and the drone trajectory planning problem on the short - time scale through the scheme based on double - time - scale reinforcement learning, and obtain the network mode configuration and adjustment result based on double - time - scale reinforcement learning.

[0089] On the other hand, a network mode configuration and adjustment device is provided, which includes: a processor; a memory, on which computer-readable instructions are stored, and when the computer-readable instructions are executed by the processor, any of the methods in the above-mentioned network mode configuration and adjustment method based on double time-scale reinforcement learning is implemented.

[0090] On the other hand, a computer-readable storage medium is provided, in which at least one instruction is stored, and the at least one instruction is loaded and executed by a processor to implement any of the methods in the above-mentioned network mode configuration and adjustment method based on double time-scale reinforcement learning.

[0091] The beneficial effects brought by the technical solutions provided in the embodiments of the present invention at least include:

[0092] In the present invention, a network mode configuration and adjustment method based on double time-scale reinforcement learning is proposed. In order to reduce the overall cost of the system, this method decouples the resource allocation at the network mode level, the resource allocation at the user level, and the UAV trajectory planning problem, forming two optimization problems on different time scales. This separation not only helps to simplify the complexity of the problem, but also can be more effectively optimized according to their respective characteristics.

[0093] In addition, compared with the above-mentioned third type of solution, in the solution proposed in the present invention, the critic network V in reinforcement learning is used to utilize the characteristic of the expected cumulative discounted reward in a specific state s, and by comparing the reconfiguration cost and the reconfiguration benefit of the time window size on the long time scale, the dynamic adjustment of the time window size on the long time scale is realized. BRIEF DESCRIPTION OF THE DRAWINGS

[0094] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0095] Figure 1 is a flowchart of a network mode configuration and adjustment method based on double time-scale reinforcement learning provided by an embodiment of the present invention;

[0096] Figure 2 is a system model diagram provided by an embodiment of the present invention;

[0097] Figure 3 is a structural diagram of a policy network provided by an embodiment of the present invention;

[0098] Figure 4 is a timing diagram provided by an embodiment of the present invention;

[0099] Figure 5 is a flowchart of a method for network mode configuration and adjustment based on a dual-time-scale reinforcement learning network provided by an embodiment of the present invention;

[0100] Figure 6 is a block diagram of a device for network mode configuration and adjustment based on a dual-time-scale reinforcement learning network provided by an embodiment of the present invention;

[0101] Figure 7 is a schematic structural diagram of a network mode configuration and adjustment device provided by an embodiment of the present invention. Detailed implementation manners

[0102] The technical solutions in the present invention will be described below with reference to the accompanying drawings.

[0103] In the embodiments of the present invention, words such as "exemplarily" and "for example" are used to represent examples, illustrations or explanations. Any embodiment or design solution described as an "example" in the present invention should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Rather, the use of the word "example" is intended to present concepts in a specific manner. In addition, in the embodiments of the present invention, the meaning expressed by "and / or" can be both, or either one of the two.

[0104] In the embodiments of the present invention, "image" and "picture" can sometimes be used interchangeably. It should be noted that when the difference is not emphasized, the meanings they express are the same. "(of)", "corresponding", and "corresponding" can sometimes be used interchangeably. It should be noted that when the difference is not emphasized, the meanings they express are the same.

[0105] In the embodiments of the present invention, sometimes subscripts such as W 1 may be written in a non-subscript form such as W1. When the difference is not emphasized, the meanings they express are the same.

[0106] To make the technical problems, technical solutions and advantages to be solved by the present invention clearer, the following will be described in detail with reference to the accompanying drawings and specific embodiments.

[0107] The embodiments of the present invention provide a method for network mode configuration and adjustment based on a dual-time-scale reinforcement learning network. This method can be implemented by a network mode configuration and adjustment device, and the network mode configuration and adjustment device can be a terminal or a server. As Figure 1 shown in the flowchart of the method for network mode configuration and adjustment based on a dual-time-scale reinforcement learning network, the processing flow of this method can include the following steps:

[0108] S1. Build a system model for the UAV-assisted network mode.

[0109] In a feasible implementation manner, in the present invention, the multi-modal network management problem of the UAV-assisted mobile edge computing system is mainly studied. Specifically, as Figure 2 shown, the present invention considers an industrial Internet of Things scenario in a square area with a side length of L, where there is a UAV with an edge server, providing computing offloading services for a large number of Internet of Things devices distributed on the ground. Let represent the set of Internet of Things devices.

[0110] Furthermore, the present invention assumes that these Internet of Things devices have heterogeneous computing tasks, and their task data sizes, task complexities, and quality of service requirements are different. In order to provide differentiated services for these Internet of Things devices with heterogeneous computing tasks, the present invention uses multi-modal network technology to divide the UAV-assisted physical network into multiple virtual and independent logical network modes. In addition, the present invention assumes that the computing tasks of the Internet of Things devices are not only heterogeneous, but also their task types change dynamically over time, so the entire decision time period is discretized into multiple equal time slots, that is The present invention assumes that the Internet of Things devices do not have computing capabilities, so the Internet of Things devices need to offload the generated computing tasks to the edge server of the UAV for processing in each time slot.

[0111] In the system architecture considered by the present invention, the computing tasks generated by the Internet of Things devices exhibit significant heterogeneity and dynamic characteristics. These tasks vary in terms of computing resource requirements, data volume scale, and latency sensitivity, and these requirements change over time. In order to effectively cope with this diversity and dynamicity, the present invention adopts multi-modal network technology, which allows multiple virtual logical networks to operate on the same physical network infrastructure.

[0112] Multi-modal network technology is a technology that divides the physical network into multiple virtual and independent logical network modes. Each network mode can be customized according to different service requirements (such as different types of task offloading requirements), including configurations in terms of network bandwidth, latency, reliability, etc.

[0113] Specifically, the present invention assumes that the computing tasks generated by the Internet of Things can be divided into three major categories: compute-intensive tasks, which require a large amount of computing resources to process; data-intensive tasks, which involve the transmission of a large volume of data; and latency-sensitive tasks, which have extremely high requirements for the response time of processing. To precisely match these different types of computing tasks, the present invention divides the physical network resources (including computing resources and bandwidth resources, etc.) assisted by drones into three independent virtual network modes. Each network mode is designed for a specific type of task: a network mode NS CI is used to process compute-intensive tasks, a network mode NS DI is used to process data-intensive tasks, and a network mode NS DS is used to process latency-sensitive tasks. Assuming that the total available computing resources of the drone are F, let and respectively represent the proportion of computing resources allocated to the three network modes NS CI 、NS DI and NS DS in time slot t. The computing resources allocated to the three network modes cannot exceed the total bandwidth resources, that is

[0114] Similarly, the total available bandwidth resources of the drone are B, let and respectively represent the bandwidth resources allocated to the three network modes NS CI 、NS DI and NS DS in time slot t. The bandwidth resources allocated to the three network modes cannot exceed the total bandwidth resources, that is

[0115] Through multi-modal network technology, each logically isolated network mode dynamically allocates corresponding network resources according to the requirements of computing tasks, effectively reducing the resource competition and preemption between different computing tasks, ensuring that various Internet of Things devices obtain the most suitable computing offloading services in their corresponding network modes, thereby optimizing the performance and efficiency of the entire system of Internet of Things devices.

[0116] Optionally, the system model in S1 includes a drone and a set of Internet of Things devices.

[0117] Among them, the drone includes multiple independent virtual network modes; the Internet of Things devices have heterogeneous computing tasks, and the types of computing tasks change dynamically over time.

[0118] The above step S1 may include the following steps S11-S13:

[0119] S11. Define the channel gain between the UAV and the m-th Internet of Things device at time slot t.

[0120] In a feasible implementation, in order to represent the relative position relationship among the UAV, the Internet of Things device, and the sensed target, the present invention establishes a three-dimensional Cartesian coordinate system. Let represent the coordinates of the UAV's position at time slot t, represent the position coordinates of the m-th Internet of Things device. It should be noted that the present invention assumes that the position of the Internet of Things device is fixed and known to the UAV, and the UAV can fly within the height range from Z min to Z max above the square target area.

[0121] The channel between the UAV and the Internet of Things device is an air-to-ground (A2G) channel. The present invention defines the channel gain between the UAV and the m-th Internet of Things device at time slot t as:

[0122]

[0123] In the formula, represents the channel gain between the UAV and the m-th Internet of Things device at time slot t, g T represents the transmitting antenna gain of the Internet of Things device, g R represents the receiving antenna gain of the UAV, λ C represents the wavelength of the signal, represents the coordinates of the UAV's position at time slot t, represents the position coordinates of the m-th Internet of Things device, and p represents the path loss exponent.

[0124] S12. Construct the total delay of the computing task generated by the m-th Internet of Things device at time slot t.

[0125] In a feasible implementation, at time slot t, the computing task generated by the m-th Internet of Things device is defined as where is the data size of the computing task, is the number of CPU cycles required to process the computing task, is the delay constraint of the computing task, is the type of the computing task. Let and respectively represent the sets of Internet of Things devices with computing-intensive, data-intensive, and delay-sensitive computing task types at time slot t. Therefore, there is

[0126] Since IoT devices do not have computing capabilities, they need to offload the entire computing task to the server of the drone for processing. Therefore, at time slot t, the delay caused by the m-th IoT device offloading the computing task to the drone edge server is:

[0127]

[0128] where, is the data transmission rate of the m-th IoT device.

[0129] This invention assumes that all IoT devices use Orthogonal Frequency Division Multiple Access (OFDMA) technology when transmitting computing tasks to the edge server of the drone, that is, they transmit data through different orthogonal subcarriers, so there is no inter-user interference between IoT devices. Therefore, the data transmission rate can be expressed as:

[0130]

[0131] where, is the ratio of the bandwidth allocated by the drone to the m-th IoT device to the total available bandwidth of its corresponding network mode, is the transmission power of the m-th IoT device, and σ is the Gaussian noise power.

[0132] After the computing task is offloaded to the drone, the edge server of the drone allocates corresponding computing resources according to the information of the task to process the task within its delay constraint. At time slice t, the delay caused by the edge server of the drone processing the computing task offloaded by the m-th IoT device is:

[0133]

[0134] where, is the ratio of the computing resources allocated by the drone to the m-th IoT device to the total available computing resources of its corresponding network mode.

[0135] Since the amount of data after the computing task is processed by the drone is usually much smaller than the original data, this paper, like most existing studies, ignores the delay caused by returning the computing result from the drone to the IoT device. Therefore, at time slice t, the total delay of the computing task generated by the m-th IoT device includes the delay of the IoT device offloading the computing task to the drone and the delay of the edge server of the drone processing the computing task, that is

[0136] S13. Construct the total cost of the drone-assisted network mode system at time slot t.

[0137] In a feasible implementation, the total cost of the UAV-assisted network mode system may include resource consumption cost, network mode reconfiguration cost, and computing task processing failure cost.

[0138] Resource consumption cost:

[0139]

[0140] Among them, ω comm and ω comp are weight factors.

[0141] Network mode reconfiguration cost:

[0142]

[0143] Among them, the inherent cost of network mode reconfiguration is P pn , and are weighting coefficients. Since the release of resources is fast and low-cost, the present invention only considers the inherent cost of network mode reconfiguration and the cost of network mode reconfiguration under the influence of resource increase.

[0144] Computing task processing failure cost:

[0145]

[0146] Among them, P comp is the cost (penalty) of processing computing task failure.

[0147] The total cost of the system within time slot t can be expressed as:

[0148]

[0149] S2. According to the system model, construct an optimization problem model.

[0150] In a feasible implementation, the objective of the present invention is to minimize the total cost of the system by optimizing network mode cycle management and resource allocation. Therefore, the optimization problem can be formulated as:

[0151]

[0152]

[0153] In the formula, L max is the side length of the square target area, is the maximum displacement of the UAV.

[0154] S3. Reconstruct the optimization problem model as a two - time - scale Markov decision process within the framework of reinforcement learning.

[0155] In a feasible implementation, P1 cannot be solved by traditional optimization methods, or the latency brought by the solution is unacceptable. On the one hand, it is a long - term optimization problem, which makes solving this problem challenging. On the other hand, due to the dynamic changes in wireless channels and network modality requests, exhaustive search for the optimal solution will bring huge computational complexity and time complexity, making the problem difficult to solve. In addition, since the computing tasks of IoT devices are heterogeneous and dynamic, in order to match the resources required by the computing tasks as much as possible, reduce the resource usage cost, and improve the execution success rate of computing tasks and user experience, the system needs to frequently adjust the network modality resources, which will result in additional network modality overhead.

[0156] Therefore, to reduce the overall cost of the system, the present invention proposes an intelligent two - time - scale optimization scheme, which decouples the resource allocation at the network modality level from the resource allocation at the user level and the UAV trajectory planning problem, forming two optimization problems on different time scales. This separation not only helps to simplify the complexity of the problem but also enables more effective optimization according to their respective characteristics.

[0157] Network modality - level resource allocation:

[0158] The resource allocation at the network modality level is regarded as a long - time - scale optimization problem. At this level, decisions are made within a relatively long time frame with the aim of reasonably allocating resources for each network modality. Considering the high cost of network modality reconfiguration, this long - term perspective optimization helps to reduce unnecessary resource adjustments and reconfigurations, thus reducing the overall cost. In the optimization process, congestion management techniques in network modalities, such as load balancing and traffic steering, as well as new traffic engineering techniques, such as network coding and multipath transmission, can be used to improve the resilience and efficiency of the network.

[0159] User - level resource allocation and UAV trajectory planning:

[0160] In contrast, the resource allocation at the user level and the UAV trajectory planning problem are regarded as a short - time - scale optimization problem. This means that within a short time interval, the system needs to accurately allocate resources for each user and simultaneously plan the flight path of the UAV. This refined management can maximize the utilization efficiency of resources while maintaining the quality of service. The resource allocation strategy can be optimized by real - time monitoring and dynamic adjustment of resource allocation, as well as by using algorithms such as reinforcement learning and game theory.

[0161] By decoupling the resource allocation at the network modality level and the user level from the UAV trajectory planning problem and optimizing them on different time scales respectively, the method of the present invention aims to reduce the total cost of the system. This method not only considers the cost of network modality reconfiguration but also takes into account the requirements of user service quality, achieving a balance between cost-effectiveness and performance optimization.

[0162] To solve this problem, the present invention first reconstructs the optimization problem P1 into a two-time-scale Markov decision process under the framework of reinforcement learning, and then proposes a two-time-scale reinforcement learning-based scheme.

[0163] Specifically, first, the present invention reconstructs the problem P1 into a two-time-scale Markov decision process, which can be expressed as It includes is the state space, is the long-time-scale action space for the network modality-level resource allocation problem, is the short-time-scale action space for the user-level resource allocation and UAV trajectory planning problem, is the reward function, and P is the state transition probability.

[0164] State space The state of the system includes the position of the UAV, the allocation ratios of the computing resources and bandwidth resources of each network modality, the positions of the IoT devices, and the information of the computing tasks. Therefore, at time slot t, the global state of the system can be expressed as:

[0165]

[0166] Long-time-scale action space The long-time-scale action of the system is the decision for the network modality-level resource allocation problem, which includes the allocation variables of the computing resources and bandwidth resources of each network modality. Therefore, at time slot t, the long-time-scale action of the system can be expressed as:

[0167]

[0168] In the formula,

[0169] At time slot t, the long-time-scale action of the system is expressed as:

[0170]

[0171] In the formula, represents the ratio of the computing resources allocated by the UAV to the m-th IoT device to the total available computing resources of its corresponding network modality, represents the ratio of the bandwidth allocated by the UAV to the m-th Internet of Things device to the total available bandwidth of its corresponding network mode. represents the coordinates of the UAV at time slot t.

[0172] Short-term scale action space The short-term scale action of the system is a decision for the user-level resource allocation and UAV trajectory planning problem, which includes the allocation variables of computing resources and bandwidth resources for each user and the UAV trajectory planning variables. Therefore, at time slot t, the long-term scale action of the system can be expressed as:

[0173]

[0174] In the formula, represents the total cost, represents the resource consumption cost, represents the network mode reconfiguration cost, represents the cost of computing task processing failure.

[0175] State transition probability P: The state transition probability refers to the probability of jumping from a Markov state s(t) to the successor state s(t + 1), that is:

[0176] P: = P(s(t + 1)|s(t)) (15)

[0177] It should be noted that due to the dynamics of the network environment and computing tasks, the state transition probability is unknown.

[0178] S4. Through the scheme based on double-time-scale reinforcement learning, optimize the resource allocation problem at the network mode level on the long time scale, optimize the resource allocation at the user level and the UAV trajectory planning problem on the short time scale, solve the double-time-scale Markov decision process, and obtain the network mode configuration and adjustment results based on double-time-scale reinforcement learning.

[0179] In a feasible implementation, to solve the above double-time-scale Markov decision process, the present invention proposes a double-time-scale reinforcement learning scheme based on PPO.

[0180] In the double-time-scale PPO, the agent has a policy network π with parameter θ and a critic network V with parameter φ, where the policy network π takes the state s as input and outputs the joint action a = {a S ,a U} The critic network V is used to approximate the expected cumulative discounted reward of the agent in a specific state s, that is, the state-value function V(s), which takes the global state s as input during the training phase.

[0181] As Figure 3 shown, the policy network π consists of a long-time-scale policy network π with parameters θ S and a short-time-scale policy network π with parameters θ S The long-time-scale policy network π takes the state s as input and outputs the long-time-scale action a U ; The short-time-scale policy network π takes the state s and the long-time-scale action a U as input and outputs the short-time-scale action a S ; S U S U .

[0182] In addition, there is also an experience replay buffer D in the double-time-scale PPO, which is used to store the experience tuples (s, a, r) generated by the interaction between the agent and the environment. The experience tuples in the experience replay pool are used to update the policy network and the critic network of the agent. The algorithm optimization goal is to maximize its cumulative discounted reward where γ ∈ [0, 1) is the discount factor.

[0183] In the double-time-scale PPO algorithm, the importance sampling technique is used to reuse the experience tuples to update the network parameters to speed up the training. The policy network optimizes the parameters by maximizing the truncated objective function:

[0184]

[0185] In the formula, clip(X, 1 - ∈, 1 + ∈) truncates X between 1 - ∈ and 1 + ∈, and ∈ is a (small) hyperparameter indicating how far the new policy is allowed to deviate from the old policy. d θ is the policy probability ratio, that is where θ old is the parameter of the policy network when the agent generates the experience tuples in D through the interaction with the environment. is the estimated advantage function, which can be calculated using the truncated version of the generalized advantage estimation: where δ(t) = r(t) + γV(s(t + 1)) - V(s(t)) is the time difference residual, and λ is used to adjust the bias-variance trade-off.

[0186] The evaluation network can update its parameters by minimizing the mean square error, as shown in Equation (13) below:

[0187]

[0188] wherein represents the cumulative discounted reward starting from state s(t).

[0189] In existing research, it is usually assumed that the computing tasks generated by IoT devices are heterogeneous, but their task types do not change dynamically over time. Then, based on double-time-scale reinforcement learning, the resource allocation at the long-time-scale network modality level and the short-time-scale user level are optimized respectively, where the window size of the network modality is usually fixed.

[0190] However, in the present invention, a more realistic and general situation is considered, that is, the computing tasks generated by IoT devices are not only heterogeneous, but also their types change dynamically over time. Therefore, the drone needs to reconfigure the parameters of each network modality in a timely and dynamic manner according to the dynamically changing computing task types of IoT devices and their requirements for computing resources and bandwidth resources, allocate appropriate resources for them, so as to achieve an effective balance between resource utilization and task processing, improve the processing success rate of computing tasks while avoiding waste of resources. However, reconfiguring the network modality requires cost, and frequent reconfiguration of the network modality parameters will undoubtedly increase the operating cost. Therefore, the present invention proposes an intelligent network modality life cycle management scheme to dynamically adjust the window size of the network modality and reconfigure the network modality at the appropriate time.

[0191] Specifically, the present invention utilizes the characteristic of the critic network V in reinforcement learning, which is used to calculate the expected cumulative discounted reward in a specific state s, to achieve the dynamic adjustment of the window size of the network modality.

[0192] The specific steps are as follows:

[0193] S41. Based on the new long-time-scale action output by the long-time-scale policy network π S (i.e., the variable for reconfiguring the network modality resources calculate the new network modality resource configuration:

[0194] S42. Based on the new network modality resource configuration, obtain the new state

[0195]

[0196] S43. Based on the critic network V, calculate the cumulative discounted reward V(s(t)) of the old state s(t) and the cumulative discounted reward of the new state respectively. And calculate the corresponding network mode reconfiguration cost:

[0197]

[0198] S44. Based on the cumulative discounted reward V(s(t)) of the old state s(t), the new state of the cumulative discounted reward and the network mode reconfiguration cost, determine whether to reconfigure the network mode.

[0199] In a feasible implementation, when the benefit brought by the network mode reconfiguration (i.e., the difference between the cumulative discounted rewards of the new and old states) is greater than the corresponding network mode reconfiguration cost, that is: Reconfigure the network mode, and reconfigure the resources of the network mode based on the long-term scale action Otherwise, the network mode is not reconfigured, and the resources of the network mode remain unchanged.

[0200] Through the above method, the present invention realizes the intelligent and dynamic management of the network mode window size, as Figure 4 shown. The overall process of the present invention is as Figure 5 shown.

[0201] The specific algorithm is as follows:

[0202]

[0203]

[0204] To solve the problems of the prior art, the present invention proposes a new double-time-scale reinforcement learning scheme based on the third type of solution, and illustrates the solution with the scenario of a multi-modal network system as an example.

[0205] In the multi-modal network system considered by the present invention, the drone-assisted physical network is virtualized into multiple network modes, providing personalized computing offloading services for Internet of Things devices with heterogeneous and dynamic task types. The system needs to reconfigure the parameters of each mode in a timely and dynamic manner according to the dynamically changing computing task types of the Internet of Things devices and their requirements for computing resources and bandwidth resources, allocate appropriate resources for them, so as to achieve an effective balance between resource utilization and task processing, improve the processing success rate of computing tasks while avoiding waste of resources. However, the reconfiguration of the network mode requires costs, and frequent reconfiguration of the parameters of the network mode will undoubtedly increase the operating costs. Therefore, the present invention uses a double-time-scale reinforcement learning scheme to dynamically adjust the time window size of the network model and reconfigure the network mode at an appropriate time.

[0206] Specifically, the present invention proposes a dual-time-scale reinforcement learning scheme based on PPO (Proximal Policy Optimization). In the above multi-modal network system, the resource allocation at the network modality level is decoupled from the resource allocation at the user level and the UAV trajectory planning problem, and they are optimized separately on different time scales. The resource allocation problem at the network modality level is optimized on a long time scale, and the resource allocation at the user level and the UAV trajectory planning problem are optimized on a short time scale. By utilizing the characteristic of the critic network V in reinforcement learning, which is used for the expected cumulative discounted reward in a specific state s, the dynamic adjustment of the network modality window size (i.e., the time window size on the long time scale) is achieved by comparing the reconfiguration cost and reconfiguration benefit of the network modality (i.e., comparing the reconfiguration cost and reconfiguration benefit of the time window on the long time scale).

[0207] In an embodiment of the present invention, a method for network modality configuration and adjustment based on a dual-time-scale reinforcement learning is proposed. To reduce the overall cost of the system, the resource allocation at the network modality level is decoupled from the resource allocation at the user level and the UAV trajectory planning problem, forming two optimization problems on different time scales. This separation not only helps to simplify the complexity of the problem but also enables more effective optimization according to their respective characteristics.

[0208] In addition, compared with the above-mentioned third type of solution, in the solution proposed by the present invention, by utilizing the characteristic of the critic network V in reinforcement learning, which is used for the expected cumulative discounted reward in a specific state s, the dynamic adjustment of the time window size on the long time scale is achieved by comparing the reconfiguration cost and reconfiguration benefit of the time window size on the long time scale.

[0209] Figure 6 It is a block diagram of a device for network modality configuration and adjustment based on a dual-time-scale reinforcement learning network shown according to an exemplary embodiment. This device is used for the method of network modality configuration and adjustment based on a dual-time-scale reinforcement learning. Referring to Figure 6 , the device includes a system construction module 310, an optimization problem construction module 320, a reconstruction module 330, and an output module 340. Among them:

[0210] The system construction module 310 is used to construct a system model of the UAV-assisted network modality.

[0211] The optimization problem construction module 320 is used to construct an optimization problem model according to the system model.

[0212] The reconstruction module 330 is used to reconstruct the optimization problem model into a dual-time-scale Markov decision process under the framework of reinforcement learning.

[0213] The output module 340 uses a scheme based on double-time-scale reinforcement learning to optimize the resource allocation problem at the network mode level on a long time scale, and optimize the resource allocation at the user level and the UAV trajectory planning problem on a short time scale. It solves the double-time-scale Markov decision process to obtain the network mode configuration and adjustment results based on double-time-scale reinforcement learning.

[0214] In the embodiment of the present invention, a network mode configuration and adjustment method based on a double-time-scale reinforcement learning is proposed. To reduce the overall cost of the system, this method decouples the resource allocation at the network mode level from the resource allocation at the user level and the UAV trajectory planning problem, forming two optimization problems on different time scales. This separation not only helps to simplify the complexity of the problem, but also enables more effective optimization according to their respective characteristics.

[0215] In addition, compared with the above-mentioned third type of scheme, in the scheme proposed in the present invention, the critic network V in reinforcement learning is used to utilize the characteristic of the expected cumulative discounted reward in a specific state s, and by comparing the reconfiguration cost and reconfiguration benefit of the time window size on the long time scale, the dynamic adjustment of the time window size on the long time scale is realized.

[0216] Figure 7 It is a schematic structural diagram of a network mode configuration and adjustment device provided by an embodiment of the present invention. As Figure 7 shown, the network mode configuration and adjustment device may include the above-mentioned Figure 6 network mode configuration and adjustment device based on double-time-scale reinforcement learning shown. Optionally, the network mode configuration and adjustment device 410 may include a first processor 2001.

[0217] Optionally, the network mode configuration and adjustment device 410 may further include a memory 2002 and a transceiver 2003.

[0218] Among them, the first processor 2001, the memory 2002 and the transceiver 2003 may be connected through a communication bus.

[0219] Next, in combination with Figure 7 each component of the network mode configuration and adjustment device 410 will be specifically introduced:

[0220] Among them, the first processor 2001 is the control center of the network mode configuration and adjustment device 410, which can be a single processor or a collective term for multiple processing elements. For example, the first processor 2001 is one or more central processing units (CPUs), or can be an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present invention. For example: one or more digital signal processors (DSPs), or one or more field programmable gate arrays (FPGAs).

[0221] Optionally, the first processor 2001 can execute various functions of the network mode configuration and adjustment device 410 by running or executing software programs stored in the memory 2002 and calling data stored in the memory 2002.

[0222] In a specific implementation, as an embodiment, the first processor 2001 can include one or more CPUs, such as Figure 7 the CPU0 and CPU1 shown in

[0223] In a specific implementation, as an embodiment, the network mode configuration and adjustment device 410 can also include multiple processors, such as Figure 7 the first processor 2001 and the second processor 2004 shown in

[0224] Here, each of these processors can be a single-core processor (single-CPU) or a multi-core processor (multi-CPU). The processor here can refer to one or more devices, circuits, and / or processing cores for processing data (such as computer program instructions).

[0225] Optionally, the memory 2002 may be a read-only memory (ROM) or other type of static storage device that can store static information and instructions, a random access memory (RAM) or other type of dynamic storage device that can store information and instructions, or may also be an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM), or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media, or other magnetic storage devices, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory 2002 may be integrated with the first processor 2001 or may exist independently and be coupled to the first processor 2001 through an interface circuit of the network mode configuration and adjustment device 410 ( Figure 7 not shown in the figure), and the embodiments of the present invention do not make specific limitations in this regard.

[0226] The transceiver 2003 is used to communicate with a network device or communicate with a terminal device.

[0227] Optionally, the transceiver 2003 may include a receiver and a transmitter ( Figure 7 not shown separately in the figure). Among them, the receiver is used to implement the receiving function, and the transmitter is used to implement the transmitting function.

[0228] Optionally, the transceiver 2003 may be integrated with the first processor 2001 or may exist independently and be coupled to the first processor 2001 through an interface circuit of the network mode configuration and adjustment device 410 ( Figure 7 not shown in the figure), and the embodiments of the present invention do not make specific limitations in this regard.

[0229] It should be noted that Figure 7 the structure of the network mode configuration and adjustment device 410 shown in the figure does not constitute a limitation on the router. The actual knowledge structure recognition device may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements.

[0230] In addition, the technical effects of the network mode configuration and adjustment device 410 may refer to the technical effects of the network mode configuration and adjustment method based on the double-time-scale reinforcement learning described in the above method embodiments, and will not be elaborated here.

[0231] It should be understood that the first processor 2001 in the embodiments of the present invention may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0232] It should also be understood that the memory in the embodiments of the present invention may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM) or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of random access memory (RAM) are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced SDRAM (ESDRAM), synchlink DRAM (SLDRAM), and direct rambus RAM (DR RAM).

[0233] The above embodiments can be implemented in whole or in part by software, hardware (such as circuits), firmware, or any other combination. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions described in the embodiments of the present invention are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (such as infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or a data center that contains one or more collections of available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, or magnetic tape), an optical medium (such as a DVD), or a semiconductor medium. The semiconductor medium can be a solid-state drive.

[0234] It should be understood that the term "and / or" in this document is merely a description of the association relationship between associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. Here, A and B can be singular or plural. In addition, the character " / " in this document generally represents an "or" relationship between the associated objects before and after, but it may also represent an "and / or" relationship, which can be specifically understood by referring to the context before and after.

[0235] In the present invention, "at least one" means one or more, and "a plurality" means two or more. "At least one of the following" or its similar expressions refer to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b, or c can represent: a, b, c, a - b, a - c, b - c, or a - b - c, where a, b, and c can be single or multiple.

[0236] It should be understood that in various embodiments of the present invention, the magnitudes of the sequence numbers of the above processes do not mean the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.

[0237] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or by a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.

[0238] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the devices, apparatuses, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0239] In several embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.

[0240] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place, or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0241] In addition, the functional units in each embodiment of the present invention can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit.

[0242] When the above-mentioned function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs.

[0243] As described above, the above are only specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.

Claims

1. A method for configuring and adjusting network modalities based on dual-time-scale reinforcement learning, characterized in that: The method comprises: S1. Construct the system model of UAV-assisted network mode; S2. constructing an optimization problem model according to the system model; S3, reconstructing the optimization problem model into a dual-time-scale Markov decision process under the framework of reinforcement learning; S4. Through a solution based on dual-time-scale reinforcement learning, the resource allocation problem at the network modality level is optimized on a long time scale, and the resource allocation and UAV trajectory planning problems at the user level are optimized on a short time scale. The dual-time-scale Markov decision process is solved to obtain the network modality configuration and adjustment results based on dual-time-scale reinforcement learning.

2. The method for configuring and adjusting network modalities based on dual-time-scale reinforcement learning according to claim 1, characterized in that: The system model in S1 includes a drone and a set of IoT devices; The drone includes multiple independent virtual network modes; the IoT device has heterogeneous computing tasks, and the types of computing tasks change dynamically over time; The system model of constructing the drone-assisted network mode in S1 includes: S11, define the channel gain between the UAV and the mth IoT device in time slot t; S12, construct the total delay of the computing task generated by the mth IoT device in time slot t; S13. Construct the total cost of the UAV-assisted network modal system under time slot t.

3. The method for configuring and adjusting network modalities based on dual-time-scale reinforcement learning according to claim 2, characterized in that: The channel gain is shown in the following formula (1): In the formula, represents the channel gain between the UAV and the mth IoT device at time slot t, g T represents the transmitting antenna gain of the IoT device, g R represents the receiving antenna gain of the UAV, λ C represents the wavelength of the signal, represents the coordinates of the drone at time slot t, represents the location coordinates of the mth IoT device, and p represents the path loss index; The total delay is shown in the following formula (2): Where, T m,t represents the total delay, represents the delay caused by the mth IoT device offloading the computation task to the drone edge server at time slot t, represents the delay caused by the UAV edge server processing the computational task offloaded from the mth IoT device at time slot t; The total cost is shown in the following formula (3): In the formula, represents the total cost, represents the resource consumption cost, represents the network modality reconfiguration cost, Indicates the cost of computing task processing failure.

4. The method for configuring and adjusting network modalities based on dual-time-scale reinforcement learning according to claim 1, characterized in that: The optimization problem model in S2 is shown in the following formula (4): In the formula, Indicates that time slot t is allocated to network mode NS CI The proportion of computing resources, Indicates that time slot t is allocated to network mode NS DI The proportion of computing resources, Indicates that time slot t is allocated to network mode NS DS The proportion of computing resources, Indicates that time slot t is allocated to network mode NS CI bandwidth resources, Indicates that time slot t is allocated to network mode NS DI bandwidth resources, Indicates that time slot t is allocated to network mode NS DS bandwidth resources, The time slot set is shown as follows: It represents the ratio of the computing resources allocated by the drone to the mth IoT device to the total available computing resources of its corresponding network mode. It represents the ratio of the bandwidth allocated by the drone to the mth IoT device to the total available bandwidth of its corresponding network mode. represents the IoT device set, T represents the decision time period, Represents the total cost.

5. The method for configuring and adjusting network modalities based on dual-time-scale reinforcement learning according to claim 1, characterized in that: The dual-time-scale Markov decision process in S3 is shown in the following equation (5): in, is the state space, is a long-time-scale action space for the network modality-level resource allocation problem, A short-time-scale action space for user-level resource allocation and UAV trajectory planning problems. is the reward function, P is the state transition probability; At time slot t, the global state of the system It is expressed as: In the formula, represents the coordinates of the drone at time slot t, Indicates that time slot t is allocated to network mode NS CI The proportion of computing resources, Indicates that time slot t is allocated to network mode NS DI The proportion of computing resources, Indicates that time slot t is allocated to network mode NS DS The proportion of computing resources, Indicates that time slot t is allocated to network mode NS CI bandwidth resources, Indicates that time slot t is allocated to network mode NS DI bandwidth resources, Indicates that time slot t is allocated to network mode NS DS bandwidth resources, Represents the location coordinates of the mth IoT device, represents the data size of the computing task generated by the mth IoT device at time slot t, represents the number of CPU cycles required by the mth IoT device to process the computing task at time slot t, represents the delay constraint of the computation task generated by the mth IoT device at time slot t, represents the type of computing task generated by the mth IoT device at time slot t, Represents a collection of IoT devices; At time slot t, the long-time scale action of the system It is expressed as: In the formula, At time slot t, the long-time scale action of the system It is expressed as: In the formula, It represents the ratio of the computing resources allocated by the drone to the mth IoT device to the total available computing resources of its corresponding network mode. It represents the ratio of the bandwidth allocated by the drone to the mth IoT device to the total available bandwidth of its corresponding network mode. represents the coordinates of the drone’s position at time slot t; At time slot t, the system's reward It is expressed as: In the formula, represents the total cost, represents the resource consumption cost, represents the network modality reconfiguration cost, Indicates the cost of calculating task processing failure; The state transition probability is expressed as: P: =P(s(t+1)|s(t)) (10) Where P represents the state transition probability, s(t) represents the Markov state, and s(t+1) represents the subsequent Markov state.

6. The method for configuring and adjusting network modalities based on dual-time-scale reinforcement learning according to claim 1, characterized in that: The dual-time-scale reinforcement learning scheme in S4 includes a policy network π with a parameter θ and a critic network V with a parameter φ; The policy network π includes parameters θ S The long time scale strategy network π S and the parameter is θ U The short-time-scale strategy network π U ; The optimization goal of the dual-time-scale reinforcement learning scheme is to maximize the cumulative discounted reward, as shown in the following equation (11): Where T represents the decision time period, γ t-1 represents the discount factor, r(t) represents the reward of the system; The policy network optimizes parameters by maximizing the truncation objective function, as shown in equation (12): Where J(θ) represents the truncated objective function, d θ represents the strategy probability ratio, represents the estimated advantage function, ∈ represents the hyperparameter; The evaluation network updates parameters by minimizing the mean square error, as shown in the following equation (13): Where L(φ) represents the mean square error, represents the cumulative discounted reward starting from state s(t), V φ (s) represents the critic network V with parameters φ.

7. The method for configuring and adjusting network modalities based on dual-time-scale reinforcement learning according to claim 6, characterized in that: The dual-time-scale reinforcement learning-based solution in S4 further includes: dynamically adjusting the window size of the network modality; The dynamically adjusting the window size of the network modality includes: S41. Long-term strategy network π S The new long-time scale action output Calculate new network modality resource configuration; S42. Based on the new network modality resource configuration, obtain a new state S43. Based on the critic network V, calculate the cumulative discounted reward V(s(t)) of the old state s(t) and the new state Cumulative discount rewards And calculate the corresponding network modality reconfiguration cost; S44, based on the cumulative discount reward V(s(t)) of the old state s(t), the new state Cumulative discount rewards As well as the network modality reconfiguration cost, determine whether to reconstruct the network modality and reconfigure the resources of the network modality.

8. A dual-time-scale reinforcement learning network modality configuration and adjustment device, the dual-time-scale reinforcement learning network modality configuration and adjustment device is used to implement the dual-time-scale reinforcement learning network modality configuration and adjustment method according to any one of claims 1 to 7, characterized in that: The device comprises: System building module, used to build the system model of UAV-assisted network mode; An optimization problem construction module is used to construct an optimization problem model according to the system model; A reconstruction module, for reconstructing the optimization problem model into a dual-time scale Markov decision process under the framework of reinforcement learning; The output module uses a solution based on dual-time-scale reinforcement learning to optimize the resource allocation problem at the network modality level on a long time scale, and optimizes the resource allocation and drone trajectory planning problems at the user level on a short time scale. The dual-time-scale Markov decision process is solved to obtain the network modality configuration and adjustment results based on dual-time-scale reinforcement learning.

9. A network mode configuration and adjustment device, characterized in that: The network mode configuration and adjustment device includes: processor; A memory having computer-readable instructions stored thereon, wherein when the computer-readable instructions are executed by the processor, the method according to any one of claims 1 to 7 is implemented.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores program codes, which can be called by a processor to execute the method according to any one of claims 1 to 7.