Wireless mobile communication network resource scheduling method and system based on DGQN

By introducing improved reinforcement learning algorithm DGQN and dual-path architecture into wireless mobile communication networks, the concave and additive metrics are solved, and the problem that the existing technology is difficult to apply to wireless mobile network resource scheduling is achieved, and efficient and intelligent resource scheduling is achieved.

CN120111615AActive Publication Date: 2025-06-06CHINESE PEOPLES LIBERATION ARMY UNIT 61905
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510206309.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-25
Publication Date
2025-06-06
Estimated Expiration
2045-02-25

AI Technical Summary

Technical Problem

Existing reinforcement learning algorithms, such as DQN, are difficult to directly apply to resource scheduling of wireless mobile networks, especially in the face of frequent changing network environments and limited computing power.

Method used

A wireless mobile communication network resource scheduling method based on the improved reinforcement learning algorithm DGQN is proposed. The concave metrics and additive metrics are processed through a dual-path architecture, and the lightweight network design based on GhostNet is adopted to dynamically adapt to differentiated services and network environments.

Benefits of technology

It realizes efficient and intelligent scheduling of communication network resources under diversified business scenarios, can dynamically adapt to changes in the network environment and differences in business demands, and solves the resource scheduling problem when computing power is limited.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120111615A_ABST
    Figure CN120111615A_ABST
Patent Text Reader

Abstract

The invention discloses a wireless mobile communication network resource scheduling method and system based on a DGQN, and relates to the technical field of communication network resource scheduling, and the method comprises the steps: firstly, mapping a resource scheduling process of a wireless mobile communication network into a Markov decision process; a concavity measurement index and an additive measurement index in the resource scheduling process are processed in a double-path architecture mode, then weighted summation is carried out on q values of the two types of indexes obtained through processing, and the q values of the two types of indexes are synthesized for representation to obtain a comprehensive model q value; and finally, firstly determining a final strategy, and then determining a final routing path of the service flow, thereby realizing resource scheduling of the wireless mobile communication network. Aiming at a wireless mobile communication network with high-frequency change of a topological structure, the invention provides the improved scheme by considering real-time dynamic change of a network environment and differentiated requirements of various service flows in the network on different performance indexes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of communication network resource scheduling, and in particular to a method and system for wireless mobile communication network resource scheduling based on DGQN. Background Art

[0002] Communication network resource scheduling technology refers to a technology that optimizes network resource configuration through scheduling strategies and algorithms under the condition of limited network resources. In wireless mobile networks, it usually involves multiple business modes such as voice, video, and messaging. Since different types of business flows have significant differences in their requirements for performance indicators such as network bandwidth, latency, and packet loss rate, resource scheduling strategies must be able to dynamically adapt to the needs of various business flows, and from the perspective of global network resource scheduling, various types of parameter indicators need to be classified and processed. At the same time, the wireless mobile network environment is susceptible to real-time changes due to factors such as the external environment and the electromagnetic environment, and the scheduling strategy must also be able to adapt to these environmental changes. In addition, since the computing power of devices in wireless mobile networks is limited, the design of the scheduling algorithm should fully consider the situation of limited computing power.

[0003] At present, a variety of intelligent algorithms have been applied to solve the problem of network resource scheduling, including neural network algorithms, ant colony algorithms, and reinforcement learning algorithms. However, the parameter adjustment of neural network algorithms is relatively complex, and usually requires a large amount of historical and real-time data to train the model to achieve ideal performance. Its training process not only requires a lot of computing resources and time costs, but also the neural network has less interaction with the environment during the training process, which may lead to a significant decrease in the algorithm effect when facing a frequently changing network environment. The ant colony algorithm takes a long time in the search process and is prone to fall into a local optimal solution. The reinforcement learning algorithm interacts with the environment in real time through a reward mechanism, but the setting of its learning rules is relatively demanding. Although the reinforcement learning algorithm has been more widely used after combining it with deep learning such as convolutional neural networks, it also brings more computing power loss, so it is not suitable for direct application in resource-constrained wireless mobile networks. In the existing research on the application of reinforcement learning to network resource scheduling, most of them use additive metrics such as delay and hop count as input parameters, while paying insufficient attention to other important metrics such as concave metrics such as remaining bandwidth. In fact, the q value of the remaining bandwidth is related to the entire transmission path of the service flow, and it is impossible to directly extract rewards from each action. Therefore, traditional deep reinforcement learning methods, such as DQN (Deep Q Learning), cannot be directly applied to resource scheduling in wireless mobile networks. Summary of the invention

[0004] The purpose of this application is to provide a wireless mobile communication network resource scheduling method and system based on DGQN, which can realize efficient and intelligent scheduling of communication network resources in diversified business scenarios.

[0005] To achieve the above objectives, this application provides the following solutions:

[0006] In a first aspect, the present application provides a wireless mobile communication network resource scheduling method based on DGQN, comprising the following steps:

[0007] The topological structure of the wireless mobile communication network and the index data of each node in the wireless mobile communication network are obtained; the index data include the delay parameters and bandwidth parameters of each node through different links.

[0008] The index data of each node in the wireless mobile communication network is preprocessed to obtain the preprocessed index data.

[0009] According to the topological structure of the wireless mobile communication network, the resource scheduling process of the wireless mobile communication network is mapped into a Markov decision process in a reinforcement learning algorithm.

[0010] A dual-path architecture is adopted to process the concave metric and additive metric of the resource scheduling process respectively, and the concave metric q value and the additive metric q value are obtained; the concave metric is the remaining bandwidth on the routing path of the business flow, and the additive metric is the transmission delay on the routing path of the business flow; the improved reinforcement learning algorithm DGQN is used to process the additive metric to obtain the additive metric q value; the improved reinforcement learning algorithm DGQN is a network improved by GhostNet on the traditional reinforcement learning algorithm DQN.

[0011] The concavity metric q value and the additive metric q value are weightedly summed to obtain the comprehensive model q value; the comprehensive model q value is a q value representation that combines the two types of indicators.

[0012] The final strategy is determined according to the q value of the comprehensive model, and the final routing path of any business flow is determined based on the final strategy.

[0013] The final routing path of the service flow is sent to each node of the wireless mobile communication network to perform resource scheduling of the wireless mobile communication network.

[0014] Optionally, preprocessing the indicator data of each node in the wireless mobile communication network to obtain the preprocessed indicator data specifically includes the following steps:

[0015] For any missing values ​​in the indicator data of a node, they are filled based on the same indicator data of the previous time frame.

[0016] All indicator data after the missing values ​​are filled are normalized to obtain preprocessed indicator data.

[0017] Optionally, a dual-path architecture is adopted to process the concave metric index and the additive metric index of the resource scheduling process respectively to obtain the concave metric index q value and the additive metric index q value, which specifically includes the following steps:

[0018] The concavity metric index is processed by the reinforcement learning algorithm Q-Learning to obtain the concavity metric index q value.

[0019] An improved reinforcement learning algorithm DGQN is used to process the additive metric index to obtain the additive metric index q value; the improved reinforcement learning algorithm DGQN is a network based on GhostNet, which uses a Ghost bottleneck structure to replace the convolution kernel in the traditional reinforcement learning algorithm DQN.

[0020] Optionally, the comprehensive model q value is calculated according to the following formula:

[0021]

[0022] Among them, q(s,a * ) is the comprehensive model q value, s is the state, a * is the action taken, q w (s,a * ) is the concavity metric q value, w thr is the bandwidth requirement of the service flow, q d (s,a * ) is the additive metric q value, qMin d =min(q d (s,a * )), μ is the weight of the additive metric q value, and A is the action set.

[0023] Alternatively, the concavity metric q-value can be calculated by the following formula:

[0024]

[0025] Among them, q w (s,a) is the concavity metric q value of taking action a in the current state s, q w (s -1 ,a * ) is the concavity metric q value of taking action a* in the state s at the previous moment, A is the action set, w ss' for…….

[0026] Optionally, the remaining bandwidth on the routing path of the service flow can be expressed by the following formula:

[0027] M W (E ρ )=min(w ij),ij∈E ρ .

[0028] Among them, M W (E ρ ) is the remaining bandwidth on the routing path of service flow ρ, w ij is the bandwidth parameter of link ij, E ρ is the routing path of service flow ρ.

[0029] Optionally, the transmission delay on the routing path of the service flow can be expressed by the following formula:

[0030]

[0031] Among them, M D (E ρ ) is the transmission delay on the routing path of the service flow ρ, d ij is the delay parameter of link ij, E ρ is the routing path of service flow ρ.

[0032] Optionally, the delay parameter of link ij can be determined by the following formula:

[0033]

[0034] Among them, d ij is the delay parameter of link ij, α T-1 ,…,α is the forgetting factor, T is the preset time length, is the initial delay parameter of link ij at the current moment, is the delay parameter of link ij at the previous moment.

[0035] Optionally, the forgetting factor α T-1 ,…,α increases gradually to reduce the impact of delay parameters far away from the current moment.

[0036] In a second aspect, the present application provides a wireless mobile communication network resource scheduling system based on DGQN, comprising the following modules:

[0037] The communication network parameter acquisition module is used to acquire the topological structure of the wireless mobile communication network and the index data of each node in the wireless mobile communication network; the index data includes the delay parameters and bandwidth parameters of each node passing through different links.

[0038] The index data preprocessing module is used to preprocess the index data of each node in the wireless mobile communication network to obtain the preprocessed index data.

[0039] The decision scheduling process mapping module is used to map the resource scheduling process of the wireless mobile communication network into a Markov decision process in a reinforcement learning algorithm according to the topological structure of the wireless mobile communication network.

[0040] The dual-path indicator processing module is used to process the concave metric indicator and the additive metric indicator of the resource scheduling process respectively in a dual-path architecture manner to obtain the concave metric indicator q value and the additive metric indicator q value; the concave metric indicator is the remaining bandwidth on the routing path of the business flow, and the additive metric indicator is the transmission delay on the routing path of the business flow; the improved reinforcement learning algorithm DGQN is used to process the additive metric indicator to obtain the additive metric indicator q value; the improved reinforcement learning algorithm DGQN is a network obtained by improving the traditional reinforcement learning algorithm DQN based on GhostNet.

[0041] The comprehensive model q value calculation module is used to perform weighted summation of the concavity measurement index q value and the additive measurement index q value to obtain the comprehensive model q value; the comprehensive model q value is a q value representation that combines the two types of indicators.

[0042] The final routing path determination module is used to determine the final strategy according to the q value of the comprehensive model, and determine the final routing path of any business flow based on the final strategy.

[0043] The final routing path sending module is used to send the final routing path of the service flow to each node of the wireless mobile communication network to perform resource scheduling of the wireless mobile communication network.

[0044] According to the specific embodiments provided in this application, this application discloses the following technical effects:

[0045] The present application provides a wireless mobile communication network resource scheduling method and system based on DGQN. In the method, after obtaining parameters and preprocessing in the wireless mobile communication network, the resource scheduling process of the wireless mobile communication network is first mapped to a Markov decision process in a reinforcement learning algorithm according to the topological structure of the wireless mobile communication network, and then a dual-path architecture is adopted to process the concave metric index and the additive metric index of the resource scheduling process respectively, wherein a network obtained by improving the traditional reinforcement learning algorithm DQN based on GhostNet is used to process the additive metric index, and then the processed concave metric index q value and the additive metric index q value are weightedly summed to obtain a comprehensive model q value, which integrates the q value representation of the two types of indicators; finally, the final strategy is determined first, and then the final routing path of the service flow is determined, and it is sent to each node of the wireless mobile communication network to realize resource scheduling of the wireless mobile communication network. This application targets wireless mobile communication networks with high-frequency changes in topology structures, takes into account the real-time dynamic changes of the network environment and the differentiated requirements of various business flows in the network for different performance indicators, comprehensively considers two types of indicators, namely concavity metric and additive metric, in the communication network, and designs a dual-path architecture for the characteristics of the two types of data in order to apply both types of indicators to network resource scheduling calculations, so that there are constraints and connections between them. In addition, the traditional reinforcement learning algorithm DQN is improved, and a lightweight improved reinforcement learning algorithm DGQN is designed to dynamically adapt to differentiated services and network environments, and to a certain extent solve the problem of communication network resource scheduling under limited computing power. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0047] Figure 1 A flowchart of a wireless mobile communication network resource scheduling method based on DGQN is provided in an embodiment of the present application.

[0048] Figure 2 A schematic diagram of a dual-path architecture in a wireless mobile communication network resource scheduling method based on DGQN provided in an embodiment of the present application.

[0049] Figure 3 A schematic diagram of an improved DGQN algorithm in a DGQN-based wireless mobile communication network resource scheduling method provided in an embodiment of the present application.

[0050] Figure 4A schematic diagram of functional modules of a wireless mobile communication network resource scheduling system based on DGQN provided in one embodiment of the present application.

[0051] Figure 5 A schematic diagram of the structure of a computer device provided in one embodiment of the present application. DETAILED DESCRIPTION

[0052] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0053] In order to make the above-mentioned objects, features and advantages of the present application more obvious and easy to understand, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.

[0054] In an exemplary embodiment, Figure 1 As shown, a wireless mobile communication network resource scheduling method based on DGQN is provided, which can be executed by a computer device and includes the following steps:

[0055] S1. Obtain the topological structure of the wireless mobile communication network and the index data of each node in the wireless mobile communication network; the index data includes the delay parameters and bandwidth parameters of each node passing through different links. In this embodiment, the index data of each node is obtained by reporting through network monitoring software.

[0056] The topological structure of the communication network is represented by a graph G(V,E), where V and E are the point set and edge set in the communication network respectively, i∈V represents the Mesh node in the network, and is defined as is the set of neighbor nodes of node i, ij∈E represents the wireless link from node i to node j.

[0057] In this application, it is assumed that the service flow is inseparable. To meet the differentiated QoS (Quality of Service) requirements of the service flow, two types of QoS indicators of the routing path are mainly considered. M is defined D (E ρ ) represents the transmission delay on the routing path of the service flow ρ, M W (E ρ ) represents the remaining bandwidth on the routing path of service flow ρ. The bandwidth parameter is a concave measure and is the remaining bandwidth w of all links on the routing path. ij The remaining bandwidth on the routing path of the service flow can be expressed by the following formula:

[0058] MW (E ρ )=min(w ij ),ij∈E ρ .

[0059] Among them, M W (E ρ ) is the remaining bandwidth on the routing path of service flow ρ, w ij is the bandwidth parameter of link ij, E ρ is the routing path of service flow ρ.

[0060] The delay parameter is an additive measure and is the delay d of all links on the routing path. ij The transmission delay on the routing path of the service flow can be expressed by the following formula:

[0061]

[0062] Among them, M D (E ρ ) is the transmission delay on the routing path of the service flow ρ, d ij is the delay parameter of link ij, E ρ is the routing path of service flow ρ.

[0063] In a dynamic network environment, in order to ensure that the routing path has a certain degree of anti-fluctuation and stability, the parameters of the link need to maintain high performance for a certain period of time. Therefore, the average value of the delay parameter within the preset time length of the link is used as the representation of the delay at the current moment. The delay parameter of link ij can be determined by the following formula:

[0064]

[0065] Among them, d ij is the delay parameter of link ij, α T-1 ,…,α is the forgetting factor, T is the preset time length, is the initial delay parameter of link ij at the current moment, is the delay parameter of link ij at the previous moment.

[0066] Since the reference value of information far from the current time is low, its contribution ratio needs to be reduced, and the forgetting factor α is introduced to reduce its interference effect. T-1 ,…,α increases gradually to reduce the impact of delay parameters far away from the current moment.

[0067] The resource scheduling problem of the communication network is modeled as a multi-objective optimization problem. The bandwidth parameter is expressed as a constraint. When the remaining bandwidth of the routing path is greater than the bandwidth requirement of the service flow, a link with lower latency and higher remaining bandwidth is selected; otherwise, the service request is rejected. ij=1 represents the path ij in the routing path E of the service ρ On, otherwise x ij =0. The mathematical model is expressed as follows:

[0068]

[0069] maximize M W (E ρ )-w thr

[0070] subject to M W (E ρ )≥w thr

[0071]

[0072] Where t represents the service life cycle, E[ε] t Represents the expectation within time t. In the model, target 1 represents the obtained routing path E ρ The expected transmission delay is minimized, and goal two represents maximizing the remaining bandwidth.

[0073] Convert objective 2 into a minimization form, and add one to the denominator to avoid it being zero:

[0074]

[0075] The weight μ combines the two objectives and optimizes them simultaneously.

[0076]

[0077] S2, preprocessing the index data of each node in the wireless mobile communication network to obtain the preprocessed index data. Specifically in this embodiment, step S2 includes the following steps:

[0078] S21. Fill the missing values ​​in the indicator data of any node with the same indicator data of the previous time frame.

[0079] S22, normalizing all the index data after the vacancy value filling is completed to obtain the pre-processed index data. In this step, the index data is normalized to facilitate the subsequent calculation.

[0080] S3. According to the topological structure of the wireless mobile communication network, the resource scheduling process of the wireless mobile communication network is mapped to the Markov decision process in the reinforcement learning algorithm. Specifically in this embodiment, the topological structure graph G(V,E) is mapped to the Markov decision process E=(S,A,P,R), where S represents the state set, A represents the action set, P represents the probability transfer matrix, and R represents the reward function. The corresponding relationship between the routing problem of the communication network and the Markov decision process is shown in Table 1:

[0081] Table 1 The correspondence between the routing problem of the communication network and the Markov decision process

[0082] Markov Decision Process Routing issues in communication networks State set S Each node in the graph G(V,E) Dynamic Set A Which neighbor node is selected as the next hop node Transition probability P For the selected action transition probability p = 1 Reward R <![CDATA[Link parameter index d ij , w ij related]]>

[0083] In the reinforcement learning algorithm, the agent and the environment interact in a series of discrete times to complete an overall goal task. The interaction between the agent and the environment includes the following processes:

[0084] (a) At time t, the agent observes the state s of the environment i ∈S, the agent follows the strategy π(s i )Select action a i .

[0085] (b) Action a i Acting on the environment, the environmental state changes to s i+1 , the environment gives the agent a corresponding reward r i .

[0086] (c) The agent will then make new decisions based on the new environment state and strategy.

[0087] (d) The above interaction process is repeated until the agent completes the corresponding target task.

[0088] In the reinforcement learning algorithm, the goal of the agent is to maximize the cumulative reward. It can be defined as:

[0089]

[0090] Q under T-step cumulative reward π (s,a) can be defined as:

[0091]

[0092] Perform full probability expansion to get T-step cumulative rewards

[0093]

[0094]

[0095] S4. A dual-path architecture is adopted to process the concave metric and additive metric of the resource scheduling process respectively to obtain the concave metric q value and the additive metric q value; the concave metric is the remaining bandwidth on the routing path of the business flow, and the additive metric is the transmission delay on the routing path of the business flow; the improved reinforcement learning algorithm DGQN is used to process the additive metric to obtain the additive metric q value; the improved reinforcement learning algorithm DGQN is a network obtained by improving the traditional reinforcement learning algorithm DQN based on GhostNet.

[0096] Considering that the remaining bandwidth is a concave metric and the transmission delay is an additive metric, the two types of indicators are updated in different ways. Therefore, this application adopts different calculation methods for the two types of indicators when designing the Q network. The dual-path architecture is adopted to process the concave metric and the additive metric respectively, and there are certain constraints between the two, such as Figure 2 A schematic diagram of the dual-path architecture is shown.

[0097] In this embodiment, step S4 specifically includes the following steps:

[0098] S41. The concavity metric is processed using the reinforcement learning algorithm Q-Learning to obtain the concavity metric q value. Assuming the action set is A, in state s, action a selects s' as the next hop node. For the remaining bandwidth as a concavity metric, its reward r has no definite pattern and cannot be directly obtained, because the q value is the minimum value of the remaining bandwidth of the entire transmission link. As an example, the concavity metric q value can be calculated by the following formula:

[0099]

[0100] Among them, q w (s,a) is the concavity metric q value of taking action a in the current state s, q w (s -1 ,a * ) is the concavity metric q value of taking action a* in the state s at the previous moment, A is the action set, w ss' for…….

[0101] Existence constraint: When q w (s,a)>w thr Then, for additive metrics such as link delay and hop count, the reward r is calculated based on the delay from s to s', taking -d ij , when s' is the target node (s is the current node, s' is the next hop node state), r is 0.

[0102] S42. An improved reinforcement learning algorithm DGQN is used to process the additive metric to obtain an additive metric q value; the improved reinforcement learning algorithm DGQN is a network based on Ghost Net, which uses a Ghost bottleneck structure to replace the convolution kernel in the traditional reinforcement learning algorithm DQN.

[0103] In Ghostnet, linear mapping is used to replace part of the convolution calculation, which effectively reduces the computing power burden. In this embodiment, DQN is improved based on GhostNet, and the Ghostbottleneck structure is used to replace the traditional convolution kernel to design a lightweight DGQN (Deep Ghost Q Network) algorithm. The improved DGQN schematic diagram is shown in the figure. Figure 3 The specific algorithm flow of DGQN is as follows:

[0104] (I). Input state space S, action space A, and target network parameter update frequency F.

[0105] (II). Initialize the experience replay library D with a capacity of N.

[0106] (III) Randomly initialize the actual Q network parameters and target The network parameters are initialized to be equal. The Q network adopts the ghostbottleneck structure with a lightweight design, which effectively saves computing power.

[0107] (IV). For the starting routing node i, the starting routing node state is s i , select action a=π(s i ), execute action a, the agent interacts with the environment, and obtains a single-step immediate reward r and a new state s'.

[0108] (V). i ,a,r,s' is put into the experience replay library D.

[0109] (VI). Sample new ss, aa, rr, ss' data sets from D for training.

[0110] (VII). Train the Q network for the loss function , where γ is the discount rate.

[0111] (VIII). Status update: s←s'.

[0112] (VIIII). Every F steps, The network parameters are updated.

[0113] (X). When state s is the destination node, return to step IV.

[0114] (XI). Training ends when the Q network parameters converge or other termination conditions are met.

[0115] S5. Perform weighted summation of the concavity metric index q value and the additive metric index q value to obtain the comprehensive model q value; the comprehensive model q value is a q value representation that combines the two types of indicators. In an exemplary embodiment of the present application, the comprehensive model q value is calculated according to the following formula:

[0116]

[0117] Among them, q(s,α * ) is the comprehensive model q value, s is the state, a * is the action taken, q w (s,a * ) is the concavity metric q value, w thr is the bandwidth requirement of the service flow, q d (s,a * ) is the additive metric q value, qMin d =min(q d (s,a * )), μ is the weight of the additive metric q value, and A is the action set.

[0118] S6. Determine the final strategy based on the comprehensive model q value, and determine the final routing path of any service flow based on the final strategy. * ) converges or reaches the set number of iterations, the complete q value representation of the two types of indicators can be obtained, where each value represents the optimal reward value that can be obtained by selecting each action in the current state. q(s,a) is the normalized weighted value of the two types of objectives. According to q(s,a), the smallest q value is selected in turn to obtain the final strategy π. According to π, the optimal path for each service transmission can be obtained to achieve efficient scheduling of communication network resources.

[0119] S7, send the final routing path of the service flow to each node of the wireless mobile communication network to schedule the resources of the wireless mobile communication network. Send the final routing path calculated based on the above process to the relevant nodes for execution according to the signaling, so as to realize the efficient transmission of diversified services and the intelligent scheduling of communication network resources.

[0120] This application scheme applies reinforcement learning technology to the frequently changing topology of wireless mobile communication networks, fully considering the real-time dynamic change characteristics of the wireless mobile communication network environment and the differentiated needs of services for different performance indicators. By modeling the resource scheduling problem of the wireless mobile communication network as a multi-objective optimization problem, two types of indicators, namely concavity measurement and additive measurement, in the communication network are comprehensively considered. In order to enable both types of indicators to be effectively applied to the calculation process of network resource scheduling, a dual-path architecture is designed for different data characteristics, and the two paths in the architecture constrain each other and maintain contact. In addition, this scheme improves the traditional reinforcement learning algorithm DQN and proposes a lightweight reinforcement learning algorithm DGQN to dynamically adapt to changes in different services and network environments, thereby solving the problem of communication network resource scheduling under limited computing power to a certain extent.

[0121] Based on the same inventive concept, the embodiment of the present application also provides a system for implementing the wireless mobile communication network resource scheduling method based on DGQN. The implementation solution provided by the system is similar to the implementation solution recorded in the above method. In an exemplary embodiment, Figure 4 As shown, a wireless mobile communication network resource scheduling system based on DGQN is provided, comprising the following modules:

[0122] The communication network parameter acquisition module is used to acquire the topological structure of the wireless mobile communication network and the index data of each node in the wireless mobile communication network; the index data includes the delay parameters and bandwidth parameters of each node passing through different links.

[0123] The index data preprocessing module is used to preprocess the index data of each node in the wireless mobile communication network to obtain the preprocessed index data.

[0124] The decision scheduling process mapping module is used to map the resource scheduling process of the wireless mobile communication network into a Markov decision process in a reinforcement learning algorithm according to the topological structure of the wireless mobile communication network.

[0125] The dual-path indicator processing module is used to process the concave metric indicator and the additive metric indicator of the resource scheduling process respectively in a dual-path architecture manner to obtain the concave metric indicator q value and the additive metric indicator q value; the concave metric indicator is the remaining bandwidth on the routing path of the business flow, and the additive metric indicator is the transmission delay on the routing path of the business flow; the improved reinforcement learning algorithm DGQN is used to process the additive metric indicator to obtain the additive metric indicator q value; the improved reinforcement learning algorithm DGQN is a network obtained by improving the traditional reinforcement learning algorithm DQN based on GhostNet.

[0126] The comprehensive model q value calculation module is used to perform weighted summation of the concavity measurement index q value and the additive measurement index q value to obtain the comprehensive model q value; the comprehensive model q value is a q value representation that combines the two types of indicators.

[0127] The final routing path determination module is used to determine the final strategy according to the q value of the comprehensive model, and determine the final routing path of any business flow based on the final strategy.

[0128] The final routing path sending module is used to send the final routing path of the service flow to each node of the wireless mobile communication network to perform resource scheduling of the wireless mobile communication network.

[0129] certainly, Figure 4 The architecture shown is only exemplary and can be omitted according to actual needs when implementing different functions. Figure 4 One or at least two components of the system shown.

[0130] In an exemplary embodiment, a computer device is provided. The computer device may be a server or a terminal. The internal structure diagram thereof may be as follows: Figure 5 As shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, referred to as I / O) and a communication interface. The processor, the memory and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a wireless mobile communication network resource scheduling method based on DGQN provided in the above embodiment can be implemented.

[0131] Those skilled in the art will understand that Figure 5 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.

[0132] In an exemplary embodiment, a computer device is further provided, including a memory and a processor, wherein a computer program is stored in the memory, and the processor implements the steps in the above-mentioned method embodiments when executing the computer program.

[0133] In an exemplary embodiment, a computer-readable storage medium is provided, storing a computer program, and when the computer program is executed by a processor, the steps in the above method embodiments are implemented.

[0134] In an exemplary embodiment, a computer program product is provided, including a computer program, and when the computer program is executed by a processor, the steps in the above method embodiments are implemented.

[0135] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant regulations.

[0136] Those of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to the memory, database or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM may be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM).

[0137] The database involved in each embodiment provided in this application may include at least one of a relational database and a non-relational database. The non-relational database may include a distributed database based on blockchain, etc., but is not limited thereto. The processor involved in each embodiment provided in this application may be a general-purpose processor, a central processing unit, a graphics processor, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, etc., but is not limited thereto.

[0138] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0139] This article uses specific examples to illustrate the principles and implementation methods of this application. The description of the above embodiments is only used to help understand the method and core ideas of this application. At the same time, for those skilled in the art, according to the ideas of this application, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as limiting this application.

Claims

1. A wireless mobile communication network resource scheduling method based on DGQN, characterized in that: include: Acquire the topological structure of the wireless mobile communication network and the index data of each node in the wireless mobile communication network; the index data includes the delay parameters and bandwidth parameters of each node passing through different links; Preprocessing the index data of each node in the wireless mobile communication network to obtain preprocessed index data; According to the topological structure of the wireless mobile communication network, the resource scheduling process of the wireless mobile communication network is mapped into a Markov decision process in the reinforcement learning algorithm; A dual-path architecture is adopted to process the concave metric and the additive metric of the resource scheduling process respectively, and obtain the concave metric q value and the additive metric q value; the concave metric is the remaining bandwidth on the routing path of the service flow, and the additive metric is the transmission delay on the routing path of the service flow; the improved reinforcement learning algorithm DGQN is adopted to process the additive metric to obtain the additive metric q value; the improved reinforcement learning algorithm DGQN is a network obtained by improving the traditional reinforcement learning algorithm DQN based on Ghost Net; Performing a weighted summation of the concavity metric index q value and the additive metric index q value to obtain a comprehensive model q value; The q value of the comprehensive model is a q value representation that combines two types of indicators; Determine a final strategy according to the q value of the comprehensive model, and determine a final routing path of any service flow based on the final strategy; The final routing path of the service flow is sent to each node of the wireless mobile communication network to perform resource scheduling of the wireless mobile communication network.

2. The wireless mobile communication network resource scheduling method based on DGQN according to claim 1 is characterized in that: Preprocessing the index data of each node in the wireless mobile communication network to obtain preprocessed index data, specifically including: For any missing values ​​in the indicator data of a node, fill them with the same indicator data of the previous time frame; All indicator data after the missing values ​​are filled are normalized to obtain preprocessed indicator data.

3. The wireless mobile communication network resource scheduling method based on DGQN according to claim 1, characterized in that: The dual-path architecture is adopted to process the concave metric and additive metric of the resource scheduling process respectively, and obtain the concave metric q value and the additive metric q value, which specifically include: The concavity metric index is processed by using a reinforcement learning algorithm Q-Learning to obtain a q value of the concavity metric index; The additive metric index is processed by using an improved reinforcement learning algorithm DGQN to obtain the additive metric index q value; the improved reinforcement learning algorithm DGQN is a network based on GhostNet, which uses a Ghost bottleneck structure to replace the convolution kernel in the traditional reinforcement learning algorithm DQN.

4. The wireless mobile communication network resource scheduling method based on DGQN according to claim 1, characterized in that: The comprehensive model q value is calculated according to the following formula: Among them, q(s,a * ) is the comprehensive model q value, s is the state, a * is the action taken, q w (s,a * ) is the concavity metric q value, w thr is the bandwidth requirement of the service flow, q d (s,a * ) is the additive metric q value, qMin d =min(q d (s,a * )), μ is the weight of the additive metric q value, and A is the action set.

5. The wireless mobile communication network resource scheduling method based on DGQN according to claim 4 is characterized in that: The concavity metric q value can be calculated by the following formula: Among them, q w (s,a) is the concavity metric q value of taking action a in the current state s, q w (s -1 ,a * ) is the concavity metric q value of taking action a* in the state s at the previous moment, A is the action set, w ss' for…….

6. The wireless mobile communication network resource scheduling method based on DGQN according to claim 1, characterized in that: The remaining bandwidth on the routing path of the service flow can be expressed by the following formula: M W (IT ρ )=min(w ij ),ij∈E ρ 4 Among them, m W (E ρ ) is the remaining bandwidth on the routing path of service flow ρ, w ij is the bandwidth parameter of link ij, E ρ is the routing path of service flow ρ.

7. The wireless mobile communication network resource scheduling method based on DGQN according to claim 1, characterized in that: The transmission delay on the routing path of the service flow can be expressed by the following formula: Among them, M D (E ρ ) is the transmission delay on the routing path of the service flow ρ, d ij is the delay parameter of link ij, E ρ is the routing path of service flow ρ.

8. The wireless mobile communication network resource scheduling method based on DGQN according to claim 7, characterized in that: The delay parameter of link ij can be determined by the following formula: Among them, d ij is the delay parameter of link ij, α T-1 ,…,α is the forgetting factor, T is the preset time length, is the initial delay parameter of link ij at the current moment, is the delay parameter of link ij at the previous moment.

9. The wireless mobile communication network resource scheduling method based on DGQN according to claim 8, characterized in that: Forgetting Factor α T-1 ,…,α increases gradually to reduce the impact of delay parameters far away from the current moment.

10. A wireless mobile communication network resource scheduling system based on DGQN, characterized in that: include: A communication network parameter acquisition module, used to acquire the topological structure of the wireless mobile communication network and the index data of each node in the wireless mobile communication network; the index data includes the delay parameters and bandwidth parameters of each node passing through different links; The index data preprocessing module is used to preprocess the index data of each node in the wireless mobile communication network to obtain the preprocessed index data; A decision scheduling process mapping module is used to map the resource scheduling process of the wireless mobile communication network into a Markov decision process in a reinforcement learning algorithm according to the topological structure of the wireless mobile communication network; A dual-path indicator processing module is used to process the concave metric indicator and the additive metric indicator of the resource scheduling process respectively in a dual-path architecture manner to obtain a concave metric indicator q value and an additive metric indicator q value; the concave metric indicator is the remaining bandwidth on the routing path of the business flow, and the additive metric indicator is the transmission delay on the routing path of the business flow; the additive metric indicator is processed by using an improved reinforcement learning algorithm DGQN to obtain the additive metric indicator q value; the improved reinforcement learning algorithm DGQN is a network obtained by improving the traditional reinforcement learning algorithm DQN based on GhostNet; A comprehensive model q value calculation module, used for performing weighted summation of the concavity metric index q value and the additive metric index q value to obtain a comprehensive model q value; The q value of the comprehensive model is a q value representation that combines two types of indicators; A final routing path determination module, used to determine a final strategy according to the q value of the comprehensive model, and determine a final routing path of any service flow based on the final strategy; The final routing path sending module is used to send the final routing path of the service flow to each node of the wireless mobile communication network to perform resource scheduling of the wireless mobile communication network.

Citation Information

Patent Citations

  • Network resource division and path planning joint optimization method based on bilevel planning

    CN116915622A

  • Subway pedestrian anomaly detection method and system based on improved yolov8 network and clip model

    CN119478529A

  • Method for constructing wireless network resource allocation system, and resource management method

    WO2024207564A1