Intelligent traffic scheduling method and system based on attention mechanism and deep reinforcement learning

By combining global and local attention mechanisms with deep reinforcement learning, an intelligent traffic scheduling method is developed, which solves the problem of inflexible scheduling in dynamic network environments by traditional algorithms and achieves efficient network resource utilization and load balancing.

CN120957183APending Publication Date: 2025-11-14GUIZHOU POWER GRID CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510850377.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-24
Publication Date
2025-11-14

AI Technical Summary

Technical Problem

Traditional traffic scheduling algorithms struggle to effectively handle sudden traffic changes, network topology shifts, and congestion in dynamic and complex network environments, and are difficult to adjust according to network conditions and performance requirements.

Method used

By combining global and local attention mechanisms with deep reinforcement learning, a centrally coordinated-distributed decision-making system is constructed. This system optimizes traffic scheduling using policy networks and value networks, extracts multi-scale feature vectors, and achieves global optimum through iterative optimization of policy networks and value networks.

Benefits of technology

It improves the accuracy of traffic scheduling and adaptability to dynamic environments, achieves more efficient network resource utilization and performance optimization, can flexibly respond to complex changes, and realize efficient resource allocation and load balancing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120957183A_ABST
    Figure CN120957183A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of wireless communication, and discloses an intelligent traffic scheduling method and system based on an attention mechanism and deep reinforcement learning, and the method comprises the steps: capturing the fine-grained traffic features through a local attention mechanism, carrying out the deep understanding of a global attention mechanism on a network overall traffic mode, and carrying out the real-time traffic scheduling of the network. Local details and global context information are integrated and combined to extract flow characteristics of the network, a central coordination-distributed decision-making system of multiple collaborative optimization units is adopted, the flow characteristics are used as input, and strategy actions obtained by each collaborative optimization unit through a strategy network are subjected to strategy optimization. The value network fits a global state value function to evaluate the current action and environment state and feed back the current action and environment state to the strategy network, a random gradient strategy is adopted to update a strategy, and a scheduling decision is continuously optimized. Through multiple times of training and iteration of central coordination-distributed decision making, optimization and coordination of a collaborative optimization unit are realized, the network state is continuously monitored, and a real-time decision is made for analyzing the network flow.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of wireless communication technology, and in particular to an intelligent traffic scheduling method and system based on attention mechanisms and deep reinforcement learning. Background Technology

[0002] Traditional traffic scheduling algorithms typically rely on predefined rules or simple network metrics, such as bandwidth utilization or latency, to allocate traffic in a static or semi-static manner. While these algorithms perform well in fixed and predictable network environments, they often struggle to effectively handle sudden traffic fluctuations, network topology changes, and congestion in dynamic and complex environments. Furthermore, once a scheduling strategy is set, traditional algorithms are difficult to adjust according to network conditions and performance requirements. Therefore, there is an urgent need to introduce new theories and technologies to advance this field and achieve more innovative and efficient design methods.

[0003] Introducing global and local attention mechanisms into intelligent traffic scheduling algorithms can significantly improve model performance. Local attention mechanisms focus on the detailed features of neighboring nodes, capturing local traffic patterns and network topology characteristics, which helps address traffic peaks and bottlenecks. Global attention mechanisms, on the other hand, focus on the overall network information, capturing dependencies between long-distance nodes and global traffic distribution trends, thus optimizing overall network performance. By combining these two mechanisms, the model not only enhances its ability to represent fine-grained local and macro-global information but also improves the accuracy of traffic scheduling decisions and its adaptability to dynamic environments. Simultaneously, attention mechanisms help filter out irrelevant information, enhancing the information processing capabilities of collaborative optimization units, ultimately achieving more efficient network resource utilization and performance optimization.

[0004] Reinforcement learning is an intelligent decision-making technique that combines deep learning and reinforcement learning. Introducing reinforcement learning into traffic scheduling algorithms enables collaborative optimization units to select scheduling actions based on real-time traffic changes and network conditions in a dynamic network environment. These actions are then updated with immediate rewards to achieve global optimum. The optimization process involves iterative optimization through trial and error and feedback, gradually improving the efficiency and adaptability of the scheduling strategy. The introduction of reinforcement learning allows traffic scheduling algorithms to more flexibly respond to complex network changes, achieving efficient resource allocation and load balancing. Summary of the Invention

[0005] In view of the aforementioned existing problems, the present invention is proposed.

[0006] Therefore, this invention provides an intelligent traffic scheduling method based on attention mechanisms and deep reinforcement learning, which can overcome the bottleneck of traditional traffic scheduling algorithms that struggle to capture key traffic and adjust it to the optimal global reward based on network conditions and performance requirements by combining global and local attention mechanisms with reinforcement learning.

[0007] To address the aforementioned technical problems, this invention provides the following technical solution: an intelligent traffic scheduling method based on attention mechanisms and deep reinforcement learning, comprising: extracting multi-scale feature vectors of network traffic by fusing local and global attention mechanisms; constructing a centrally coordinated distributed decision-making system, including a collaborative optimization unit composed of a policy network and a value network; inputting the multi-scale feature vectors into the policy network to generate scheduling actions, evaluating the value of the actions through the value network and feeding it back to the policy network; iteratively optimizing the policy network and the value network until the target reward converges to the global optimum.

[0008] As a preferred embodiment of the intelligent traffic scheduling method based on attention mechanism and deep reinforcement learning described in this invention, wherein: the local attention mechanism captures fine-grained traffic features;

[0009] The global attention mechanism captures the overall network traffic pattern, and the two are spliced ​​and fused to form the multi-scale feature vector.

[0010] As a preferred embodiment of the intelligent traffic scheduling method based on attention mechanism and deep reinforcement learning described in this invention, the value network includes: calculating the time difference target value based on experience buffer data, outputting the value function prediction value through dual online value networks respectively, and selecting the minimum value of the output value of the dual online value networks to update the target value network parameters.

[0011] As a preferred embodiment of the intelligent traffic scheduling method based on attention mechanism and deep reinforcement learning described in this invention, the policy network includes an online policy network and a target policy network. The parameters of the target policy network are softly updated by importance sampling, and the online policy network is updated by random policy gradient.

[0012] As a preferred embodiment of the intelligent traffic scheduling method based on attention mechanism and deep reinforcement learning described in this invention, the local attention mechanism includes dividing the traffic data into non-overlapping windows.

[0013] Calculate the request tensor, identifier tensor, and value tensor for each window;

[0014] After performing pooling operations on the request tensor and the identifier tensor, the similarity score matrix is ​​calculated.

[0015] The normalized similarity score matrix is ​​then weighted and a weighted tensor is used to generate local attention features.

[0016] As a preferred embodiment of the intelligent traffic scheduling method based on attention mechanism and deep reinforcement learning described in this invention, the global attention mechanism includes selecting the first i maximum values ​​from the local attention similarity matrix to generate a correlation matrix.

[0017] Extract the identifier sub-tensor and value sub-tensor corresponding to the correlation matrix;

[0018] Global attention features are generated by combining the request tensor and subtensors.

[0019] As a preferred embodiment of the intelligent traffic scheduling method based on attention mechanism and deep reinforcement learning described in this invention, the method comprises: before extracting the multi-scale feature vector of network traffic, constructing a traffic distribution matrix dataset and a link matrix D within a time period t;

[0020] In this matrix, the rows and columns represent time points, sensor data, and network data, respectively. Based on the time points, each sensor data point and network environment data point is used as a parameter to fill the corresponding column of the matrix.

[0021] Data is collected once per second, and a traffic distribution matrix is ​​constructed over three seconds. The traffic distribution matrix dataset is obtained within a time period t.

[0022] The rows and columns of the link matrix D represent the numbers of the sending and receiving nodes, respectively. If the nodes are connected, the parameter of the element in D is 1; otherwise, it is 0.

[0023] This invention provides an intelligent traffic scheduling system based on attention mechanisms and deep reinforcement learning.

[0024] As a preferred embodiment of the intelligent traffic scheduling system based on attention mechanism and deep reinforcement learning described in this invention, it includes a feature extraction module, a decision generation module, a value evaluation module, and an optimization control module.

[0025] The feature extraction module integrates local and global attention mechanisms to extract multi-scale feature vectors from traffic data.

[0026] The decision generation module generates a probability distribution of scheduling actions based on feature vectors and then executes the decision.

[0027] The value assessment module evaluates the value of an action through a dual-value network and updates the target network with the minimum value.

[0028] The optimization control module controls the iterative optimization of the policy network and the value network, driving the reward to converge to the global optimum.

[0029] The present invention provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of an intelligent traffic scheduling method based on attention mechanism and deep reinforcement learning.

[0030] The present invention provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the steps of an intelligent traffic scheduling method based on attention mechanism and deep reinforcement learning.

[0031] The beneficial effects of this invention are as follows: This invention utilizes global and local attention mechanisms to extract network features. The local attention mechanism focuses on capturing fine-grained information of specific regions or nodes, enabling detailed analysis of the features and micro-patterns of local data. The global attention mechanism can capture the overall flow patterns and structural information of the network, taking into account the relationships between all input data and providing an overall view. Combining the two provides a multi-level and multi-perspective capability for network feature extraction. The extracted features can capture key information while filtering out irrelevant features.

[0032] By combining reinforcement learning, we define the network state, action space, and reward function, and initialize the collaborative optimization unit. Then, each unit selects an action based on the observed environment and probability distribution, and receives state transitions and reward feedback after execution. The collaborative optimization unit updates its policy using a stochastic gradient strategy, continuously optimizing the scheduling decision. Finally, after multiple training and iterations, the target reward reaches the global optimum. Attached Figure Description

[0033] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0034] Figure 1 This diagram illustrates the centralized training and distributed execution of an intelligent traffic scheduling method based on attention mechanism and deep reinforcement learning, as provided in an embodiment of the present invention.

[0035] Figure 2 This is a schematic diagram illustrating the extraction of network features in an intelligent traffic scheduling method based on attention mechanism and deep reinforcement learning, provided as an embodiment of the present invention.

[0036] Figure 3 This is a schematic diagram of the strategy-value network of an intelligent traffic scheduling method based on attention mechanism and deep reinforcement learning, provided as an embodiment of the present invention. Detailed Implementation

[0037] To make the above-mentioned objects, features, and advantages of the present invention more apparent and understandable, specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of them. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the protection scope of the present invention.

[0038] Example 1, referring to Figure 1 This is the first embodiment of the present invention, which provides an intelligent traffic scheduling method based on attention mechanisms and deep reinforcement learning, including:

[0039] S1: Extract multi-scale feature vectors of network traffic by fusing local attention mechanisms and global attention mechanisms.

[0040] It should be noted that before extracting the multi-scale feature vectors of network traffic, a traffic distribution matrix dataset is constructed within the time period t to determine the link matrix D;

[0041] Construct a traffic distribution matrix dataset within a time period t, and determine the link matrix D. The rows of the traffic distribution matrix represent time nodes per second, and the columns include sensor data and network status data. The data collection interval is determined to be once per second, with a duration of three seconds. Fill the collected data into the corresponding columns of the matrix according to time to obtain a traffic distribution matrix dataset. The rows and columns of the link matrix D represent the numbers of the sending and receiving nodes, respectively. If nodes are connected, the element parameter in D is 1; otherwise, it is 0.

[0042] Furthermore, a feature extraction network is constructed based on local and global attention mechanisms. This feature extraction network includes both local and global attention mechanisms.

[0043] The input is X∈R N×H×W Where N represents the batch size, H represents the row of the flow distribution matrix, W represents the column of the flow matrix, and X is a tensor;

[0044] This includes a traffic distribution matrix and a link matrix (N=2). First, the input data is divided into blocks X∈R. H×W×N Remodeled as:

[0045]

[0046] Where α is an identifier, it is divided into p×p non-overlapping windows, such that each window carries Given a feature vector, a window-level request R, identifier I, and value V are derived in tensor form through linear transformation:

[0047] R = X α W q +S q I = X α W k +S k V = X α W v +S v

[0048] Among them W q W k W v ∈R represent the linear transformation weights of request R, identifier I, and value V, respectively. These are the linear transformation bias values ​​of request R, identifier I, and value V, respectively.

[0049] It should be noted that, for local attention capture, after obtaining three tensors through the above steps, the request R and the identifier I are averaged in the spatial dimension to obtain a tensor request R with lower spatial resolution but containing window-level information. t and I identifier t Then, the similarity score matrix H is obtained by calculating the similarity between the two tensors that have undergone average pooling. t :

[0050] H t =R t (I t ) T

[0051] Where T is the transpose matrix; then, after normalizing the similarity score matrix, the local attention feature vector O is calculated. P :

[0052]

[0053] Where O∈R H×W×N , As a scalar factor, the local attention weight O is obtained. P Focus on information interaction within a local window and handle local details and spatial relationships.

[0054] For global attention capture, in the similarity score matrix H t Select the first "i" maximum values ​​to obtain the correlation matrix. J t The most relevant identifiers I for each request R:

[0055] J t =First i (H t )

[0056] Then, the parts corresponding to the identifier I and value V in the correlation matrix are extracted:

[0057] I f =Extract(I,J t ),V f =Extract(V,J t )

[0058] The extracted tensor is flattened, and then the request R and the flattened tensor R are combined. f V f Calculate the global attention feature vector O G :

[0059]

[0060] Finally, the global and local attention features are concatenated and fused to obtain the final output feature vector O:

[0061] O = O P +O G

[0062] It should be noted that by integrating local attention mechanisms and global attention mechanisms, and simultaneously capturing micro-level sudden traffic changes and macro-level traffic distribution trends, the accuracy of identifying network bottlenecks and congestion points is improved, and feature extraction errors are reduced.

[0063] S2: Construct a centrally coordinated, distributed decision-making system, which includes collaborative optimization units composed of a strategy network and a value network.

[0064] Furthermore, an experience buffer is constructed, in which the policy network in each collaborative optimization unit is independent. First, the network parameters are initialized, and the network traffic feature vector generated in the above steps is input into the online policy network. The online policy network observes the probability distribution. Select Action Actions Environmental conditions award and the environmental state S at the next moment t+1 Temporarily store in the buffer used for sampling input to the value network.

[0065] In this embodiment, the experience buffer is constructed by initializing network parameters and uniformly sampling by storing (state, action, reward, next state) tuples.

[0066] In an optional embodiment, the construction of the experience buffer can also be achieved through priority experience replay. Specifically, a priority (such as the absolute value of the TD error) is calculated for each stored experience sample. The larger the error, the higher the priority. Samples are sampled according to the priority ratio. Samples with high error have a higher probability of being selected. Weights are calculated during sampling to compensate for priority bias. The priority of the samples is updated after each training to reflect the latest network error.

[0067] In another alternative embodiment, the experience buffer can also be constructed by a clustering buffer based on traffic features. Specifically, the stored experience is clustered into multiple sub-buffers according to traffic features (such as matrix similarity). Samples are drawn from each cluster subset in equal amounts each time. When new samples are added, they are automatically classified into the nearest neighbor cluster. Similar clusters are merged periodically, and independent subsets are set up for burst traffic samples to avoid being overwhelmed by regular samples.

[0068] A value network is constructed. The value network in a single collaborative optimization unit includes an online value network 1, a target value network 2, and their corresponding Adam optimizers and target evaluation networks. The value network comprises an input layer, hidden layers, and an output layer. The input layer merges the input network feature vectors, actions, and rewards into a feature vector representation, which is then activated using a ReLU function. The hidden layers, after obtaining the merged feature vectors, output the evaluation value for the current round through a fully connected layer. The single value network is then fitted to a global value function to evaluate the environmental state.

[0069] S3: Input the multi-scale feature vector into the policy network to generate scheduling actions, evaluate the value of the actions through the value network and feed it back to the policy network.

[0070] Furthermore, we first initialize the network parameters and then randomly sample the environmental conditions. Reward r t i ,action and the environmental state at the next moment Input the target value network and perform a time difference algorithm to obtain the target value y. t :

[0071]

[0072] Where, r t Q represents the instant reward at each time step, γ is the discount factor, and Q is the instant reward at each time step. i Action value function prediction, w Q The target value network weight parameters;

[0073] Set the loss function L for the online value network:

[0074]

[0075] in, For batch sample size, For the target value of sample k, Here, d represents the online network parameters, and d represents the action space.

[0076] Update the online network parameters w using the Adam optimizer by minimizing the loss function. Q1i ,w Q2i

[0077] Action value function prediction value Choose the smaller of the two online value networks as the value judgment for the current value network:

[0078]

[0079] The current value judgment is used as input to the online policy network for processing, and simultaneously based on... And the Adam optimizer smoothly updates the target value network parameters

[0080]

[0081] Where τ is the soft update coefficient, The target value network parameters are those of the previous time step.

[0082] In this embodiment, the training process is stabilized through a dual-value network architecture (online network + target network), and the minimum value of the two outputs is selected as the final action value to suppress the risk of overestimation.

[0083] In an optional embodiment, the value network update mechanism can also be implemented through average value network integration. Specifically, three independent online value networks are constructed, sharing the same input but with different initializations. The action value is taken as the average of the outputs of the three networks. Each network updates its parameters asynchronously with different data batches, and the average of the online network parameters is periodically assigned to the target network.

[0084] In another optional embodiment, the value network update mechanism can also be implemented through a sliding target value constraint. Specifically, after calculating the target value, the deviation between it and the current online value is limited to no more than a threshold δ. Initially, a larger deviation is allowed (δ is larger), and it is gradually tightened to a stable value later. If the target value exceeds the range of the current value ± δ, it is truncated to the boundary value.

[0085] S4: Iteratively optimize the policy network and value network until the target reward converges to the global optimum.

[0086] Furthermore, the policy network generates scheduling decisions. The entire policy network includes an online policy network, a target policy network, and an Adam optimizer. The online policy network includes hidden layers and an output layer. The hidden layers process the input feature vectors into different representation domains and connect them using a fully connected approach with ReLU as the activation function. The output layer outputs the actions, states, and policies. Figure 3 A schematic diagram of the gradient update for the policy-value network.

[0087] The gradient of the target policy network stochastic policy of the coordinated optimization unit i is calculated using the fractional function estimator.

[0088]

[0089] in, Let θ' be the gradient of the objective function of the target policy network parameters. To distribute π from the empirical buffer d Expectation of sampling, π θ’ For the target strategy, π θj For online strategies, Let θ be the multi-agent advantage function, and θ be the current parameters of the online policy network.

[0090] Update the online policy network parameters θ based on the objective function gradient and the Adam optimizer. i Simultaneously, the online policy network is updated to achieve the optimal global reward solution after the gradient stabilizes.

[0091]

[0092] in, The gradient of the objective function for the updated target policy network parameters θ';

[0093] Set the objective function of the target policy network as follows:

[0094]

[0095] in It is by The integral is obtained. The modified objective function is defined by δ, which is the KL divergence constraint coefficient. KL divergence is added as a constraint term to maintain the policy's variation range and control stability. Subsequently, the target network parameters are smoothly updated based on the objective function and the Adam optimizer.

[0096] The update process is as follows:

[0097] in, The online network parameters from the previous time step. These are the target network parameters for the previous time step;

[0098] Through continuous gradient updates, the system gradually converges to the optimal strategy to achieve intelligent traffic scheduling, while eliminating the risk of strategy collapse, ensuring long-term optimization stability, optimizing bandwidth utilization, and balancing network load.

[0099] Example 2 is the second embodiment of the present invention, which differs from the previous embodiment in that:

[0100] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0101] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a processor-included system, or other system that can fetch and execute instructions from, an instruction execution system, apparatus, or device). For the purposes of this specification, "computer-readable medium" can be any means that can contain, store, communicate, propagate, or transmit programs for use by, or in conjunction with, an instruction execution system, apparatus, or device.

[0102] More specific examples of computer-readable media (a non-exhaustive list) include: electrical connections (electronic devices) having one or more wires, portable computer disk drives (magnetic devices), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic devices, and portable optical disc read-only memory (CDROM). Furthermore, computer-readable media can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, interpreting, or otherwise processing as necessary, and then stored in computer memory.

[0103] It should be understood that various parts of the present invention can be implemented in hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented in software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.

[0104] Example 3, an embodiment of the present invention, provides an intelligent traffic scheduling system based on attention mechanism and deep reinforcement learning, including: a feature extraction module, a decision generation module, a value evaluation module, and an optimization control module;

[0105] The feature extraction module integrates local and global attention mechanisms to extract multi-scale feature vectors from traffic data;

[0106] The decision generation module generates a probability distribution of scheduling actions based on feature vectors and then executes the decision.

[0107] The value assessment module evaluates the value of actions through a dual value network and updates the target network with the minimum value.

[0108] The control module is optimized, and the iterative optimization of the control policy network and value network drives the reward to converge to the global optimum.

[0109] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.

Claims

1. An intelligent traffic scheduling method based on attention mechanism and deep reinforcement learning, characterized by: include, Multi-scale feature vectors of network traffic are extracted by fusing local attention mechanisms and global attention mechanisms; Construct a centrally coordinated, distributed decision-making system, which includes collaborative optimization units composed of a strategy network and a value network; The multi-scale feature vector is input into the policy network to generate scheduling actions, and the value of the actions is evaluated by the value network and fed back to the policy network. Iteratively optimize the policy network and value network until the target reward converges to the global optimum.

2. The intelligent traffic scheduling method based on attention mechanism and deep reinforcement learning as described in claim 1, characterized in that: The local attention mechanism captures fine-grained flow characteristics; The global attention mechanism captures the overall network traffic pattern, and the two are spliced ​​and fused to form the multi-scale feature vector.

3. The intelligent traffic scheduling method based on attention mechanism and deep reinforcement learning as described in claim 2, characterized in that: The value network includes calculating a time difference target value based on experience buffer data, outputting value function prediction values ​​through dual online value networks, and selecting the minimum value of the dual online value network outputs to update the target value network parameters.

4. The intelligent traffic scheduling method based on attention mechanism and deep reinforcement learning as described in claim 3, characterized in that: The policy network includes an online policy network and a target policy network. The parameters of the target policy network are soft-updated by importance sampling, and the online policy network is updated using stochastic policy gradient.

5. The intelligent traffic scheduling method based on attention mechanism and deep reinforcement learning as described in claim 4, characterized in that: The local attention mechanism includes dividing the traffic data into non-overlapping windows; Calculate the request tensor, identifier tensor, and value tensor for each window; After performing pooling operations on the request tensor and the identifier tensor, the similarity score matrix is ​​calculated. The normalized similarity score matrix is ​​then weighted and a weighted tensor is used to generate local attention features.

6. The intelligent traffic scheduling method based on attention mechanism and deep reinforcement learning as described in claim 5, characterized in that: The global attention mechanism includes selecting the top i maximum values ​​from the local attention similarity matrix to generate a correlation matrix; Extract the identifier sub-tensor and value sub-tensor corresponding to the correlation matrix; Global attention features are generated by combining the request tensor and subtensors.

7. The intelligent traffic scheduling method based on attention mechanism and deep reinforcement learning as described in claim 6, characterized in that: Before extracting the multi-scale feature vector of network traffic, a traffic distribution matrix dataset and a link matrix D are constructed within a time period t. In this matrix, the rows and columns represent time points, sensor data, and network data, respectively. Based on the time points, each sensor data point and network environment data point is used as a parameter to fill the corresponding column of the matrix. Data is collected once per second, and a traffic distribution matrix is ​​constructed over three seconds. The traffic distribution matrix dataset is obtained within a time period t. The rows and columns of the link matrix D represent the numbers of the sending and receiving nodes, respectively. If the nodes are connected, the element parameter in D is 1; otherwise, it is 0.

8. A system for intelligent traffic scheduling based on attention mechanisms and deep reinforcement learning, employing the intelligent traffic scheduling method based on attention mechanisms and deep reinforcement learning as described in any one of claims 1 to 7, characterized in that, include: Feature extraction module, decision generation module, value assessment module, and optimization control module; The feature extraction module integrates local and global attention mechanisms to extract multi-scale feature vectors from traffic data. The decision generation module generates a probability distribution of scheduling actions based on feature vectors and then executes the decision. The value assessment module evaluates the value of an action through a dual-value network and updates the target network with the minimum value. The optimization control module controls the iterative optimization of the policy network and the value network, driving the reward to converge to the global optimum.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the intelligent traffic scheduling method based on attention mechanism and deep reinforcement learning as described in any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the intelligent traffic scheduling method based on attention mechanism and deep reinforcement learning as described in any one of claims 1 to 7.

Citation Information

Cited By

  • Novel flexible manufacturing system dynamic scheduling method based on multi-agent reinforcement learning

    CN121764010A