Network traffic scheduling optimization method based on deep learning

By constructing a network traffic scheduling method that integrates the retrieval enhancement diffusion model and a hybrid linear expert model, combined with reinforcement learning and multi-objective optimization, the problems of low prediction accuracy and rigid strategy of traditional scheduling methods in complex network environments are solved, and efficient and adaptive network resource management is achieved.

CN120281665AActive Publication Date: 2025-07-08北京领雾科技有限公司

Patent Information

Application Number
CN202510759008.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-09
Publication Date
2025-07-08
Estimated Expiration
2045-06-09

AI Technical Summary

Technical Problem

Traditional network traffic scheduling methods are difficult to meet the needs of dynamic, efficient and intelligent management. Especially in network environments where multiple business types are concurrent, the existing deep learning-based traffic prediction methods have weak generalization capabilities, slow response to emergencies, rigid scheduling strategies and difficult to update in time according to changes in the network environment.

Method used

Build a traffic prediction method of a fusion retrieval enhancement diffusion model and a hybrid linear expert model, combine reinforcement learning and multi-objective optimization mechanisms, and realize dynamic scheduling strategy optimization and resource management through real-time data acquisition, multiple rounds of online reinforcement learning and reward bonus item design.

Benefits of technology

It improves traffic prediction accuracy and adaptive capabilities of scheduling strategies, improves network resource utilization and service quality, reduces latency and energy consumption, and adapts to complex dynamic network environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120281665A_ABST
    Figure CN120281665A_ABST
Patent Text Reader

Abstract

The invention provides a network traffic scheduling optimization method based on deep learning. The method is applied to the technical field of communication networks, and comprises the following steps: S1, collecting flow data, link state, delay and packet loss rate indexes deployed at network nodes in real time, and constructing a sample database; s2, according to the sample database, constructing a traffic prediction model fusing a retrieval enhancement diffusion model and a mixed linear expert model to perform multi-modal traffic prediction; s3, constructing a state space of a reinforcement learning agent according to a real-time network state and the future traffic prediction result; s4, performing strategy optimization based on a multi-objective optimization mechanism; and S5, deploying the trained and optimized model to a network controller or an edge computing node to realize real-time sensing and dynamic scheduling control of network resources. According to the method, the traffic prediction precision is remarkably improved, the scheduling strategy is updated and optimized in time according to the network environment change, and the comprehensive performance of the network is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of communication networks, and particularly to a method for optimizing network traffic scheduling based on deep learning. Background Art

[0002] With the continuous growth of network scale and complexity, traditional traffic scheduling methods are gradually difficult to meet the dynamic, efficient, and intelligent management requirements. Especially in a network environment where multiple service types (such as voice, video, big data synchronization, etc.) coexist, how to achieve efficient and refined scheduling according to real-time traffic status, network topology, and service requirements has become a key problem in network optimization.

[0003] In recent years, deep learning has made remarkable progress in fields such as time series modeling and decision optimization, showing great potential in network scheduling. However, existing deep learning-based traffic prediction methods generally have problems such as weak generalization ability and slow response to emergencies, and cannot fully adapt to the complex situations in real networks; at the same time, scheduling strategies often rely on static training models and are difficult to update and optimize in a timely manner according to network environment changes. In addition, in multi-objective optimization scenarios, there are conflicts between objectives such as delay, bandwidth utilization, and energy consumption, and a strategy for coordinating various objectives needs to be designed to improve the comprehensive performance.

[0004] Therefore, there is an urgent need for a deep learning method with intelligent perception, accurate prediction, and adaptive scheduling capabilities to address the problem of optimizing traffic scheduling in complex network environments. Summary of the Invention

[0005] To solve the above problems, the present invention proposes a method for optimizing network traffic scheduling based on deep learning to solve technical problems such as low prediction accuracy, rigid strategies, and resource waste in traditional network traffic scheduling.

[0006] To achieve the above object, it is realized through the following technical solutions: The present invention provides a method for optimizing network traffic scheduling based on deep learning, including the following steps: S1. Real-time collect traffic data, link status, delay, and packet loss rate metrics deployed on network nodes, and construct a sample database; S2. According to the sample database, construct a traffic prediction model that combines a retrieval enhanced diffusion model and a hybrid linear expert model for multi-modal traffic prediction, and output a future traffic prediction result through weighted fusion; S3. Construct the state space of a reinforcement learning agent according to the real-time network state and the future traffic prediction result, train the model using a multi-round online reinforcement learning mechanism, optimize the scheduling strategy in stages, guide the strategy improvement through a reward bonus item, and dynamically adjust the action space according to the state space; S4. Optimize the strategy based on the multi-objective optimization mechanism. Combine the weighted linear normalization method and the Pareto front approximation algorithm to screen out the Pareto front solution set that meets the multi-objective trade-off from the candidate solutions formed by the action space for actual deployment selection; S5. Deploy the trained and optimized model to the network controller or edge computing node to achieve real-time perception and dynamic scheduling control of network resources, and update the parameters of the traffic prediction model online according to the scheduling effect data.

[0007] Further, among them, the construction of the retrieval enhanced diffusion model includes: Intercept the sample time series in the sample database into fixed-length segments according to a unified time window to form a candidate sample set D; Map the currently input sample time series into a first embedding vector through a pre-trained encoder; For each sample in the sample database, intercept its first π time steps and generate a second embedding vector through the same encoder; Calculate the similarity between the first embedding vector and all the second embedding vectors, and select the k sample indices with the smallest similarity to form a retrieval sample set; Construct a diffusion probability model framework, and encode and guide denoising for the retrieval sample set through a mechanism that destroys the data distribution through forward diffusion and reconstructs the target sequence through reverse diffusion to generate a denoised prediction result.

[0008] Further, among them, the construction of a diffusion probability model framework, encoding and guiding denoising for the retrieval sample set through a mechanism that destroys the data distribution through forward diffusion and reconstructs the target sequence through reverse diffusion to generate a denoised prediction result includes: In the forward diffusion stage: Starting from the real traffic sequence x0, gradually add Gaussian noise through T time steps to generate a noise vector T containing the noise-added sequences x1, x2, …, x , which approximately follows the standard normal distribution N(0, I) and has the same dimension as the prediction target; In the reverse diffusion prediction stage: Map each reference sample in the retrieval sample set into an embedding vector through a pre-trained time series encoder, extract its deep time series features, and obtain the retrieval sample encoding features ; The diffusion state feature at the current diffusion time step t, the feature after encoding the input time series , and the retrieval sample encoding features And the current diffusion time step t is input into the reference modulation attention module (RMA), and the dynamic weights between features are calculated through the multi-head attention mechanism, dynamically adjusting the contribution ratio of the reference sample to the denoising process. Denoising is performed at each step during the reverse diffusion from t = T to t = 1 in reverse, gradually generating denoising predictions. After T steps of denoising, the first traffic prediction result is obtained. 。

[0009] Furthermore, the structure of the hybrid linear expert model includes: a multi-expert parallel prediction layer, a dynamic weight router, and a weighted fusion and post-processing layer. Among them, the multi-expert parallel prediction layer contains n independent linear expert models, and each expert is a linear regression model with different structures or parameters, respectively used to capture the long-term trend, periodicity, or sudden characteristics of the time series. The dynamic weight router includes an encoder for timestamp embedding encoding of the input time series and a hybrid layer composed of two fully connected networks; the hybrid layer is used to map the timestamp embedding encoding features into a channel-specific weight matrix. The weighted fusion and post-processing layer includes a channel weighted fusion unit and a post-processing layer; the channel weighted fusion unit is used to calculate the weights of the prediction results of each expert channel by channel according to the channel-specific weight matrix to obtain a fusion result; the post-processing layer is used to perform non-linear correction on the fusion result and output the final prediction result.

[0010] Furthermore, the construction of the hybrid linear expert model includes: Set n linear expert models , each expert independently predicts the input sample time series and outputs a prediction result with a length of p for the future time window ; where c is the number of channels for traffic monitoring. Based on the timestamp embedding features of the input sequence , m represents the dimension of the timestamp embedding features; through a hybrid layer composed of two MLPs perform dynamic weight routing to generate a channel-specific weight matrix , and the calculation formula is: ; weight the prediction results of each expert model by the channel-corresponding specific weights and sum them, and through the post-processing layer adjust the output trend and output the final second prediction result 。

[0011] Furthermore, through the post-processing layer Adjust the output trend and output the final second prediction result , including: ; among which, : defined as a channel-specific multiplication operation; : weighted-combine the outputs of each expert model to form a comprehensive prediction representation; : the post-processing layer includes residual connection and layer normalization operations, used to compensate for the prediction error of weighted fusion, for further fitting the error and adjusting the trend, : the final second prediction result; The timestamp embedding feature is encoded as: concatenate the normalized day of the week, hour number, and holiday flag into a vector respectively.

[0012] Furthermore, the weighted fusion formula of the traffic prediction model is:

[0013] where Z is the predicted future traffic prediction result; α, β are dynamically adjusted fusion coefficients, and satisfy α + β = 1.

[0014] Furthermore, the phased optimization scheduling strategy includes: Phase 1: Decoupling attempt initialization Policy generation; generate the action policy for the first scheduling attempt based on the initial state, including route selection and bandwidth allocation; improve the action policy for the second scheduling attempt according to the performance metrics feedback after the first scheduling attempt; Optimization objective: maximize the reward of the second scheduling attempt, and at the same time, constrain the policy distribution of the first scheduling attempt through the KL divergence regularization term to prevent the policy from deviating from the reasonable distribution of the basic model; Phase 2: Joint optimization of reward shaping Based on the optimization results of the first phase, enter the joint optimization of the second phase. Through the reward shaping mechanism, the optimization process is driven by the feedback difference between the two scheduling attempts to iteratively optimize the policy, encouraging the model to effectively improve the first policy in the second scheduling attempt.

[0015] Furthermore, the S4, perform policy optimization based on the multi-objective optimization mechanism, combine the weighted linear normalization method and the Pareto front approximation algorithm, and screen out the Pareto front solution set that meets the multi-objective trade-off from the candidate solutions for actual deployment selection, including: In the training stage, dynamically allocate the weight coefficients of each optimization objective according to the current network scenario requirements, and construct the joint objective function : ; among which, is the weight coefficient of the i-th index; is the normalized score for the q-th indicator; Generate a Pareto optimal solution set through a multi-objective optimization algorithm, which is used to provide an unbiased scheduling reference set during the deployment phase, ensuring that each scheduling strategy in the solution set cannot be further optimized on a certain objective without harming other objectives; According to the real-time network status, select the optimal strategy from the Pareto solution set. The priority rules are as follows: in high-load scenarios, prefer low-latency solutions; in energy-saving modes, prefer high-energy-efficiency solutions; in balanced modes, select the solution with the highest comprehensive score.

[0016] Furthermore, the online update mechanism in step S5 includes: (1) Store historical state-action-reward samples through an experience replay pool and periodically sample to fine-tune model parameters; (2) Use the PPO algorithm for policy gradient update and combine it with the target network to stabilize the training process; (3) Deploy multi-model redundancy and gradually update online model parameters through soft replacement.

[0017] Compared with the prior art, the present invention has the following beneficial effects: 1. By constructing a retrieval-enhanced diffusion model, combining the reference sample retrieval mechanism with the generation ability of the diffusion model, and introducing a reference-guided reverse diffusion process in the prediction stage, the present invention significantly enhances the modeling stability and generalization ability of the model for future traffic trends, improves the prediction accuracy, and is particularly suitable for network traffic scenarios with strong sudden changes.

[0018] 2. By constructing a hybrid linear expert model: designing a lightweight and scalable multi-expert system, using multiple parallel linear experts to capture long-term trend features, and improving the fitting effect on periodic and trend traffic through dynamic weight fusion, effectively complementing the deficiencies of non-linear models in long-term modeling.

[0019] 3. By constructing a multi-round online reinforcement learning method, aiming at problems such as high state-action space dimensions and feedback delays in the process of formulating scheduling strategies, introducing a multi-round online reinforcement learning mechanism, dynamically adjusting strategies based on prediction results, and realizing continuous optimization and rapid adaptation of strategies. At the same time, self-reinforcement is combined with historical scheduling feedback to improve strategy stability and long-term benefits.

[0020] 4. By constructing a design mechanism for reward bonus items, which plays an auxiliary signal role in the process of optimizing reinforcement learning strategies, guiding the learning process to focus on key decision points by modeling the relationship between intermediate states (such as sub-strategies or transitional traffic states) and optimal objectives, effectively alleviating the sparse reward problem, and improving the strategy convergence speed and performance.

[0021] It should be understood that the content described in the Summary of the Invention section is not intended to limit the key or important features of the embodiments of the present invention, nor to limit the scope of the present invention. Other features of the present invention will become easily understandable through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] In conjunction with the accompanying drawings and with reference to the following detailed description, the above and other features, advantages, and aspects of the embodiments of the present invention will become more apparent. The drawings are used to better understand the solution and do not constitute a limitation to the present invention. In the drawings, the same or similar reference numerals denote the same or similar elements, where: Figure 1 FIG. shows a schematic flow chart of a method for optimizing network traffic scheduling based on deep learning according to an embodiment of the present invention; Figure 2 FIG. shows a schematic architecture diagram of a method for optimizing network traffic scheduling based on deep learning according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0023] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.

[0024] In addition, the term "and / or" in this article is merely a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this article generally represents an "or" relationship between the associated objects before and after.

[0025] Figure 1 FIG. shows a schematic flow chart of a method for optimizing network traffic scheduling based on deep learning according to an embodiment of the present invention; Figure 2 FIG. shows a schematic architecture diagram of a method for optimizing network traffic scheduling based on deep learning according to an embodiment of the present invention. As Figure 1 and Figure 2 shown, a method 100 for optimizing network traffic scheduling based on deep learning includes the following steps: S1. Real-time collect traffic data, link status, delay, and packet loss rate metrics deployed on network nodes, and construct a sample database; In this step S1, first, key metrics such as traffic data, link status, latency, and packet loss rate are collected in real time through monitoring modules deployed on each node of the network. Then, for time series data sets with different characteristics, two database construction strategies are proposed. For data sets with insufficient scale and difficult to label a single category label, the entire training set is defined as the database; for data sets with complete category labels but category imbalance, a sample subset containing all categories is selected as the database.

[0026] S2. Based on the sample database, construct a traffic prediction model that combines a retrieval-enhanced diffusion model and a mixture of linear experts model for multimodal traffic prediction, and output the future traffic prediction result through weighted fusion; In this step S2, a traffic prediction model integrating two different modeling paradigms is constructed: a retrieval-enhanced diffusion model and a mixture of linear experts model, which are used to enhance the stability and multi-mode modeling ability of the model respectively. Moreover, the fusion of the mixture of linear experts model (lightweight plug-in architecture) and the retrieval-enhanced diffusion model (pre-trained encoder) can reduce the computational overhead and adapt to the resource limitations of edge devices; by performing weighted fusion on the prediction results of the two, the overall prediction performance and robustness are effectively improved.

[0027] This step S2 specifically includes the following steps: Step S2.1: Construct a retrieval-enhanced diffusion model To address the problems of large performance fluctuations and weak generalization ability faced by traditional time series prediction models in traffic prediction tasks, the present invention introduces a retrieval-enhanced diffusion model to model and predict future traffic. Based on the diffusion probability modeling framework, the retrieval-enhanced diffusion model introduces an auxiliary mechanism of reference samples, and realizes enhanced modeling of traffic / network flow time series features through two modules: embedding retrieval and reference-guided diffusion, effectively improving prediction stability and accuracy.

[0028] Furthermore, the construction of the retrieval-enhanced diffusion model includes: intercepting the sample time series in the sample database into fixed-length segments according to a unified time window to form a candidate sample set D; mapping the currently input sample time series into a first embedding vector through a pre-trained encoder; for each sample in the sample database, intercepting its first π time steps and generating a second embedding vector through the same encoder; calculating the similarity between the first embedding vector and all the second embedding vectors, selecting the k sample indices with the smallest similarity to form a retrieval sample set; constructing a diffusion probability model framework, and encoding and guiding denoising for the retrieval sample set through the mechanism of forward diffusion to destroy the data distribution and backward diffusion to reconstruct the target sequence, generating a denoised prediction result.

[0029] Step S2.1 specifically includes the following sub-steps: Step S2.1.1: Embedding Retrieval Mechanism The retrieval-enhanced diffusion model first passes through a sample database containing a large number of historical traffic time series in Step S1. Each time series segment is intercepted into a sequence of fixed length according to a unified time window as the candidate set for subsequent retrieval. Then, the input historical traffic time series X (such as the traffic in the past 24 hours or 7 days) is encoded. Specifically, each time component in the timestamp is encoded. Taking the encoding of the day of the week as an example, the encoder extracts the embedding vector of the current input sequence and matches it with the sequences in the database: ; where ; where, : The input historical traffic time series represents the observed values of network traffic data of c channels (such as multiple network interface, sensor channels) within the past s time steps (such as the number of observations in the past 24 hours, 7 days, etc.), serving as the main modeling object.; or : The embedding representation of the input historical traffic time series X after passing through the encoder (Embedding function), extracting its temporal semantic information. : The i-th historical sample sequence in the database. D represents the set of all stored samples, used as the retrieval reference. : The first π time steps of the i-th sample, used to calculate the embedding (i.e., only taking the initial part of the sample as the reference basis). or : The pre-trained time series encoder function, used to map the database sample sequence into the embedding space representation , facilitating similarity comparison. Sim(a,b) represents the similarity metric function between two embedding vectors, commonly used such as the negative Euclidean distance, the opposite of the cosine similarity, etc. (the smaller the value, the more similar). : Represents the set composed of the k reference time series samples retrieved from the database D that are most similar to the current input historical time series X. : Represents the set of sample indices selected with the smallest similarity (i.e., the closest) to select the l-th reference sample from the database.

[0030] The above Step S2.1.1 uses the pre-trained encoder to perform embedding mapping on the traffic sequence within the current time window, obtaining a low-dimensional feature vector ; The distance metric (such as the Euclidean distance) is performed between this vector and the embedding representation of each candidate sample in the database, and the k most similar reference samples are retrieved from it , for subsequent diffusion guidance.

[0031] Step S2.1.2: Construct a diffusion probability model framework Furthermore, in the construction of the diffusion probability model framework, the retrieval sample set is encoded and guided for denoising through a mechanism of forward diffusion to destroy the data distribution and reverse diffusion to reconstruct the target sequence, generating a denoised prediction result, including: In the forward diffusion stage: Starting from the real traffic sequence x0, Gaussian noise is gradually added through T diffusion time steps to generate a noise vector containing the noise sequences x1, x2, …, x T of the initial noise vector , which approximately follows the standard normal distribution N(0, I) and has the same dimension as the prediction target; To prevent the model from overly relying on reference samples, the retrieval enhanced diffusion model also introduces a tuning parameter in the reference guidance to limit the weight of the reference information and maintain the robustness and generalization of the prediction. Assuming there are T diffusion time steps in total, predictions are generated through the model at each step:

[0032] : The initial state of the diffusion process, also known as the "noise vector" or "reverse denoising starting point". It is the input at T diffusion time steps (i.e., the state furthest from the target prediction), usually a random variable with the same dimension as the target prediction sequence. T: The maximum number of time steps in the diffusion process, representing the number of denoising stages required from the most chaotic state to a clear prediction. : Represents is randomly sampled from a multi-dimensional standard normal distribution with a mean of 0 and a covariance of the identity matrix I. This distribution defines the starting point of the model in the uninformative state.

[0033] In the reverse diffusion prediction stage: Each reference sample in the retrieval sample set is mapped to an embedding vector through a pre-trained time series encoder, and its deep temporal features are extracted to obtain the retrieval sample encoding features ; The retrieval enhanced diffusion model adopts a diffusion model framework in the prediction stage, where: the forward diffusion stage is the same as the traditional diffusion model, that is, the target prediction traffic sequence is destroyed by gradually adding noise to construct a noise distribution; in the reverse diffusion process, the retrieval enhanced diffusion model uses the previously retrieved reference samples and guides the denoising process by designing a reference modulation attention module (RMA). Specifically, the RMA module fuses the features of the current prediction state, the features of the reference samples, and the side information, and dynamically adjusts the contribution weights of each feature through an adaptive modulation attention mechanism.

[0034] The diffusion state features at the current diffusion time step t and the features after encoding the input time series , the encoded features of the retrieved samples and the current diffusion time step t are input into the reference modulation attention module (RMA). The dynamic weights between features are calculated through the multi-head attention mechanism, and the contribution ratio of the reference samples to the denoising process is dynamically adjusted. Denoising is performed at each step during the reverse diffusion from t = T to t = 1, and the denoising prediction is gradually generated:

[0035]

[0036]

[0037] : The predicted mean of the reverse diffusion denoising function, which is a denoising function that fuses the current state, input sequence features, and reference sample features. : The diffusion state at the current diffusion time step t, which is the pre-noised predicted flow sequence and is an intermediate variable gradually denoised during the reverse diffusion process. RMA: Reference Modulation Attention Module, which is an attention mechanism or variant structure that dynamically generates a denoising prediction result that fuses context and external references. t: The current diffusion time step, which is used to control the diffusion or denoising progress, and the RMA module will use it to implement time modulation (such as time embedding). : The k reference time series samples retrieved from the database through the embedding retrieval mechanism and most similar to X, which are used to provide external "experience" guidance for the prediction. : Represents the embedded representation obtained after encoding the input time series X (usually through a pre-trained time series encoder), which is used to extract the sequence feature vector. : Represents the set of embeddings obtained after encoding the retrieved k reference samples, which enters the RMA module as external knowledge input and participates in the denoising guidance. : The covariance (or standard deviation) term corresponding to the current diffusion time step t, which is used to control the amplitude of the added noise. Usually, it is the output of a predefined time scheduling function, such as linear, cosine, or a learned noise scale table. : The noise vector sampled from the standard normal distribution, which is the core random variable in the diffusion model for generating diversity and simulating uncertainty.

[0038] After multiple rounds of reverse diffusion iterations, after T steps of denoising, the diffusion model finally outputs a denoised future flow sequence , which is the prediction result of the flow for a future period of time, that is, the first flow prediction result : ; The first flow prediction result , p is the length of the future time window predicted by the model, is the number of future time steps predicted, and represents the time dimension. For example: if the input historical traffic is the past 24 hours (s=24), and the traffic is predicted for the next 6 hours, then p=6; if the traffic is predicted for the next hour (sampled every minute), then p=60. Each time step corresponds to a fixed time interval (such as 1 minute, 1 hour, etc.), which is determined by the data collection frequency. c is the number of channels for traffic monitoring (such as multi-channel network interfaces, different sensors, etc.), representing different data sources, interfaces or paths in the network that need to be independently monitored or scheduled, representing the spatial dimension, for example: multi-channel network interfaces: such as multiple network cards (NICs) on a server, each network card corresponds to a channel; sensor nodes: different monitoring points distributed in the network (such as routers, switch ports); virtual channels: logically divided independent data transmission paths (such as VPN tunnels, QoS priority queues). The traffic data of each channel is collected independently, and the model needs to predict the future status of all channels at the same time. For example , indicating that the model predicts the traffic data of the three channels for the next 24 time steps (such as 24 hours).

[0039] According to the retrieval enhancement diffusion model constructed in step S2.1 above, through historical similar sample retrieval ( ) and the reference modulation attention (RMA) module can significantly improve the prediction accuracy of burst traffic and abnormal fluctuations; the multi-step denoising mechanism of the diffusion model combined with external reference guidance reduces the dependence on the distribution of training data and can enhance the generalization ability of the model in unknown network scenarios.

[0040] Step S2.2: Constructing a mixed linear expert model The hybrid linear expert model is a linear-centric multi-expert system, a lightweight, scalable plug-in architecture suitable for long-term traffic trend modeling. It uses multiple parallel linear expert models to learn time series features of different dimensions / periods, and fuses them through a dynamic weight router to finally output the predicted value.

[0041] Furthermore, the structure of the hybrid linear expert model includes: a multi-expert parallel prediction layer, a dynamic weight router, and a weighted fusion and post-processing layer; The multi-expert parallel prediction layer contains n independent linear expert models, each of which is a linear regression model with different structures or parameters, which is used to capture the long-term trend, periodicity or burst characteristics of the time series. The dynamic weight router includes an encoder for timestamp embedding encoding of the input time series and a mixing layer consisting of two layers of fully connected networks (MLPs); the mixing layer is used to map the timestamp embedding encoding features into a channel-specific weight matrix; Weighted fusion and post-processing layer, including a channel weighted fusion unit and a post-processing layer; the channel weighted fusion unit is used to calculate the weighted sum of the prediction results of each expert channel by channel according to the channel-specific weight matrix to obtain a fusion result; the post-processing layer is used to perform non-linear correction on the fusion result and output the final prediction result.

[0042] Further, the construction of the mixture of linear experts model includes: Given the input historical traffic time series , preprocess its start timestamp information, that is, encode each time component in the timestamps in the historical traffic time series X to obtain multi-dimensional timestamp embedding features , where m is the dimension of the timestamp embedding feature, then each expert model outputs: ; : X is the input historical traffic time series; c is the number of channels for traffic monitoring (such as multi-channel network traffic, traffic at different intersections, etc.); s is the historical time step, such as the past 24 hours or 7 days, etc. : The i'-th expert model, usually a linear model with different structures or parameters; set n linear expert models , each expert independently predicts the historical traffic time series X and outputs the traffic prediction results for a future time window of length p and c channels .

[0043] Based on the timestamp embedding features of the input historical traffic time series X , a mixture layer composed of two-layer MLP is used for dynamic weight routing to generate a channel-specific weight matrix , and the calculation formula is: ; where : A mixture layer (also called a router) composed of two-layer MLP, which is used to generate the weighting coefficients of each channel for each expert model; : W is the channel-specific weight matrix; n is the number of expert models.

[0044] Among them, the start timestamp information of the input historical traffic time series X is encoded to obtain multi-dimensional timestamp embedding features The encoding method is: for example, normalize the day of the week, hour, holiday flag, etc. and splice them into a vector. It usually includes a combination of the following time components: day of the week (1-dimensional, normalized encoding); time period (such as hours, normalized to a scalar of 0,1); holiday flag (1-dimensional, 0 / 1 indicates whether it is a holiday); season or month (optional, 1-dimensional, normalized encoding of 12 months). After each time component is converted into a numerical feature through encoding, it is concatenated into a vector Taking the encoding of the day of the week as an example, the formula is Among them, "index" represents the index of the day of the week, with Monday's index being 0 and Sunday's index being 6. Through this formula, Monday is encoded as -0.5, Tuesday is encoded as approximately -0.33, Wednesday is encoded as approximately -0.17, and so on, with Sunday being encoded as 0.5. This encoding method converts time information into numerical form, making it easier for the model to learn and utilize the periodic features in the time series.

[0045] The prediction results of each expert model The weighted sum is performed according to the specific weights corresponding to the channels, and the output trend is adjusted through the post-processing layer P(⋅). After weighted fusion, the final second flow prediction result is output : ;in, : Defined as channel-specific multiplication operation; : The outputs of each expert model are weighted and combined to form a comprehensive prediction representation; : Post-processing layer, used to further fit the error, adjust the trend, and improve the prediction accuracy (such as residual connection, regularization, etc.).

[0046] This method can dynamically adjust the contribution of different experts to the final prediction, has good interpretability and stability, and is particularly suitable for long-term sequence modeling tasks.

[0047] The hybrid linear expert model constructed in step S2.2 combines historical experience knowledge with current context features through the process of "first search, then guide, and finally predict", and shows stronger time series modeling capabilities and prediction stability in traffic sequence prediction tasks. Multiple experts capture long-term trends, periodicity, and burst characteristics in parallel, and the weighted fusion mechanism (dynamic router) adapts to different time scenarios (such as working days / holidays) to improve long-term prediction stability; the post-processing layer compensates for prediction errors through residual connections to further optimize the output quality.

[0048] Step S2.3: Weighted fusion of prediction results In order to take into account the prediction advantages of the two models, a simple and effective weighted fusion mechanism is proposed: The traffic prediction model obtains the future traffic prediction result Z through the following weighted fusion formula:

[0049] Among them, is the first traffic prediction result obtained through step S2.1; is the second traffic prediction result obtained through step S2.2; the weighting coefficients are α and β respectively, and α + β = 1.

[0050] Among them, α and β can be dynamically adjusted according to the performance on the validation set. For example, the optimal combination coefficients can be learned through indicators such as MSE / MAE, or an attention mechanism can be introduced for adaptive weighting.

[0051] S3. Construct the state space of the reinforcement learning agent according to the real-time network state and the future traffic prediction result, adopt a multi-round online reinforcement learning mechanism to train the model, optimize the scheduling strategy in stages, guide the policy improvement through the reward bonus item, and dynamically adjust the action space according to the state space; Based on the multi-round online reinforcement learning method, through two-stage training and reward shaping mechanism, solve the common distribution shift and behavior collapse problems in the network environment, enable the scheduling strategy to self-correct using the data generated by itself, and improve the robustness and generalization ability. This method does not rely on external feedback, and only through multiple attempts and optimizations of the model itself, improves the correctness of the final response. The agent responds according to the current network state, selects the optimal scheduling strategy to maximize the long-term benefits (such as throughput, load balancing, etc.). Through continuous interaction and learning, the system can adapt to the changes in the network environment and achieve self-optimization.

[0052] Furthermore, the phased optimization of the scheduling strategy in step S3 includes: Phase 1: Decoupled attempt initialization Input: The current network state (such as traffic load, link delay, topology information).

[0053] Policy generation; based on the initial state Generate the action policy for the first scheduling attempt ( ), including route selection and bandwidth allocation; improve the action policy for the second scheduling attempt according to the performance metric p1 feedback after the first scheduling attempt, and obtain the second scheduling attempt ( ); Optimization objective: Maximize the reward of the second scheduling attempt , and at the same time, constrain the policy distribution of the first scheduling attempt ( ) through the KL divergence regularization term to prevent the policy from deviating from the reasonable distribution of the basic model; During the above-mentioned stage-one decoupling attempt initialization process, first, the basic scheduling policy model is fine-tuned. Using a reinforcement learning framework, while optimizing the accuracy of the second scheduling attempt, a Kullback-Leibler divergence regularization term is introduced to constrain the behavior distribution of the first attempt to be close to the output distribution of the basic model, thereby avoiding the policy deviating from the original reasonable region and reducing the risk of behavior collapse. The optimization objective is as follows: ; where θ: the parameters of the current scheduling policy model; : the current network state, such as characteristics like user demand volume, link load, channel state, etc.; : the action policy output by the first scheduling attempt, such as resource block allocation, task migration, etc.; : the feedback information after the first attempt, such as performance metrics (such as real-time data like latency, packet loss rate, bandwidth utilization, energy consumption, etc.) or the reward signal, i.e., the immediate reward value of the first scheduling action (such as load balancing score, quality of service score) or the update of the network state after executing (such as link load change, node resource occupancy rate); : the improved policy output by the second scheduling attempt; : the probability distribution of the current model outputting the scheduling policy under the input state x; : the output distribution of the reference policy model (basic policy) in the state ; : the second scheduling result and the similarity score or performance score between the ideal scheduling target ; : the Kullback-Leibler divergence, used to measure the difference between two policy distributions; : the regularization coefficient of the Kullback-Leibler divergence, used to control whether the behavior output by the first attempt deviates from the reference model.

[0054] Stage Two: Joint Optimization of Reward Shaping Based on the optimization results of the first stage, enter the joint optimization of the second stage. Through the reward shaping mechanism, encourage the model to effectively improve the first policy in the second attempt and avoid the learning from falling into a non-self-correcting path. The original optimization objective is: ; where : the input state of the first scheduling attempt, usually the original network state (such as traffic load, link state, etc.). : the input state of the second scheduling attempt, usually the combined state after splicing the first attempt state and its partial feedback or response, helping the agent adjust the subsequent policy based on the effects of historical actions. : The input state of the i-th attempt , which is used to guide the generation of the scheduling strategy for each round. : The output of the action strategy for the first scheduling attempt, that is, the scheduling decision generated by the model based on the initial state . : The output of the action strategy for the second scheduling attempt, which is a new strategy after self-correction considering the state feedback information of the first attempt . : The strategy output of the i-th attempt and the performance matching score between the ideal target strategy , which can be comprehensively defined based on indicators such as throughput, load balancing, and energy consumption. : The probability distribution of the strategy generated by the current strategy model in the state . : The probability distribution of the strategy generated by the reference strategy (such as the base model or the previous strategy) in the state , which is used as the reference target for the Kullback-Leibler divergence. : The Kullback-Leibler divergence between the current strategy distribution and the reference strategy distribution in the i-th attempt, which is used to prevent the strategy from deviating too far. : Controls the weight of the Kullback-Leibler regularization term in the overall loss. The larger the value, the more emphasis is placed on the strategy behavior not deviating from the reference strategy to avoid "policy collapse" or overfitting.

[0055] To further encourage the "progress" of the second attempt in performance, a reward bonus term is introduced on the basis of the original reward, which is used to quantify the performance difference between the two scheduling attempts and is defined as: ; : The reward bonus, which is used to measure the performance improvement degree of the second scheduling compared to the first. : The reward amplification factor , which is used to enhance the incentive intensity of the model's self-correction behavior. , : Respectively represent the scores of the second and the first scheduling results.

[0056] In step S3 above, in phase one (decoupling attempt initialization), for the first scheduling: Based on the initial state x1x1, a policy y1y1 is generated, and after execution, feedback p1 is obtained. For the second scheduling: The input state becomes x2 = [x1, p1]x2 = [x1, p1], and an improved policy y2y2 is generated. Optimization objective: Maximize the reward of y2y2 while constraining the behavior distribution of y1y1 (through the KL divergence regularization term).

[0057] The traffic prediction result Z in step S2 is the core input feature for the reinforcement learning scheduling decision, directly affecting policy generation: The prediction result serves as the state representation, that is, the future traffic trend (Z) to be predicted needs to be included in the state space (x1, x2) of the reinforcement learning agent. For example: x1 = [real-time network state, Z]. The agent generates scheduling actions (such as routing adjustment, bandwidth allocation) based on this state.

[0058] Dynamically adjust the action space: If the prediction shows that the future traffic will surge (such as a significant increase in Z), the scheduling policy may expand the link or migrate tasks in advance to avoid congestion. For example, if the prediction accuracy in step S2 is insufficient (such as failure to capture bursty traffic), the policy in step S3 may generate suboptimal scheduling actions (such as insufficient resource allocation) due to "misjudging" the network state. Conversely, high-precision prediction can provide a reliable basis for policy optimization and improve scheduling efficiency.

[0059] Step S3 generates multiple possible scheduling actions (such as y1, y2) through multiple rounds of attempts as candidate solutions for multi-objective optimization.

[0060] S4. Based on the multi-objective optimization mechanism, perform policy optimization. Combine the weighted linear normalization method and the Pareto front approximation algorithm to screen out the Pareto front solution set that meets multi-objective trade-offs from the candidate solutions formed by the action space for actual deployment selection; The network scheduling problem is essentially a complex decision problem with multiple performance constraints and optimization objectives. The involved metrics include but are not limited to communication delay, link bandwidth utilization, system throughput, energy consumption, load balancing, and quality of service, etc. In order to achieve a reasonable balance among multiple performance objectives, the present invention introduces a multi-objective optimization mechanism, models the optimization of the scheduling policy as a multi-objective problem, and constructs a joint objective function containing multiple loss terms.

[0061] Specifically, this system adopts a strategy that combines the weighted linear normalization method with the Pareto front approximation algorithm, allowing the system to dynamically balance multiple objectives such as latency, energy consumption, and throughput: The weighted method is used to dynamically adjust the weight coefficients of various performance indicators according to scenario preferences during the training phase to achieve flexible control of different objectives; while the Pareto optimal solution set is used to provide an unbiased scheduling reference set during the deployment phase for the agent to select the most suitable scheduling strategy for the current network state from it. This step S4 screens out the Pareto front solution set that meets the multi-objective trade-off from the candidate solutions for actual deployment selection.

[0062] Furthermore, a normalization scoring mechanism is introduced during the scheduling strategy evaluation process to unify indicators with different dimensions (such as ms-level latency and kW-level energy consumption) to ensure the fairness and convergence stability of the optimization process. This mechanism allows the system to switch the optimization focus according to actual needs when facing scenarios with high reliability and low latency (such as the Internet of Vehicles) and scenarios with high energy efficiency requirements (such as low-power edge devices), thereby improving the system's adaptability and service elasticity. By eliminating the differences in indicator dimensions through the normalization scoring mechanism, the fairness of optimization is ensured. For example, in a high-load scenario, the low-latency strategy is automatically prioritized; in the case of low power consumption requirements, the energy efficiency optimal strategy is prioritized.

[0063] Furthermore, in this step S4 during the training phase, the weight coefficients of each optimization objective are dynamically allocated according to the current network scenario requirements to construct a joint objective function : ; where is the weight coefficient of the i-th indicator; is the normalized score of the q-th indicator; The Pareto optimal solution set is generated through a multi-objective optimization algorithm to provide an unbiased scheduling reference set during the deployment phase, ensuring that each scheduling strategy in the solution set cannot be further optimized in a certain objective without sacrificing other objectives; According to the real-time network state, the optimal strategy is selected from the Pareto solution set, and the priority rules include: in a high-load scenario, the low-latency solution is preferentially selected; in the energy-saving mode, the high-energy-efficiency solution is preferentially selected; in the balanced mode, the solution with the highest comprehensive score is selected.

[0064] S5. Deploy the trained and optimized model to the network controller or edge computing node to achieve real-time perception and dynamic scheduling control of network resources, and online update the parameters of the traffic prediction model according to the scheduling effect data.

[0065] The trained reinforcement learning scheduling model is deployed in a network controller (such as an SDN controller) or an edge computing node to achieve real-time perception and dynamic scheduling control of network resources. In actual deployment, the model continuously perceives the current network state (including link load, node status, traffic pattern, etc.), automatically generates scheduling decisions, and realizes intelligent scheduling response with millisecond-level latency.

[0066] The scheduling effect data (such as real latency, energy consumption) collected in actual deployment can be fed back to step S2 for updating the prediction model (such as online fine-tuning of diffusion model parameters); meanwhile, the scheduling feedback data can also be used to correct the multi-objective weights, forming a closed-loop optimization of "prediction → decision → feedback → re-prediction".

[0067] In some embodiments, to cope with the continuous changes in the network environment, such as dynamic business requirements, fluctuating link states, node online / offline, etc., the present invention supports the online learning and incremental update mechanism of the model. The specific methods are as follows: introduce an experience replay mechanism and a policy soft update mechanism, extract samples from the new experiences accumulated during operation, and periodically fine-tune the model parameters to maintain the long-term adaptability of the model. Use reinforcement learning algorithms (such as DQN, PPO, or Actor-Critic, etc.) to perform lightweight iterative updates on edge nodes to ensure the self-evolution ability under low computing resources. On the premise of ensuring uninterrupted service, achieve a smooth transition between old and new policies through a model hot replacement mechanism or a multi-model redundancy switching mechanism, effectively preventing model degradation or catastrophic forgetting.

[0068] In summary, the above embodiments of the present invention provide a network traffic scheduling optimization method based on deep learning, which solves problems such as low prediction accuracy, rigid policies, and resource waste in traditional network traffic scheduling through a prediction-decision-optimization closed-loop system, and realizes high-precision prediction, adaptive decision-making, low-energy consumption operation, and high-reliability service in a complex dynamic environment. It is applicable to scenarios such as 5G, Internet of Things, and cloud computing, significantly improving network service quality and operation economy.

[0069] Furthermore, an embodiment of the present application also provides an electronic device, including: a processor, a memory, and a system bus; the processor and the memory are connected through the system bus; the memory is used to store one or more programs, and the one or more programs include instructions, and when the instructions are executed by the processor, the processor executes any of the above methods.

[0070] Furthermore, an embodiment of the present application also provides a computer program product, which, when running on a terminal device, causes the terminal device to execute any of the above methods.

[0071] As can be seen from the description of the above embodiments, those skilled in the art can clearly understand that all or part of the steps in the above method embodiments can be implemented by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art can be embodied in the form of a software product, and this computer software product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., including several instructions for causing a computer device (which can be a personal computer, a server, or a network communication device such as a media gateway, etc.) to execute the methods described in each embodiment or some parts of the embodiments of the present application.

[0072] It should be noted that the various embodiments in this specification are described in a progressive manner, and the key point of each embodiment is to illustrate the differences from other embodiments. The same or similar parts among the various embodiments can be referred to each other. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the description of the method part.

[0073] It should also be noted that in the embodiments of the present application, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, the element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article or device including the said element.

[0074] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present application. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined in the embodiments of the present application can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to these embodiments shown in the embodiments of the present application, but will conform to the widest scope consistent with the principles and novel features disclosed in the embodiments of the present application.

Claims

1. A method for optimizing network traffic scheduling based on deep learning, characterized in that, It includes the following steps: S1. Collect traffic data, link status, latency, and packet loss rate metrics deployed on network nodes in real time, and construct a sample database; S2. According to the sample database, construct a traffic prediction model that combines a retrieval-enhanced diffusion model and a hybrid linear expert model for multimodal traffic prediction, and output the future traffic prediction result through weighted fusion; S3. Construct the state space of the reinforcement learning agent according to the real-time network state and the future traffic prediction result, train the model using a multi-round online reinforcement learning mechanism, optimize the scheduling strategy in stages, guide the policy improvement through the reward bonus item, and dynamically adjust the action space according to the state space; S4. Optimize the policy based on the multi-objective optimization mechanism, combine the weighted linear normalization method and the Pareto front approximation algorithm, and screen out the Pareto front solution set that meets the multi-objective trade-off from the candidate solutions formed by the action space for actual deployment selection; S5. Deploy the trained and optimized model to the network controller or edge computing node to achieve real-time perception and dynamic scheduling control of network resources, and update the parameters of the traffic prediction model online according to the scheduling effect data.

2. The network traffic scheduling optimization method based on deep learning according to claim 1, wherein, Among them, The construction of the retrieval-enhanced diffusion model includes: Intercept the sample time series in the sample database into fixed-length segments according to a unified time window to form a candidate sample set D; Map the currently input sample time series to a first embedding vector through a pre-trained encoder; For each sample in the sample database, intercept its first π time steps and generate a second embedding vector through the same encoder; Calculate the similarity between the first embedding vector and all the second embedding vectors, and select the k sample indices with the smallest similarity to form a retrieval sample set; Construct a framework based on the diffusion probability model, and encode and guide denoising for the retrieval sample set through the mechanism of forward diffusion to destroy the data distribution and backward diffusion to reconstruct the target sequence, and generate a denoised prediction result.

3. The method for optimizing network traffic scheduling based on deep learning according to claim 2, wherein Among them, The construction of a framework based on the diffusion probability model, encoding and guiding denoising for the retrieval sample set through the mechanism of forward diffusion to destroy the data distribution and backward diffusion to reconstruct the target sequence, and generating a denoised prediction result includes: In the forward diffusion stage: Starting from the real traffic sequence x0, Gaussian noise is gradually added through T time steps to generate a noisy sequence x1, x2, …, x T of the initial noise vector , approximately follows the standard normal distribution N(0, I) and has the same dimension as the prediction target; In the reverse diffusion prediction stage: Each reference sample in the retrieval sample set is mapped to an embedding vector through a pre-trained time series encoder, and its deep time series features are extracted to obtain the retrieval sample coding features ; The diffusion state features at the current diffusion time step t , the features after encoding the input time series , the retrieval sample encoding features and the current diffusion time step t are input into the reference modulation attention module (RMA). The dynamic weights between the features are calculated through the multi-head attention mechanism, and the contribution ratio of the reference sample to the denoising process is dynamically adjusted. Perform denoising at each step in the reverse diffusion from t = T to t = 1, and gradually generate a denoised prediction; After T steps of denoising, the first traffic prediction result is obtained .

4. The optimization method for network traffic scheduling based on deep learning according to claim 3, characterized in that, Among them, The structure of the hybrid linear expert model includes: a multi-expert parallel prediction layer, a dynamic weight router, and a weighted fusion and post-processing layer; Among them, the multi-expert parallel prediction layer contains n independent linear expert models, and each expert is a linear regression model with different structures or parameters, which are respectively used to capture the long-term trend, periodicity, or sudden characteristics of the time series; The dynamic weight router includes an encoder for timestamp embedding encoding of the input time series and a hybrid layer composed of two fully connected networks; the hybrid layer is used to map the timestamp embedding encoding features to a channel-specific weight matrix; The weighted fusion and post-processing layer includes a channel weighted fusion unit and a post-processing layer; the channel weighted fusion unit is used to calculate the weighted sum of each expert prediction result channel by channel according to the channel-specific weight matrix to obtain a fusion result; the post-processing layer is used to perform non-linear correction on the fusion result and output the final prediction result.

5. The network traffic scheduling optimization method based on deep learning according to claim 4, characterized in that Among them, The construction of the hybrid linear expert model includes: Set up n linear expert models , each expert independently predicts the input sample time series and outputs the prediction result with a future time window length of p ; where c is the number of channels for traffic monitoring; Timestamp embedding features based on the input sequence , where m represents the dimension of the timestamp embedding features; a hybrid layer composed of two layers of MLP performs dynamic weight routing to generate a channel-specific weight matrix , and the calculation formula is as follows: ; Weight the prediction results of each expert model by the channel-corresponding specific weights and adjust the output trend through the post-processing layer to output the second traffic prediction result .

6. The method for optimizing network traffic scheduling based on deep learning according to claim 5, wherein The post-processing layer adjusts the output trend and outputs a second traffic prediction result , including: ; wherein, : is defined as a channel-specific multiplication operation; : combines the outputs of each expert model by weighting to form a comprehensive prediction representation; : the post-processing layer contains residual connections and layer normalization operations, which are used to compensate for the prediction errors of weighted fusion, further fit the errors, and adjust the trend, : the second traffic prediction result; The timestamp embedding feature is encoded as follows: The day of the week, the hour number, and the holiday flag are normalized respectively and then concatenated into a vector.

7. The method according to claim 1, wherein The weighted fusion formula of the traffic prediction model is: ; where Z is the predicted future traffic prediction result; α and β are dynamically adjusted fusion coefficients, and satisfy α + β = 1.

8. The method for optimizing network traffic scheduling based on deep learning according to claim 1, characterized in that The phased optimization scheduling strategy includes: Phase 1: Decoupling attempt initialization Policy generation; generating an action policy for the first scheduling attempt based on the initial state, including route selection and bandwidth allocation; improving the action policy for the second scheduling attempt according to the performance metrics feedback after the first scheduling attempt; Optimization objective: maximizing the reward of the second scheduling attempt, while constraining the policy distribution of the first scheduling attempt through the KL divergence regularization term to prevent the policy from deviating from the reasonable distribution of the base model; Phase 2: Joint optimization of reward shaping Based on the optimization results of the first phase, enter the joint optimization of the second phase. Through the reward shaping mechanism, the optimization process is driven by the feedback difference between the two scheduling attempts to iteratively optimize the policy, encouraging the model to effectively improve the first policy in the second scheduling attempt.

9. The method for optimizing network traffic scheduling based on deep learning according to claim 1, wherein In step S4, based on the multi-objective optimization mechanism for policy optimization, combining the weighted linear normalization method and the Pareto front approximation algorithm, screening out the Pareto front solution set that meets the multi-objective trade-off from the candidate solutions for actual deployment selection, including: During the training phase, the weight coefficients of each optimization objective are dynamically allocated according to the current network scenario requirements to construct a joint objective function : ; among them, is the weight coefficient of the i-th index; is the normalized score of the q-th index; Generating a Pareto optimal solution set through a multi-objective optimization algorithm to provide an unbiased scheduling reference set during the deployment phase, ensuring that each scheduling policy in the solution set cannot be further optimized in a certain objective without compromising other objectives; According to the real-time network status, selecting the optimal policy from the Pareto solution set, and the priority rules include: preferentially selecting low-latency solutions in high-load scenarios; preferentially selecting energy-efficient solutions in energy-saving modes; selecting the solution with the highest comprehensive score in the balanced mode.

10. The method for optimizing network traffic scheduling based on deep learning according to claim 1, wherein The online update mechanism in step S5 includes: (1) Storing historical state-action-reward samples through an experience replay pool and periodically sampling to fine-tune the model parameters; (2) Using the PPO algorithm for policy gradient update and combining the target network to stabilize the training process; (3) Deploying multi-model redundancy and gradually updating the online model parameters through soft replacement.

Citation Information

Patent Citations

  • Private domain traffic peak identification and route scheduling method and system based on time sequence prediction

    CN119449630A

  • Database adaptive data flow acquisition optimization method and system based on reinforcement learning

    CN119719783A

  • Generative AI endogenous communication network architecture and hierarchical collaborative scheduling method

    CN119815557A

  • Network traffic optimization scheduling method based on deep reinforcement learning and suitable for periodic traffic characteristics

    CN120090989A

  • Method for intelligent traffic scheduling based on deep reinforcement learning

    US20230362095A1

Cited By

  • 5G network self-optimization method and device based on artificial intelligence

    CN120602965A

  • Traffic scheduling method and electronic equipment

    CN120896909A

  • Data stream real-time scheduling method based on intelligent agent

    CN121000684A

  • Flow scheduling method and device based on machine room multi-index real-time prediction

    CN121098800A