AI-driven real-time network optimization algorithm

Through the AI-driven real-time network optimization algorithm, using technologies such as generative adversarial networks, graph neural networks and deep reinforcement learning, the problems of low computing efficiency and unstable convergence in real-time network optimization are solved, and efficient and stable network resource scheduling is achieved.

CN120128494AActive Publication Date: 2025-06-10北京思普艾斯科技有限公司

Patent Information

Application Number
CN202510340113.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-21
Publication Date
2025-06-10
Estimated Expiration
2045-03-21

AI Technical Summary

Technical Problem

Traditional gradient optimization methods have low computational efficiency, strong hyperparameter dependence, and unstable convergence in real-time network optimization scenarios.

Method used

Using AI-driven real-time network optimization algorithm, the global resource allocation is dynamically optimized by generating adversarial networks, graph neural networks, deep reinforcement learning and multi-objective optimization methods, and combined with Lagrangian relaxation method and graph optimization technology, the global resource scheduling strategy is realized.

Benefits of technology

It improves computing efficiency, reduces dependence on hyperparameters, ensures the stability of the optimization process, and realizes fast and accurate regression analysis and network resource scheduling.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120128494A_ABST
    Figure CN120128494A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of network optimization, and discloses an AI-driven real-time network optimization algorithm, which comprises the following steps: data optimization: preprocessing network state data by using a generative adversarial network, extracting optimization features and generating enhanced data; topology modeling: modeling network topology by adopting a graph neural network, and updating node features; resource optimization: dynamically adjusting calculation, storage and bandwidth resources by using deep reinforcement learning; global scheduling: optimizing resource allocation based on a multi-objective optimization method; secondary optimization: combining a Lagrangian relaxation method and a graph optimization technology to further optimize a scheduling strategy; and adaptive learning: improving the adaptive capacity of the system in a dynamic environment through self-supervised learning and meta-learning. Through a matrix operation technology based on a regular equation and in cooperation with an efficient data preprocessing process, the effect of improving the solving speed of the regression model is achieved, the problem that calculation is slow on a big data set in a traditional method is solved, and rapid and accurate regression analysis is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of network optimization, specifically an AI-driven real-time network optimization algorithm. Background Art

[0002] Currently, the application of artificial intelligence (AI) technology in the field of network optimization is becoming increasingly widespread. Many optimization algorithms are used in key tasks such as traffic scheduling, resource allocation, and data transmission to improve the performance and stability of network systems. Among them, algorithms based on gradient optimization, such as gradient descent and its variants, are widely used in parameter adjustment and optimization calculations in complex network environments. These methods gradually optimize network performance by continuously iteratively adjusting parameters and have achieved good results in many application scenarios. At the same time, mathematical tools such as matrix calculation and regression analysis also play an important role in network optimization, helping the model better adapt to the changing network environment.

[0003] However, with the continuous expansion of network scale and the explosion of data traffic, traditional gradient optimization methods face certain challenges in terms of computational efficiency and stability. First of all, iterative optimization methods usually require multiple training rounds and have a large amount of calculation, which will affect the response speed in application scenarios with high real-time requirements. Secondly, such methods are sensitive to hyperparameters. For example, the choice of learning rate directly affects the convergence speed and optimization effect. If the parameters are set improperly, it will lead to oscillations or slow convergence in the training process. In addition, in complex data environments, gradient optimization methods are affected by gradient vanishing or gradient explosion, reducing the stability of the algorithm. Therefore, how to improve computational efficiency, reduce dependence on hyperparameters, and at the same time ensure the stability of the optimization process has become a direction that urgently needs to be broken through in current network optimization technology. Summary of the Invention

[0004] Aiming at the deficiencies of the prior art, the present invention provides an AI-driven real-time network optimization algorithm, which solves the problems of low computational efficiency, strong hyperparameter dependence, and unstable convergence of traditional optimization algorithms in real-time network optimization scenarios.

[0005] To achieve the above objectives, the present invention is realized through the following technical solutions: An AI-driven real-time network optimization algorithm includes the following steps: S1. Preprocess network state data through a generative adversarial network, extract optimized network feature data, and generate high-quality enhanced data; S2. Based on the enhanced data, use a graph neural network to model the network topology structure, capture the dependencies between nodes, and update node features through a graph convolutional network; S3. Based on node features, adopt a deep reinforcement learning algorithm to dynamically optimize and adjust computing, storage, and bandwidth resources, and adaptively allocate resources; S4. Input the resource allocation scheme into the multi-objective optimization method to further minimize the latency, packet loss rate, maximize the throughput, and optimize the energy efficiency, and finally output the global optimized resource scheduling strategy; S5. Combine the Lagrangian relaxation method and graph optimization technology to optimize the global optimized resource scheduling strategy again, or the final scheduling scheme; S6. The final scheduling scheme improves the adaptability of the network model in a dynamic environment through self-supervised learning and meta-learning, ensuring the efficient operation of the optimized system in different network states.

[0006] Preferably, in step S1, the generative adversarial network is trained using a conditional generative adversarial network, specifically including: The generator generates enhanced data based on the input of network states such as network traffic, node load, and bandwidth utilization; The discriminator evaluates the quality of the generated data and feeds back to optimize the generator to improve the authenticity and effectiveness of the data.

[0007] Preferably, in step S2, the graph neural network performs multi-layer feature propagation through a graph convolutional network to update node features and fuse adjacent node information to optimize the resource allocation strategy, specifically including: Use a K-layer graph convolutional network to update node features layer by layer, so that the computing power, bandwidth utilization, and storage resource information of each node are gradually propagated in the adjacency relationship, enhancing the perception of the global network state; Assign different weights to adjacent nodes through an attention mechanism to increase the accuracy of resource allocation decisions; Generate an optimized node embedding vector and use it as input for the reinforcement learning algorithm to improve the decision-making ability of resource allocation.

[0008] Preferably, in step S3, the deep reinforcement learning optimizes network resource allocation specifically including: The reward function is calculated based on network performance metrics such as latency, throughput, packet loss rate, and energy efficiency, and the weights are dynamically adjusted according to the current network state; Combine deep Q-network for offline training and proximal policy optimization for online optimization to improve the policy adjustment ability; The state input of the reinforcement learning includes node computing power, storage utilization, bandwidth load, and topology structure information, and the action space includes resource allocation policy adjustment, and uses the features generated by the graph neural network for policy decision-making.

[0009] Preferably, in step S4, the multi-objective optimization adopts multi-task learning, specifically including: Balance the optimization objectives of latency, throughput, packet loss rate, and energy efficiency in different network states through dynamic weight adjustment; For the conflicts between the optimization objectives, the Pareto optimization method is used to dynamically adjust the weights, and combined with reinforcement learning to adaptively adjust the weight coefficients of different optimization objectives.

[0010] Preferably, in the step S5, the following steps are further included: Perform global optimization on resource scheduling through graph optimization, and achieve global resource allocation coordination through local information transmission in a distributed environment; Use Lagrange multipliers to optimize local resource allocation so that the overall resource scheduling meets the global constraint conditions; The local information includes but is not limited to: Computing resource status: CPU load, GPU load, memory occupancy, and storage resources; Network resource status: bandwidth occupancy rate, link load, packet loss rate, and network latency; Task scheduling information: current task queue, task execution priority, and task execution status; Energy consumption information: power consumption of a single computing node, heat dissipation management, and battery power supply status; Load balancing information: load condition of the local node, available computing resources, and task migration information; The global constraint conditions include but are not limited to computing resource constraints, network resource constraints, delay constraints, energy consumption constraints, global optimization of task scheduling, and reliability and fault tolerance constraints.

[0011] Preferably, in the step S6, the following steps are further included: Self-supervised learning automatically learns network state features through unsupervised pre-training; Meta-learning quickly adjusts model parameters through a small amount of sample data so that the optimization strategy can adapt to different network environments; The sample data includes computing resource status, network status, task scheduling information, energy consumption data, and user demand data; The model parameters include computing resource allocation weights, task priority adjustment parameters, reinforcement learning reward functions, and deep learning optimization parameters.

[0012] Preferably, in the step S6, the network status is analyzed in real time through the self-attention mechanism and the Transformer model, specifically: Improve the calculation efficiency of the policy through the self-attention mechanism, accelerate the Q-value calculation and policy update, and improve the real-time performance of network optimization.

[0013] Preferably, the algorithm has a fault tolerance mechanism and performs dynamic adjustment in case of network failure, specifically including: Monitor bandwidth utilization, response time, and packet loss rate indicators through fault detection to identify faulty nodes; Automatically perform load transfer to ensure that tasks are not interrupted, and select the optimal node for takeover based on resource availability; Perform dynamic resource reallocation to optimize the allocation of bandwidth, computing resources, and storage space, and reduce the impact of faults on network performance.

[0014] Preferably, the fault tolerance mechanism further includes a preventive fault tolerance strategy based on predictive analysis, specifically including: Analyze historical data through a machine learning model to predict potential faulty nodes, and calculate the probability of fault occurrence based on long-term trends; Pre-adjust the resource allocation within 5-10 time windows before the fault occurs to reduce the impact of future faults; When the faulty node recovers, adaptively adjust the resource allocation, optimize the global resource utilization rate, and update the prediction model using historical data to improve the accuracy of future predictions.

[0015] The present invention provides an AI-driven real-time network optimization algorithm. It has the following beneficial effects: 1. Through the matrix operation technology based on the normal equation and in cooperation with an efficient data preprocessing process, the present invention achieves the effect of improving the solution speed of the regression model. Compared with traditional iterative algorithms, by reducing the amount of calculation and avoiding multiple optimization steps, the present invention solves the problem of slow calculation of traditional methods on large datasets and realizes fast and accurate regression analysis.

[0016] 2. By combining matrix inverse operation with an optimization algorithm, the present invention successfully solves the training instability caused by manual parameter tuning in the prior art. This technical solution not only has a high degree of automation but also ensures the adaptability and reliability of the regression model on different datasets through stable mathematical derivations, improving the accuracy of the training process.

[0017] 3. By optimizing the regression solution strategy, the regression analysis of the present invention is not only more efficient but also can operate stably in a variety of complex data environments. Compared with the existing gradient descent-dependent solutions, the present invention avoids the problems of gradient disappearance and hyperparameter adjustment, realizes a smoother convergence process, and effectively improves the processing ability of large-scale datasets.

[0018] 4. Through the matrix operation optimization scheme, the present invention reduces redundant calculations and improves data utilization rate, solves the deficiencies of high memory occupancy and low efficiency existing in traditional algorithms, and in cooperation with automated model training, provides a new method for regression analysis that is both efficient and accurate, greatly improving work efficiency and analysis quality. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 It is a flowchart of the algorithm optimization steps of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0020] The technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0021] Please refer to the attached Figure 1 , the embodiments of the present invention provide an AI-driven real-time network optimization algorithm, including the following steps: S1. Preprocess the network state data through a generative adversarial network, extract the optimized network feature data, and generate high-quality enhanced data; S2. Based on the enhanced data, use a graph neural network to model the network topology structure, capture the dependencies between nodes, and update the node features through a graph convolutional network; S3. Based on the node features, adopt a deep reinforcement learning algorithm to dynamically optimize and adjust the computing, storage, and bandwidth resources, and adaptively allocate resources; S4. Input the resource allocation scheme into a multi-objective optimization method to further minimize the latency, packet loss rate, maximize the throughput, and optimize the energy efficiency, and finally output a global optimized resource scheduling strategy; S5. Combine the Lagrangian relaxation method with graph optimization technology to optimize the global optimized resource scheduling strategy again, or the final scheduling scheme; S6. The final scheduling scheme improves the adaptability of the network model in a dynamic environment through self-supervised learning and meta-learning, and ensures the efficient operation of the optimization system under different network states; The algorithm has a fault tolerance mechanism and makes dynamic adjustments in case of network failures, specifically including: Monitor the bandwidth utilization rate, response time, and packet loss rate indicators through fault detection to identify faulty nodes; Automatically perform load transfer to ensure that the tasks are not interrupted, and select the optimal node for takeover based on resource availability; Perform dynamic resource reallocation to optimize the allocation of bandwidth, computing resources, and storage space, and reduce the impact of faults on network performance; The fault tolerance mechanism further includes a preventive fault tolerance strategy based on predictive analysis, specifically including: Analyze historical data through a machine learning model to predict potential faulty nodes, and calculate the probability of fault occurrence based on long-term trends; Pre-adjust the resource allocation within 5-10 time windows before the fault occurs to reduce the impact of future faults; After the faulty node recovers, adaptively adjust resource allocation, optimize the global resource utilization rate, and update the prediction model using historical data to improve future prediction accuracy; In step S1, the generative adversarial network is trained using a conditional generative adversarial network, specifically including: The generator generates enhanced data based on the input of network status such as network traffic, node load, and bandwidth utilization rate; The discriminator evaluates the quality of the generated data and feeds back to optimize the generator to improve the authenticity and effectiveness of the data.

[0022] Specifically, the main objective of step S1 is to preprocess the network status data through a generative adversarial network to extract optimized network feature data and generate high-quality enhanced data. This step lays the foundation for the subsequent modeling process based on the enhanced data. In particular, the generative adversarial network can generate realistic data samples through adversarial training, thereby improving the performance and robustness of the network model when dealing with complex network statuses. It should be noted that step S1 is closely related to the subsequent steps, and the optimized network feature data will directly affect the modeling effect of the graph neural network and the optimization of the resource allocation strategy in step S2.

[0023] In this embodiment, the generative adversarial network is trained using a conditional generative adversarial network. As a possible implementation, the conditional generative adversarial network helps generate data samples that are more in line with the actual network environment by introducing network traffic, node load, and bandwidth utilization rate, etc. as conditional inputs. In a specific implementation, the conditional generative adversarial network consists of a generator and a discriminator, where the generator generates enhanced data based on the network status input (e.g., network traffic, node load, bandwidth utilization rate), and the discriminator is responsible for evaluating the quality of the generated data and optimizing the generator through feedback, thereby improving the authenticity and effectiveness of the data.

[0024] Specifically, when the generator generates network status data, it first extracts feature information from parameters such as network traffic, node load, and bandwidth utilization rate, and generates enhanced data based on these features. Here, the network traffic can be expressed as , the node load can be expressed as , and the bandwidth utilization rate can be expressed as . The generator generates enhanced data based on these conditional data through a deep neural network. The goal of the generator is to minimize the gap between the generated data and the actual network data, thereby improving the authenticity of the generated data.

[0025] While the generator generates data, the discriminator receives the generated data and the actual network data as inputs and evaluates its authenticity. The output of the discriminator is a scalar , indicating the authenticity of the generated data. The discriminator continuously optimizes its judgment ability. The goal during the training process is to maximize the recognition rate of the discriminator for real data and minimize its misjudgment rate for generated data. During the training process, the discriminator will feedback its scores to the generator, and the generator will improve the data generation quality by adjusting parameters.

[0026] In some embodiments, the optimization process of the generative adversarial network can be achieved by minimizing the following loss function: ; where is the actual network data distribution; is the loss function, representing the total loss of the generative adversarial network; is the real data sampled from the real data distribution ; ; is the expectation operator, representing the expected value of the data distribution; is the discriminator 's output for the real data ; is the discriminator 's output for the generated data ; is the generated data sampled from the generated data distribution ; ; is the generated data distribution, and respectively represent the discriminant results of the discriminator for real data and generated data. The goal of the generator is to minimize the gap between the generated data and the real data, while the goal of the discriminator is to maximize the discriminant ability for real data and the misjudgment rate for generated data.

[0027] The training of the generative adversarial network can, to a certain extent, make up for the lack of data in reality and generate high-quality enhanced data through adversarial training. These enhanced data can effectively supplement the scarce data samples that may exist in the real network environment, thereby improving the adaptability of the network optimization model under various different network states.

[0028] In some implementation manners, the output enhanced data of the generative adversarial network is further used for the network topology modeling based on the graph neural network in the subsequent step S2. The optimized network feature data will be provided to the graph neural network as input to capture the dependencies between network nodes and update the node features through the graph convolutional network, thereby providing strong support for subsequent dynamic optimization and adjustment.

[0029] Through the preprocessing in step S1, the network state data is optimized, and enhanced data is generated. These data not only improve the accuracy of network state modeling but also effectively enhance the efficiency of subsequent resource scheduling optimization.

[0030] In step S2, the graph neural network performs multi-layer feature propagation through the graph convolutional network, updates the node features, and fuses the adjacent node information to optimize the resource allocation strategy, specifically including: Adopt a K-layer graph convolutional network to update the node features layer by layer, enabling the computational power, bandwidth utilization, and storage resource information of each node to gradually spread in the adjacency relationship, enhancing the perception of the global network state; Assign different weights to adjacent nodes through the attention mechanism to increase the accuracy of resource allocation decisions; Generate optimized node embedding vectors and use them as inputs for the reinforcement learning algorithm to improve the decision-making ability of resource allocation; Specifically, in step S2, based on the enhanced data generated in the previous step S1, further topological modeling of the network is performed, and the resource allocation strategy is optimized. Through the graph neural network, we can effectively capture the complex dependencies between network nodes, improving the accuracy and flexibility of the model in the network optimization process. This process provides strong support for subsequent dynamic resource scheduling and network traffic management. It should be noted that step S2 directly depends on the enhanced data generated in step S1, and the quality of these data determines the performance of the graph neural network in network topological modeling.

[0031] In this embodiment, the input of the graph neural network is the enhanced data processed by the generative adversarial network. These data include the state information of network nodes, link bandwidth, traffic demand, etc., and are processed by the graph neural network to capture the connection relationships and mutual influences between nodes. Specifically, in the structure of the graph neural network, each node represents a physical or logical node in the network (such as switches, routers, etc.), and the edges represent the connections between nodes. In order to improve the effect of the graph neural network, in some embodiments, the input features of each node in the network not only include the basic information of the node but also incorporate the enhanced data generated in the previous step S1, which helps to make the network topological modeling more accurate.

[0032] In a possible implementation, each layer of the graph neural network updates the state information of nodes through graph convolution operations. The update rule of the graph convolutional layer can be expressed as: ; where, represents the feature of node at the th layer; is the adjacent node at the The characteristics of the layer; is the set of adjacent nodes of the node ; is the normalization coefficient; and represent the weight and bias of the graph convolutional layer respectively; is the activation function; represents an adjacent node of the node in the graph . The graph neural network updates the characteristics of the nodes through this process, enabling each node in the network topology to learn richer information from its adjacent nodes.

[0033] During the processing of the graph neural network, the generated data will be passed through the convolutional layer to gradually update the characteristics of the nodes. The graph convolution operation of each layer takes into account the connection relationship between the nodes, which helps to capture complex network dependencies. In some embodiments, this process can further enhance the model's ability to learn long-range dependencies through multiple layers of graph convolution.

[0034] Specifically, after the calculation of multiple layers of graph convolution, the finally obtained node characteristics will be used for subsequent resource allocation and optimization decisions. These characteristics will affect the formulation of the network resource scheduling strategy, and further optimize the network traffic control, bandwidth allocation, etc. As an option, the output of the graph neural network can be used to construct a decision-making model for network resource scheduling. This model makes optimal decisions based on the state information of the network, thereby improving the overall efficiency and resource utilization rate of the network.

[0035] It should be emphasized that during the training process of the graph neural network, the parameters and are updated through the backpropagation algorithm. By continuously optimizing these parameters, the model can better fit the data in the actual network environment and improve the robustness and adaptability of the model; In step S3, the deep reinforcement learning to optimize network resource allocation specifically includes: The reward function is calculated based on network delay, throughput, packet loss rate, and energy efficiency performance metrics, and the weights are dynamically adjusted according to the current network state; Combined with the deep Q network for offline training, and using proximal policy optimization for online optimization to improve the policy adjustment ability; The state input of the reinforcement learning includes node computing power, storage utilization rate, bandwidth load, and topology structure information, and the action space includes resource allocation strategy adjustment, and uses the characteristics generated by the graph neural network for policy decision-making.

[0036] Specifically, in step S3, based on the network topology model constructed in step S2, resource scheduling optimization is further performed. Through a reinforcement learning algorithm, the resource allocation strategy in the network is dynamically adjusted to achieve the optimal utilization level of network resources under different load conditions. The dynamic nature of the network environment makes it difficult for static optimization methods to adapt, so a method based on deep reinforcement learning is adopted to achieve adaptive optimization. The key to this step is to use a reinforcement learning agent to learn the optimal strategy during continuous interaction to improve network performance.

[0037] In this embodiment, the state input of the reinforcement learning agent is provided by the graph neural network model in step S2. The state information includes the traffic load of nodes, bandwidth occupancy, link availability, etc., and this information constitutes the environmental state. . In some embodiments, the agent needs to extract key features from the environmental state and select an action according to the current policy , that is, adjust the resource allocation of a certain link or optimize the data flow routing.

[0038] In a possible implementation, the agent uses a policy gradient-based method for learning, and the loss function can be expressed as: ; where represents the trajectory experienced by the agent when executing the policy at time step ; is the reward obtained by the agent at time step ; represents a final time step; are the parameters of the policy network. Generally, the reward is calculated comprehensively from indicators such as the overall throughput of the network, link utilization rate, and delay.

[0039] Specifically, the policy network is implemented through a deep neural network (DNN), and a parameterized function is used to approximate the optimal policy. During the training process, the agent continuously samples state information from the environment, selects actions based on the policy network, and then updates the parameters according to the reward function. In some embodiments, the agent adopts an Actor-Critic architecture, where the Actor is responsible for selecting actions, and the Critic evaluates the value of the current state through a value function to improve the stability of policy optimization.

[0040] As an option, the value function Updated recursively through the following Bellman equation: ; where is the state at time step ; is the desired state; is the reward obtained at time step ; is the value function of the state at the next time step; is the discount factor, used to control the influence of future rewards on the current decision; a larger makes the agent more concerned about long-term rewards, while a smaller favors short-term optimization.

[0041] During the training process, the agent optimizes the policy network through the policy gradient method, and the direction of gradient update is given by the following formula: ; where is the objective function; is the expectation of the trajectory generated by the policy ; is the trajectory, usually containing a series from the starting state to the ending state; is the conditional probability of taking action at state , given by the policy controlled by the parameter ; is the sum from time step from 0 to ; is the gradient of with respect to ; is the reward obtained at time step .

[0042] In some possible implementations, the agent can adopt the Proximal Policy Optimization (PPO) algorithm, which improves the training stability and convergence speed by restricting the magnitude of policy updates and avoiding drastic changes in the policy network during training.

[0043] Specifically, the objective function of proximal policy optimization is as follows: ; where is the objective function of the PPO algorithm; represents the expectation over time steps or samples; is the minimum operation; is usually the probability ratio of the current policy to the old policy, used to measure the magnitude of policy updates; is the advantage function, representing the additional reward of an action in the current state relative to the baseline value; then the ratio is limited to within the interval to prevent over-updating of the policy; is a hyperparameter used to control the clipping range; are the upper and lower limits of clipping.

[0044] In a possible implementation, the training data of the reinforcement learning model is generated by a simulation environment, which is constructed based on the actual network topology and simulates different traffic patterns to improve the generalization ability of the model. After training, the reinforcement learning agent can adaptively adjust network resources in the actual environment, thereby optimizing network performance under different load conditions.

[0045] In step S4, multi-objective optimization adopts multi-task learning, specifically including: By dynamically adjusting weights, balance the optimization objectives of delay, throughput, packet loss rate, and energy efficiency under different network states; For the conflicts between optimization objectives, the Pareto optimization method is used to dynamically adjust the weights, and combined with reinforcement learning to adaptively adjust the weight coefficients of different optimization objectives.

[0046] Specifically, the core of step S4 is to execute the optimized resource scheduling policy and dynamically adjust the operation of the entire system. This step directly depends on the policy model obtained in step S3 and is verified in the actual network environment. The optimization objectives not only include the efficient utilization of network resources but also consider the adaptability under different network states. Due to the uncertainty of network states, relying solely on static allocation is difficult to meet the actual needs, so real-time evaluation and adjustment of the scheduling policy are required.

[0047] In this embodiment, the scheduling system performs resource allocation based on the output of the reinforcement learning model. Specifically, the system periodically samples the current network state , and through the trained policy network selects a suitable scheduling scheme. As an option, the scheduling decision is not only based on the current state but also combines the historical state sequence to enhance the perception ability of long-term trends.

[0048] In some embodiments, the system evaluates the execution effect of the scheduling decision and calculates the reward value of the current policy , which is weighted and calculated by multiple metrics such as throughput, delay, and packet loss rate: ; where, represents the network throughput; represents the average delay; Indicates the packet loss rate; are hyperparameters used to balance the impacts of different performance metrics. Generally, adjusting these weight parameters can adapt to different optimization goals, such as low-latency applications or high-throughput scenarios.

[0049] In one possible implementation, the system adopts a rolling window method to perform online adjustment of the scheduling policy, and the window size is , and the cumulative reward is calculated at the end of each window: ; where is the sum of the time steps from to ; is the discount factor, which is used to measure the weight of future rewards in the current value; indicates that the reward at time step is discounted in a manner proportional to the time difference ; is the time step obtained reward; As the overall benefit metric, it determines whether to perform policy adjustment; when is lower than the set threshold, the system re-executes policy update to adapt to the new network state.

[0050] Specifically, the system maintains an experience buffer pool to record the scheduling execution data of the past rounds . In some embodiments, these data are used to perform incremental training on the policy model, and the parameters are optimized using Stochastic Gradient Descent (SGD): ; where is the learning rate; is the optimization objective function. Generally, reinforcement learning methods such as PPO and DQN can be used to update the parameters; indicates the parameter to be optimized; indicates the update operation, which assigns the value on the right side to the left side; indicates the gradient of the objective function with respect to the parameter .

[0051] In different application scenarios, the system can dynamically adjust the scheduling policy. For example, in high-load situations, the system gives priority to low-latency routing, while in low-load situations, the system may adopt an equal distribution policy to improve resource utilization. As an option, multi-task learning can be combined to enable the scheduling policy to have stronger generalization ability under different application requirements.

[0052] In step S5, the following steps are further included: Global optimization of resource scheduling is carried out through graph optimization, and global resource allocation coordination is achieved through local information transmission in a distributed environment; Lagrange multipliers are used to optimize local resource allocation so that the overall resource scheduling meets the global constraint conditions; Local information includes but is not limited to: Computing resource status: CPU load, GPU load, memory occupancy, and storage resources; Network resource status: bandwidth occupancy rate, link load, packet loss rate, and network latency; Task scheduling information: current task queue, task execution priority, and task execution status; Energy consumption information: power consumption of a single computing node, heat dissipation management, and battery power supply status; Load balancing information: load conditions of local nodes, available computing resources, and task migration information; Global constraint conditions include but are not limited to computing resource constraints, network resource constraints, latency constraints, energy consumption constraints, global optimization of task scheduling, and reliability and fault tolerance constraints; Specifically, step S5 is mainly responsible for the comprehensive analysis of the processing results of the foregoing S1 to S4 and subsequent optimization processing tasks, and its goal is to improve the robustness and accuracy of the overall system. This step has clear designs in data fusion, anomaly detection, and filtering optimization to ensure seamless docking of data streams when passing between modules, meeting the expected technical effects.

[0053] In this embodiment, the implementation method of step S5 is described in detail as follows. Generally, step S5 includes preprocessing of the output data of each module, outlier detection, data fusion, and signal filtering optimization. As an option, step S5 first performs outlier detection on the input data, using the mean and standard deviation judgment method, and its judgment formula is: ; Among them, represents the original data of the th module; is the corrected data after outlier detection; is the mean of the data set; is the standard deviation; is a preset constant (such as 2 or 3) to control the tolerance range; is the absolute difference between the data point and the average value.

[0054] Specifically, in a possible implementation manner, the preprocessed data is fused using the weighted average method, and the formula is as follows: ; Among them, is the final output after fusion; represents the weight coefficient of the module data; is the number of modules participating in the fusion process; is the sum of all weighted terms; is the sum of all weights; is the number of samples. The weight coefficient can be dynamically calculated according to the signal-to-noise ratio or historical error of each module, and its calculation formula is: ; wherein, represents the variance of the module data; is a small positive constant to prevent division by zero.

[0055] In some embodiments, in order to further reduce high-frequency noise interference, a frequency-domain filtering module is further added after data fusion. As an option, the filtering process can be described by the following formula: ; wherein, represents the representation of the filtered signal in the frequency domain;

[0056] is the spectrum of the fused signal;

[0057] is the designed filter function. Generally, a low-pass filter can be selected.

[0058] In a possible implementation manner, to further improve the system adaptive performance, step S5 can also introduce a weight adaptive learning mechanism based on error backpropagation. Specifically, the update iteration formula for each weight coefficient is: ; wherein, is the value of the th weight at the th iteration; represents the weight of the module in the th iteration; is the learning rate; is the objective function with respect to the weight partial derivative, indicating the influence degree of weight change on the loss.

[0059] Generally, each processing link in step S5 closely cooperates with the data output and processing processes of the aforementioned S1 to S4, which not only ensures data continuity but also enhances the overall stability and robustness of the system. As an option, in the embodiment, each link can also be modularly extended according to actual needs, such as adding data preprocessing algorithms, introducing multiple filter combinations, or adopting other adaptive learning algorithms to meet the requirements of different application scenarios.

[0060] In step S6, the following steps are further included: Self-supervised learning automatically learns network state features through unsupervised pre-training; Meta-learning quickly adjusts model parameters through a small amount of sample data, enabling the optimization strategy to adapt to different network environments; The sample data includes computing resource status, network status, task scheduling information, energy consumption data, and user demand data; The model parameters include computing resource allocation weights, task priority adjustment parameters, reinforcement learning reward functions, and deep learning optimization parameters; In step S6, the network state is analyzed in real time through the self-attention mechanism and the Transformer model. Specifically: The self-attention mechanism is used to improve the efficiency of policy calculation, accelerate Q-value calculation and policy update, and improve the real-time performance of network optimization.

[0061] Specifically, step S6 plays a role in final decision correction and feedback control. It receives the fused data from step S5 and forms a closed-loop connection with the previous processing results. This link performs a final verification of the data, adjusts system parameters, and ensures the overall accuracy and stability.

[0062] In this embodiment, step S6 analyzes the fused output. Generally, step S6 adopts a method combining threshold judgment and error correction. As an option, step S6 uses the fused result and the preset reference value to determine the decision based on the deviation therebetween. The judgment formula is: ; where represents the final decision output; is the data fused in step S5; is the expected reference value; is the allowable deviation threshold.

[0063] Specifically, in a possible implementation manner, step S6 adopts a dynamic threshold adjustment strategy. In some embodiments, the threshold parameter is updated by the following formula: ; where is the threshold value at the moment; is the threshold value at the moment; is the learning rate, and its value range is 0 < < 10.

[0064] In another implementation, step S6 generates a feedback control signal for correcting the front-end module. Specifically, the calculation formula for the feedback signal F is: ; where is the feedback control signal; is the feedback gain coefficient, which is set according to system requirements.

[0065] In some embodiments, in order to further correct the model error, the least squares method is used for parameter update.

[0066] In this case, the model parameter correction formula is: ; where represents the corrected parameter vector; is the design matrix, which contains the values of all input features; is the design matrix transpose matrix; is the inverse matrix of the matrix obtained by multiplying the transpose of the design matrix by the original matrix; is the error vector, which is defined as , where and respectively represent the actual output and the expected output of the th acquisition.

[0067] Generally, the data correction and feedback control in step S6 are closely connected with the data processing in the foregoing steps S1 to S5.

[0068] As an option, after receiving the fusion result, step S6 can also compare with the output data of the preprocessing and feature extraction module to verify the overall consistency of the system. In a possible implementation, fuzzy logic control is also introduced to adaptively adjust the threshold and feedback.

[0069] Specifically, the inputs of the fuzzy controller are and its change rate , and the output is the feedback correction value .

[0070] In addition, the design of the fuzzy rules and membership functions refers to the system requirements and can be determined according to empirical values or statistical analysis.

[0071] Although embodiments of the present invention have been shown and described, those of ordinary skill in the art will appreciate that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. AI-driven real-time network optimization algorithm, characterized by: The following steps are involved: S1. Preprocess the network state data through the generative adversarial network, extract the optimized network feature data, and generate high-quality enhanced data; S2. Based on the enhanced data, the graph neural network is used to model the network topology structure, capture the dependencies between nodes, and update the node features through the graph convolutional network; S3, based on node characteristics, uses deep reinforcement learning algorithms to dynamically optimize and adjust computing, storage, and bandwidth resources, and adaptively allocate resources; S4, input the resource allocation plan into the multi-objective optimization method to further minimize the delay and packet loss rate, maximize the throughput, and optimize the energy efficiency, and finally output the global optimization resource scheduling strategy; S5. Combine Lagrangian relaxation method and graph optimization technology to further optimize the global optimal resource scheduling strategy or the final scheduling plan; S6. The final scheduling scheme improves the adaptability of the network model in a dynamic environment through self-supervised learning and meta-learning, ensuring that the optimization system operates efficiently under different network conditions.

2. The AI-driven real-time network optimization algorithm according to claim 1, characterized in that: In the step S1, the generative adversarial network is trained using a conditional generative adversarial network, specifically including: The generator generates enhanced data based on the network status input of network traffic, node load and bandwidth utilization; The discriminator evaluates the quality of generated data and provides feedback to optimize the generator, improving the authenticity and effectiveness of the data.

3. The AI-driven real-time network optimization algorithm according to claim 1, characterized in that: In the S2 step, the graph neural network performs multi-layer feature propagation through the graph convolutional network, updates node features, and integrates adjacent node information to optimize the resource allocation strategy, specifically including: A K-layer graph convolutional network is used to update node features layer by layer, so that the computing power, bandwidth utilization and storage resource information of each node are gradually propagated in the neighbor relationship, enhancing the perception of the global network status; The attention mechanism is used to assign different weights to adjacent nodes, thus increasing the accuracy of resource allocation decisions. Generate optimized node embedding vectors and provide them as input to the reinforcement learning algorithm to improve the decision-making ability of resource allocation.

4. The AI-driven real-time network optimization algorithm according to claim 1, characterized in that: In the step S3, deep reinforcement learning to optimize network resource allocation specifically includes: The reward function is calculated based on network latency, throughput, packet loss rate, and energy efficiency performance indicators, and the weights are dynamically adjusted according to the current network status; Combine the deep Q network for offline training and use proximal policy optimization for online optimization to improve the policy adjustment capability; The state input of the reinforcement learning includes node computing power, storage utilization, bandwidth load and topological structure information. The action space includes resource allocation strategy adjustment, and strategy decisions are made using the features generated by the graph neural network.

5. The AI-driven real-time network optimization algorithm according to claim 1, characterized in that: In the step S4, multi-objective optimization adopts multi-task learning, which specifically includes: Through dynamic weight adjustment, the optimization goals of latency, throughput, packet loss rate and energy efficiency are balanced under different network conditions; In order to solve the conflicts among the optimization objectives, the Pareto optimization method is used to dynamically adjust the weights, and the weight coefficients of different optimization objectives are adaptively adjusted in combination with reinforcement learning.

6. The AI-driven real-time network optimization algorithm according to claim 1, characterized in that: The step S5 further comprises the following steps: Globally optimize resource scheduling through graph optimization, and achieve global resource allocation coordination through local information transmission in a distributed environment; Lagrange multipliers are used to optimize local resource allocation so that the overall resource scheduling meets global constraints; The local information includes but is not limited to: Computing resource status: CPU load, GPU load, memory usage, and storage resources; Network resource status: bandwidth utilization, link load, packet loss rate, and network delay; Task scheduling information: current task queue, task execution priority and task execution status; Energy consumption information: power consumption, thermal management, and battery power status of individual computing nodes; Load balancing information: local node load, available computing resources, and task migration information; Global constraints include, but are not limited to, computing resource constraints, network resource constraints, latency constraints, energy consumption constraints, global optimization of task scheduling, and reliability and fault tolerance constraints.

7. The AI-driven real-time network optimization algorithm according to claim 1, characterized in that: The step S6 further comprises the following steps: Self-supervised learning automatically learns network state features through unsupervised pre-training; Meta-learning quickly adjusts model parameters through a small amount of sample data, allowing the optimization strategy to adapt to different network environments; The sample data includes computing resource status, network status, task scheduling information, energy consumption data and user demand data; The model parameters include computing resource allocation weights, task priority adjustment parameters, reinforcement learning reward functions, and deep learning optimization parameters.

8. The AI-driven real-time network optimization algorithm according to claim 1, characterized in that: In step S6, the network state is analyzed in real time through the self-attention mechanism and the Transformer model, specifically: The self-attention mechanism is used to improve the efficiency of strategy calculation, accelerate Q-value calculation and strategy update, and improve the real-time performance of network optimization.

9. The AI-driven real-time network optimization algorithm according to claim 1, characterized in that: The algorithm has a fault-tolerant mechanism and makes dynamic adjustments when a network failure occurs, including: Monitor bandwidth utilization, response time, and packet loss rate through fault detection to identify faulty nodes. Automatically transfer load to ensure uninterrupted tasks and select the best node to take over based on resource availability; Perform dynamic resource reallocation to optimize bandwidth, computing resources, and storage space allocation, and reduce the impact of failures on network performance.

10. The AI-driven real-time network optimization algorithm according to claim 9, characterized in that: The fault tolerance mechanism further includes a preventive fault tolerance strategy based on predictive analysis, specifically including: Analyze historical data through machine learning models, predict potential failure nodes, and calculate the probability of failure based on long-term trends; Pre-adjust resource allocation within a 5-10 time window before a failure occurs to reduce the impact of future failures; When the failed node is restored, resource allocation is adaptively adjusted to optimize global resource utilization, and the prediction model is updated using historical data to improve the accuracy of future predictions.

Citation Information

Patent Citations

  • Implementation method of inter-domain flow engineering achieving double optimization of operating cost and transmission performance

    CN103200113A

  • Container cluster online deployment method fusing graph neural network and reinforcement learning in edge computing

    CN115686846A

  • Super-fusion expansion AI performance optimization system based on adaptive network topology algorithm

    CN118747518A

  • Routing optimization method based on graph neural network and deep reinforcement learning

    CN118784547A

  • Multi-objective reinforcement learning method and system for adaptive network flow optimization

    CN119071236A

Cited By

  • AI optimization system and method for cloud-side integrated distributed computing architecture

    CN120768902A

  • An ai optimization system and method of cloud-edge integrated distributed computing architecture

    CN120768902B

  • Diabetes health management method and system based on AI big data

    CN121075563A

  • Electric power communication network construction method and device based on deep learning technology

    CN121418310A