Routing policy adjustment method and device, storage medium and electronic equipment

By building a state vector and inputting it into the target network model, adjusting the model parameters using the reward function, and dynamically updating the routing strategy, the problem of being unable to dynamically adjust the network routing strategy in the existing technology is solved, and the stability and adaptability of system performance are achieved.

CN120200955APending Publication Date: 2025-06-24JINAN INSPUR DATA TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510404888.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-01
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

The existing technology cannot dynamically adjust network routing policies and cannot effectively adapt to dynamically changing network environments and load needs.

Method used

By determining the current network status and historical network status of the target cluster, a status vector is built and input into the target network model, the reward function is used to calculate the reward value, adjust the model parameters, and dynamically update the routing strategy.

Benefits of technology

It realizes dynamic adjustment of network routing strategies to ensure the stability and adaptability of system performance, and can effectively respond to changes in network environment and load.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120200955A_ABST
    Figure CN120200955A_ABST
Patent Text Reader

Abstract

The invention discloses a routing strategy adjustment method and device, a storage medium and electronic equipment, and relates to the technical field of computers, and the method comprises the steps: constructing a first state vector through employing a current network state and a historical network state, inputting the first state vector into a reward function, and obtaining a reward value determined by the reward function. The reward value and the current network state can be utilized to adjust the model parameters of the target network model, an updated network model can be obtained, and the routing strategy of the current target cluster is adjusted through the updated network model. The reward value of the reward function can be determined according to the first state vector determined according to the current network state and the historical network state, and the model parameters of the target network model are dynamically learned and adjusted online based on the reward value and the current network state, so that the stability and adaptability of system performance can be ensured. Therefore, the technical problem that the network routing strategy cannot be dynamically adjusted can be solved, and the technical effect of dynamically adjusting the network routing strategy is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and particularly to a method, apparatus, storage medium, and electronic device for adjusting routing policies. Background Art

[0002] In the related art, network routing policies cannot effectively adapt to dynamic network environments and load requirements.

[0003] It can be seen therefrom that there is a problem in the related art that network routing policies cannot be dynamically adjusted.

[0004] In view of the above problems in the related art, no effective solution has been proposed yet. Summary of the Invention

[0005] This application provides a method, apparatus, storage medium, and electronic device for adjusting routing policies, so as to at least solve the problem in the related art that network routing policies cannot be dynamically adjusted.

[0006] This application provides a method for adjusting a routing policy, including: determining the current network state of a target cluster and the historical network state within a historical time period; determining a first state vector based on the current network state and the historical network state; inputting the first state vector into a target network model to obtain a reward value determined based on a reward function included in the target network model, where the reward function is constructed based on the weights, exponential function, and linear function of each sub-state included in the network state of the target cluster; adjusting the model parameters of the target network model based on the reward value and the current network state to obtain an updated network model, where the model parameters include the weights of each sub-state; and adjusting the routing policy of the target cluster through the updated network model.

[0007] This application also provides an apparatus for adjusting a routing policy, including: a first determination module for determining the current network state of a target cluster and the historical network state within a historical time period; a second determination module for determining a first state vector based on the current network state and the historical network state; an input module for inputting the first state vector into a target network model to obtain a reward value determined based on a reward function included in the target network model, where the reward function is constructed based on the weights, exponential function, and linear function of each sub-state included in the network state of the target cluster; a first adjustment module for adjusting the model parameters of the target network model based on the reward value and the current network state to obtain an updated network model, where the model parameters include the weights of each sub-state; and a second adjustment module for adjusting the routing policy of the target cluster through the updated network model.

[0008] The present application also provides an electronic device, including: a memory for storing a computer program; and a processor for implementing the steps of any one of the above routing policy adjustment methods when executing the computer program.

[0009] The present application also provides a computer-readable storage medium storing a computer program, wherein the computer program, when executed by a processor, implements the steps of any one of the above routing policy adjustment methods.

[0010] The present application also provides a computer program product including a computer program, which, when executed by a processor, implements the steps of any one of the above routing policy adjustment methods.

[0011] Through the present application, after determining the current network state of the target cluster and the historical network state within the historical time period, a first state vector can be constructed using the current network state and the historical network state, and the first state vector is input into a reward function included in the target network model and constructed by the weights of each sub-state included in the network state of the target cluster, the exponential function, and the linear function, and a reward value determined by the reward function can be obtained. The model parameters of the target network model, that is, the weights of each sub-state, can be adjusted using the reward value and the current network state, and an updated network model can be obtained, and the routing policy of the current target cluster can be adjusted through the updated network model. Since the reward value of the reward function can be determined based on the first state vector determined according to the current network state and the historical network state, and the model parameters of the target network model are dynamically and online learned and adjusted based on the reward value and the current network state, the stability and adaptability of the system performance can be ensured. Therefore, the technical problem of being unable to dynamically adjust the network routing policy can be solved, and the technical effect of dynamically adjusting the network routing policy can be achieved. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] To more clearly illustrate the embodiments of the present application, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application, and those of ordinary skill in the art can obtain other drawings based on these drawings without creative efforts.

[0013] Figure 1 It is a flowchart of the routing policy adjustment method according to an embodiment of the present invention;

[0014] Figure 2 It is a structural diagram of the routing policy adjustment system according to an embodiment of the present invention;

[0015] Figure 3 It is a schematic diagram of the reinforcement learning model according to an embodiment of the present invention;

[0016] Figure 4Schematic diagram of reward function design according to an embodiment of the present invention;

[0017] Figure 5 Block diagram of the structure of the adjustment device for the routing strategy according to an embodiment of the present invention. Detailed implementation manners

[0018] Next, the technical solutions in the embodiments of the present application will be clearly and completely described with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0019] It should be noted that in the description of the present application, the terms "include", "comprise" or any other variant thereof are intended to cover a non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. The terms "first", "second", etc. in the present application are used to distinguish similar objects and are not used to describe a specific order or sequence.

[0020] To enable those skilled in the art of the present technology to better understand the solutions of the present application, the present application will be further described in detail below with reference to the accompanying drawings and specific implementation manners.

[0021] Combined with the specific application environment architecture or specific hardware architecture on which the execution of the routing strategy adjustment method depends, the specific application environment architecture or specific hardware architecture will be described herein.

[0022] An embodiment of the present application provides a method for adjusting a routing strategy. The method will be described in detail in combination with the execution process of the routing strategy adjustment method. Figure 1 It is a flowchart of the method for adjusting the routing strategy according to an embodiment of the present invention, as Figure 1 shown. The method includes the following steps:

[0023] Step S102, determining the current network state of the target cluster and the historical network state within a historical time period;

[0024] Step S104, determining a first state vector based on the current network state and the historical network state;

[0025] Step S106, inputting the first state vector into the target network model to obtain a reward value determined based on the reward function included in the target network model, where the reward function is constructed based on the weights, exponential function, and linear function of each sub-state included in the network state of the target cluster;

[0026] Step S108: Adjust the model parameters of the target network model based on the reward value and the current network state to obtain an updated network model, where the model parameters include the weights of each sub-state.

[0027] Step S110: Adjust the routing policy of the target cluster through the updated network model.

[0028] In the above embodiments, the target cluster can be understood as a system composed of multiple computers (nodes) connected through a network. The nodes cooperate to complete computing tasks, which can provide higher and more reliable performance than a single computer. Examples include Kubernetes (k8s cluster), Apache Mesos, Hadoop, etc. In the target cluster, resources such as processors, memory, storage, and network bandwidth can be shared among nodes, and the overall computing efficiency and data processing capacity can be improved through parallel processing or load sharing.

[0029] In the above embodiments, the real-time status information of the target cluster nodes can be obtained by using the Kubernetes API (Application Programming Interface). Figure 2 It is a structural diagram of a routing policy adjustment system according to an embodiment of the present invention. As Figure 2 shown, the data acquisition module can collect the network status data of the target cluster in real time. At the same time, network monitoring tools (such as Prometheus, Istio, etc.) can also be deployed to monitor the network status data between nodes. Among them, the network status data can include latency d, bandwidth utilization b, node load l, and traffic characteristics. The latency can be understood as the network latency between nodes; the bandwidth utilization can be understood as the bandwidth usage between nodes; the node load can be understood as the load information such as the CPU and memory usage of the node; the traffic characteristics can be understood as traffic characteristics such as the packet size and transmission frequency. The collected raw data can be preprocessed, which can include data cleaning, data normalization, and time series data processing. Data cleaning can be understood as removing noise data and outliers to ensure the accuracy and consistency of the data; data normalization can be understood as normalizing data with different dimensions, for example, scaling features such as latency, bandwidth utilization, and node load to the range of [0, 1]; time series data processing can be understood as integrating historical status data into the current state to form a time series state vector.

[0030] In the above embodiments, the preprocessed data can be integrated into a unified state vector (i.e., the above first state vector). First, the data of the current network state can be obtained, and features such as latency, bandwidth utilization, and node load can be arranged in sequence to form a basic state vector: S = [d1, d2,..., b1, b2,..., l1, l2,...], which can provide direct network state information. Secondly, in order to capture the dynamic changes of the network state, historical network state information (i.e., the data of the above historical network state) can be introduced to form a temporal state vector: S t = [S t-1 , S t-2 ,..., S t-n , where n can represent the window size of the network historical state, which can be 5 minutes, 7 minutes, 10 minutes, etc., and this is not limited in this article. Finally, a deep neural network (DNN, Deep Neural Network) can be used to perform non-linear transformation on multi-dimensional features. The temporal state vector and the basic state vector are input into the DNN, and non-linear transformation is performed on the input features through multiple hidden layers to extract more complex features, and a higher-level feature vector, that is, the first state vector, can be generated. Among them, the deep neural network DNN can be understood as constructing a more complex model by increasing the number of layers (i.e., depth) of the network to achieve high-level abstraction of data and more accurate prediction or decision-making. That is, DNN can be understood as a feedforward neural network, and the input data is transformed through a series of hidden layers and finally generates an output.

[0031] In the above embodiments, the target network model can be understood as a reinforcement learning model, which can include a state space, an action space, and a reward function. Figure 3 It is a schematic diagram of the reinforcement learning model according to an embodiment of the present invention. As Figure 3 shown, the state space (i.e., the state observation in the figure) can be understood as defining the network state, which can include parameters such as latency between nodes, bandwidth utilization, and node load; the action space (i.e., the above action selection) can be understood as defining possible routing adjustment actions, which can include selecting different network paths, adjusting bandwidth allocation policies, etc.; the reward function can be understood as a reward function designed according to the optimization goal, aiming at reducing latency and increasing bandwidth utilization, and dynamically calculating the reward value. By optimizing the routing path in real time through the reinforcement learning model, the latency of data transmission can be effectively reduced, and the response speed and user experience of the application can be improved, especially in scenarios sensitive to latency such as real-time video transmission.

[0032] In the above embodiments, the design of the action space can optimize the data transmission path by selecting basic routing actions, that is, different network paths can be selected, the path selection can be dynamically adjusted according to the current network state, weights can be assigned to each path, and priorities can be defined. Among them, the weights can be dynamically adjusted according to factors such as latency and bandwidth utilization. The design of the action space can also optimize network performance by dynamically adjusting the bandwidth allocation ratio between nodes, that is, the bandwidth allocation ratio can be dynamically adjusted according to the real-time network state (such as node load and bandwidth utilization), and multiple bandwidth allocation strategies can be designed for selection, such as equal allocation and priority allocation. The design of the action space can also optimize network performance through different traffic scheduling strategies, that is, the routing strategy can be dynamically adjusted according to the priority of traffic (such as real-time video and file transfer), and the traffic can be dynamically allocated according to the load of nodes to achieve load balancing. For traffic characteristic analysis, different traffic types can be identified to provide a basis for the scheduling strategy. The design of the action space can also dynamically generate the optimal routing action according to the real-time network traffic characteristics, that is, the characteristics of network traffic, such as packet size, transmission frequency, and traffic type, can be extracted, and the optimal routing action can be dynamically generated according to the extracted traffic characteristics. For example, a low-latency path can be selected for real-time video traffic, and a high-bandwidth path can be selected for file transfer traffic. According to the real-time changes in traffic characteristics, the priority of routing actions can be dynamically adjusted to ensure that the optimal decision can be made according to the current network environment. The design of the action space can also implement routing actions and bandwidth allocation strategies through the network policies and bandwidth management tools of Kubernetes. The bandwidth management function of the integrated bandwidth management tool (Calico) can be combined with the network policies of Kubernetes to achieve more refined network control. Through the calls of Kubernetes API and Calico API, the network configuration can be dynamically updated to ensure that the network state can be adjusted in real time.

[0033] In the above embodiments, the reward function in the target network model can be constructed by the weights, exponential functions, and linear functions of each sub-state included in the network state of the target cluster. The combined first state vector is input into the reward function included in the target network model, and then the reward value R calculated by the reward function can be obtained. The reward value R is fed back to the reinforcement learning agent, which can guide it to select the optimal routing action. That is, the reinforcement learning agent will update the policy parameters (i.e., the above model parameters) of the model using the reinforcement learning algorithm according to the current network state and the reward value R to optimize the routing selection policy, which may include: recording the current network state, the actions taken, and the obtained reward value R, updating the policy parameters of the reinforcement learning model based on the collected historical network state data. Through continuous iteration, the model gradually optimizes its policy to maximize the cumulative reward, and then the updated network model can be obtained. Then, the routing policy of the target cluster is adjusted through the updated network model. Among them, the reinforcement learning agent can be understood as the core component in the Reinforcement Learning (RL) framework, which can learn and improve its behavior policy through interaction with the environment to maximize a certain reward signal.

[0034] In the above embodiments, a dynamic routing policy can be generated according to the output of the updated network model and applied to the target cluster. That is, a clustering algorithm can be used to classify network traffic, analyze the time series characteristics of the data (such as arrival interval time, burst traffic frequency, etc.) and identify different traffic types (such as real-time video, file transfer, etc.). Among them, the clustering algorithm can be understood as an unsupervised learning method that can group or cluster the samples in the dataset so that the samples within the same group are similar to each other. It can be the K-Means algorithm, the hierarchical clustering algorithm, or the DBSCAN (Density-Based Spatial Clustering of Applications with Noise) algorithm, but not limited to this. According to different types of traffic, different routing strategies can be selected. For example, for real-time video traffic, a low-latency path can be selected; for file transfer traffic, a high-bandwidth path can be selected. In this embodiment, fusing the traffic characteristics with the network state characteristics can form a comprehensive state representation, which can assist the updated network model to better understand the network environment. After generating the routing policy, the generated routing policy can be converted into a configuration format understandable by the Kubernetes API, and the routing policy configuration is sent to the target cluster nodes through the Kubernetes API, and the new routing policy starts to be applied.

[0035] Through the present application, after determining the current network state of the target cluster and the historical network state within a historical time period, a first state vector can be constructed using the current network state and the historical network state, and the first state vector is input into a reward function included in the target network model and constructed by the weights of each sub-state included in the network state of the target cluster, the exponential function, and the linear function, and a reward value determined by the reward function can be obtained. The model parameters of the target network model, that is, the weights of each sub-state, can be adjusted using the reward value and the current network state, and an updated network model can be obtained, and the routing policy of the current target cluster can be adjusted through the updated network model. Since the reward value of the reward function can be determined based on the first state vector determined according to the current network state and the historical network state, and the model parameters of the target network model are dynamically and online learned and adjusted based on the reward value and the current network state, the stability and adaptability of the system performance can be ensured. Therefore, the technical problem of being unable to dynamically adjust the network routing policy can be solved, and the technical effect of dynamically adjusting the network routing policy can be achieved.

[0036] In an exemplary embodiment, before inputting the first state vector into the target network model to obtain a reward value determined based on the reward function included in the target network model, the method further includes: constructing a first exponential function based on a first sub-state included in the network state; constructing a linear function based on a second sub-state included in the network state; constructing a second exponential function based on a third sub-state included in the network state; determining a first weight of the first sub-state, a second weight of the second sub-state, and a third weight of the third sub-state; and constructing a reward function based on the first exponential function, the linear function, the second exponential function, the first weight, the second weight, and the third weight.

[0037] In the above embodiment, the key network performance indicators that need to be optimized may include: network latency d (i.e., the above-mentioned first sub-state), that is, reducing latency to improve real-time performance; bandwidth utilization b (i.e., the above-mentioned second sub-state), that is, improving bandwidth utilization to optimize resource usage; node load l (i.e., the above-mentioned third sub-state), that is, achieving node load balancing to avoid overload. Figure 4 is a schematic diagram of the reward function design according to an embodiment of the present invention, as Figure 4 shown, a first exponential function f delay (d) can be constructed through the first sub-state network latency, a linear function f bandwidth (b) can be constructed through the second sub-state bandwidth utilization, and a second exponential function f load(l), assign three different weights α, β, γ (i.e., the above-mentioned first weight, second weight, and third weight) to the three functions, which can be used to adjust the contributions of the above three sub-states to the reward function. By subdividing the network state into different sub-states (such as latency, bandwidth utilization, node load, etc.) and constructing the first exponential function, linear function, and second exponential function respectively, the reward function can evaluate the current network state from multiple perspectives, which helps to understand the network environment more comprehensively rather than relying solely on a single metric.

[0038] In an exemplary embodiment, constructing a reward function based on the first exponential function, linear function, second exponential function, first weight, second weight, and third weight includes: determining a first product of the first exponential function and the first weight; determining a second product of the linear function and the second weight; determining a third product of the second exponential function and the third weight; and determining the sum value of the first product, second product, and third product as the reward function.

[0039] In the above embodiment, in order to comprehensively consider the above three metrics, a hybrid reward function can be constructed in the form of a combination of a linear combination and an exponential function, that is, R = α·f delay (d)+β·f bandwidth (b)+γ·f load (l), where α·f delay (d) is the above-mentioned first product, β·f bandwidth (b) is the above-mentioned second product, γ·f load (l) is the above-mentioned third product. The sum value of the first product, second product, and third product can be determined as the reward function. By mapping different network performance metrics (such as latency, bandwidth utilization, node load) to exponential functions and linear functions respectively and assigning different weights, the reward function can consider and optimize these metrics simultaneously, enabling the reinforcement learning model to find the best balance among multiple objectives. In addition, the first weight, second weight, and third weight can be dynamically adjusted according to the actual situation, which means that in different network scenarios, the model can prioritize optimizing a certain or certain performance metrics. For example, when the network is congested, the weight of bandwidth utilization can be increased, and when the network load is unbalanced, the weight of node load can be increased.

[0040] In an exemplary embodiment, constructing the first exponential function based on the first sub-state included in the network state includes: determining a fourth product of the first sub-state and the first hyperparameter, determining the fourth product power of the natural constant to obtain a first function, and determining the negative of the first function as the first exponential function; constructing the second exponential function based on the third sub-state included in the network state includes: determining a fifth product of the third sub-state and the second hyperparameter, determining the fifth product power of the natural constant to obtain a second function, and determining the negative of the second function as the second exponential function.

[0041] In the above embodiments, the network latency d is a key performance metric, and the smaller d is, the better, that is, the lower the latency, the better. Since the increase in latency has a non-linear impact on network performance, an exponential function form can be used to penalize high latency, that is, f delay (d) = -e λd , where λ > 0 can be understood as a hyperparameter (i.e., the above first hyperparameter), which can control the steepness of the exponential function. When the latency increases, f delay (d) will decrease sharply. λd is the above fourth product, and e is the natural constant, and e λd is the first function.

[0042] In the above embodiments, the node load l is a negative metric, that is, the lower the load, the better. Since the increase in load also has a non-linear impact on network performance, an exponential function form can also be used to penalize high load, that is, f load (l) = -e μ·l , where μ > 0 can be understood as a hyperparameter (i.e., the above second hyperparameter), which can control the steepness of the exponential function. When the load increases, f load (l) = -e μ·l will decrease sharply. μ·l is the above fifth product, and e is the natural constant, and e μ·l is the second function.

[0043] In the above embodiments, the bandwidth utilization rate d is a positive metric, that is, the higher the utilization rate, the better. Therefore, a linear function form can be used to directly reflect the contribution of the bandwidth utilization rate, that is, f bandwidth (b) = b, where the value range of b can be [0, 1], which can represent the percentage of the bandwidth utilization rate.

[0044] In the above embodiments, by constructing an exponential function, non-linear penalties can be imposed on network latency and node load, and a linear function is constructed to impose a linear penalty on the bandwidth utilization rate. The construction of the exponential function depends on the product with hyperparameters, and these hyperparameters can be dynamically adjusted according to changes in the network environment. That is, when the network conditions change, the model can respond quickly. By adjusting the hyperparameters, it can more accurately reflect the impact degree of latency and load changes on network performance, and thus make more reasonable strategy selections.

[0045] In an exemplary embodiment, adjusting the model parameters of the target network model based on the reward value and the current network state includes: when the reward value does not meet the predetermined conditions, updating the reward function based on the current network state; determining the first partial derivative of the reward function with respect to the first sub-state included in the network state; determining the second partial derivative of the reward function with respect to the second sub-state included in the network state; determining the third partial derivative of the reward function with respect to the third sub-state included in the network state; adjusting the first weight of the first sub-state included in the reward function based on the first partial derivative; adjusting the second weight of the second sub-state included in the reward function based on the second partial derivative; adjusting the third weight of the third sub-state included in the reward function based on the third partial derivative.

[0046] In the above embodiment, when the reward value does not meet the predetermined conditions (such as not achieving the performance target, slow model convergence speed, etc.), the reward function can be updated based on the current network state. That is, in order to make the reward function adapt to the changes in the network environment, it is necessary to dynamically adjust the weight coefficients. The adjustment of the weight coefficients is based on the changes in the network state. When the network delay increases significantly, increase the weight of α to make the model pay more attention to delay optimization. When the bandwidth utilization rate is low, increase the weight of β to make the model pay more attention to the improvement of bandwidth utilization. When the node load is too high, increase the weight of γ to make the model pay more attention to load balancing.

[0047] In the above embodiment, the partial derivatives of the reward function with respect to the first sub-state (network delay), the second sub-state (bandwidth utilization rate), and the third sub-state (node load) in the network state can be calculated first. The partial derivatives measure the influence degree of each sub-state on the reward value, and it can be understood how the changes in the network state affect the total reward of the model, thereby guiding the adjustment of the weights (i.e., the above first weight, second weight, and third weight). By calculating the partial derivatives of the reward function with respect to different sub-states in the network state, the change rate of the contribution of each sub-state to the total reward can be determined, and then the weights can be dynamically adjusted. Under specific network conditions, the network performance indicators that need to be concerned can be automatically identified and preferentially optimized, improving the pertinence and flexibility of the optimization strategy.

[0048] In an exemplary embodiment, adjusting the first weight of the first sub-state included in the reward function based on the first partial derivative includes: determining the first historical weight of the first sub-state, determining the learning rate of the target network model, determining the sixth product of the first partial derivative and the learning rate, and determining the sum value of the sixth product and the first historical weight as the first weight; adjusting the second weight of the second sub-state included in the reward function based on the second partial derivative includes: determining the second historical weight of the second sub-state, determining the learning rate of the target network model, determining the seventh product of the second partial derivative and the learning rate, and determining the sum value of the seventh product and the second historical weight as the second weight; adjusting the third weight of the third sub-state included in the reward function based on the third partial derivative includes: determining the third historical weight of the third sub-state, determining the learning rate of the target network model, determining the eighth product of the third partial derivative and the learning rate, and determining the sum value of the eighth product and the third historical weight as the third weight.

[0049] In the above embodiment, the dynamic adjustment of the weight coefficient can be achieved through the following formula: Wherein, can be understood as the gradient of the reward function with respect to the weight coefficient (i.e., the above first partial derivative, second partial derivative, and third partial derivative). α0, β0, and γ0 can be understood as the initial weight coefficients (i.e., the above first historical weight, second historical weight, and third historical weight), and η is the above learning rate, which can control the step size of weight adjustment. is the above sixth product, is the above seventh product, is the above eighth product. Combining the above adjustments, the complete form of the mixed reward function is R = α·(-e λ·d ) + β·b + γ·(-e μ·l ). By continuously performing dynamic updates of the weight coefficients through the partial derivatives and the learning rate, it is possible to guide the model to converge to the optimal policy faster. The partial derivatives provide gradient information about the reward function, while the learning rate controls the speed of weight updates. The combination of the two can help the model efficiently explore the network optimization strategy space and reduce the time required to reach the stable state. In addition, through the combination of the historical weights and the partial derivatives, the model can perform real-time policy adjustments according to changes in the network environment. Even when facing sudden network conditions or unpredictable load changes, it can quickly respond through the dynamic weight update mechanism and maintain the robustness and flexibility of the policy.

[0050] In an exemplary embodiment, adjusting the routing policy of the target cluster by updating the network model includes: recalculating the routing path when the number of nodes included in the target cluster changes; adjusting the bandwidth allocation ratio of the target cluster when the load of the nodes included in the target cluster changes.

[0051] In the above embodiments, the reinforcement learning agent can be enabled to quickly adapt to changes in the network environment through an adaptive learning mechanism, such as the addition or deletion of nodes, load fluctuations, etc. Through online learning, the routing policy can be continuously optimized to ensure the stability of system performance. That is, when a change in the network environment is detected, the online learning mechanism can be triggered, and the learning frequency and intensity can be dynamically adjusted according to the severity of the change, and the detected environmental change information can be transmitted to the reinforcement learning agent. During this process, the model parameters are continuously updated to adapt to the new network environment. Among them, during the online learning process, new network state data can be collected as the input of online learning, and the parameters of the reinforcement learning model are updated with the new data to make it adapt to the new network environment. The model update method that combines progressive update and experience replay mechanism can be used, that is, the model parameters are updated step by step to avoid performance fluctuations caused by one-time updates; the model parameters are updated by combining the old and new network state change data to avoid forgetting the data learned before. After online learning, the performance of the model can also be evaluated, including indicators such as latency, bandwidth utilization, and node load balancing. If the performance is improved, the updated model is applied to the actual routing policy; if the performance decreases, roll back to the previous model. The entire process forms a closed loop, that is, network state data can be collected during the monitoring stage to detect environmental changes; during the learning stage, when a change in the environment is detected, the online learning mechanism is triggered to update the model parameters; during the optimization stage, the routing policy can be adjusted according to the updated model to optimize network performance; during the evaluation stage, the optimized performance can be evaluated, and the parameters and policies of adaptive learning can be dynamically adjusted to ensure system stability. At the same time, a feedback mechanism can also be used to calculate network performance indicators in real time, including latency, bandwidth utilization, and load balancing conditions, generate feedback signals according to changes in the performance indicators, be used to adjust the model parameters and weight coefficients, and use the optimized performance indicators as the basis for subsequent monitoring and learning.

[0052] In the above embodiments, it can be dynamically adjusted according to the type of environmental change. For example, when the type of environmental change is the addition or deletion of nodes, the state space and action space related to nodes in the model can be updated, and the optimal routing path can be recalculated; when the type of environmental change is load fluctuation, the weight coefficient in the reward function can be adjusted to prioritize the optimization of load balancing, that is, the bandwidth allocation policy is adjusted to balance the node load.

[0053] In the above embodiments, when there are multiple objectives to be optimized, the weights of multi-objective optimization can be dynamically adjusted. For example, when nodes are added or deleted, latency and bandwidth utilization are preferentially optimized; when the load fluctuates, node load balancing is preferentially optimized. The Deep Q-Network (DQN) can be used as the infrastructure, including an input layer, multiple hidden layers (such as LSTM or CNN), and an output layer. Among them, the input layer design can use the comprehensive reward function as the reward signal of the model to guide the learning process of the model. By maximizing the comprehensive reward, multi-objective optimization is achieved; the output layer is designed for multi-objective output, corresponding to latency, bandwidth utilization, and node load balancing respectively. Each output corresponds to the optimization result of one objective. Through the multi-objective optimization mechanism, while optimizing latency and bandwidth utilization, an even distribution of node load can be achieved. This not only avoids the problem of node overload but also improves the overall stability and reliability of the system.

[0054] In the above embodiments, a verification and testing phase can be set up, that is, different network loads, topology changes, and fault scenarios can be simulated to test the performance of the model in various environments. In the simulated environment, performance metrics such as network latency, bandwidth utilization, node load, service response time, and packet loss rate are evaluated to verify the optimization effect of the reward function. According to the test results, the hyperparameters and learning rate in the reward function are adjusted to optimize the performance of the model. The model is deployed in the actual target cluster, and the network performance is monitored in real time to further verify and optimize the design of the reward function.

[0055] In the foregoing embodiments, the design of the reward function can be based on the optimization objectives, comprehensively considering factors such as latency reduction and bandwidth utilization improvement, and dynamically calculating the reward value to guide the model to converge to the optimal strategy. When a change in the environment is detected, the model parameters can be dynamically adjusted through online learning to ensure that the agent can quickly adapt to the new network environment and maintain the stability of the system performance. By dynamically adjusting the weights of each objective according to the change of the network environment, the model can find the optimal balance between different optimization objectives, enhancing the flexibility and adaptability of the system. In addition, by dynamically adjusting the bandwidth allocation strategy and making full use of network resources, the bandwidth utilization can be significantly improved. This effect is particularly prominent in high-load scenarios, which can effectively avoid bandwidth waste and improve network throughput.

[0056] In an exemplary embodiment, adjusting the model parameters of the target network model based on the reward value and the current network state includes: determining a fourth partial derivative of the reward function with respect to a first sub-state included in the network state, determining a first initial weight of the first sub-state, determining a ninth product of the learning rate and the fourth partial derivative, and determining a sum value of the first initial weight and the ninth product as the first updated weight; determining a fifth partial derivative of the reward function with respect to a second sub-state included in the network state, determining a second initial weight of the second sub-state, determining a tenth product of the learning rate and the fifth partial derivative, and determining a sum value of the second initial weight and the tenth product as the second updated weight; determining a sixth partial derivative of the reward function with respect to a third sub-state included in the network state, determining a third initial weight of the third sub-state, determining an eleventh product of the learning rate and the sixth partial derivative, and determining a sum value of the third initial weight and the eleventh product as the third updated weight; in a case where the first sub-state exceeds a first preset threshold, determining a first difference between the first sub-state and a delay threshold, determining a twelfth product of the first difference and an adjustment coefficient, determining a first sum value of the twelfth product and a target constant, and determining a product of the first sum value and the first updated weight as a first weight. In a case where the second sub-state exceeds a second preset threshold, determining a second difference between the second sub-state and a bandwidth utilization threshold, determining a thirteenth product of the second difference and the adjustment coefficient, determining a second sum value of the thirteenth product and the target constant, and determining a product of the second sum value and the second updated weight as a second weight. In a case where the third sub-state exceeds a third preset threshold, determining a third difference between the third sub-state and a node load balancing threshold, determining a fourteenth product of the third difference and the adjustment coefficient, determining a third sum value of the fourteenth product and the target constant, and determining a product of the third sum value and the third updated weight as a third weight.

[0057] In the above embodiment, a weight coefficient α d , β b , γ l (i.e., the above first initial weight, second initial weight, and third initial weight) can be initialized for each state feature, and the weight coefficient can be updated in real time according to the change of network performance using the online gradient descent algorithm.

[0058] In the above embodiment, taking delay as an example: the first updated weight coefficient can be calculated by the following formula: Where is the above fourth partial derivative, is the above ninth product. When the delay exceeds the first preset threshold (e.g., 100 ms), the weight can be updated, i.e., α = αd1 ·(1 + η·(d - t)), where t is the above-mentioned delay threshold.

[0059] In the above embodiments, taking the bandwidth utilization rate as an example: The second updated weight coefficient can be calculated by the following formula: where is the fifth partial derivative, is the above-mentioned tenth product. When the bandwidth utilization rate exceeds the second preset threshold (e.g., 80%), the weight can be updated, that is, β = β b1 ·(1 + η·(b - t1)), where t1 is the above-mentioned bandwidth utilization threshold.

[0060] In the above embodiments, taking the node load rate as an example: The third updated weight coefficient can be calculated by the following formula: where is the sixth partial derivative, is the above-mentioned eleventh product. When the degree of node load imbalance exceeds the third preset threshold (e.g., 20%), the weight can be updated, that is, γ = γ l1 ·(1 + η·(l - t2)), where t2 is the above-mentioned node load balancing threshold.

[0061] Through the description of the above implementation manners, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases, the former is a better implementation manner.

[0062] The embodiments of the present application further provide an adjustment device for a routing policy, Figure 5 is a structural block diagram of an adjustment device for a routing policy according to an embodiment of the present invention, as Figure 5 shown. The device includes:

[0063] A first determination module 502, configured to determine the current network state of the target cluster and the historical network state within a historical time period;

[0064] A second determination module 504, configured to determine a first state vector based on the current network state and the historical network state;

[0065] An input module 506, configured to input the first state vector into the target network model to obtain a reward value determined based on a reward function included in the target network model, where the reward function is constructed based on the weights, exponential functions, and linear functions of each sub-state included in the network state of the target cluster;

[0066] The first adjustment module 508 adjusts the model parameters of the target network model based on the reward value and the current network state to obtain an updated network model, where the model parameters include the weights of each sub-state;

[0067] The second adjustment module 510 is configured to adjust the routing policy of the target cluster through the updated network model.

[0068] In an exemplary embodiment, before the apparatus inputs the first state vector into the target network model to obtain a reward value determined based on the reward function included in the target network model: constructing a first exponential function based on the first sub-state included in the network state; constructing a linear function based on the second sub-state included in the network state; constructing a second exponential function based on the third sub-state included in the network state; determining the first weight of the first sub-state, the second weight of the second sub-state, and the third weight of the third sub-state; constructing a reward function based on the first exponential function, the linear function, the second exponential function, the first weight, the second weight, and the third weight.

[0069] In an exemplary embodiment, the apparatus can construct a reward function based on the first exponential function, the linear function, the second exponential function, the first weight, the second weight, and the third weight in the following manner: determining a first product of the first exponential function and the first weight; determining a second product of the linear function and the second weight; determining a third product of the second exponential function and the third weight; determining the sum value of the first product, the second product, and the third product as the reward function.

[0070] In an exemplary embodiment, the apparatus can also construct a first exponential function based on the first sub-state included in the network state in the following manner: determining a fourth product of the first sub-state and the first hyperparameter, determining the fourth product power of the natural constant to obtain a first function, and determining the negative of the first function as the first exponential function; constructing a second exponential function based on the third sub-state included in the network state, including: determining a fifth product of the third sub-state and the second hyperparameter, determining the fifth product power of the natural constant to obtain a second function, and determining the negative of the second function as the second exponential function.

[0071] In an exemplary embodiment, the first adjustment module 508 may adjust the model parameters of the target network model based on the reward value and the current network state in the following manner: when the reward value does not meet the predetermined condition, update the reward function based on the current network state; determine the first partial derivative of the reward function with respect to the first sub-state included in the network state; determine the second partial derivative of the reward function with respect to the second sub-state included in the network state; determine the third partial derivative of the reward function with respect to the third sub-state included in the network state; adjust the first weight of the first sub-state included in the reward function based on the first partial derivative; adjust the second weight of the second sub-state included in the reward function based on the second partial derivative; adjust the third weight of the third sub-state included in the reward function based on the third partial derivative.

[0072] In an exemplary embodiment, the first adjustment module 508 may adjust the first weight of the first sub-state included in the reward function based on the first partial derivative in the following manner: determine the first historical weight of the first sub-state, determine the learning rate of the target network model, determine the sixth product of the first partial derivative and the learning rate, and determine the sum value of the sixth product and the first historical weight as the first weight; adjusting the second weight of the second sub-state included in the reward function based on the second partial derivative includes: determining the second historical weight of the second sub-state, determining the learning rate of the target network model, determining the seventh product of the second partial derivative and the learning rate, and determining the sum value of the seventh product and the second historical weight as the second weight; adjusting the third weight of the third sub-state included in the reward function based on the third partial derivative includes: determining the third historical weight of the third sub-state, determining the learning rate of the target network model, determining the eighth product of the third partial derivative and the learning rate, and determining the sum value of the eighth product and the third historical weight as the third weight.

[0073] In an exemplary embodiment, the second adjustment module 510 may adjust the routing policy of the target cluster by updating the network model in the following manner: when the number of nodes included in the target cluster changes, recalculate the routing path; when the load of the nodes included in the target cluster changes, adjust the bandwidth allocation ratio of the target cluster.

[0074] For the description of the features in the corresponding embodiment of the routing policy adjustment device, reference may be made to the relevant description in the corresponding embodiment of the routing policy adjustment method, which will not be elaborated here one by one.

[0075] An embodiment of the present application further provides an electronic device, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any one of the above embodiments of the routing policy adjustment method.

[0076] Embodiments of the present application also provide a computer-readable storage medium storing a computer program, where the computer program is configured to execute the steps in the embodiments of any of the above routing policy adjustment methods when running.

[0077] In an exemplary embodiment, the above computer-readable storage medium may include, but is not limited to: various media that can store computer programs such as USB flash drives, read-only memories (ROM for short), random access memories (RAM for short), mobile hard disks, magnetic disks, or optical discs.

[0078] Embodiments of the present application also provide a computer program product. The above computer program product includes a computer program, and when the computer program is executed by a processor, it implements the steps in the embodiments of any of the above routing policy adjustment methods.

[0079] Embodiments of the present application also provide another computer program product, including a non-volatile computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, it implements the steps in the embodiments of any of the above routing policy adjustment methods.

[0080] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.

[0081] The above has introduced in detail a routing policy adjustment method, device, storage medium, and electronic device provided by the present application. Specific examples are used in this article to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application. It should be noted that for those of ordinary skill in the art in the technical field, without departing from the principle of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the protection scope of the claims of the present application.

Claims

1. A method for adjusting a routing strategy, characterized in that: include: Determine the current network status of the target cluster and the historical network status over a historical time period; Determine a first state vector based on the current network state and the historical network state; Inputting the first state vector into a target network model to obtain a reward value determined based on a reward function included in the target network model, wherein the reward function is constructed based on a weight, an exponential function, and a linear function of each sub-state included in the network state of the target cluster; Adjusting the model parameters of the target network model based on the reward value and the current network state to obtain an updated network model, wherein the model parameters include the weight of each of the sub-states; The routing strategy of the target cluster is adjusted by updating the network model.

2. The method for adjusting the routing strategy according to claim 1, characterized in that: Before inputting the first state vector into the target network model to obtain a reward value determined based on a reward function included in the target network model, the method further includes: constructing a first exponential function based on a first sub-state included in the network state; constructing a linear function based on a second sub-state included in the network state; constructing a second exponential function based on a third sub-state included in the network state; determining a first weight of the first substate, a second weight of the second substate, and a third weight of the third substate; The reward function is constructed based on the first exponential function, the linear function, the second exponential function, the first weight, the second weight, and the third weight.

3. The method for adjusting routing strategy according to claim 2, characterized in that: The constructing the reward function based on the first exponential function, the linear function, the second exponential function, the first weight, the second weight and the third weight comprises: determining a first product of the first exponential function and the first weight; determining a second product of the linear function and the second weight; determining a third product of the second exponential function and the third weight; A sum of the first product, the second product, and the third product is determined as the reward function.

4. The method for adjusting routing strategy according to claim 2, characterized in that: The constructing of the first exponential function based on the first sub-state included in the network state includes: determining a fourth product of the first sub-state and a first hyperparameter, determining a fourth product power of a natural constant to obtain a first function, and determining a negative number of the first function as the first exponential function; The constructing of the second exponential function based on the third substate included in the network state includes: determining the fifth product of the third substate and a second hyperparameter, determining the fifth power of the product of a natural constant to obtain a second function, and determining the negative of the second function as the second exponential function.

5. The method for adjusting routing strategy according to claim 1, characterized in that: The adjusting the model parameters of the target network model based on the reward value and the current network state includes: When the reward value does not satisfy a predetermined condition, updating the reward function based on the current network state; determining a first partial derivative of the reward function with respect to a first substate included in the network state; determining a second partial derivative of the reward function with respect to a second substate included in the network state; determining a third partial derivative of the reward function with respect to a third sub-state included in the network state; adjusting a first weight of the first sub-state included in the reward function based on the first partial derivative; adjusting a second weight of the second sub-state included in the reward function based on the second partial derivative; A third weight of the third sub-state included in the reward function is adjusted based on the third partial derivative.

6. The method for adjusting routing strategy according to claim 5, characterized in that: The adjusting the first weight of the first substate included in the reward function based on the first partial derivative comprises: determining a first historical weight of the first substate, determining a learning rate of the target network model, determining a sixth product of the first partial derivative and the learning rate, and determining a sum of the sixth product and the first historical weight as the first weight; The adjusting the second weight of the second substate included in the reward function based on the second partial derivative comprises: determining a second historical weight of the second substate, determining a learning rate of the target network model, determining a seventh product of the second partial derivative and the learning rate, and determining a sum of the seventh product and the second historical weight as the second weight; The step of adjusting the third weight of the third substate included in the reward function based on the third partial derivative comprises: determining a third historical weight of the third substate, determining a learning rate of the target network model, determining an eighth product of the third partial derivative and the learning rate, and determining the sum of the eighth product and the third historical weight as the third weight.

7. The method for adjusting routing strategy according to claim 1, characterized in that: The adjusting the routing strategy of the target cluster by updating the network model includes: When the number of nodes included in the target cluster changes, recalculating the routing path; When the load of the nodes included in the target cluster changes, the bandwidth allocation ratio of the target cluster is adjusted.

8. A routing strategy adjustment device, characterized in that: include: A first determination module is used to determine the current network status of the target cluster and the historical network status within a historical time period; A second determination module, configured to determine a first state vector based on the current network state and the historical network state; An input module, configured to input the first state vector into a target network model to obtain a reward value determined based on a reward function included in the target network model, wherein the reward function is constructed based on a weight, an exponential function, and a linear function of each sub-state included in the network state of the target cluster; A first adjustment module, which adjusts the model parameters of the target network model based on the reward value and the current network state to obtain an updated network model, wherein the model parameters include the weight of each of the sub-states; The second adjustment module is used to adjust the routing strategy of the target cluster by updating the network model.

9. An electronic device, characterized in that: include: Memory for storing computer programs; A processor, configured to implement the steps of the method for adjusting the routing policy as claimed in any one of claims 1 to 7 when executing the computer program.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, wherein the computer program, when executed by a processor, implements the steps of the method for adjusting the routing policy according to any one of claims 1 to 7.