Container arrangement method and device, computer equipment, storage medium and product
By using trained policy networks and value networks to optimize container orchestration strategies in a cloud-edge collaborative environment, the problem of low resource allocation efficiency in existing technologies is solved, and container scheduling effects with high resource utilization and low latency are achieved.
Patent Information
- Application Number
- CN202510666246.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-22
- Publication Date
- 2025-09-26
AI Technical Summary
Existing container orchestration technology has difficulty responding quickly to environmental changes in a cloud-edge collaborative environment, resulting in inefficient resource allocation, especially the lack of a dynamic policy adjustment mechanism in the collaborative scheduling across cloud-edge nodes.
Through the trained policy network and value network, the initial orchestration strategy is updated based on the current system status data of the cloud-edge collaborative environment, and the reinforcement learning algorithm is used to optimize the policy network parameters to generate the current target orchestration strategy to cope with dynamic changes and improve resource utilization and the robustness of container scheduling.
It achieves high resource utilization, low-latency response, and highly robust container scheduling in dynamic environments, improving the performance and reliability of container clusters.
Smart Images

Figure CN120704794A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of resource management and scheduling, and in particular to a container orchestration method, apparatus, computer equipment, storage medium, and product. Background Art
[0002] With the rapid development of emerging applications such as the Industrial Internet of Things and smart cities, edge computing scenarios are becoming increasingly prevalent. Cloud-edge collaborative architectures have become critical infrastructure for supporting low-latency, highly reliable services. Containerization technology, with its advantages of lightweight design, high flexibility, and rapid deployment, has been widely adopted in cloud-edge environments comprised of cloud centers and edge nodes. As the core technology for unified scheduling, management, and resource allocation of massive containerized applications in cloud-edge environments, container orchestration is crucial for improving overall system performance, resource utilization, and task execution efficiency.
[0003] Current container orchestration technology is implemented by configuring static rule-based scheduling strategies for different application scenarios. However, its strategies have weak generalization in coping with the complexity and dynamism of cloud-edge collaborative environments, making it difficult to respond quickly to environmental changes, resulting in low long-term resource allocation efficiency. Summary of the Invention
[0004] Based on this, it is necessary to provide a container orchestration method, device, computer equipment, storage medium and product that can improve resource allocation efficiency to address the above technical problems.
[0005] In a first aspect, the present application provides a container orchestration method, comprising:
[0006] Obtaining current system state data of the cloud-edge collaborative environment and a current initial orchestration policy corresponding to the current system state data; the current initial orchestration policy is used to indicate a preset container orchestration action that matches the current system state data;
[0007] Update the current initial orchestration strategy based on the trained strategy network and the current system state data to obtain a current target orchestration strategy;
[0008] The policy network is obtained through periodic iterative training. In each iterative training of each training cycle, historical system state data and a historical initial orchestration strategy corresponding to the historical system state data are input into the policy network to obtain a candidate orchestration strategy corresponding to the historical system state data. The candidate orchestration strategy of the historical system state data and the expected orchestration strategy of the historical system state data are used to optimize the parameters of the policy network. The orchestration action value of the expected orchestration strategy under the historical system state is greater than the orchestration action value of the candidate orchestration strategy under the historical system state. The orchestration action value is used to indicate at least one of container startup delay, resource fragmentation rate, and node load rate.
[0009] In one embodiment, the policy network, in each iteration of each training cycle:
[0010] Inputting historical system state data and historical initial orchestration strategies corresponding to the historical system state data into a strategy network to obtain candidate orchestration strategies corresponding to the historical system state data;
[0011] Obtaining the expected orchestration strategy corresponding to the historical system state data based on the historical system state data, the candidate orchestration strategy corresponding to the historical system state data, and a trained value network; wherein the trained value network includes a mapping relationship between system state, orchestration strategy, and orchestration action value;
[0012] Parameters of a policy network are optimized using the candidate orchestration strategies and the expected orchestration strategy.
[0013] In one embodiment, the value network is obtained by the following process:
[0014] Based on the historical system state data and the historical new system state data, obtaining a historical instant reward for executing the historical target orchestration strategy corresponding to the historical system state data under the historical system state; the historical new system state data refers to the system state after executing the historical target orchestration strategy;
[0015] Based on the historical new system state data and the historical new target orchestration strategy corresponding to the historical new system state data, obtaining a historical future reward for executing the historical target orchestration strategy corresponding to the historical system state data under the historical system state;
[0016] The value network is trained using the historical system state data, the historical target orchestration strategy, and the target orchestration action value; the target orchestration action value includes the historical immediate reward and the historical future reward.
[0017] In one embodiment, obtaining current system state data of the cloud-edge collaborative environment and a current initial orchestration policy corresponding to the current system state data includes:
[0018] Obtain the current system status data of the cloud-edge collaborative environment; the current system status data includes current computing resource status data, current network status data, current container operation indicator data and current task queue load data corresponding to the cloud center node and edge node in the cloud-edge collaborative environment respectively;
[0019] Based on the current computing resource status data, the current network status data, the current container operation indicator data, and the current task queue load data, a preset container orchestration model is called to obtain the current initial orchestration strategy corresponding to the current system status data; the container orchestration model includes a container orchestration target optimization function, and the container orchestration target optimization function is used to indicate at least one of a container startup delay, a resource fragmentation rate, and a node load rate.
[0020] In one embodiment, the current system state data further includes current weight data corresponding to the optimization target in the container orchestration target optimization function;
[0021] Among them, the current weight data is obtained by obtaining the historical actual system performance value after the execution of the historical target orchestration strategy and the target system performance value in the preset service level agreement, and adjusting the historical weight data based on the difference between the historical actual system performance value and the target system performance value in the service level agreement.
[0022] In one embodiment, the current system status data includes current computing resource status data and current network status data of the edge node in the cloud-edge collaborative environment; the method further includes:
[0023] When the current computing resource status data or the current network status data of the edge node meets a preset abnormal condition, some containers of the edge node are controlled to migrate to an adjacent edge node or cloud center node of the edge node.
[0024] In a second aspect, the present application further provides a container orchestration device, comprising:
[0025] An acquisition unit, configured to acquire current system state data of the cloud-edge collaborative environment and a current initial orchestration policy corresponding to the current system state data; the current initial orchestration policy is used to indicate a pre-set container orchestration action that matches the current system state data;
[0026] An orchestration unit, configured to update the current initial orchestration strategy based on the trained strategy network and the current system state data to obtain a current target orchestration strategy;
[0027] The policy network is obtained through periodic iterative training. In each iterative training of each training cycle, historical system state data and a historical initial orchestration strategy corresponding to the historical system state data are input into the policy network to obtain a candidate orchestration strategy corresponding to the historical system state data. The candidate orchestration strategy of the historical system state data and the expected orchestration strategy of the historical system state data are used to optimize the parameters of the policy network. The orchestration action value of the expected orchestration strategy under the historical system state is greater than the orchestration action value of the candidate orchestration strategy under the historical system state. The orchestration action value is used to indicate at least one of container startup delay, resource fragmentation rate, and node load rate.
[0028] In a third aspect, the present application further provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the container orchestration method provided in the first aspect of the present application is implemented.
[0029] In a fourth aspect, the present application further provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the container orchestration method provided in the first aspect of the present application.
[0030] In a fifth aspect, the present application also provides a computer program product, including a computer program, which, when executed by a processor, implements the container orchestration method provided in the first aspect of the present application.
[0031] The above-mentioned container orchestration method, apparatus, computer equipment, storage medium and product obtain the current system state data of the cloud-edge collaborative environment and the current initial orchestration policy corresponding to the current system state data, and update the current initial orchestration policy based on the trained policy network and the current system state data to obtain the current target orchestration policy. The policy network is obtained through periodic iterative training; in each iterative training of each training cycle, the historical system state data and the historical initial orchestration policy corresponding to the historical system state data are input into the policy network to obtain the candidate orchestration policy corresponding to the historical system state data, and the candidate orchestration policy of the historical system state data and the expected orchestration policy of the historical system state data are used to optimize the parameters of the policy network; the orchestration action value of the expected orchestration policy in the historical system state is greater than the orchestration action value of the candidate orchestration policy in the historical system state; the orchestration action value is used to indicate at least one of the container startup delay, resource fragmentation rate and node load rate. This application generates a preliminary orchestration strategy through the current snapshot of the environment, and then learns from actual historical operating experience through the policy network. It can master how to make more refined and robust decisions under various dynamic changes. It is not limited to a static optimal solution given by the preliminary strategy, but according to the real-time changing situation, it fine-tunes, corrects or even boldly deviates from the preliminary strategy to effectively respond to instantaneous changes in node resources, fluctuations in network conditions and sudden task loads, achieve high resource utilization, low-latency response and high-robustness container scheduling, and improve the performance and reliability of container clusters. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following briefly introduces the drawings required for use in the embodiments or related technical descriptions. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0033] Figure 1 This is an application environment diagram of a container orchestration method in one embodiment;
[0034] Figure 2 Schematic diagram of a process of a container orchestration method in one embodiment;
[0035] Figure 3 Schematic diagram of a process of a container orchestration method according to another embodiment;
[0036] Figure 4 1 is a flow chart of a policy network training process in one embodiment;
[0037] Figure 51 is a flow chart of a value network training process in one embodiment;
[0038] Figure 6 is a structural block diagram of a container orchestration device in one embodiment;
[0039] Figure 7 FIG. 1 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION
[0040] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.
[0041] Current container orchestration technology implements static, rule-based scheduling policies configured for different application scenarios. These policies primarily focus on resource requests or limits, node affinity, instance affinity, priority, and preemption strategies. These scheduling policies are based on single or combined static rules. In scenarios with heterogeneous resource distribution, dynamic network environments, and complex task requirements, these scheduling policies lack generalizability and are unable to quickly respond to environmental changes (such as sudden task queues and resource preemption). In particular, in collaborative scheduling across cloud-edge nodes, the lack of a dynamic policy adjustment mechanism based on reinforcement learning leads to inefficient long-term resource allocation.
[0042] The container orchestration method provided in the embodiment of the present application can be applied to Figure 1 In the application environment shown, the cloud-edge collaborative environment includes cloud center nodes and edge nodes, which communicate with servers through the network. The server can be implemented as a standalone server or a server cluster consisting of multiple servers.
[0043] In an exemplary embodiment, Figure 2 As shown, a container orchestration method is provided, which is applied to Figure 1 The server in FIG. 1 is used as an example to illustrate the method, including the following steps 202 to 206. Among them:
[0044] Step 202: Obtain current system status data of the cloud-edge collaborative environment and a current initial orchestration strategy corresponding to the current system status data.
[0045] The current initial orchestration policy is used to indicate a pre-set container orchestration action that matches the current system state data. The container orchestration action may include container deployment location, container resource allocation, and container migration.
[0046] The system status data of the embodiment of the present application may include cloud edge node attribute data, cloud edge node real-time status data, cloud edge network status data, container and task requirement data.
[0047] For example, the server can be pre-configured with matching rules that represent system state data and the initial orchestration policy. When the system state changes, a new task arrives, or according to a predetermined orchestration frequency, the server obtains the current system state data of the cloud-edge collaborative environment and determines the current initial orchestration policy that matches the current system state data based on the matching rules.
[0048] Step 204: Update the current initial orchestration strategy based on the trained strategy network and the current system state data to obtain a current target orchestration strategy.
[0049] The current target orchestration strategy is used to indicate the container orchestration action that the server is about to perform.
[0050] The policy network in this embodiment is obtained through periodic iterative training. When preset trigger conditions are met, such as sufficient data collection or reaching a certain time point, the server controls the policy network to update using historical data. During each training cycle, the server obtains a training sample set based on historical data and uses this training sample set to iterate the policy network multiple times.
[0051] Exemplarily, the server trains and updates the policy network at a preset first time interval to obtain an updated policy network. The server iteratively trains the policy network multiple times within each training cycle. During each iterative training of each training cycle, the server inputs historical system state data and a historical initial orchestration policy corresponding to the historical system state data into the policy network to obtain a candidate orchestration policy corresponding to the historical system state data.
[0052] Understandably, the candidate orchestration strategy does not refer to the historical target orchestration strategy that the system has actually executed under the historical system state data. Rather, when training the current policy network, the server collects historical system state data and historical initial orchestration strategies, inputs them into the current policy network, and obtains the actions that should be executed under the historical system state data output by the current policy network, namely the candidate orchestration strategy.
[0053] In this embodiment, the candidate orchestration strategy and the expected orchestration strategy based on the historical system state data are used to optimize the parameters of the policy network. The orchestration action value of the expected orchestration strategy under the historical system state is greater than the orchestration action value of the candidate orchestration strategy under the historical system state; the orchestration action value is used to indicate at least one of container startup latency, resource fragmentation rate, and node load rate.
[0054] In the above container orchestration method, by obtaining the current system state data of the cloud-edge collaborative environment and the current initial orchestration strategy corresponding to the current system state data, the current initial orchestration strategy is updated based on the trained policy network and the current system state data to obtain the current target orchestration strategy. The policy network is obtained through periodic iterative training; in each iterative training of each training cycle, the historical system state data and the historical initial orchestration strategy corresponding to the historical system state data are input into the policy network to obtain the candidate orchestration strategy corresponding to the historical system state data, and the candidate orchestration strategy of the historical system state data and the expected orchestration strategy of the historical system state data are used to optimize the parameters of the policy network; the orchestration action value of the expected orchestration strategy in the historical system state is greater than the orchestration action value of the candidate orchestration strategy in the historical system state; the orchestration action value is used to indicate at least one of the container startup delay, resource fragmentation rate and node load rate. The embodiment of the present application generates a preliminary orchestration strategy through the current snapshot of the environment, and then learns from actual historical operating experience through the policy network, which can master how to make more refined and robust decisions under various dynamic changes. It is not limited to a static optimal solution given by the preliminary strategy, but according to the real-time changing situation, it fine-tunes, corrects or even boldly deviates from the preliminary strategy to effectively respond to instantaneous changes in node resources, fluctuations in network conditions and sudden task loads, achieve high resource utilization, low-latency response and high-robustness container scheduling, and improve the performance and reliability of the container cluster.
[0055] In an exemplary embodiment, another container orchestration method is provided, such as Figure 3 As shown, the container orchestration method includes steps 302 to 306. In which:
[0056] Step 302: Obtain the current system status data of the cloud-edge collaborative environment.
[0057] Among them, the current system status data includes the current computing resource status data, current network status data, current container operation indicator data and current task queue load data corresponding to the cloud center node and edge node in the cloud-edge collaborative environment.
[0058] For example, the server can be configured with a resource monitoring module that collects real-time computing resource status, network status, container operation indicators, and task queue load data from cloud center nodes and edge nodes. The computing resource status includes the real-time utilization of the node's CPU, memory, and storage; the network status includes network latency, bandwidth fluctuations, and packet loss rate of the cloud-edge link; the container operation indicators include at least the container startup time and the computing resources required by the container; and the task queue load data includes at least the task priority and task response time.
[0059] Step 304: Based on the current computing resource status data, the current network status data, the current container operation indicator data, and the current task queue load data, a preset container orchestration model is called to obtain a current initial orchestration strategy corresponding to the current system status data.
[0060] The container orchestration model includes a container orchestration objective optimization function, which is used to indicate at least one of container startup latency, resource fragmentation rate, and node load rate. For example, the container orchestration objective optimization function includes multiple optimization objectives such as container startup latency, resource fragmentation rate, and node load rate, and container orchestration actions such as container deployment location, container resource allocation, and container migration are decision variables of the container orchestration objective optimization function.
[0061] For example, a dynamic scheduling module can be configured within the server. Based on data collected by the resource monitoring module, the server generates the current initial orchestration strategy through a multi-objective optimization algorithm, enabling elastic resource allocation across cloud edge nodes and dynamic migration of container instances. The dynamic scheduling module's multi-objective optimization algorithm aims to minimize resource fragmentation, balance node load, and reduce container response latency.
[0062] For example, the resource fragmentation rate is calculated based on the computing resource status.
[0063] For example, the node load rate is calculated by weighting the collected computing resources, and is expressed as follows:
[0064] ;
[0065] in, For nodes The load factor, is the CPU usage, is the memory usage, The storage usage rate.
[0066] For example, the container response delay is determined by the container startup time and the network status.
[0067] For example, the container orchestration objective optimization function of the multi-objective optimization algorithm is expressed as follows:
[0068] ;
[0069] in, is the objective function, is the total number of nodes in the cloud edge, is the total number of containers in the cloud edge, For nodes The total resource capacity, is the amount of resources used, a set of response delays for the container, Indicates the maximum container response delay, For nodes The load factor, is the average load rate of the cloud edge, 、 、 is the weight coefficient.
[0070] In other embodiments, the dynamic scheduling module deployment strategy includes prioritizing delay-sensitive tasks to edge nodes, allocating computationally intensive tasks to cloud center nodes, and optimizing data transmission paths using overlay networks for cross-node container communications.
[0071] Step 306: Update the current initial orchestration strategy based on the trained strategy network and the current system state data to obtain a current target orchestration strategy.
[0072] The policy network includes the mapping relationship between system state data, initial orchestration policy and target orchestration policy.
[0073] Exemplarily, the current system state data and the previous initial orchestration strategy are input into a trained strategy network, and the trained strategy network is used to output a current target orchestration strategy to be executed under the current system state data.
[0074] Step 308: execute the current target orchestration strategy, and store the corresponding relationship between the current system state data, the current initial orchestration strategy, the current target orchestration strategy, and the current new system state data after executing the current target orchestration strategy.
[0075] Among them, the current new system state data is used to indicate the new state that the system transfers to after the server schedules and manages the containers in the cloud-edge collaborative environment according to the current target orchestration strategy.
[0076] In an optional embodiment, the server is based on the current system state data s at time step t t , determine the current system status data s t The current target orchestration strategy a to be executed t , the server controls the execution of the current target orchestration strategy a t After that, the system state is transferred to s t+1 , the server is based on the current new system status data s t+1 and current system status data s t , determine the current system status data s t Execute the current target orchestration strategy a t The instant reward r st , will s t、a t 、s t+1 、r st The corresponding relationship is stored in the experience replay pool.
[0077] The server of the embodiment of the present application repeats the above steps 302 to 306 when the system status changes, a new task arrives, or according to a predetermined scheduling frequency. It continuously perceives the environment, outputs decisions, executes decisions, obtains feedback, and stores experience.
[0078] In this example, multi-objective optimization provides a relatively efficient and robust initial solution for static snapshots, avoiding the inefficiencies of starting from scratch with the reinforcement learning policy network. Reinforcement learning then builds on this foundation by making dynamic and precise adjustments, improving both actual performance and long-term optimization in dynamic environments.
[0079] The following further illustrates the training process of the policy network in the embodiments of the present application.
[0080] In an exemplary embodiment, Figure 4 As shown in Figure 2, the policy network includes the following steps in each iteration of each training cycle:
[0081] Step 402 : Input historical system status data and historical initial orchestration strategies corresponding to the historical system status data into a strategy network to obtain candidate orchestration strategies corresponding to the historical system status data.
[0082] Exemplarily, the server collects historical system state data and the historical initial orchestration strategies corresponding to the historical system state data from the experience replay pool. The historical system state data and the historical initial orchestration strategies are input into the initial policy network (i.e., the policy network before this iteration of training), and the server outputs candidate orchestration strategies corresponding to the historical system state data.
[0083] Step 404 : Based on the historical system status data, the candidate orchestration strategies corresponding to the historical system status data, and the trained value network, obtain the expected orchestration strategy corresponding to the historical system status data.
[0084] Exemplarily, the trained value network includes mappings between system states, orchestration policies, and orchestration action values. The server uses the trained value network to determine a first gradient of the orchestration action value for executing a candidate orchestration policy under the historical system state. This first gradient represents the target change direction or trend of the candidate orchestration policy to maximize the orchestration action value under the historical system state. In embodiments of the present application, this first gradient can be used to represent the expected orchestration policy.
[0085] Step 406: Optimize parameters of the policy network using the candidate orchestration strategy and the expected orchestration strategy.
[0086] Exemplarily, the server calculates the second gradient of the initial policy network with respect to its parameters, where the second gradient indicates how a change in the policy network's parameters affects the action it outputs. Based on the first and second gradients, the server uses the chain rule to optimize the policy network's parameters.
[0087] In this embodiment, a policy network is used to learn an optimal policy so that the policy network outputs a target orchestration policy under different system states, which can maximize the expected cumulative reward. The container orchestration policy is dynamically adjusted using a reinforcement learning algorithm, and the container startup delay, resource fragmentation rate, and node load rate are optimized in combination with historical task execution data.
[0088] The following further illustrates the training process of the value network in the embodiment of the present application.
[0089] In an exemplary embodiment, Figure 5 As shown, the value network of the embodiment of the present application can also be obtained through periodic iterative training. For example, the server trains and updates the value network at a preset second time interval to obtain an updated value network. Each iterative training in each training cycle includes the following steps:
[0090] Step 502 : Based on the historical system state data and the historical new system state data, obtain a historical instant reward for executing the historical target orchestration strategy corresponding to the historical system state data under the historical system state.
[0091] The historical new system state data refers to the system state after executing the historical target orchestration strategy.
[0092] For example, the historical instant rewards in the embodiments of the present application can be calculated and stored in the experience replay pool after executing the historical target orchestration strategy. When the server trains and updates the value network according to a preset period, it obtains historical system state data, the historical new system state data corresponding to the historical system state data, and the corresponding historical instant rewards from the experience replay pool.
[0093] The historical instant rewards of the embodiments of the present application can be calculated based on a reward function, which may include cloud-edge resource utilization, container response delay, cloud-edge node load rate, network status, migration offset, etc. Specifically, based on the cloud-edge resource utilization, container response delay, cloud-edge node load rate, network status, and migration offset corresponding to the historical system state data and the historical new system state data, the historical instant rewards for executing the historical target orchestration strategy under the historical system state are obtained.
[0094] Step 504 : Based on the historical new system state data and the historical new target orchestration strategy corresponding to the historical new system state data, obtain a historical future reward for executing the historical target orchestration strategy corresponding to the historical system state data under the historical system state.
[0095] Exemplarily, the server uses the target policy network and historical new system state data to obtain a historical new target orchestration policy corresponding to the historical new system state data. This historical new target orchestration policy is used to represent the container orchestration action that should be taken under the new system state, as provided by the target policy network. The server inputs the historical new system state data and the historical new target orchestration policy into the target value network, and uses the target value network to output a new orchestration action value for executing the historical new target orchestration policy under the historical new system state. This new orchestration action value is used to represent the historical future reward for executing the historical target orchestration policy under the historical system state.
[0096] It should be noted that to improve training stability, the target policy network and the target value network can be replicas of the policy network and the value network, respectively. Taking the target value network as an example, the target value network can perform parameter updates at a third time interval that is greater than the second time interval. For example, rather than performing gradient updates at every training iteration like the value network, the target value network replicates the parameters of the value network at a lower update frequency. This allows the parameters of the target value network to change much more slowly than those of the value network, providing a more stable reference for value assessment.
[0097] Step 506: Train the value network using the historical system state data, the historical target orchestration strategy, and the target orchestration action value.
[0098] The target choreography action value includes the historical immediate reward and the historical future reward.
[0099] Exemplarily, the historical system state data and the historical target orchestration strategy are input into the initial value network, the value network is used to output an estimated orchestration action value, and then the estimated orchestration action value and the target orchestration action value are combined to perform parameter tuning on the value network.
[0100] In an exemplary embodiment, the reinforcement learning algorithm of the embodiment of the present application adopts deep deterministic policy gradient (DDPG), which includes the following parts: state space, action space, reward function, policy network, value network and soft update of target network.
[0101] Wherein, the state space is expressed as:
[0102] ;
[0103] in, is the state space, The computing resources required for the container, The remaining computing resources of the cloud edge nodes, The cloud-edge network state; the state space serves as the input of the policy network and the value network to generate actions and evaluate values;
[0104] Furthermore, the action space is expressed as:
[0105] ;
[0106] in, choreographing an action space for said container, is the selection weight of N nodes; They are the allocation ratios of cpu, memory and storage respectively; a load threshold for triggering migration and a response delay threshold for the container; The migration target offset is used to represent the container migration direction in two-dimensional coordinates, which is used to guide container deployment and migration.
[0107] Furthermore, the reward function is expressed as follows:
[0108] ;
[0109] in, is the reward function, edge cloud resource utilization; is the delay decay term, is the attenuation coefficient; is the cloud edge node load rate, ; is the network status; is the migration offset; 、 、 、 and is the dynamically adjusted reward weight.
[0110] Furthermore, the policy network parameters are updated by the gradient ascent method, which is expressed as:
[0111]
[0112]
[0113] in, are the parameters of the policy network; Strategy objective function right The gradient of ; s is the current state; a is the action, represents the probability distribution of the policy network generating action a in a given state s; choreographing action values for the container output by the value network; An experience replay pool that stores historical computing resource data; The gradient of the value network to the action, representing the action the impact on expected returns; is the gradient of the policy network to the parameter, indicating the parameter Impact on container orchestration action generation.
[0114] Furthermore, the parameters of the value network Update by minimizing the Bellman error, which is expressed as follows:
[0115]
[0116]
[0117]
[0118] in, is the Bellman error; is the current state; For the current action; For the next state; is the target, which represents the estimated value of future returns; Arranging action values for the container generated by the value network for calculating target values; is the target value network; for the target strategy network; is a discount factor used to balance the importance of immediate rewards and future rewards; is the immediate reward, calculated by the reward function.
[0119] Furthermore, the target network parameters are gradually adjusted through soft updates, as shown below:
[0120]
[0121]
[0122] in, is the value network parameter, used to calculate the target value y; The policy network parameters are used to generate the target container orchestration action; is the soft update coefficient, which controls the update speed of the target network parameters.
[0123] In an exemplary embodiment, the current system state data further includes current weight data corresponding to the optimization target in the container orchestration target optimization function.
[0124] Among them, the current weight data is obtained by obtaining the historical actual system performance value after the execution of the historical target orchestration strategy and the target system performance value in the preset service level agreement, and adjusting the historical weight data based on the difference between the historical actual system performance value and the target system performance value in the service level agreement.
[0125] Exemplarily, the server may be configured with a feedback adjustment module, which adopts a closed-loop control mechanism, ie, dynamically adjusts the weight coefficients of the multi-objective optimization algorithm according to the service level agreement.
[0126] The weight coefficients include 、 、 , and the adjustment formula is:
[0127]
[0128]
[0129]
[0130] in, is the learning rate, ranging from 0.01 to 0.1, is the actual resource fragmentation rate, is the target resource fragmentation rate; The actual maximum container response delay, is the container response delay threshold; is the actual load rate, is the target load factor.
[0131] In this embodiment, the cloud-edge scenario needs to simultaneously optimize multi-dimensional objectives such as resource fragmentation rate, load rate, and task delay. Existing methods often use fixed weight coefficients, which lack the ability to adaptively adjust to real-time load, network status, and service level agreements. This is especially prone to local optimality problems in the event of sudden traffic or node failures. This embodiment iteratively updates the multi-objective optimization weight parameters based on preset target indicators and service level agreements to form a closed-loop control mechanism. The preset target indicators include the target resource fragmentation rate, target delay threshold, and target load rate.
[0132] In an exemplary embodiment, the current system status data includes the current computing resource status data and current network status data of the edge node in the cloud-edge collaborative environment; the method also includes: when the current computing resource status data or the current network status data of the edge node meets a preset abnormal condition, controlling part of the containers of the edge node to migrate to an adjacent edge node or cloud center node of the edge node.
[0133] Exemplarily, the server can be configured with an edge collaboration module, and the task migration mechanism of the edge collaboration module includes migrating the container instance to the adjacent edge node or cloud center when it is detected that the edge node resources are overloaded or the network delay exceeds the threshold; when the target node resources of the adjacent edge node or cloud center are insufficient, dynamically releasing low-priority container resources according to the task priority; and using incremental data synchronization and breakpoint resumption technology during the migration process to reduce service interruption time.
[0134] Furthermore, the triggering conditions for task migration are:
[0135] ;
[0136] in, is the edge node load rate, is the upper threshold of node load, is the network delay, is the network delay threshold, is the packet loss rate, is the packet loss rate threshold.
[0137] In this embodiment, the edge network experiences frequent fluctuations and a high node failure rate. Existing container migration mechanisms often rely on threshold triggering, lacking coordinated optimization of migration paths, data synchronization efficiency, and service continuity. For example, improper selection of migration target nodes can lead to secondary congestion, and an imperfect incremental synchronization strategy can prolong service interruptions. This embodiment triggers a container task migration strategy and a localized rescheduling mechanism in the event of network fluctuations or edge node failures.
[0138] Compared with the prior art, the advantages of the embodiments of the present application are:
[0139] (1) Flexible resource scheduling based on dynamic weights: Existing technologies often use static resource allocation strategies, which are difficult to cope with the dynamic changes of heterogeneous cloud-edge resources. This invention uses dynamic weight calculation and multi-objective optimization algorithms to perceive node resource utilization, network status, and task priority in real time, generating a flexible deployment strategy. This improves resource utilization by more than 20% while reducing resource fragmentation.
[0140] (2) Adaptive optimization for multiple objectives: Traditional methods rely on a single scheduling rule (such as load balancing or latency priority), which makes it difficult to take into account multiple objectives. This paper designs a reinforcement learning engine based on the DDPG algorithm. It achieves multi-objective collaborative optimization through a comprehensive reward function (covering resource utilization, latency, load balancing, network stability, and migration overhead) and a dynamic weight adjustment mechanism.
[0141] (3) Intelligent disaster recovery and migration mechanism based on cloud-edge collaboration: Existing technologies are not well-suited to network fluctuations and node failures, and their migration strategies are crude and simplistic. This invention introduces an edge collaboration module that triggers a migration strategy based on localized rescheduling and incremental data synchronization when the network packet loss rate is greater than 15% or the latency exceeds a threshold. At the same time, it optimizes the migration path by using the migration target offset to reduce bandwidth usage.
[0142] (4) Closed-loop feedback-driven learning capability: Traditional systems lack real-time feedback and adaptive adjustment capabilities. This invention uses a user feedback module to dynamically update the policy network parameters and multi-objective weight coefficients of the reinforcement learning model, forming a closed-loop optimization. This allows the system to maintain stable performance in a dynamic environment, and the SLA (Service Level Agreement) compliance rate is increased to over 95%.
[0143] It should be understood that, although the various steps in the flowcharts involved in the various embodiments described above are displayed in sequence according to the instructions of the arrows, these steps are not necessarily executed in sequence in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and these steps can be executed in other orders. Moreover, at least a portion of the steps in the flowcharts involved in the various embodiments described above can include multiple steps or multiple stages, and these steps or stages are not necessarily executed and completed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a portion of steps or stages in other steps.
[0144] It should be understood that the term "based on" as used herein is used to describe one or more factors that influence a determination, and does not exclude other factors that may influence the determination. For example, the phrase "determine A based on B" means that the determination of A may be based entirely or at least partially on factor B. In other words, B is a factor that influences the determination of A, but does not exclude the determination of A being based on C as well.
[0145] Based on the same inventive concept, embodiments of the present application also provide a container orchestration device for implementing the aforementioned container orchestration method. The solution provided by this device is similar to the solution described in the aforementioned method. Therefore, the specific limitations in one or more container orchestration device embodiments provided below can be found in the above-described limitations on the container orchestration method and will not be further elaborated here.
[0146] In an exemplary embodiment, Figure 6 As shown, a container orchestration device is provided, including: an acquisition unit 602 and an orchestration unit 604, wherein:
[0147] The acquisition unit 602 is used to obtain the current system state data of the cloud-edge collaborative environment and the current initial orchestration policy corresponding to the current system state data; the current initial orchestration policy is used to indicate a preset container orchestration action that matches the current system state data.
[0148] The orchestration unit 604 is configured to update the current initial orchestration strategy based on the trained strategy network and the current system state data to obtain a current target orchestration strategy.
[0149] The policy network is obtained through periodic iterative training. In each iterative training of each training cycle, historical system state data and a historical initial orchestration strategy corresponding to the historical system state data are input into the policy network to obtain a candidate orchestration strategy corresponding to the historical system state data. The candidate orchestration strategy of the historical system state data and the expected orchestration strategy of the historical system state data are used to optimize the parameters of the policy network. The orchestration action value of the expected orchestration strategy under the historical system state is greater than the orchestration action value of the candidate orchestration strategy under the historical system state. The orchestration action value is used to indicate at least one of container startup delay, resource fragmentation rate, and node load rate.
[0150] Each module in the container orchestration device described above can be implemented in whole or in part through software, hardware, or a combination thereof. Each module can be embedded in or independent of a processor in a computer device in hardware form, or can be stored in a computer device memory in software form, so that the processor can call and execute the corresponding operations of each module.
[0151] In an exemplary embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as shown in FIG. Figure 7As shown. The computer device includes a processor, a memory, an input / output interface (I / O) and a communication interface. The processor, memory and input / output interface are connected via a system bus, and the communication interface is connected to the system bus via the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal via a network connection. When the computer program is executed by the processor, a container orchestration method is implemented.
[0152] Those skilled in the art will understand that Figure 7 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.
[0153] In an exemplary embodiment, a computer device is provided, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the container orchestration method provided in the first aspect of the embodiment of the present application.
[0154] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the container orchestration method provided in the first aspect of the embodiment of the present application is implemented.
[0155] In one embodiment, a computer program product is provided, including a computer program, which, when executed by a processor, implements the container orchestration method provided in the first aspect of the embodiment of the present application.
[0156] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant regulations.
[0157] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the above-mentioned embodiments. In particular, any reference to memory, database, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The databases involved in the various embodiments provided herein may include at least one of a relational database and a non-relational database. Non-relational databases may include, but are not limited to, distributed databases based on blockchains. The processors involved in the various embodiments provided herein may be, but are not limited to, general-purpose processors, central processing units (CPUs), graphics processing units (GPUs), digital signal processors (DSPs), programmable logic devices (PLDs), data processing logic devices based on quantum computing, and the like.
[0158] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0159] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present application. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present application, and these modifications and improvements fall within the scope of protection of the present application. Therefore, the scope of protection of the present application shall be determined by the appended claims.
Claims
1. A container orchestration method, characterized in that: The method comprises: Obtaining current system state data of the cloud-edge collaborative environment and a current initial orchestration policy corresponding to the current system state data; the current initial orchestration policy is used to indicate a preset container orchestration action that matches the current system state data; Update the current initial orchestration strategy based on the trained strategy network and the current system state data to obtain a current target orchestration strategy; The policy network is obtained through periodic iterative training. In each iterative training of each training cycle, historical system state data and a historical initial orchestration strategy corresponding to the historical system state data are input into the policy network to obtain a candidate orchestration strategy corresponding to the historical system state data. The candidate orchestration strategy of the historical system state data and the expected orchestration strategy of the historical system state data are used to optimize the parameters of the policy network. The orchestration action value of the expected orchestration strategy under the historical system state is greater than the orchestration action value of the candidate orchestration strategy under the historical system state. The orchestration action value is used to indicate at least one of container startup delay, resource fragmentation rate, and node load rate.
2. The method according to claim 1, characterized in that The policy network performs the following training in each iteration of each training cycle: Inputting historical system state data and historical initial orchestration strategies corresponding to the historical system state data into a strategy network to obtain candidate orchestration strategies corresponding to the historical system state data; Obtaining the expected orchestration strategy corresponding to the historical system state data based on the historical system state data, the candidate orchestration strategy corresponding to the historical system state data, and the trained value network; The trained value network includes mapping relationships between system states, orchestration strategies, and orchestration action values; Parameters of a policy network are optimized using the candidate orchestration strategies and the expected orchestration strategy.
3. The method according to claim 2, characterized in that The value network is obtained through the following process: Based on the historical system state data and the historical new system state data, obtaining a historical instant reward for executing the historical target orchestration strategy corresponding to the historical system state data under the historical system state; the historical new system state data refers to the system state after executing the historical target orchestration strategy; Based on the historical new system state data and the historical new target orchestration strategy corresponding to the historical new system state data, obtaining a historical future reward for executing the historical target orchestration strategy corresponding to the historical system state data under the historical system state; Training the value network using the historical system state data, the historical target orchestration strategy, and the target orchestration action value; The target orchestration action value includes the historical immediate reward and the historical future reward.
4. The method according to claim 1, wherein The obtaining of the current system state data of the cloud-edge collaborative environment and the current initial orchestration strategy corresponding to the current system state data includes: Obtain the current system status data of the cloud-edge collaborative environment; the current system status data includes current computing resource status data, current network status data, current container operation indicator data and current task queue load data corresponding to the cloud center node and edge node in the cloud-edge collaborative environment respectively; Based on the current computing resource status data, the current network status data, the current container operation indicator data, and the current task queue load data, a preset container orchestration model is called to obtain the current initial orchestration strategy corresponding to the current system status data; the container orchestration model includes a container orchestration target optimization function, and the container orchestration target optimization function is used to indicate at least one of a container startup delay, a resource fragmentation rate, and a node load rate.
5. The method according to claim 4, characterized in that The current system state data also includes current weight data corresponding to the optimization target in the container orchestration target optimization function; Among them, the current weight data is obtained by obtaining the historical actual system performance value after the execution of the historical target orchestration strategy and the target system performance value in the preset service level agreement, and adjusting the historical weight data based on the difference between the historical actual system performance value and the target system performance value in the service level agreement.
6. The method according to claim 1, wherein The current system status data includes current computing resource status data and current network status data of the edge node in the cloud-edge collaborative environment; the method further includes: When the current computing resource status data or the current network status data of the edge node meets a preset abnormal condition, some containers of the edge node are controlled to migrate to an adjacent edge node or cloud center node of the edge node.
7. A container arrangement device, characterized in that: The device comprises: An acquisition unit, configured to acquire current system state data of the cloud-edge collaborative environment and a current initial orchestration policy corresponding to the current system state data; the current initial orchestration policy is used to indicate a pre-set container orchestration action that matches the current system state data; An orchestration unit, configured to update the current initial orchestration strategy based on the trained strategy network and the current system state data to obtain a current target orchestration strategy; The policy network is obtained through periodic iterative training. In each iterative training of each training cycle, historical system state data and a historical initial orchestration strategy corresponding to the historical system state data are input into the policy network to obtain a candidate orchestration strategy corresponding to the historical system state data. The candidate orchestration strategy of the historical system state data and the expected orchestration strategy of the historical system state data are used to optimize the parameters of the policy network. The orchestration action value of the expected orchestration strategy under the historical system state is greater than the orchestration action value of the candidate orchestration strategy under the historical system state. The orchestration action value is used to indicate at least one of container startup delay, resource fragmentation rate, and node load rate.
8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.
10. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.