Power grid operation scheduling management method, system and device based on Kubernetes and storage medium

By adopting a combination solution of Kubernetes multi-cluster architecture, load saturation scheduling algorithm and deep learning model in the grid management system, the challenges of traditional grid management systems in resource scheduling and security are solved, and efficient resource allocation and system stability and security are improved.

CN120109999APending Publication Date: 2025-06-06POWER DISPATCHING CONTROL CENT OF GUANGDONG POWER GRID CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510150416.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-11
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

Traditional power grid management systems have challenges in resource scheduling and security, and it is difficult to achieve efficient allocation of resources and system stability and security.

Method used

The multi-cluster architecture design based on Kubernetes is adopted, and resource scheduling is optimized through load saturation scheduling algorithm, and a deep learning model is built for real-time analysis and adjustment of adaptive security policies, while implementing a multi-tenant isolation mechanism.

Benefits of technology

It improves the resource utilization rate of the power grid operation management system and the security of the system, enhances the reliability and stability of the system, and ensures that the resources and data of each business department are independent of each other.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120109999A_ABST
    Figure CN120109999A_ABST
Patent Text Reader

Abstract

The invention discloses a Kubernetes-based power grid operation scheduling management method, system and device and a storage medium, and the method comprises the steps: a multi-cluster architecture design: dividing a power grid operation management system into a plurality of functional modules, and deploying each module on an independent Kubernetes cluster; resource scheduling optimization design: calculating a scheduling adaptation score of a schedulable node in each cluster, and selecting the schedulable node with the highest scheduling adaptation score as a selected node of the corresponding cluster for binding a scheduled Pod to complete scheduling; and establishing an adaptive security policy, namely deploying a deep learning model into a Kubernetes cluster to receive and analyze new behavior data of the system in real time, and realizing mutual independence of resources and data of different business departments by utilizing a multi-tenant isolation mechanism. Different service functions of the power grid operation management system can be subjected to isolation management, the disaster tolerance capability of the system is improved, efficient allocation and scheduling of resources are achieved, resource configuration is dynamically optimized and adjusted, it is ensured that resources and data of all service departments are independent, safety risks can be automatically recognized and prevented, and the safety performance of the system is improved. And the reliability and the stability of the system are obviously improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of power operation management, and in particular to a power grid operation dispatching management method, system, device and storage medium based on Kubernetes. Background Art

[0002] With the continuous expansion of the power system, the complexity of power grid management and dispatching has increased rapidly, which has put forward higher requirements for efficient management and safe operation of the system. For a long time, the power system has faced problems such as uneven resource dispatching, unbalanced loads, and system failures, which have affected the reliability and stability of the system. Traditional dispatching and management methods are time-consuming and labor-intensive, and are easily affected by human factors, making it difficult to ensure the dispatching effect.

[0003] Chinese patent CN202410059955.8 discloses a multi-cluster component management method based on k8s and its related devices and equipment. The method includes obtaining sub-cluster metadata and rendering the sub-cluster metadata in the main cluster; defining the data structure of the component metadata in the main cluster, redefining and rendering it; obtaining the deployment information of the component sub-cluster based on the component metadata, and rendering it to the corresponding sub-cluster; monitoring the component deployment status of the sub-cluster, and synchronizing the status information to the main cluster component metadata. This method achieves unified management and operation lifecycle management of different components in multiple sub-clusters by maintaining component metadata in multiple sub-clusters in the main cluster, simplifying the maintenance of multiple clusters, and improving management efficiency and system reliability. However, when the number of sub-clusters is large and the distribution is complex, the synchronization and management of metadata still faces certain challenges, which may lead to delays and data inconsistency risks.

[0004] Chinese patent CN115604362A discloses a scheduling management method based on Kubernetes and its related devices. The method includes obtaining resource levels in a multi-tenant management module, a project module, a resource management module, and a node module, and performing resource scheduling and management based on these levels. The specific process is as follows: the current tenant level in the multi-tenant management module is set to level a, the corresponding project level in the project module is set to level b, the corresponding software and hardware resource level in the resource management module is set to level c, and the corresponding node level in the node module is set to level d; when all nodes are in a network interconnected state, a blacklist network isolation strategy is used to block the scheduling operation of all nodes, and nodes with a level greater than the target scheduling level k are screened out to schedule the target node. The method solves the problem of resource scheduling and management in the Kubernetes cluster through level division and scheduling strategies. However, in a complex network environment, the dynamic adjustment of resource levels and the real-time application of isolation strategies may affect the stability and scheduling efficiency of the system. Summary of the invention

[0005] Purpose of the invention: The purpose of the present invention is to provide a Kubernetes-based power grid operation and dispatching management method, system, device and storage medium to achieve efficient allocation and dispatch of resources and improve the reliability and stability of the power grid operation and management system.

[0006] Technical solution: To achieve the above purpose, the present invention provides a Kubernetes-based power grid operation dispatching management method, comprising:

[0007] Multi-cluster architecture design: According to the different business functions of power grid operation and management, the power grid operation and management system is divided into multiple functional modules, and each module is deployed on an independent Kubernetes cluster;

[0008] Resource scheduling optimization: Calculate the scheduling adaptation score of the schedulable nodes in each cluster, select the schedulable node with the highest scheduling adaptation score as the selected node of the corresponding cluster; bind the Pod to be scheduled to the selected node to complete the scheduling;

[0009] Establishing adaptive security policies: Build a deep learning model and use the collected operation logs and indicator data of the power grid operation and management system for training. Deploy the trained deep learning model to the Kubernetes cluster level to receive and analyze new behavior data in the power grid operation and management system in real time.

[0010] This also includes using the multi-tenant isolation mechanism to achieve the independence of resources and data of different business departments. The implementation method is:

[0011] First, divide tenants according to the roles and needs of business departments, and determine the resource and data isolation requirements of each tenant in the power grid operation and management system;

[0012] In Kubernetes, each tenant is assigned a separate namespace and resource quotas and limits are configured for each namespace.

[0013] The role-based access control mechanism RBAC configures access control policies for each namespace. At the same time, the Kubernetes network policy Network Policy is used to limit network communication between namespaces.

[0014] This also includes subsequent management and control of namespaces, which is achieved by using Prometheus and Grafana to monitor the resource usage of each namespace, configuring ELK Stack to collect and analyze log data of each namespace, regularly reviewing and optimizing the multi-tenant isolation mechanism, and continuously improving and perfecting the isolation strategy based on actual operating conditions and changes in business needs.

[0015] The calculation of the scheduling adaptation score of the schedulable nodes in each cluster includes using the load saturation scheduling algorithm to collect the total resources of each schedulable node and the resource usage of the Pod already running on the node, calculating the remaining amount of each resource of the node, and obtaining the resource request amount of the Pod to be scheduled, and then calculating it with the remaining amount of the corresponding resources of the node to obtain the scheduling adaptation score of the schedulable node. The scoring function is:

[0016]

[0017] Among them, R CPU 、T CPU , U CPU They represent the CPU request of the Pod to be scheduled, the total CPU of the node, and the current CPU usage of the node; R mem 、T mem , U mem They represent the memory request of the Pod to be scheduled, the total memory of the node, and the current memory usage of the node.

[0018] The calculation of the scheduling adaptation score of the schedulable nodes in each cluster includes supporting the user to customize the weights of various resources. The final node score is the weighted sum of the scores of each resource. The scheduling adaptation score of the schedulable node has a scoring function of:

[0019]

[0020] Among them, R CPU 、T CPU , U CPU Represents the GPU request amount of the Pod to be scheduled, the total number of GPUs on the node, and the current GPU usage of the node; W CPU , W mem , W GPU Represents the user-defined weights of CPU, memory, and GPU.

[0021] Among them, the operation logs and indicator data of the power grid operation management system are collected by deploying Prometheus and ELK Stack. Specifically, Prometheus is configured to collect system operation indicators such as CPU usage, memory usage, and network traffic. ELK Stack is configured to collect application logs, system logs, and security event logs. Prometheus is responsible for real-time monitoring of the performance of the power grid operation management system, collecting resource usage data, and using Elasticsearch as a storage and search engine to store all logs and monitoring data. Logstash processes and forwards log data to Elasticsearch, and Kibana is used to visualize and analyze the collected data.

[0022] The method of constructing a deep learning model and using the collected operation logs and indicator data of the power grid operation management system for training is:

[0023] Use Python to preprocess the collected data, including cleaning, removing missing values ​​and outliers, and data standardization. The preprocessed data includes system operation indicators, application logs, system logs, and security event logs;

[0024] Divide the preprocessed data into training set and test set;

[0025] Build a long short-term memory network (LSTM) deep learning model, and use the training set and test set to train and optimize the deep learning model under the TensorFlow and Keras framework.

[0026] The present invention provides a Kubernetes-based power grid operation dispatching and management system, comprising:

[0027] Multi-cluster architecture design module: According to the different business functions of power grid operation and management, the power grid operation and management system is divided into multiple functional modules, and each module is deployed on an independent Kubernetes cluster;

[0028] Resource scheduling optimization module: calculates the scheduling adaptation score of the schedulable nodes in each cluster, selects the schedulable node with the highest scheduling adaptation score as the selected node of the corresponding cluster; binds the Pod to be scheduled to the selected node to complete the scheduling;

[0029] Adaptive security policy implementation module: Build a deep learning model and use the collected operation logs and indicator data of the power grid operation management system for training. Deploy the trained deep learning model to the Kubernetes cluster to receive and analyze new behavior data of the power grid operation management system in real time.

[0030] A computing device described in the present invention includes one or more processors, one or more memories and one or more programs, wherein the one or more programs are stored in the one or more memories and are configured to be executed by the one or more processors, and the one or more programs include instructions for executing any one of the above-mentioned Kubernetes-based power grid operation and scheduling management methods.

[0031] The present invention provides a computer-readable storage medium storing one or more programs, wherein the one or more programs include instructions, and when the instructions are executed by a computing device, the computing device executes any one of the above-mentioned methods for power grid operation and scheduling management based on Kubernetes.

[0032] Beneficial effects: The present invention has the following advantages: by dividing the power grid operation and management system into multiple functional modules and deploying them on an independent Kubernetes cluster, the isolation of business functions and the high availability of the system are achieved. At the same time, the load saturation scheduling algorithm is used to optimize resource scheduling, reduce resource fragmentation, and improve resource utilization. Further, through the multi-tenant isolation mechanism, the resources and data of each department are ensured to be independent of each other. Combined with adaptive security strategies, security measures and multi-tenant isolation mechanisms are dynamically adjusted to enhance the security and efficiency of the system, thereby improving the reliability and stability of the power grid operation and management system, which is of great significance to the efficient operation of the power system. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] Figure 1 A schematic flow chart of a method for managing power grid operation and dispatching provided in one embodiment of the present invention. DETAILED DESCRIPTION

[0034] The technical solution of the present invention is described in detail below in conjunction with the embodiments and drawings.

[0035] like Figure 1 As shown, the present invention constructs a power grid operation management system through the Kubernetes platform, and provides a power grid operation dispatching management method based on Kubernetes, so as to control the power operation business. The overall process is: in the power grid operation management system based on Kubernetes, a multi-cluster architecture design is adopted, and different business functions (such as power grid management, operation plan management, operation control management, etc.) are divided into independent functional modules, and each module is deployed on an independent Kubernetes cluster to improve the disaster recovery capability of the system. The resource utilization rate is improved and the scheduling time is reduced through the load saturation scheduling algorithm, and a multi-tenant isolation mechanism is implemented to ensure that the resources and data of each department are independent of each other. The security measures and multi-tenant isolation mechanism are dynamically adjusted in combination with the adaptive security strategy to ensure the security and efficiency of the system. The specific design method is as follows:

[0036] 1. Architecture Design and Component Selection

[0037] (1) Multi-cluster architecture design

[0038] For the containers related to power grid operation and management, the Kubernetes platform is used for container orchestration and management. According to the different business functions of power grid operation and management, such as power grid management, operation plan management, operation control management, operation risk management, operation evaluation and improvement management, secondary system management, operation support management and spot market operation, the system is divided into multiple functional modules. Each functional module is deployed on an independent Kubernetes cluster, which is distributed in different physical locations or data centers to improve the disaster recovery capability of the system.

[0039] (2) Component selection

[0040] Use Helm to package and deploy applications to simplify the management and version control of complex applications. Use Istio to achieve secure communication and traffic management between services to ensure reliable communication and load balancing between services. Use Prometheus to monitor the operating status of the system, collect indicator data, and use Grafana to visualize the data to display the system performance and health status in real time. Use KubeEdge to delegate some computing tasks to edge nodes to achieve edge computing, so as to monitor and control the operating status of the power grid in real time.

[0041] 2. Resource Scheduling Optimization through Load Saturation Scheduling Algorithm

[0042] The default scheduling algorithm of Kubernetes is prone to produce too much resource fragmentation. Some user tasks will have a long waiting time due to insufficient resources. In addition, the resource status of the node GPU is not considered in the pre-selection and optimization stages, which makes scheduling unreasonable in the power grid business scenarios designed for high-performance computing and deep learning. The use of a load saturation scheduling algorithm can reduce the resource fragmentation generated during the scheduling process, improve resource utilization, and reduce the waiting time for tasks to be scheduled. Specifically, it includes:

[0043] (1) Pre-selection

[0044] The load saturation scheduling algorithm not only collects the resource status information of the node's CPU and memory, but also collects user-specified external extension resources such as GPU resources. The Nvidia Device Plugin plug-in will be deployed on each working node in the Kubernetes cluster, collect the ID information of the node's GPU device, organize it into a list, and report it to the node's kubelet regularly. If the Pod needs to use the GPU, the kubelet will select it from the GPU ID list, then obtain the device path and driver directory corresponding to the selected GPU from the Nvidia Device Plugin plug-in, and mount it to the Pod container to complete the Pod's GPU allocation.

[0045] (2) Optimization of links

[0046] The load saturation scheduling algorithm collects the total resources of each schedulable node and the resource usage of the Pods already running on the node, and calculates the remaining amount of each resource on the node. For the Pod to be scheduled, the load saturation optimization strategy calculates the remaining amount of the corresponding resources of the node after obtaining its resource request amount to obtain the node score. The node score is negatively correlated with the remaining amount of the corresponding resources of the node. The score calculation of the load saturation optimization strategy is:

[0047]

[0048] Among them, R CPU 、T CPU , U CPU They represent the CPU request of the Pod to be scheduled, the total CPU of the node, and the current CPU usage of the node; R mem 、T mem , U mem They represent the memory request of the Pod to be scheduled, the total memory of the node, and the current memory usage of the node. According to the scoring function of the load saturation optimization strategy, the scheduler will tend to schedule the Pod to be scheduled to the node with less remaining resources, so as to reduce resource fragmentation and improve resource utilization.

[0049] (3) User-defined resource ratings

[0050] The load saturation optimization strategy also supports scoring of user-defined external extension resources. For example, in the deep learning training business scenario, you can add scoring of node GPU resources. The load saturation optimization strategy believes that the importance of various resources is different in different scenarios. Therefore, the load saturation optimization strategy allows users to customize the weights of various resources, and the final node score is the weighted sum of the scores of each resource. The scoring function in this case is:

[0051]

[0052] Among them, R CPU 、T CPU , U CPU Represents the GPU request amount of the Pod to be scheduled, the total number of GPUs on the node, and the current GPU usage of the node; W CPU , W mem , W GPU Represents the user-defined weights of CPU, memory, and GPU. It ensures that some resource-intensive tasks in the power grid (such as grid-connected simulation, load forecasting, real-time control, and risk analysis) can be executed efficiently. By dynamically evaluating the GPU resource status of the node and assigning corresponding weights to different tasks, the system can schedule key tasks to the most suitable node, thereby improving the performance and response speed of the overall system, reducing resource waste, and ensuring the stability and reliability of power grid operation.

[0053] In general, the specific data transmission process of the load saturation scheduling algorithm is as follows: Each node in the Kubernetes cluster will report its own resource status information in real time, including the usage of CPU, memory, and external extended resources (such as GPU). This information is passed to the kubelet (Kubernetes node agent) of each node through plug-ins such as Nvidia Device Plugin for the scheduling algorithm to call.

[0054] Then the score calculation is performed: the scheduling algorithm will calculate the ratio of the resources required by the Pod to be scheduled (such as CPU, memory, GPU, etc.) to the current remaining resources of the node for each node, so as to evaluate the adaptability of each node. The score of each resource is assigned different weights according to the user's business scenario. For example, the GPU weight of GPU-intensive tasks may be higher, and the score will give priority to this demand.

[0055] The user-defined scoring stage allows users to assign weights to different resources based on task requirements. The system then combines the resource scores and generates a comprehensive score.

[0056] After the scoring is completed, the load saturation scheduling algorithm will select the node with the highest score to ensure optimal resource allocation. The higher the score of the node, the more its resource status matches the task requirements. Selecting this node can reduce resource fragmentation and task scheduling waiting time. The system then binds the resource request of the Pod to be scheduled to the selected node. In particular, for GPU resources, the kubelet will match the device path from the selected GPU ID list and mount it to the Pod container to complete the actual resource allocation. After the scheduling is completed, the system will monitor the resource usage of the node in real time and feed back the updated resource status to the monitoring system for optimizing the evaluation of the next scheduling cycle.

[0057] Assume that a deep learning task requires 2-core CPU, 4GB memory and 1 GPU resources, and there are two nodes in the cluster: Node A has 2-core CPU, 4GB memory and 1 GPU, and Node B has 4-core CPU, 12GB memory and 1 GPU. Since the task is highly dependent on GPU resources, the user sets the GPU weight to the highest to ensure priority allocation. When calculating the score, although Node A has limited CPU and memory resources, it has a score of 0.8 because its GPU resources meet the requirements; Node B has a score of 0.9 because its CPU and memory resources are more abundant and its GPU requirements are also met. Under this setting, the high GPU weight becomes the decisive factor, giving Node B the highest score. Therefore, the system selects Node B to perform the task, achieving priority utilization of GPU resources, reducing scheduling waiting time and ensuring optimal resource allocation.

[0058] 3. Implementation of Adaptive Security Strategy

[0059] (1) Establishing a behavior analysis model

[0060] First, deploy Prometheus and ELK Stack (Elasticsearch, Logstash, Kibana) to collect system operation logs and indicator data. Configure Prometheus to collect system operation indicators such as CPU usage, memory usage, network traffic, and configure ELK Stack to collect application logs, system logs, and security event logs. Prometheus is responsible for real-time monitoring of system performance and collecting detailed resource usage data. Elasticsearch, as a storage and search engine, stores all logs and monitoring data; Logstash processes and forwards log data to Elasticsearch; Kibana is used to visualize and analyze this data.

[0061] Use Python for data preprocessing. Clean the collected data and remove missing values ​​and outliers. At the same time, standardize the data to ensure that data with different characteristics are on the same scale, thereby improving the effect of model training. The sorted data includes system operation indicators (such as CPU usage, memory usage, network traffic) and application logs, system logs, and security event logs.

[0062] The preprocessed data is divided into training sets and test sets for subsequent model training and evaluation. A LSTM (Long Short-Term Memory Network) deep learning model is built to adapt to the characteristics of time series data. The model is trained using the TensorFlow and Keras frameworks, and the training set data is used for model training. The model parameters are optimized through repeated iterations until the model performance reaches the expected level.

[0063] (2) Model deployment and real-time analysis

[0064] Deploy the trained model to the Kubernetes cluster and run it as a standalone microservice. Configure the model service to receive and analyze all new behavior data in the power grid operation management system in real time for effective management and optimization.

[0065] Adaptive security strategies are commonly used in the field of network security. Security measures are dynamically adjusted through real-time monitoring and data analysis. The present invention combines Prometheus and ELK Stack for real-time monitoring, and dynamically adjusts security strategies based on resource allocation and operating status of the power system, thereby better adapting to the complexity and real-time nature of power grid management.

[0066] 4. Implementation of multi-tenant isolation mechanism

[0067] In the Kubernetes-based private cloud platform of the power grid operation and management system, the implementation of a multi-tenant isolation mechanism can ensure that the resources and data of different business departments are independent of each other, avoid mutual influence, and improve the security and management efficiency of the system.

[0068] (1) Tenant division and space naming

[0069] Identify the various business departments and their functions in the system, such as power distribution management, substation management, and monitoring management. According to the roles and needs of these departments, divide tenants and determine the resource and data isolation requirements of each tenant in the system. Then, in Kubernetes, each tenant is assigned an independent namespace to achieve resource and workload isolation. Namespaces are used to organize and isolate resources from each business department. Create an independent namespace for each business department. For example: power distribution management namespace: namespace-distribution-management; substation management namespace: namespace-substation-management; power dispatch namespace: namespace-power-dispatch; power monitoring namespace: namespace-power-monitoring, etc. At the same time, configure resource quotas and restrictions for each namespace to ensure fairness and isolation of resource use among tenants. According to the needs of each department, configure quotas for resources such as CPU, memory, and storage to prevent the resource consumption of one department from affecting other departments.

[0070] (2) Implement access control and network isolation strategies

[0071] Through Kubernetes's role-based access control (RBAC) mechanism, configure access control policies for each namespace. Assign appropriate roles and permissions to users and applications in each department to ensure that only authorized users and applications can access the corresponding resources. At the same time, use Kubernetes's network policy to achieve network isolation and limit network communication between namespaces. Configure network policies to ensure that only necessary communications are allowed to reduce the attack surface. For example, the services of the power grid management department can only communicate with specific external services and cannot access services of other departments at will.

[0072] The multi-tenant isolation mechanism in the present invention combines the namespace and RBAC (role-based access control) of Kubernetes to ensure resource and data isolation between different business departments, and combines with the load scheduling algorithm to ensure fair use of resources. Therefore, the present invention makes targeted improvements to the adaptive security strategy and multi-tenant isolation mechanism, further improving the security and stability of the power grid management system.

[0073] 5. Subsequent system control

[0074] Use Prometheus and Grafana to monitor resource usage in each namespace and detect anomalies in a timely manner. Configure ELK Stack (Elasticsearch, Logstash, Kibana) to collect and analyze log data in each namespace to ensure the traceability of security incidents. Regularly review and optimize the multi-tenant isolation mechanism, and continuously improve and perfect the isolation strategy based on actual operation conditions and changes in business needs. Ensure that the multi-tenant architecture can adapt to the ever-changing power grid operation and management needs, and improve the security and management efficiency of the system.

[0075] Traditional power grid operation and management systems face great challenges in resource scheduling and security. They are unable to fully optimize resource utilization and ensure system security, and the system reliability and stability are insufficient. At the same time, methods that rely on traditional scheduling algorithms and security strategies are prone to resource waste and security vulnerabilities, affecting the efficient operation of the system.

[0076] The present invention adopts a power grid operation management method based on the Kubernetes platform to isolate and manage different business functions, significantly improve the disaster recovery capability of the system, realize efficient allocation and scheduling of resources, dynamically optimize and adjust resource configuration, ensure the independence of resources and data of each business department, automatically identify and prevent security risks, significantly improve the security of the system, realize continuous optimization and upgrading of the system, and ensure the optimal performance of the power grid operation management system under different operating conditions.

[0077] Based on the above-mentioned power grid operation dispatching management method, the present invention provides a power grid operation dispatching management system based on Kubernetes, including:

[0078] Multi-cluster architecture design module: According to the different business functions of power grid operation and management, the power grid operation and management system is divided into multiple functional modules, and each module is deployed on an independent Kubernetes cluster;

[0079] Resource scheduling optimization module: calculates the scheduling adaptation score of the schedulable nodes in each cluster, selects the schedulable node with the highest scheduling adaptation score as the selected node of the corresponding cluster; binds the Pod to be scheduled to the selected node to complete the scheduling;

[0080] Adaptive security policy implementation module: Build a deep learning model and use the collected operation logs and indicator data of the power grid operation management system for training. Deploy the trained deep learning model to the Kubernetes cluster to receive and analyze new behavior data of the power grid operation management system in real time.

[0081] The present invention further provides a computing device, comprising one or more processors, one or more memories, and one or more programs, wherein the one or more programs are stored in the one or more memories and are configured to be executed by the one or more processors, and the one or more programs include instructions for executing any of the above-mentioned Kubernetes-based power grid operation and dispatching management methods.

[0082] The present invention further provides a computer-readable storage medium storing one or more programs, wherein the one or more programs include instructions, and when the instructions are executed by a computing device, the computing device executes any one of the above-mentioned Kubernetes-based power grid operation and dispatching management methods.

Claims

1. A power grid operation dispatching management method based on Kubernetes, characterized in that: include: Multi-cluster architecture design: According to the different business functions of power grid operation and management, the power grid operation and management system is divided into multiple functional modules, and each module is deployed on an independent Kubernetes cluster; Resource scheduling optimization: Calculate the scheduling adaptation score of the schedulable nodes in each cluster, select the schedulable node with the highest scheduling adaptation score as the selected node of the corresponding cluster; bind the Pod to be scheduled to the selected node to complete the scheduling; Establishing adaptive security policies: Build a deep learning model and use the collected operation logs and indicator data of the power grid operation and management system for training. Deploy the trained deep learning model to the Kubernetes cluster level to receive and analyze new behavior data in the power grid operation and management system in real time.

2. The power grid operation dispatching management method based on Kubernetes according to claim 1 is characterized in that: It also includes using the multi-tenant isolation mechanism to achieve the independence of resources and data of different business departments. The implementation method is: First, divide tenants according to the roles and needs of business departments, and determine the resource and data isolation requirements of each tenant in the power grid operation and management system; In Kubernetes, each tenant is assigned a separate namespace and resource quotas and limits are configured for each namespace. The role-based access control mechanism RBAC configures access control policies for each namespace. At the same time, the Kubernetes network policy NetworkPolicy is used to limit network communication between namespaces.

3. The power grid operation dispatching management method based on Kubernetes according to claim 1 is characterized in that: It also includes subsequent management and control of the namespace, which is achieved by using Prometheus and Grafana to monitor the resource usage of each namespace, configuring ELKStack to collect and analyze the log data of each namespace, regularly reviewing and optimizing the multi-tenant isolation mechanism, and continuously improving and perfecting the isolation strategy based on actual operating conditions and changes in business needs.

4. The power grid operation dispatching management method based on Kubernetes according to claim 1 is characterized in that: The calculation of the scheduling adaptation score of the schedulable nodes in each cluster includes using the load saturation scheduling algorithm to collect the total resources of each schedulable node and the resource usage of the Pod already running on the node, calculating the remaining amount of each resource of the node, and obtaining the resource request amount of the Pod to be scheduled, and then calculating it with the remaining amount of the corresponding resources of the node to obtain the scheduling adaptation score of the schedulable node. The scoring function is: Among them, R CPU , T CPU , U CPU They represent the CPU request of the Pod to be scheduled, the total CPU of the node, and the current CPU usage of the node; R mem , T mem , U mem They represent the memory request of the Pod to be scheduled, the total memory of the node, and the current memory usage of the node.

5. The power grid operation dispatching management method based on Kubernetes according to claim 1 is characterized in that: The calculation of the scheduling adaptation score of the schedulable nodes in each cluster includes supporting the user to customize the weights of various resources. The final node score is the weighted sum of the scores of each resource. The scheduling adaptation score of the schedulable node has a scoring function of: Among them, R CPU , T CPU , U CPU Represents the GPU request amount of the Pod to be scheduled, the total number of GPUs on the node, and the current GPU usage of the node; W CPU , W mem , W GPU Represents the user-defined weights of CPU, memory, and GPU.

6. The power grid operation dispatching management method based on Kubernetes according to claim 1 is characterized in that: Prometheus and ELKStack are deployed to collect the operation logs and indicator data of the power grid operation management system. Specifically, Prometheus is configured to collect system operation indicators such as CPU usage, memory usage, and network traffic. ELKStack is configured to collect application logs, system logs, and security event logs. Prometheus is responsible for real-time monitoring of the performance of the power grid operation management system and collecting resource usage data. Elasticsearch is used as a storage and search engine to store all logs and monitoring data. Logstash processes and forwards log data to Elasticsearch. Kibana is used to visualize and analyze the collected data.

7. The power grid operation dispatching management method based on Kubernetes according to claim 1 is characterized in that: The method of constructing a deep learning model and using the collected operation logs and indicator data of the power grid operation management system for training is: Use Python to preprocess the collected data, including cleaning, removing missing values ​​and outliers, and data standardization. The preprocessed data includes system operation indicators, application logs, system logs, and security event logs; Divide the preprocessed data into training set and test set; Build a long short-term memory network (LSTM) deep learning model, and train and optimize the deep learning model using training sets and test sets under the TensorFlow and Keras frameworks.

8. A Kubernetes-based power grid operation dispatching management system, characterized in that: include: Multi-cluster architecture design module: According to the different business functions of power grid operation and management, the power grid operation and management system is divided into multiple functional modules, and each module is deployed on an independent Kubernetes cluster; Resource scheduling optimization module: calculates the scheduling adaptation score of the schedulable nodes in each cluster, selects the schedulable node with the highest scheduling adaptation score as the selected node of the corresponding cluster; binds the Pod to be scheduled to the selected node to complete the scheduling; Adaptive security policy implementation module: Build a deep learning model and use the collected operation logs and indicator data of the power grid operation management system for training. Deploy the trained deep learning model to the Kubernetes cluster to receive and analyze new behavior data of the power grid operation management system in real time.

9. A computing device, characterized in that The method comprises one or more processors, one or more memories and one or more programs, wherein the one or more programs are stored in the one or more memories and are configured to be executed by the one or more processors, and the one or more programs include instructions for executing any one of the methods according to claims 1 to 7.

10. A computer-readable storage medium storing one or more programs, characterized in that: The one or more programs include instructions which, when executed by a computing device, cause the computing device to perform any one of the methods according to claims 1 to 7.

Citation Information

Patent Citations

  • Scheduling management method and device based on Kubernetes

    CN115604362A

  • Multi-cluster component management method and device based on k8s and computer equipment

    CN117573295A