Method and system for calling temporary resources by micro-service container arrangement platform

By deploying a Kubernetes cluster on a microservices orchestration platform, fixing the number of resources and instances, and dynamically adjusting resources using Prometheus and LSTM models, we solved the problem of inefficient resource management under dynamic loads on traditional microservices orchestration platforms, achieving efficient and flexible resource management and improved stability.

CN120670142APending Publication Date: 2025-09-19珠海盈米基金销售有限公司

Patent Information

Application Number
CN202510658978.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-21
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

Traditional microservice container orchestration platforms are inefficient when handling dynamic loads, static configuration leads to insufficient or wasted resources, and resources cannot be managed efficiently.

Method used

Deploy a Kubernetes microservice orchestration cluster, fix the cluster resource scale and the number of application instances, combine Prometheus monitoring data with the LSTM model to predict the expansion threshold, and dynamically adjust the cluster scale and the number of application deployments through fuzzy logic and Bayesian inference.

Benefits of technology

Improve resource utilization efficiency, reduce operating costs, enhance system scalability and stability, simplify operation and maintenance processes, and improve service quality and user satisfaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120670142A_ABST
    Figure CN120670142A_ABST
Patent Text Reader

Abstract

The invention provides a temporary resource calling method and system for a micro-service container arrangement platform, and the method comprises the steps: deploying a micro-service arrangement cluster kubernetes, and fixing the scale of cluster resources; deploying business applications in the micro-service orchestration cluster, and fixing the number of application examples; acquiring application monitoring data, and predicting a micro-service capacity expansion threshold value; according to the real-time monitoring data and the capacity expansion threshold value, the cluster scale is dynamically adjusted, the application deployment number is dynamically adjusted, capacity expansion is conducted when the threshold value is higher than the threshold value, and capacity shrinkage is conducted when the threshold value is lower than the threshold value, through the method and the corresponding system, service interruption or performance bottlenecks caused by insufficient resources can be reduced, the stability and reliability of the system are enhanced, and the resource utilization efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention provides a method and system for a microservice container orchestration platform to call temporary resources, and relates to the field of information technology. Background Art

[0002] Traditional microservice container orchestration and resource management usually rely on static configuration, that is, presetting a fixed amount of resources (such as CPU and memory) when deploying the application. This method performs well when handling static loads, but is inefficient when facing dynamically changing loads. When the load suddenly increases, static configuration may lead to insufficient resources and affect service quality. When the load decreases, too many idle resources lead to resource waste. These problems have given rise to the demand for more efficient and sophisticated resource management methods for microservice container orchestration platforms. Summary of the Invention

[0003] The present invention provides a method and system for a microservice container orchestration platform to call temporary resources, to solve the above-mentioned problems:

[0004] The present invention proposes a method for a microservice container orchestration platform to call temporary resources, the method comprising: deploying a microservice orchestration cluster Kubernetes and fixing the cluster resource scale;

[0005] Deploy business applications in the microservice orchestration cluster and fix the number of application instances;

[0006] Obtain application monitoring data and predict microservice expansion thresholds;

[0007] Based on real-time monitoring data and expansion thresholds, the cluster size and the number of application deployments are dynamically adjusted, with capacity expanded when the threshold is exceeded and capacity reduced when the threshold is exceeded.

[0008] Furthermore, we deployed a microservice orchestration cluster, Kubernetes, and fixed the cluster resource scale, including:

[0009] Obtain the target software package, upload the target software package to the master node virtual machine, create a cluster on the master node virtual machine, use resource requests and limits when creating pods in the cluster, and add the slave node virtual machines to the cluster;

[0010] After the cluster is successfully built, the image corresponding to the slave node virtual machine is pulled by the Dockerpull command on the slave node virtual machine, and then the image corresponding to the master node virtual machine is pulled by the Dockerpull command on the master node virtual machine;

[0011] After the image is successfully pulled, start the container corresponding to the image through the Dockerrun command, obtain the container ID that was successfully started, and then access the visual interface of the container orchestration tool through the browser.

[0012] Furthermore, business applications are deployed in the microservice orchestration cluster, with a fixed number of application instances, including:

[0013] Create a Docker image for the business application and push the Docker image to a private container registry using Docker commands;

[0014] Create a Kubernetes deployment configuration, that is, create a YAML file and define the Deployment resource in the YAML file;

[0015] Use the kubectl command to apply the YAML file to deploy the business application;

[0016] Verify that the Pods are running successfully. If so, check that the number of Pods created by the Deployment matches the specified number of replicas, as specified by the spec.replicas field in the YAML file.

[0017] Create a Service, apply the Service configuration file, and check whether the Service has been successfully created and is running;

[0018] Find the external IP address assigned to the Service and access it in a browser or using the curl command to verify that the application is running and can be load balanced.

[0019] Furthermore, application monitoring data is obtained to predict microservice expansion thresholds, including:

[0020] Deploy the monitoring tool Prometheus in the Kubernetes microservice orchestration cluster;

[0021] Prometheus captures monitoring data from Pods deployed in the microservice orchestration cluster, including CPU usage, memory usage, network traffic, number of requests, and response time.

[0022] Obtain the monitoring history data captured by the monitoring tool Prometheus on all Pods;

[0023] Processing missing values, outliers and noise data in the monitoring historical data;

[0024] extracting features from the processed data, the features including a timestamp, a request type, and a node status;

[0025] Divide the data after feature extraction into training set, validation set and test set;

[0026] Use the training set data to train the LSTM model, adjust the model parameters through grid search, and use cross-validation based on the validation set to avoid overfitting;

[0027] The trained model is deployed on the test set for testing to obtain the final predicted expansion threshold model.

[0028] Furthermore, based on real-time monitoring data and expansion thresholds, the cluster size and the number of deployed applications are dynamically adjusted, with capacity expanded when above the threshold and reduced when below the threshold, including:

[0029] Collect real-time monitoring data from the cluster, including CPU usage, memory usage, network traffic, number of requests, and response time;

[0030] Inputting the real-time monitoring data into the predicted capacity expansion threshold model to obtain the predicted capacity expansion threshold;

[0031] Dynamically adjust the cluster size based on the predicted expansion threshold, dynamically adjust

[0032] The number of application deployments. When the cluster size and the number of application deployments exceed the threshold, the capacity is expanded; when it falls below the threshold, the capacity is reduced, including:

[0033] Use fuzzy logic to evaluate the load of each resource and obtain the fuzzy score M(t).

[0034]

[0035] Where x is the current resource usage (CPU, memory, network traffic, number of requests, and response time), k and c are parameters for adjusting the fuzzy logic characteristics, and M(t) represents the fuzzy score at the current time t;

[0036] Introduce dynamic weight W and weighted Bayesian inference formula:

[0037]

[0038] Among them, P(S|D) represents the probability that the system needs to be expanded given the data D, and W prior represents the preset dynamic weight, P(D|S) represents the probability of observing the current data when capacity expansion is required, P(S) represents the prior probability of event S (capacity expansion is required), and P(D) represents the overall probability of observing the current data D;

[0039]

[0040] Among them, S(t) is the final scaling decision score, which reflects the comprehensive changes of different indicators over time, T d (τ) = e -λ(t-τ) ,τ represents the time window length;

[0041] When S(t) is greater than a first preset threshold, it is decided to expand the capacity; when S(t) is less than a second preset threshold, it is decided to reduce the capacity.

[0042] The present invention proposes a system for a microservice container orchestration platform to call temporary resources, the system comprising:

[0043] Fixed cluster resource scale module, used to deploy microservice orchestration cluster Kubernetes and fix cluster resource scale;

[0044] The fixed application instance number module is used to deploy business applications in the microservice orchestration cluster and fix the number of application instances;

[0045] The prediction module is used to obtain application monitoring data and predict microservice expansion thresholds;

[0046] The adjustment module is used to dynamically adjust the cluster size and the number of application deployments based on real-time monitoring data and expansion thresholds, expanding the capacity when the threshold is exceeded and reducing the capacity when the threshold is exceeded.

[0047] Furthermore, the fixed cluster resource scale module includes:

[0048] Create a cluster module, obtain the target software package, upload the target software package to the master node virtual machine, create a cluster on the master node virtual machine, and add the slave node virtual machines to the cluster;

[0049] The image pulling module is used to pull the image corresponding to the slave node virtual machine through the Docker pull command on the slave node virtual machine after the cluster is successfully built, and then pull the image corresponding to the master node virtual machine through the Docker pull command on the master node virtual machine;

[0050] Start the container module, which is used to start the container corresponding to the image through the Docker run command after the image is successfully pulled. After obtaining the container ID of the successfully started container, you can access the visual interface of the container orchestration tool through the browser.

[0051] Furthermore, the module for fixing the number of application instances includes:

[0052] The push image module is used to create a Docker image for the business application and push the Docker image to the private container registry using Docker commands;

[0053] Create a YAML file module for creating Kubernetes deployment configurations, that is, create a YAML file and define the Deployment resource in the YAML file;

[0054] The deployment application module is used to apply the YAML file using the kubectl command to deploy business applications;

[0055] The verification module verifies that the Pods are running successfully. If the verification passes, it checks that the number of Pods created by the Deployment matches the specified number of replicas, which is implemented by the spec.replicas field in the YAML file.

[0056] Verify the Service module, which is used to create a Service, apply the Service configuration file, and check whether the Service has been successfully created and run;

[0057] The Verify Run module is used to find the external IP address assigned to the Service, and then access the address in a browser or using the curl command to verify whether the application can run and load balance.

[0058] Furthermore, the prediction module includes:

[0059] Deployment monitoring tool module, used to deploy monitoring tool Prometheus in Kubernetes microservice orchestration cluster;

[0060] The data capture module, Prometheus, captures monitoring data from the Pod deployed in the microservice orchestration cluster;

[0061] A historical data acquisition module is used to obtain monitoring historical data captured by the monitoring tool Prometheus on all Pods;

[0062] A feature extraction module is used to process missing values, abnormal values ​​and noise data in the monitoring historical data;

[0063] a segmentation dataset module for extracting features from the processed data, the features including timestamp, request type, and node status;

[0064] The dataset segmentation module divides the feature-extracted data into training set, validation set, and test set;

[0065] The training and validation module uses the training set data to train the LSTM model, adjusts the model parameters through grid search, and uses cross-validation based on the validation set to avoid overfitting;

[0066] The testing module is used to deploy the trained model to the test set for testing to obtain the final prediction expansion threshold model.

[0067] Furthermore, the adjustment module includes:

[0068] Get real-time monitoring data module, used to collect real-time monitoring data from the cluster, including CPU usage, memory usage, network traffic, number of requests, and response time;

[0069] A prediction threshold acquisition module is used to input real-time monitoring data into the prediction expansion threshold model to obtain a prediction expansion threshold;

[0070] The scaling module is used to dynamically adjust the cluster size and the number of deployed applications based on the predicted expansion threshold. When the cluster size and the number of deployed applications are above the threshold, the cluster size is expanded, and when they are below the threshold, the cluster size is reduced. This module includes:

[0071] Use fuzzy logic to evaluate the load of each resource and obtain the fuzzy score M(t).

[0072]

[0073] Where x is the current resource usage (CPU, memory, network traffic, number of requests, and response time), k and c are parameters for adjusting the fuzzy logic characteristics, and M(t) represents the fuzzy score at the current time t;

[0074] Introduce dynamic weight W and weighted Bayesian inference formula:

[0075]

[0076] Among them, P(S|D) represents the probability that the system needs to be expanded given the data D, and W prior represents the preset dynamic weight, P(D|S) represents the probability of observing the current data when capacity expansion is required, P(S) represents the prior probability of event S (capacity expansion is required), and P(D) represents the overall probability of observing the current data D;

[0077] The final scaling decision score of the current microservice container orchestration platform is calculated based on the fuzzy score of the resources of the current microservice container orchestration platform, the calculated probability that the system needs to be expanded, and the calculated value of the time decay function:

[0078]

[0079] Among them, S(t) is the final scaling decision score, which reflects the comprehensive changes of different indicators over time, T d (τ) = e -λ(t-τ) ,τ represents the time window length, T d (τ) represents the calculated value of the time decay function within τ, and P(S|D(τ)) represents the probability that the system needs to be expanded under the given data D within τ;

[0080] When S(t) is greater than a first preset threshold, it is decided to expand the capacity; when S(t) is less than a second preset threshold, it is decided to reduce the capacity.

[0081] The beneficial effects of the present invention are: improving resource utilization efficiency, ensuring that a large amount of resources will not be idle by dynamically adjusting the number of containers, thereby improving overall resource utilization efficiency; reducing operating costs, adjusting resources according to actual needs, avoiding over-investment in underutilized hardware resources, thereby reducing IT operating costs; enhancing the scalability and flexibility of the system, the use of container orchestration tools, and dynamic resource adjustment based on real-time data improve the flexibility and scalability of the system to cope with different load requirements; improving system stability and reliability, predicting system load and optimizing resource allocation accordingly, which can reduce service interruptions or performance bottlenecks caused by insufficient resources and enhance system stability and reliability; automated operation and maintenance processes, automated container deployment and management reduce the need for manpower operation and maintenance, alleviate the burden on administrators, and reduce the possibility of human errors; these mechanisms not only help enterprises use cloud resources more efficiently, but also can quickly adapt when demand changes, thereby improving service quality and user satisfaction. BRIEF DESCRIPTION OF THE DRAWINGS

[0082] Figure 1 A schematic diagram of a method for a microservice container orchestration platform to call temporary resources according to the present invention. DETAILED DESCRIPTION

[0083] In order to more clearly understand the above-mentioned objects, features and advantages of the present invention, the present invention is described in detail below with reference to the accompanying drawings and specific embodiments. It should be noted that, in the absence of conflict, the embodiments of the present application and the features therein may be combined with each other.

[0084] The following description sets forth numerous specific details to facilitate a thorough understanding of the present invention. The embodiments described are merely a portion of the embodiments of the present invention, not all of them. All other embodiments derived by persons of ordinary skill in the art based on the embodiments of the present invention without inventive effort are intended to fall within the scope of protection of the present invention.

[0085] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art of the present invention. The terms used in this specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention.

[0086] One embodiment of the present invention provides a method for a microservice container orchestration platform to call temporary resources, characterized in that the method includes:

[0087] Deploy a microservice orchestration cluster, Kubernetes, and fix the cluster resource scale;

[0088] Deploy business applications in the microservice orchestration cluster and fix the number of application instances;

[0089] Obtain application monitoring data and predict microservice expansion thresholds;

[0090] Based on real-time monitoring data and expansion thresholds, the cluster size and the number of application deployments are dynamically adjusted, with capacity expanded when the threshold is exceeded and capacity reduced when the threshold is exceeded.

[0091] The working principle and effect of the above technical solution are as follows: fixed initial resources and instances, fixed cluster resource scale, at the beginning, deploy a Kubernetes cluster with fixed resource configuration, which involves setting the number of nodes and the resources of each node (such as CPU, memory); deploy application instances, deploy business applications on the cluster, and initialize a certain number of application instances (i.e., Pods), which is specified by the spec.replicas of the Kubernetes Deployment; obtain and monitor data, using monitoring tools such as Prometheus to collect the operating status and performance data of applications and clusters, including CPU, memory utilization, response time, etc.; use analytical tools to process monitoring data, and predict system load changes based on historical data and real-time data; use machine learning algorithms or rule-based policies to analyze monitoring data, set or adjust expansion / contraction thresholds for applications and clusters, and repeat learning and adjustment to adapt the thresholds to load changes; compare monitoring data with set thresholds in real time. If the expansion threshold is above the expansion threshold, cluster resources (such as adding nodes) and application instances (adding Pods) are automatically expanded to ensure performance; if the expansion threshold is below the contraction threshold, cluster resources (such as reducing nodes) and application instances are reduced to conserve resources. Improve resource utilization efficiency, dynamically adjust cluster and application scale to match actual load requirements, and optimize resource usage to the greatest extent possible; improve system performance and user experience, and ensure that applications always have sufficient resources and instances to support requests and workloads by quickly responding to load changes; automatically scale down when the load decreases to avoid unnecessary resource occupation, thereby reducing operating costs; enhance elasticity and scalability, the system can automatically respond to load changes, has good elasticity and scalability, and is suitable for highly volatile business scenarios; simplify operation and maintenance management, and reduce the need for manual operation and maintenance intervention through automated expansion and contraction strategies, making management simpler and more efficient. Through the combination of fixed and dynamic strategies, it can provide a stable foundation while flexibly adapting to changing workload requirements.

[0092] One embodiment of the present invention deploys a microservice orchestration cluster Kubernetes and fixes the cluster resource scale, including:

[0093] Obtain the target software package, upload the target software package to the master node virtual machine, create a cluster on the master node virtual machine, use resource requests and limits when creating pods in the cluster, and add the slave node virtual machines to the cluster;

[0094] After the cluster is successfully built, the image corresponding to the slave node virtual machine is pulled by the Dockerpull command on the slave node virtual machine, and then the image corresponding to the master node virtual machine is pulled by the Dockerpull command on the master node virtual machine;

[0095] After the image is successfully pulled, start the container corresponding to the image through the Docker run command, obtain the container ID that was successfully started, and then access the visual interface of the container orchestration tool through the browser.

[0096] The working principle and effect of the above technical solution are as follows: software package preparation and upload, obtain the target software package from the software source or warehouse, and upload it to the Kubernetes master node virtual machine, which is usually a necessary installation package or configuration file for building and initializing the cluster environment; cluster initialization and configuration, on the master node virtual machine, initialize the Kubernetes cluster through software packages and installation tools (such as kubeadm), and in the process of defining Pods, specify resource requests and restrictions, so that the resource usage of each Pod can be controlled to ensure fairness and resource utilization; join the cluster from the node, by generating and using the join command, the slave node virtual machine will be added to the Kubernetes cluster where the master node is located, thereby expanding the computing power of the cluster; use the Docker pull command to pull the corresponding Docker images on the slave node and master node virtual machines respectively. These images usually contain application services or tools; use the Docker run command to start the container built from the image, and obtain the container ID after startup to ensure the normal operation of the container; access the visual interface of the application or tool running in the container through the browser. This interface is used to monitor, manage and orchestrate the cluster and containers, usually a container orchestration platform such as Kubernetes Dashboard. Simplify the installation and deployment process. The automated cluster initialization and node joining process reduces the complexity of manual configuration, allowing clusters to be deployed and run more quickly. Efficiently manage resources. By setting resource requests and limits at the Pod level, this solution ensures better resource management from the start and optimizes resource utilization. Improve system scalability and allow dynamic addition of nodes to expand cluster capabilities to meet the needs of larger business loads and rapid growth. Enhance management and monitoring capabilities. The visual interface provides convenient monitoring and management methods, allowing operation and maintenance personnel to intuitively understand cluster status, resource usage and application performance. Improve service availability and reliability. Through clustered management, applications have high availability and can automatically handle node or container failures to improve service continuity. Through automation and tooling, efficient and flexible cluster management and application deployment are achieved, meeting the needs of modern applications for agility and reliability.

[0097] In one embodiment of the present invention, a business application is deployed in a microservice orchestration cluster, and the number of application instances is fixed, including:

[0098] Create a Docker image for the business application and push the Docker image to a private container registry using Docker commands;

[0099] Create a Kubernetes deployment configuration, that is, create a YAML file and define the Deployment resource in the YAML file;

[0100] Use the kubectl command to apply the YAML file to deploy the business application;

[0101] Verify that the Pods are running successfully. If so, check that the number of Pods created by the Deployment matches the specified number of replicas, as specified by the spec.replicas field in the YAML file.

[0102] Create a Service, apply the Service configuration file, and check whether the Service has been successfully created and is running;

[0103] Find the external IP address assigned to the Service and access it in a browser or using the curl command to verify that the application is running and can be load balanced.

[0104] The working principle and effect of the above technical solution are as follows: Create a Docker image and use the Docker build tool to create a Docker image for the business application. This step typically includes writing a Dockerfile to define the environment required to build and run the application; pushing the image to a private registry using the docker push command to upload the created Docker image to the private container registry so that the Kubernetes cluster can pull and use the image for application deployment; and creating a Kubernetes deployment configuration by writing a Kubernetes YAML file and defining a Deployment resource. In this file, specify the image, resource requests and limits, and the number of replicas (spec.replicas) required for the application to control the horizontal scaling of the application; use the kubectl apply-f command to apply the defined YAML file to the Kubernetes cluster to create and deploy business applications; verify the running status of Pods and confirm that the Pods are running successfully. You can use the kubectl get pods command to view its status and check whether the number of Pods created by the Deployment matches the number of replicas specified by the spec.replicas field in the YAML file to ensure that the required number of application instances are deployed correctly; create and configure a Service. Use another YAML file to define and apply a Kubernetes Service, which will expose the application so that external traffic can access it. Use the kubectl apply-f command to create a Service and verify that the Service is running successfully through kubectl get svc; access the application, find the external IP address assigned to the Service, and use a browser or curl command to access the IP address to verify that the application is running properly and can distribute traffic through the load balancer. Automating application building and deployment through Docker and Kubernetes configuration files ensures consistency and repeatable environment setup and application deployment; using the fixed number of replicas defined by the spec.replicas field controls the number of application instances, ensuring that the application has enough instances to handle the load; Kubernetes' self-healing capabilities ensure that the application can automatically recover in the event of a failure, thereby improving system reliability and high availability; Service distributes traffic through a load balancer to ensure that each application instance handles requests evenly, improving resource utilization efficiency and user experience; easy to access and test, Service provides a fixed entry point to access the application, simplifying application testing and user access paths; through containerization, orchestration and network service configuration, stable, efficient and highly available microservice environment deployment is achieved.

[0105] In one embodiment of the present invention, obtaining application monitoring data and predicting a microservice expansion threshold includes:

[0106] Deploy the monitoring tool Prometheus in the Kubernetes microservice orchestration cluster;

[0107] Prometheus captures monitoring data from Pods deployed in the microservice orchestration cluster, including CPU usage, memory usage, network traffic, number of requests, and response time.

[0108] Obtain the monitoring history data captured by the monitoring tool Prometheus on all Pods;

[0109] Processing missing values, abnormal values ​​and noise data in the monitoring historical data;

[0110] extracting features from the processed data, the features including a timestamp, a request type, and a node status;

[0111] Divide the data after feature extraction into training set, validation set and test set;

[0112] Use the training set data to train the LSTM model, adjust the model parameters through grid search, and use cross-validation based on the validation set to avoid overfitting;

[0113] The trained model is deployed on the test set for testing to obtain the final predicted expansion threshold model.

[0114] The working principle and effect of the above technical solution are as follows: deploy the Prometheus monitoring tool, deploy Prometheus in the Kubernetes cluster to collect Pod performance and resource usage data in real time; define the Prometheus configuration file so that it can discover and capture the monitoring indicators of each Pod in the microservice cluster; Prometheus captures monitoring data from each Pod through HTTP polling. The main indicators include CPU usage, memory usage, network traffic, number of requests and response time, etc.; obtain historical data, Prometheus stores the collected data and provides a query interface to obtain complete monitoring history data; data preprocessing, in the acquired historical monitoring data, perform data cleaning, process missing values, outliers and noise data to ensure data quality; feature extraction, extract useful features from the cleaned data, such as timestamps, request types, node status, etc., for subsequent model training; data set segmentation, divide the processed data into training sets, validation sets and test sets in proportion to prepare for model training and evaluation; train LSTM model, use training set data to train LSTM (long short-term memory) model, process time series prediction tasks, adjust the model's hyperparameters through grid search, and use Cross-validation is performed on the validation set to optimize model performance and avoid overfitting. Model testing and deployment involves applying the tuned model to the test set to evaluate its performance and predict scaling thresholds. The validated model is considered the final version and can be used for real-time monitoring and automatic scaling decision support. Prometheus provides efficient real-time data collection, ensuring that microservice performance metrics are readily available, laying the foundation for subsequent predictive analysis. The data preprocessing phase effectively improves the quality of monitoring data, making subsequent analysis more accurate. To ensure the accuracy of the predictive model, the LSTM model is used for time series prediction, effectively capturing temporal dependencies in the data and providing more accurate predictions for scaling decisions. Efficient resource management, through effective scaling threshold prediction, optimizes resource usage and reduces unnecessary resource waste while maintaining application performance. Grid search is used to adjust model parameters to ensure the model can dynamically adapt to varying load conditions. System stability and availability are enhanced, and scaling-in and scaling-out measures based on the predictive model can effectively cope with traffic peaks, ensuring stable and efficient system operation. This solution combines monitoring and predictive analysis to improve resource management efficiency and system stability for microservice applications in Kubernetes clusters through automated and intelligent means.

[0115] One embodiment of the present invention dynamically adjusts the cluster size and the number of deployed applications based on real-time monitoring data and expansion thresholds, expanding capacity when capacity exceeds the threshold and shrinking capacity when capacity falls below the threshold, including:

[0116] Collect real-time monitoring data from the cluster, including CPU usage, memory usage, network traffic, number of requests, and response time;

[0117] Inputting the real-time monitoring data into the predicted capacity expansion threshold model to obtain the predicted capacity expansion threshold;

[0118] Dynamically adjust the cluster size based on the predicted expansion threshold, dynamically adjust

[0119] The number of application deployments. When the cluster size and the number of application deployments exceed the threshold, the capacity is expanded; when it falls below the threshold, the capacity is reduced, including:

[0120] Use fuzzy logic to evaluate the load of each resource and obtain the fuzzy score M(t).

[0121]

[0122] Where x is the current resource usage (CPU, memory, network traffic, number of requests, and response time), k and c are parameters for adjusting the fuzzy logic characteristics, and M(t) represents the fuzzy score at the current time t;

[0123] Introduce dynamic weight W and weighted Bayesian inference formula:

[0124]

[0125] Among them, P(S|D) represents the probability that the system needs to be expanded given the data D, and W prior represents the preset dynamic weight, P(D|S) represents the probability of observing the current data when capacity expansion is required, P(S) represents the prior probability of event S (capacity expansion is required), and P(D) represents the overall probability of observing the current data D;

[0126]

[0127] Among them, S(t) is the final scaling decision score, which reflects the comprehensive changes of different indicators over time, T d (τ) = e -λ(t-τ) ,τ represents the time window length;

[0128] When S(t) is greater than a first preset threshold, it is decided to expand the capacity; when S(t) is less than a second preset threshold, it is decided to reduce the capacity.

[0129] The first preset threshold and the second preset threshold are obtained by a predicted expansion threshold model. The first preset threshold is obtained by adding a preset value to the value predicted by the predicted expansion threshold model, and the second preset threshold is obtained by subtracting a preset value from the value predicted by the predicted expansion threshold model. The preset values ​​are obtained by calculating the standard deviation of the relevant monitoring data using historical monitoring data.

[0130] The working principle and effect of the above technical solution are as follows: A fuzzy logic function is used to evaluate the utilization of each resource, resulting in a fuzzy score M(t), which is a standard Sigmoid function used to quantify the severity of resource load; dynamic weights W and weighted Bayesian inference are introduced to evaluate the need for system expansion. The Bayesian inference formula is applied to calculate the probability of expansion under the current state; dynamic response capabilities are enhanced, and through the combination of real-time monitoring and prediction models, dynamic response to resource demand is achieved, optimizing resource utilization and performance; intelligent decision support, combining fuzzy logic with Bayesian inference, improves the intelligent understanding of complex load conditions and supports more accurate decision-making; efficient resource management and automated scaling mechanisms reduce human intervention and misjudgment, improving the overall efficiency and stability of the system; using the time decay factor T_d(τ) to weight data at different time points, making the system more sensitive to recent conditions. The S-curve is used to convert resource load into a fuzzy score, which smooths out decision changes at critical points and improves fault tolerance and flexibility. The weighted Bayesian formula allows for more robust probability estimates in uncertainty, and dynamic weights enhance model adaptability. By integrating multi-dimensional indicators and time decay, a time- and load-sensitive comprehensive score is generated, providing a more forward-looking basis for decision-making. The rational integration of real-time data analysis, prediction, and decision-making mechanisms provides an automated and intelligent solution for dynamic scaling, effectively improving the resource utilization and operational efficiency of the Kubernetes microservice system.

[0131] One embodiment of the present invention provides a system for a microservice container orchestration platform to call temporary resources, the system comprising:

[0132] Fixed cluster resource scale module, used to deploy microservice orchestration cluster Kubernetes and fix cluster resource scale;

[0133] The fixed application instance number module is used to deploy business applications in the microservice orchestration cluster and fix the number of application instances;

[0134] The prediction module is used to obtain application monitoring data and predict microservice expansion thresholds;

[0135] The adjustment module is used to dynamically adjust the cluster size and the number of application deployments based on real-time monitoring data and expansion thresholds, expanding the capacity when the threshold is exceeded and reducing the capacity when the threshold is exceeded.

[0136] The working principle and effect of the above technical solution are as follows: fixed initial resources and instances, fixed cluster resource scale, at the beginning, deploy a Kubernetes cluster with fixed resource configuration, which involves setting the number of nodes and the resources of each node (such as CPU, memory); deploy application instances, deploy business applications on the cluster, and initialize a certain number of application instances (i.e., Pods), which is specified by the spec.replicas of the Kubernetes Deployment; obtain and monitor data, using monitoring tools such as Prometheus to collect the operating status and performance data of applications and clusters, including CPU, memory utilization, response time, etc.; use analytical tools to process monitoring data, and predict system load changes based on historical data and real-time data; use machine learning algorithms or rule-based policies to analyze monitoring data, set or adjust expansion / contraction thresholds for applications and clusters, and repeat learning and adjustment to adapt the thresholds to load changes; compare monitoring data with set thresholds in real time. If the expansion threshold is above the expansion threshold, cluster resources (such as adding nodes) and application instances (adding Pods) are automatically expanded to ensure performance; if the expansion threshold is below the contraction threshold, cluster resources (such as reducing nodes) and application instances are reduced to conserve resources. Improve resource utilization efficiency, dynamically adjust cluster and application scale to match actual load requirements, and optimize resource usage to the greatest extent possible; improve system performance and user experience, and ensure that applications always have sufficient resources and instances to support requests and workloads by quickly responding to load changes; automatically scale down when the load decreases to avoid unnecessary resource occupation, thereby reducing operating costs; enhance elasticity and scalability, the system can automatically respond to load changes, has good elasticity and scalability, and is suitable for highly volatile business scenarios; simplify operation and maintenance management, and reduce the need for manual operation and maintenance intervention through automated expansion and contraction strategies, making management simpler and more efficient. Through the combination of fixed and dynamic strategies, it can provide a stable foundation while flexibly adapting to changing workload requirements.

[0137] In one embodiment of the present invention, the fixed cluster resource scale module includes:

[0138] Create a cluster module, obtain the target software package, upload the target software package to the master node virtual machine, create a cluster on the master node virtual machine, and add the slave node virtual machines to the cluster;

[0139] The image pulling module is used to pull the image corresponding to the slave node virtual machine through the Docker pull command on the slave node virtual machine after the cluster is successfully built, and then pull the image corresponding to the master node virtual machine through the Docker pull command on the master node virtual machine;

[0140] Start the container module, which is used to start the container corresponding to the image through the Docker run command after the image is successfully pulled. After obtaining the container ID of the successfully started container, you can access the visual interface of the container orchestration tool through the browser.

[0141] The working principle and effect of the above technical solution are as follows: fixed initial resources and instances, fixed cluster resource scale, at the beginning, deploy a Kubernetes cluster with fixed resource configuration, which involves setting the number of nodes and the resources of each node (such as CPU, memory); deploy application instances, deploy business applications on the cluster, and initialize a certain number of application instances (i.e., Pods), which is specified by the spec.replicas of the Kubernetes Deployment; obtain and monitor data, using monitoring tools such as Prometheus to collect the operating status and performance data of applications and clusters, including CPU, memory utilization, response time, etc.; use analytical tools to process monitoring data, and predict system load changes based on historical data and real-time data; use machine learning algorithms or rule-based policies to analyze monitoring data, set or adjust expansion / contraction thresholds for applications and clusters, and repeat learning and adjustment to adapt the thresholds to load changes; compare monitoring data with set thresholds in real time. If the expansion threshold is above the expansion threshold, cluster resources (such as adding nodes) and application instances (adding Pods) are automatically expanded to ensure performance; if the expansion threshold is below the contraction threshold, cluster resources (such as reducing nodes) and application instances are reduced to conserve resources. Improve resource utilization efficiency, dynamically adjust cluster and application scale to match actual load requirements, and optimize resource usage to the greatest extent possible; improve system performance and user experience, and ensure that applications always have sufficient resources and instances to support requests and workloads by quickly responding to load changes; automatically scale down when the load decreases to avoid unnecessary resource occupation, thereby reducing operating costs; enhance elasticity and scalability, the system can automatically respond to load changes, has good elasticity and scalability, and is suitable for highly volatile business scenarios; simplify operation and maintenance management, and reduce the need for manual operation and maintenance intervention through automated expansion and contraction strategies, making management simpler and more efficient. Through the combination of fixed and dynamic strategies, it can provide a stable foundation while flexibly adapting to changing workload requirements.

[0142] In one embodiment of the present invention, the module for fixing the number of application instances includes:

[0143] The push image module is used to create a Docker image for the business application and push the Docker image to the private container registry using Docker commands;

[0144] Create a YAML file module for creating Kubernetes deployment configurations, that is, create a YAML file and define the Deployment resource in the YAML file;

[0145] The deployment application module is used to apply the YAML file using the kubectl command to deploy business applications;

[0146] The verification module verifies that the Pods are running successfully. If the verification passes, it checks that the number of Pods created by the Deployment matches the specified number of replicas, which is implemented by the spec.replicas field in the YAML file.

[0147] Verify the Service module, which is used to create a Service, apply the Service configuration file, and check whether the Service has been successfully created and run;

[0148] The Verify Run module is used to find the external IP address assigned to the Service, and then access the address in a browser or using the curl command to verify whether the application can run and load balance.

[0149] The working principle and effect of the above technical solution are as follows: Create a Docker image and use the Docker build tool to create a Docker image for the business application. This step typically includes writing a Dockerfile to define the environment required to build and run the application; pushing the image to a private registry using the docker push command to upload the created Docker image to the private container registry so that the Kubernetes cluster can pull and use the image for application deployment; and creating a Kubernetes deployment configuration by writing a Kubernetes YAML file and defining a Deployment resource. In this file, specify the image, resource requests and limits, and the number of replicas (spec.replicas) required for the application to control the horizontal scaling of the application; use the kubectl apply-f command to apply the defined YAML file to the Kubernetes cluster to create and deploy business applications; verify the running status of Pods and confirm that the Pods are running successfully. You can use the kubectl get pods command to view its status and check whether the number of Pods created by the Deployment matches the number of replicas specified by the spec.replicas field in the YAML file to ensure that the required number of application instances are deployed correctly; create and configure a Service. Use another YAML file to define and apply a Kubernetes Service, which will expose the application so that external traffic can access it. Use the kubectl apply-f command to create a Service and verify that the Service is running successfully through kubectl get svc; access the application, find the external IP address assigned to the Service, and use a browser or curl command to access the IP address to verify that the application is running properly and can distribute traffic through the load balancer. Automating application building and deployment through Docker and Kubernetes configuration files ensures consistency and repeatable environment setup and application deployment; using the fixed number of replicas defined by the spec.replicas field controls the number of application instances, ensuring that the application has enough instances to handle the load; Kubernetes' self-healing capabilities ensure that the application can automatically recover in the event of a failure, thereby improving system reliability and high availability; Service distributes traffic through a load balancer to ensure that each application instance handles requests evenly, improving resource utilization efficiency and user experience; easy to access and test, Service provides a fixed entry point to access the application, simplifying application testing and user access paths; through containerization, orchestration and network service configuration, stable, efficient and highly available microservice environment deployment is achieved.

[0150] In one embodiment of the present invention, the prediction module includes:

[0151] Deployment monitoring tool module, used to deploy monitoring tool Prometheus in Kubernetes microservice orchestration cluster;

[0152] The data capture module, Prometheus, captures monitoring data from the Pod deployed in the microservice orchestration cluster;

[0153] A historical data acquisition module is used to obtain monitoring historical data captured by the monitoring tool Prometheus on all Pods;

[0154] A feature extraction module is used to process missing values, abnormal values ​​and noise data in the monitoring historical data;

[0155] a segmentation dataset module for extracting features from the processed data, the features including timestamp, request type, and node status;

[0156] The dataset segmentation module divides the feature-extracted data into training set, validation set, and test set;

[0157] The training and validation module uses the training set data to train the LSTM model, adjusts the model parameters through grid search, and uses cross-validation based on the validation set to avoid overfitting;

[0158] The testing module is used to deploy the trained model to the test set for testing to obtain the final prediction expansion threshold model.

[0159] The working principle and effect of the above technical solution are as follows: deploy the Prometheus monitoring tool, deploy Prometheus in the Kubernetes cluster to collect Pod performance and resource usage data in real time; define the Prometheus configuration file so that it can discover and capture the monitoring indicators of each Pod in the microservice cluster; Prometheus captures monitoring data from each Pod through HTTP polling. The main indicators include CPU usage, memory usage, network traffic, number of requests and response time, etc.; obtain historical data, Prometheus stores the collected data and provides a query interface to obtain complete monitoring history data; data preprocessing, in the acquired historical monitoring data, perform data cleaning, process missing values, outliers and noise data to ensure data quality; feature extraction, extract useful features from the cleaned data, such as timestamps, request types, node status, etc., for subsequent model training; data set segmentation, divide the processed data into training sets, validation sets and test sets in proportion to prepare for model training and evaluation; train LSTM model, use training set data to train LSTM (long short-term memory) model, process time series prediction tasks, adjust the model's hyperparameters through grid search, and use Cross-validation is performed on the validation set to optimize model performance and avoid overfitting. Model testing and deployment involves applying the tuned model to the test set to evaluate its performance and predict scaling thresholds. The validated model is considered the final version and can be used for real-time monitoring and automatic scaling decision support. Prometheus provides efficient real-time data collection, ensuring that microservice performance metrics are readily available, laying the foundation for subsequent predictive analysis. The data preprocessing phase effectively improves the quality of monitoring data, making subsequent analysis more accurate. To ensure the accuracy of the predictive model, the LSTM model is used for time series prediction, effectively capturing temporal dependencies in the data and providing more accurate predictions for scaling decisions. Efficient resource management, through effective scaling threshold prediction, optimizes resource usage and reduces unnecessary resource waste while maintaining application performance. Grid search is used to adjust model parameters to ensure the model can dynamically adapt to varying load conditions. System stability and availability are enhanced, and scaling-in and scaling-out measures based on the predictive model can effectively cope with traffic peaks, ensuring stable and efficient system operation. This solution combines monitoring and predictive analysis to improve resource management efficiency and system stability for microservice applications in Kubernetes clusters through automated and intelligent means.

[0160] In one embodiment of the present invention, the adjustment module includes:

[0161] Get real-time monitoring data module, used to collect real-time monitoring data from the cluster, including CPU usage, memory usage, network traffic, number of requests, and response time;

[0162] A prediction threshold acquisition module is used to input real-time monitoring data into the prediction expansion threshold model to obtain a prediction expansion threshold;

[0163] The scaling module is used to dynamically adjust the cluster size and the number of deployed applications based on the predicted expansion threshold. When the cluster size and the number of deployed applications are above the threshold, the cluster size is expanded, and when they are below the threshold, the cluster size is reduced. This module includes:

[0164] Use fuzzy logic to evaluate the load of each resource and obtain the fuzzy score M(t).

[0165]

[0166] Where x is the current resource usage (CPU, memory, network traffic, number of requests, and response time), k and c are parameters for adjusting the fuzzy logic characteristics, and M(t) represents the fuzzy score at the current time t;

[0167] Introduce dynamic weight W and weighted Bayesian inference formula:

[0168]

[0169] Among them, P(S|D) represents the probability that the system needs to be expanded given the data D, and W prior represents the preset dynamic weight, P(D|S) represents the probability of observing the current data when capacity expansion is required, P(S) represents the prior probability of event S (capacity expansion is required), and P(D) represents the overall probability of observing the current data D;

[0170] The final scaling decision score of the current microservice container orchestration platform is calculated based on the fuzzy score of the resources of the current microservice container orchestration platform, the calculated probability that the system needs to be expanded, and the calculated value of the time decay function:

[0171]

[0172] Among them, S(t) is the final scaling decision score, which reflects the comprehensive changes of different indicators over time, T d (τ) = e -λ(t-τ) ,τ represents the time window length, T d (τ) represents the calculated value of the time decay function within τ, and P(S|D(t)) represents the probability that the system needs to expand capacity under given data D within τ;

[0173] When S(t) is greater than a first preset threshold, it is decided to expand the capacity; when S(t) is less than a second preset threshold, it is decided to reduce the capacity.

[0174] The first preset threshold and the second preset threshold are obtained by a predicted expansion threshold model. The first preset threshold is obtained by adding a preset value to the value predicted by the predicted expansion threshold model, and the second preset threshold is obtained by subtracting a preset value from the value predicted by the predicted expansion threshold model. The preset values ​​are obtained by calculating the standard deviation of the relevant monitoring data using historical monitoring data.

[0175] The working principle and effect of the above technical solution are as follows: A fuzzy logic function is used to evaluate the utilization of each resource, resulting in a fuzzy score M(t), which is a standard Sigmoid function used to quantify the severity of resource load; dynamic weights W and weighted Bayesian inference are introduced to evaluate the need for system expansion. The Bayesian inference formula is applied to calculate the probability of expansion under the current state; dynamic response capabilities are enhanced, and through the combination of real-time monitoring and prediction models, dynamic response to resource demand is achieved, optimizing resource utilization and performance; intelligent decision support, combining fuzzy logic with Bayesian inference, improves the intelligent understanding of complex load conditions and supports more accurate decision-making; efficient resource management and automated scaling mechanisms reduce human intervention and misjudgment, improving the overall efficiency and stability of the system; using the time decay factor T_d(τ) to weight data at different time points, making the system more sensitive to recent conditions. The S-curve is used to convert resource load into a fuzzy score, which smooths out decision changes at critical points and improves fault tolerance and flexibility. The weighted Bayesian formula allows for more robust probability estimates in uncertainty, and dynamic weights enhance model adaptability. By integrating multi-dimensional indicators and time decay, a time- and load-sensitive comprehensive score is generated, providing a more forward-looking basis for decision-making. The rational integration of real-time data analysis, prediction, and decision-making mechanisms provides an automated and intelligent solution for dynamic scaling, effectively improving the resource utilization and operational efficiency of the Kubernetes microservice system.

[0176] Obviously, those skilled in the art may make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if such changes and modifications fall within the scope of the claims and their equivalents, the present invention is intended to include such changes and modifications.

Claims

1. A method for calling temporary resources on a microservice container orchestration platform, characterized in that: The method includes: deploying a microservice orchestration cluster Kubernetes and fixing the cluster resource scale; Deploy business applications in the microservice orchestration cluster and fix the number of application instances; Obtain application monitoring data and predict microservice expansion thresholds; Based on real-time monitoring data and expansion thresholds, the cluster size and the number of application deployments are dynamically adjusted, with capacity expanded when the threshold is exceeded and capacity reduced when the threshold is exceeded.

2. According to the method of claim 1, a microservice container orchestration platform calls temporary resources, characterized in that: Deploy a microservice orchestration cluster, Kubernetes, and fix the cluster resource scale, including: Obtain the target software package, upload the target software package to the master node virtual machine, create a cluster on the master node virtual machine, use resource requests and limits when creating pods in the cluster, and add the slave node virtual machines to the cluster; After the cluster is successfully built, the image corresponding to the slave node virtual machine is pulled by the Dockerpull command on the slave node virtual machine, and then the image corresponding to the master node virtual machine is pulled by the Dockerpull command on the master node virtual machine; After the image is successfully pulled, start the container corresponding to the image through the Docker run command, obtain the container ID that was successfully started, and then access the visual interface of the container orchestration tool through the browser.

3. According to the method of a microservice container orchestration platform calling temporary resources according to claim 1, it is characterized in that: Deploy business applications in a microservices orchestration cluster and fix the number of application instances, including: Create a Docker image for the business application and push the Docker image to a private container registry using Docker commands; Create a Kubernetes deployment configuration, that is, create a YAML file and define the Deployment resource in the YAML file; Use the kubectl command to apply the YAML file to deploy the business application; Verify that the Pods are running successfully. If so, check that the number of Pods created by the Deployment matches the specified number of replicas, as specified by the spec.replicas field in the YAML file. Create a Service, apply the Service configuration file, and check whether the Service has been successfully created and is running; Find the external IP address assigned to the Service and access it in a browser or using the curl command to verify that the application is running and can be load balanced.

4. According to the method of a microservice container orchestration platform calling temporary resources according to claim 1, it is characterized in that: The feature is to obtain application monitoring data and predict microservice expansion thresholds, including: Deploy the monitoring tool Prometheus in the Kubernetes microservice orchestration cluster; Prometheus captures monitoring data from Pods deployed in the microservice orchestration cluster, including CPU usage, memory usage, network traffic, number of requests, and response time. Obtain the monitoring history data captured by the monitoring tool Prometheus on all Pods; Processing missing values, abnormal values ​​and noise data in the monitoring historical data; extracting features from the processed data, the features including a timestamp, a request type, and a node status; Divide the data after feature extraction into training set, validation set and test set; Use the training set data to train the LSTM model, adjust the model parameters through grid search, and use cross-validation based on the validation set to avoid overfitting; The trained model is deployed on the test set for testing to obtain the final predicted expansion threshold model.

5. According to the method of a microservice container orchestration platform calling temporary resources according to claim 1, it is characterized in that: Based on real-time monitoring data and expansion thresholds, the cluster size and the number of deployed applications are dynamically adjusted. When the capacity exceeds the threshold, the capacity is expanded, and when it falls below the threshold, the capacity is reduced, including: Collect real-time monitoring data from the cluster, including CPU usage, memory usage, network traffic, number of requests, and response time; Inputting the real-time monitoring data into the predicted capacity expansion threshold model to obtain the predicted capacity expansion threshold; Dynamically adjust the cluster size and the number of deployed applications based on the predicted expansion threshold. When the cluster size and the number of deployed applications exceed the threshold, the cluster size is expanded; when they fall below the threshold, the cluster size is reduced. This includes: Use fuzzy logic to evaluate the load of each resource and obtain the fuzzy score M(t). Where x is the current resource usage (CPU, memory, network traffic, number of requests, and response time), k and c are parameters for adjusting the fuzzy logic characteristics, and M(t) represents the fuzzy score at the current time t; Introduce dynamic weight W and weighted Bayesian inference formula: Among them, P(S|D) represents the probability that the system needs to be expanded given the data D, and W prior represents the preset dynamic weight, P(D|S) represents the probability of observing the current data when capacity expansion is required, P(S) represents the prior probability of event S (capacity expansion is required), and P(D) represents the overall probability of observing the current data D; Among them, S(t) is the final scaling decision score at time t, which reflects the comprehensive changes of different indicators over time. d (τ) = e -λ(t-τ) ,τ represents the time window length, M CPU (τ) represents the fuzzy score of CPU usage within τ, M Memory (τ) represents the fuzzy score of memory usage within τ, M Network (τ) represents the fuzzy score of network traffic within τ, M Latency (τ) represents the fuzzy score of response time within τ; When S(t) is greater than a first preset threshold, it is decided to expand the capacity; when S(t) is less than a second preset threshold, it is decided to reduce the capacity.

6. A system for calling temporary resources on a microservice container orchestration platform, characterized in that: The system comprises: Fixed cluster resource scale module, used to deploy microservice orchestration cluster Kubernetes and fix cluster resource scale; The fixed application instance number module is used to deploy business applications in the microservice orchestration cluster and fix the number of application instances; The prediction module is used to obtain application monitoring data and predict microservice expansion thresholds; The adjustment module is used to dynamically adjust the cluster size and the number of application deployments based on real-time monitoring data and expansion thresholds, expanding the capacity when the threshold is exceeded and reducing the capacity when the threshold is exceeded.

7. A system for calling temporary resources on a microservice container orchestration platform according to claim 6, characterized in that: The fixed cluster resource scale module includes: Create a cluster module, obtain the target software package, upload the target software package to the master node virtual machine, create a cluster on the master node virtual machine, and add the slave node virtual machines to the cluster; The image pulling module is used to pull the image corresponding to the slave node virtual machine through the Docker pull command on the slave node virtual machine after the cluster is successfully built, and then pull the image corresponding to the master node virtual machine through the Docker pull command on the master node virtual machine; Start the container module, which is used to start the container corresponding to the image through the Docker run command after the image is successfully pulled. After obtaining the container ID of the successfully started container, you can access the visual interface of the container orchestration tool through the browser.

8. A system for calling temporary resources on a microservice container orchestration platform according to claim 6, characterized in that: The module for fixing the number of application instances includes: The push image module is used to create a Docker image for the business application and push the Docker image to the private container registry using Docker commands; Create a YAML file module for creating Kubernetes deployment configurations, that is, create a YAML file and define the Deployment resource in the YAML file; The deployment application module is used to apply the YAML file using the kubectl command to deploy business applications; The verification module verifies that the Pods are running successfully. If the verification passes, it checks that the number of Pods created by the Deployment matches the specified number of replicas, which is implemented by the spec.replicas field in the YAML file. Verify the Service module, which is used to create a Service, apply the Service configuration file, and check whether the Service has been successfully created and run; The Verify Run module is used to find the external IP address assigned to the Service, and then access the address in a browser or using the curl command to verify whether the application can run and load balance.

9. A system for calling temporary resources on a microservice container orchestration platform according to claim 6, characterized in that: The prediction module includes: Deployment monitoring tool module, used to deploy monitoring tool Prometheus in Kubernetes microservice orchestration cluster; The data capture module, Prometheus, captures monitoring data from the Pod deployed in the microservice orchestration cluster; A historical data acquisition module is used to obtain monitoring historical data captured by the monitoring tool Prometheus on all Pods; A feature extraction module is used to process missing values, abnormal values ​​and noise data in the monitoring historical data; a segmentation dataset module for extracting features from the processed data, the features including timestamp, request type, and node status; The dataset segmentation module divides the feature-extracted data into training set, validation set, and test set; The training and validation module uses the training set data to train the LSTM model, adjusts the model parameters through grid search, and uses cross-validation based on the validation set to avoid overfitting; The testing module is used to deploy the trained model to the test set for testing to obtain the final prediction expansion threshold model.

10. A system for calling temporary resources on a microservice container orchestration platform according to claim 6, characterized in that: The adjustment module includes: Get real-time monitoring data module, used to collect real-time monitoring data from the cluster, including CPU usage, memory usage, network traffic, number of requests, and response time; A prediction threshold acquisition module is used to input real-time monitoring data into the prediction expansion threshold model to obtain a prediction expansion threshold; The scaling module is used to dynamically adjust the cluster size and the number of deployed applications based on the predicted expansion threshold. When the cluster size and the number of deployed applications are above the threshold, the cluster size is expanded, and when they are below the threshold, the cluster size is reduced. This module includes: Use fuzzy logic to evaluate the load of each resource and obtain the fuzzy score M(t). Where x is the current resource usage (CPU, memory, network traffic, number of requests, and response time), k and c are parameters for adjusting the fuzzy logic characteristics, and M(t) represents the fuzzy score at the current time t; Introduce dynamic weight W and weighted Bayesian inference formula: Among them, P(S|D) represents the probability that the system needs to be expanded given the data D, and W prior represents the preset dynamic weight, P(D|S) represents the probability of observing the current data when capacity expansion is required, P(S) represents the prior probability of event S (capacity expansion is required), and P(D) represents the overall probability of observing the current data D; The final scaling decision score of the current microservice container orchestration platform is calculated based on the fuzzy score of the resources of the current microservice container orchestration platform, the calculated probability that the system needs to be expanded, and the calculated value of the time decay function: Among them, S(t) is the final scaling decision score at time t, which reflects the comprehensive changes of different indicators over time. d (τ) = e -λ(t-τ) ,τ represents the time window length, M CPU (τ) represents the fuzzy score of CPU usage within τ, M Memory (τ) represents the fuzzy score of memory usage within τ, M Network (τ) represents the fuzzy score of network traffic within τ, M Latency (τ) represents the fuzzy score of the response time within τ, tD(τ) represents the calculated value of the time decay function within τ, and P(S|D(τ)) represents the probability that the system needs to be expanded under the given data D within τ; When S(t) is greater than a first preset threshold, it is decided to expand the capacity; when S(t) is less than a second preset threshold, it is decided to reduce the capacity.

Citation Information

Patent Citations

  • Online application dynamic capacity expansion and shrinkage method based on micro-service call dependence perception

    CN112199150A

  • Dynamic capacity expansion and contraction method and device for satellite edge computing service and storage medium

    CN117573339A

  • Automatic operation and maintenance method based on container and big data

    CN117971384A

  • Container energy-saving elastic capacity expansion and contraction method and system based on time sequence prediction and medium

    CN118260021A

  • Service return-to-zero scaling method and system based on cloud native

    CN118819742A

Cited By

  • Collaborative method and system for containerized mixed deployment of database and financial software

    CN121743392A