Database container network optimization scheduling method and device and electronic equipment
By obtaining the network and connection status data of the database container, using pre-trained models for resource evaluation and prediction, and dynamically adjusting resource allocation, the problem of untimely resource allocation in the Kubernetes network scheduling mechanism is solved, and the processing efficiency of high concurrent services is improved.
Patent Information
- Application Number
- CN202510588138.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-08
- Publication Date
- 2025-07-29
AI Technical Summary
The existing Kubernetes network scheduling mechanism cannot accurately identify the high concurrent business needs of database containers, resulting in untimely resource allocation and affecting business efficiency, especially in financial transaction systems, transaction request responses are slow.
By obtaining the network status and connection status data of the database container, using pre-trained target requirements to evaluate the network and traffic prediction model, dynamically adjust resource allocation strategies, including network bandwidth and connection number, and combining reinforcement learning and deep learning technologies for resource scheduling.
It realizes accurate prediction of database container requirements and traffic, improves resource allocation efficiency, ensures stable operation of high concurrent services, reduces response time and improves business processing efficiency of Kubernetes clusters.
Smart Images

Figure CN120389954A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of big data processing, and in particular, to a method, device, and electronic device for optimizing the scheduling of database container networks. Background Art
[0002] The Kubernetes cluster is a distributed system composed of multiple nodes, used for automating the deployment, scaling, and management of containerized applications. The Kubernetes network scheduling mechanism is an important factor affecting the task processing efficiency in the Kubernetes cluster.
[0003] The existing Kubernetes network scheduling mechanism mainly allocates resources based on simple network metrics. For example, it only focuses on the overall usage of network bandwidth. In scenarios such as financial trading systems where extremely high requirements for data real-time are imposed, database containers generate a large number of short connection requests. The traditional scheduling mechanism cannot accurately identify and meet these special requirements, resulting in insufficient network connection resources for critical services, slow response to transaction requests, and seriously affecting business efficiency. Facing the network requirements of database containers that change with business loads, the existing scheduling mechanism also lacks the ability to dynamically adjust. This situation of untimely resource allocation greatly affects the processing efficiency of each service in the Kubernetes cluster. Summary of the Invention
[0004] In view of this, embodiments of the present invention provide a method, device, and electronic device for optimizing the scheduling of database container networks to improve the efficiency of network resource allocation and further improve the business processing efficiency of the Kubernetes cluster.
[0005] According to one aspect of the present invention, there is provided a method for optimizing the scheduling of database container networks. The method is applied to a Kubernetes cluster and includes:
[0006] Obtain target metric data of a database container, where the target metric data includes: network status data and connection status data;
[0007] Input the target metric data into a pre-trained target demand evaluation network, so that the target demand evaluation network outputs a target demand level of the database container based on the target metric data;
[0008] Input the target metric data into a pre-trained target traffic prediction model, so that the target traffic prediction model outputs a target traffic prediction result of the database container based on the target metric data;
[0009] Determine the target resource quantity that the database container needs to increase based on the current load data of the database container, the target demand level, and the target traffic prediction result, where the target resource quantity includes network bandwidth and the number of connections;
[0010] Perform resource scheduling for the database container according to the target resource quantity of the database container and the task type executed by the database container, where the task type is used to define the allocation weights of the network bandwidth and the number of connections during the resource scheduling process.
[0011] In a possible embodiment, the target demand assessment network is pre-trained through the following steps:
[0012] Obtain a demand training data set, where the demand training data set contains multiple pieces of demand training data, and each piece of demand training data includes historical database network status data, historical database connection status data, and historical database operation types;
[0013] Input the demand training data set into an initial demand assessment network, where the initial demand assessment network is a deep neural network built based on the TensorFlow framework;
[0014] Obtain the predicted demand levels output by the demand assessment network based on each piece of demand training data;
[0015] Train the initial demand assessment network based on the difference between the actual demand level of each piece of demand training data and the predicted demand level until the difference converges, and obtain the target demand assessment network.
[0016] In a possible embodiment, the target traffic prediction model is pre-trained through the following steps:
[0017] Construct a state space based on historical database network status data, historical traffic data, historical database load data, and historical business load data;
[0018] Construct an action space based on different network resource allocation strategies;
[0019] Input the state space and the action space into an initial traffic prediction model, so that the traffic prediction model selects a prediction allocation strategy in the action space based on the state space;
[0020] Calculate the reward function value based on the predicted allocation strategy. In the case where the reward function value does not reach the maximum value of the reward function, modify the parameters of the traffic prediction model, and return to the step of the traffic prediction model selecting a predicted allocation strategy in the action space for the state space until the reward function value reaches the maximum; wherein, the reward function is a business performance improvement metric, and the business performance improvement metric at least includes the amount of response time reduction and throughput.
[0021] Determine the traffic prediction model corresponding to when the reward function value reaches the maximum as the target traffic prediction model.
[0022] In a possible embodiment, the Kubernetes cluster includes multiple nodes, and an Ethereum client is deployed in each of the nodes. The method further includes:
[0023] In the case of a faulty node, the faulty node broadcasts fault information to other nodes in the Kubernetes cluster.
[0024] The master node in the Kubernetes cluster collects the fault information of each faulty node and broadcasts pre-preparation information to other nodes. The pre-preparation information includes a fault identifier and a fault type; so that other nodes return preparation information to the master node in response to the pre-preparation information.
[0025] In the case where the number of received preparation messages by the master node reaches a preset number, perform fault handling on the faulty node. The fault handling includes task migration of the faulty node and network resource allocation.
[0026] According to another aspect of the present invention, there is provided a database container network optimization scheduling device. The device is applied to a Kubernetes cluster. The device includes:
[0027] An acquisition module, configured to acquire target metric data of a database container. The target metric data includes: network status data and connection status data.
[0028] An input module, configured to input the target metric data into a pre-trained target demand evaluation network, so that the target demand evaluation network outputs a target demand level of the database container based on the target metric data.
[0029] Input the target metric data into a pre-trained target traffic prediction model, so that the target traffic prediction model outputs a target traffic prediction result of the database container based on the target metric data.
[0030] A determination module, configured to determine a target resource quantity that the database container needs to increase based on the current load data of the database container, the target demand level, and the target traffic prediction result, where the target resource quantity includes network bandwidth and the number of connections;
[0031] A scheduling module, configured to perform resource scheduling for the database container according to the target resource quantity of the database container and the type of task executed by the database container, where the type of task is used to define the allocation weights of the network bandwidth and the number of connections during the resource scheduling process.
[0032] In a possible embodiment, the target demand evaluation network is pre-trained through the following steps:
[0033] Obtain a demand training data set, where the demand training data set contains multiple pieces of demand training data, and each piece of demand training data includes historical database network status data, historical database connection status data, and historical database operation types;
[0034] Input the demand training data set into an initial demand evaluation network, where the initial demand evaluation network is a deep neural network built based on the TensorFlow framework;
[0035] Obtain the predicted demand levels output by the demand evaluation network based on each piece of demand training data;
[0036] Train the initial demand evaluation network based on the difference between the actual demand level of each piece of demand training data and the predicted demand level until the difference converges, to obtain a target demand evaluation network.
[0037] In a possible embodiment, the target traffic prediction model is pre-trained through the following steps:
[0038] Based on historical database network status data, historical traffic data, historical database load data, and historical service load data, construct a state space;
[0039] Based on different network resource allocation strategies, construct an action space;
[0040] Input the state space and the action space into an initial traffic prediction model, so that the traffic prediction model selects a predicted allocation strategy in the action space based on the state space;
[0041] Calculate the reward function value based on the predicted allocation strategy. When the reward function value does not reach the maximum value of the reward function, modify the parameters of the traffic prediction model, and return to the step of the traffic prediction model selecting the predicted allocation strategy in the action space for the state space until the reward function value reaches the maximum; wherein, the reward function is a business performance improvement index, and the business performance improvement index at least includes the reduction in response time and throughput.
[0042] Determine the traffic prediction model corresponding to when the reward function value reaches the maximum as the target traffic prediction model.
[0043] In a possible embodiment, the Kubernetes cluster includes multiple nodes, and an Ethereum client is deployed in each of the nodes. The apparatus further includes:
[0044] A fault response module, configured to, when a faulty node appears, broadcast fault information from the faulty node to other nodes in the Kubernetes cluster;
[0045] The master node in the Kubernetes cluster collects the fault information of each faulty node and broadcasts pre-preparation information to other nodes. The pre-preparation information includes a fault identifier and a fault type; so that other nodes return preparation information to the master node in response to the pre-preparation information;
[0046] When the number of received preparation messages at the master node reaches a preset number, perform fault handling on the faulty node. The fault handling includes task migration of the faulty node and network resource allocation.
[0047] According to another aspect of the present invention, there is provided an electronic device, including:
[0048] A processor; and
[0049] A memory storing a program,
[0050] wherein, the program includes instructions, and when the instructions are executed by the processor, the processor executes any one of the above-mentioned database container network optimization scheduling methods.
[0051] According to another aspect of the present invention, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to execute any one of the above-mentioned database container network optimization scheduling methods.
[0052] One or more technical solutions provided in the embodiments of the present invention collect target metric data of a database container including network status data and connection status data, and input the target metric data into a pre-trained target demand evaluation network and a target traffic prediction model to perform demand level evaluation and adaptive traffic prediction of the database container, and determine the target resource quantity that the database container needs to increase based on the demand level evaluation result, the traffic prediction result, and the current load data of the database container, and perform resource allocation for the database container according to the target resource quantity and the current task type executed by the database container. Applying the embodiments of the present invention, using the pre-trained model to output the demand evaluation level and the traffic prediction result based on the multi-modal metric data of the database container realizes the in-depth mining of the processing status of the database container, and further realizes the accurate prediction of the demand and traffic of the database container, so that resource allocation can be quickly responded to business needs, improving the resource allocation efficiency, and further improving the business processing efficiency of the Kubernetes cluster. In addition, through traffic prediction, the traffic peak period can be predicted in advance, and resource allocation can be performed in advance to ensure the stable operation of high-concurrency services. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] In the following description of exemplary embodiments with reference to the accompanying drawings, more details, features, and advantages of the present invention are disclosed, in which:
[0054] Figure 1 It is a schematic flowchart of a method for optimizing network scheduling of a database container provided by an embodiment of the present invention;
[0055] Figure 2 It is another schematic flowchart of a method for optimizing network scheduling of a database container provided by an embodiment of the present invention;
[0056] Figure 3 It is a schematic structural diagram of a device for optimizing network scheduling of a database container provided by an embodiment of the present invention;
[0057] Figure 4 It shows a schematic block diagram of an exemplary electronic device that can be used to implement the embodiments of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0058] The embodiments of the present invention will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present invention are shown in the drawings, it should be understood that the present invention can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present invention. It should be understood that the drawings and embodiments of the present invention are only for exemplary purposes and are not used to limit the protection scope of the present invention.
[0059] It should be understood that the various steps described in the method embodiments of the present invention can be executed in different orders and / or executed in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present invention is not limited in this regard.
[0060] As used herein, the term "including" and its variants are open-ended, that is, "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". The relevant definitions of other terms will be given in the following description. It should be noted that the concepts such as "first" and "second" mentioned in the present invention are only used to distinguish different devices, modules or units, and are not used to limit the order or mutual dependence relationship of the functions performed by these devices, modules or units.
[0061] It should be noted that the modifications of "one" and "multiple" mentioned in the present invention are illustrative rather than restrictive. Those skilled in the art should understand that unless otherwise clearly specified in the context, it should be understood as "one or more".
[0062] The names of the messages or information exchanged between multiple devices in the embodiments of the present invention are only for illustrative purposes and are not used to limit the scope of these messages or information.
[0063] In order to improve the accuracy of network resource allocation, the present invention provides a method, device, electronic device and storage medium for optimizing and scheduling a database container network. The method for optimizing and scheduling a database container network provided by the embodiments of the present invention can be applied to any electronic device with the function of optimizing and scheduling a database container network. The electronic device can be a server, a computer or a mobile terminal, etc. The solution of the present invention will be described below with reference to the drawings:
[0064] Figure 1 FIG. is a schematic flow chart of a method for optimizing and scheduling a database container network provided by an embodiment of the present invention, which may include the following steps:
[0065] S101. Obtain target index data of a database container, where the target index data includes: network status data and connection status data;
[0066] S102. Input the target index data into a pre-trained target demand evaluation network, so that the target demand evaluation network outputs a target demand level of the database container based on the target index data;
[0067] S103. Input the target metric data into a pre-trained target traffic prediction model, so that the target traffic prediction model outputs a target traffic prediction result of the database container based on the target metric data;
[0068] S104. Determine the target resource quantity that the database container needs to increase based on the current load data, target demand level, and target traffic prediction result of the database container. The target resource quantity includes network bandwidth and the number of connections;
[0069] S105. Perform resource scheduling for the database container according to the target resource quantity of the database container and the task type executed by the database container. The task type is used to define the allocation weights of the network bandwidth and the number of connections during the resource scheduling process.
[0070] In the embodiments of the present invention, target metric data including network status data and connection status data of the database container is collected, and the target metric data is input into a pre-trained target demand evaluation network and a target traffic prediction model to perform demand level evaluation and adaptive traffic prediction of the database container. Based on the demand level evaluation result, traffic prediction result, and current load data of the database container, the target resource quantity that the database container needs to increase is determined, and resource allocation is performed for the database container according to the target resource quantity and the task type currently executed by the database container. By applying the embodiments of the present invention, a pre-trained model is used to output a demand evaluation level and a traffic prediction result based on the multi-modal metric data of the database container, realizing in-depth mining of the processing status of the database container, and further realizing accurate prediction of the demand and traffic of the database container. Thus, resource allocation can be quickly responded to business needs, improving the resource allocation efficiency and further improving the business processing efficiency of the Kubernetes cluster. In addition, through traffic prediction, the traffic peak period can be predicted in advance, and resource allocation can be performed in advance to ensure the stable operation of high-concurrency services.
[0071] The following is an exemplary description of the above S101 - S105:
[0072] A database container refers to packaging database software and its dependencies into an independent and portable container so that it can run in any environment that supports container technology. The target metric data of the database container can be obtained according to preset metrics, which can be set according to the actual application scenario. Exemplarily, the preset metrics can include network status metrics of the database container, and the network status metrics can include metrics reflecting the network conditions of the database container such as network bandwidth, network latency, and packet loss rate. The above preset metrics can also include database connection status metrics, specifically, the status metrics of the database connection pool. The database connection pool is the "bridge" between the application and the database. By pre-creating and managing a certain number of database connections, it improves the efficiency of the application to obtain database connections. The connection status metrics can include the current number of connections, the number of waiting connections, etc. The above preset metrics can also include metrics related to the operation characteristics of the database. For example, when executing complex query statements, multi-modal data such as query execution time, the amount of data returned, and the network bandwidth and time interval occupied by transmitting this data are recorded.
[0073] In a possible embodiment, the target metric data of the database container can be obtained through a preset hook function. A hook function is a function that is automatically triggered and executed when a specific event occurs. As a possible implementation, a hook function can be set to be triggered when the number of data writes exceeds a preset write quantity threshold, so as to collect the network status data of the database container through this hook function, such as network bandwidth occupancy, network latency, and packet loss rate.
[0074] Exemplarily, on each node of the Kubernetes cluster, a data collection proxy module can be set between the kernel state and the user state with the help of eBPF (extended Berkeley Packet Filter) technology. Among them, the kernel state, also known as the kernel mode or privileged mode, is the state in which the operating system kernel runs and has full control over all key resources and functions such as system hardware, memory management, and process scheduling. The user state, also known as the user mode or user space, refers to the running state of the application when it is executing. An application in the user state can only access its own memory space and limited system resources and cannot directly access hardware resources (such as CPU, memory, disk, etc.) and execute privileged operations. eBPF is a technology that can dynamically load and run bytecode in the Linux kernel. It allows developers to insert custom code logic into the kernel without modifying the kernel source code.
[0075] For the characteristics of database containers with extremely high requirements for network real-time performance, eBPF collects data closely related to the network layer during database operations in real time by implanting hook functions at the network device driver layer. For example, when a large amount of data is written to the database, eBPF can accurately capture the network bandwidth occupancy, network latency, and changes in packet loss rate at this time. These data are crucial for analyzing the impact of database operations on the network.
[0076] The connection status data of database containers can be collected inside the database containers. Exemplarily, a lightweight probe based on Prometheus can be deployed inside the database containers. Prometheus is an open-source system monitoring and alerting tool that adopts a pull-based monitoring model and realizes comprehensive monitoring of the system by regularly obtaining metric data from the HTTP endpoints exposed by the target application. The above lightweight probe can collect multi-modal data such as the initiation frequency of database connection requests, data transfer size, and time interval inside the database containers.
[0077] Taking the MySQL database container as an example, the probe can query the connection pool status of the database at a set time interval of 5 seconds, obtain conventional information such as the current number of connections and the number of waiting connections, and can also deeply collect data related to the characteristics of database operations. For example, when executing a complex query statement, multi-modal data such as the query execution time, the amount of returned data, and the network bandwidth and time interval occupied by transmitting this data are recorded to comprehensively understand the network usage of the database in different operation scenarios.
[0078] Since the collected target metric data comes from a wide range of sources and has various formats, the target metric data can be preprocessed. This preprocessing process can include format unification, abnormal data filtering, and data normalization.
[0079] In a possible embodiment, the format of each target metric data can be initially unified in JSON format. Specifically, the network status data collected by eBPF and the database connection status data obtained by the Prometheus probe can both be converted into JSON objects containing fields such as timestamps, data types, and data values. During the format conversion process, the network data corresponding to database operations can be associated. For example, the start time of a database query operation is matched with the timestamp of the peak network bandwidth occupancy for subsequent analysis of the impact of different database operations on network performance.
[0080] After that, the data with unified format can be filtered. Specifically, data points with obvious errors or anomalies in the target metric data can be filtered. For example, for bandwidth data, if a value that instantaneously exceeds the hardware theoretical upper limit appears, it can be determined that this value is abnormal data. For instance, if the bandwidth monitoring data of a certain node shows 100 Gbps while the actual network card bandwidth of this node is only 10 Gbps, it is then determined as abnormal data and excluded.
[0081] As a possible implementation, the interquartile range (IQR) method can be used to detect outliers in various types of target metric data. The interquartile range is a robust statistic for measuring the central tendency of data, mainly used to describe the dispersion degree of data and identify outliers. The IQR is the difference between the upper quartile (Q3) and the lower quartile (Q1), that is, IQR = Q3 - Q1. After sorting a set of data from small to large, the value in the middle position is called the median Q2, and the median Q2 divides the data set into two parts of equal size. Q1 is the median of the part on the left side of the median, representing the middle value of the first half of the data set; Q3 is the median of the part on the right side of the median, representing the middle value of the second half of the data set.
[0082] Data points below Q1 - 1.5IQR or above Q3 + 1.5IQR can be further reviewed and processed. During the review, the rationality of the data is judged in combination with the database operation scenario. For example, during database backup, the change ranges of network bandwidth and data transfer volume are relatively large, and it is necessary to judge whether the data is abnormal according to the actual situation. For example, the business information executed by the database can be collected simultaneously when collecting the target metric data, and then it can be determined whether the target metric data is abnormal or incorrect data according to the network status data range corresponding to this business information.
[0083] During the process of normalizing the target metric data, the Min - Max normalization algorithm can be adopted to convert data of different magnitudes into the 0 - 1 interval. Its formula is: Xnorm = (X - Xmin) / (Xmax - Xmin), where X is the original data, and Xmin and Xmax are respectively the minimum and maximum values in the data set, and Xnorm is the normalized data. For example, for network latency data, assuming the minimum value in the data set is 1 ms and the maximum value is 100 ms, when the latency data collected at a certain moment is 20 ms, after normalization calculation, Xnorm = (20 - 1) / (100 - 1). During the normalization process, the tolerance of the database operation to network latency can be considered, and the network latency data can be combined with the response time requirements of the database operation to ensure that the data processing meets the database business requirements.
[0084] After preprocessing the target metric data, the target metric data can be input into a pre-trained model to obtain the demand level corresponding to the target metric data and the traffic prediction result. Specifically, the target demand level of the database container can be obtained through a pre-trained target demand evaluation network, and the target traffic prediction result of the database container can be obtained through a pre-trained target traffic prediction model.
[0085] In a possible embodiment, the target demand evaluation network can be pre-trained through the following steps:
[0086] S21. Obtain a demand training data set, where the demand training data set contains multiple pieces of demand training data, and each piece of demand training data includes historical database network status data, historical database connection status data, and historical database operation types.
[0087] Each piece of demand training data in this demand training data set can be obtained through an open-source data set or through Kubernetes cluster historical data. The historical data can include database network status data, database connection status data, database operation types, etc. Among them, the database network status data can include data such as bandwidth, network latency, packet loss rate, number of network connections, and network traffic direction. The database connection status data can include connection request frequency, data transfer size, and data transfer time interval, etc. The database operation types can include query, insert, update, delete, etc.
[0088] S22. Input the demand training data set into an initial demand evaluation network, where the initial demand evaluation network is a deep neural network built based on the TensorFlow framework.
[0089] S23. Obtain the predicted demand level output by the demand evaluation network based on each piece of demand training data;
[0090] TensorFlow is an open-source deep learning framework that can run on CPU / GPU / TPU, making full use of the computing resources of different hardware platforms to accelerate the model training and inference processes. The above demand evaluation network can be a network with a structure such as CNN or RNN. The demand evaluation network can include an input layer, a hidden layer, and an output layer. Among them, the input layer is used to receive data input. The input layer can include multiple nodes, and the number of nodes in the input layer can be determined according to the input data type. Exemplarily, if the input data includes 5 types of network status data and 3 types of connection status data, the number of nodes in the input layer is 8.
[0091] The hidden layer is used to extract data features. There can be multiple such hidden layers, and the specific number can be flexibly set according to the actual application scenario. For example, it can be set that the requirement evaluation network includes three hidden layers. Exemplarily, this hidden layer can adopt a convolutional neural network (CNN) structure. The CNN network structure uses a convolutional kernel to extract data features, and the size of this convolutional kernel can be set to 3x3 with a stride of 1. This setting can effectively reduce the computational amount while retaining the local features of the data.
[0092] The output of each hidden layer can increase the non-linear expression ability of the model through the ReLU activation function. The expression of the ReLU function is f(x) = max(0, x). It can solve the gradient vanishing problem and accelerate the convergence speed of the model. After convolution and ReLU activation in one hidden layer, some key features in the data are highlighted, providing more valuable information for the processing of subsequent layers. When processing data related to the database, the model can learn the complex relationships between different database operations and network performance metrics.
[0093] The output layer maps the features to the evaluation result of the database container network requirements through a fully connected layer. The fully connected layer integrates the features extracted by the previous hidden layers and obtains the final evaluation result through the calculation of the weight matrix and bias term. Both the weight matrix and the bias term are trainable parameters in the requirement evaluation network. For example, the requirement levels are divided into three levels: low, medium, and high. When the calculation result of the output layer is greater than 0.8, it is determined as a high requirement level; between 0.3 and 0.8 is a medium requirement level; less than 0.3 is a low requirement level. The evaluation result is directly related to the network resource allocation strategy of the database container, and corresponding network bandwidth, number of connections, etc. resources are provided for the database container according to different requirement levels.
[0094] S24. Train the initial requirement evaluation network based on the difference between the actual requirement level and the predicted requirement level of each of the said requirement training data until the difference converges, and obtain the target requirement evaluation network.
[0095] This difference can be calculated based on a preset loss function. The preset loss function can be a cross-entropy function, mean square error function, etc. After obtaining the difference between the predicted requirement level and the actual requirement level, methods such as gradient descent method and gradient ascent method can be used to adjust the parameters of the requirement evaluation model based on this difference to train the requirement evaluation model. After each round of training, the difference between the actual requirement level and the predicted requirement level is calculated, and when this difference converges, it can be determined that the model training is completed. The convergence of this difference can be that the difference is less than a preset difference threshold or the difference between the differences obtained from two rounds of training is less than a preset difference threshold.
[0096] Based on the target metric data, the trained target demand assessment network can output the target demand level of the database container, and different demand levels correspond to different network bandwidth resources, connection number resources, etc.
[0097] The above-mentioned target traffic prediction model can be constructed based on the principle of reinforcement learning and the Deep Q-Network (DQN) algorithm is used for model training. The DQN algorithm combines the advantages of deep learning and Q-learning and is used to solve the problems of high-dimensional state and action spaces in reinforcement learning. It uses a neural network to approximate the Q-value function, and through continuous iterative training, enables the agent to learn the optimal behavior strategy. Reinforcement learning is a machine learning paradigm where the agent interacts with the environment and learns the optimal behavior strategy based on the reward signal feedback from the environment. In a possible embodiment, the target traffic prediction model is pre-trained through the following steps:
[0098] S31. Construct a state space based on historical database network status data, historical traffic data, historical database load data, and historical business load data;
[0099] The state space may include multiple sets of historical database data. Each set of historical database data includes historical database network status data, historical traffic data, historical database load data, and historical business load data. Among them, the historical database network status data is the current network status data based on the collection time, which may include bandwidth utilization, network latency, etc. The historical traffic data is the traffic data within a preset time period before the collection time, such as the traffic data per minute in the past hour. The historical database load data is the number of active link libraries and the number of complex queries being executed corresponding to the collection time, etc. The historical business load data is the business type corresponding to the collection time, such as an e-commerce promotion activity identifier. By constructing the state space with the above-mentioned various data, the model can consider the operating state of the database more comprehensively.
[0100] S32. Construct an action space based on different network resource allocation strategies;
[0101] The action space is different network resource allocation strategies, including increasing or decreasing network bandwidth, adjusting network connection number limits, etc. For example, when the model determines that the network demand of the current database container increases, it can choose to increase the network bandwidth by 10% or increase the network connection number limit by 20 connections. In the decision-making process, fully consider the demand characteristics of database operations for network resources. For example, for a large number of data query operations, give priority to ensuring network bandwidth; for high-concurrency connection requests, focus on adjusting the network connection number limit.
[0102] S33. Input the state space and the action space into an initial traffic prediction model, so that the traffic prediction model selects a prediction and allocation strategy in the action space based on the state space;
[0103] S34. Calculate the reward function value based on the prediction and allocation strategy. In the case where the reward function value does not reach the maximum value of the reward function, modify the parameters of the traffic prediction model, and return to the step of the traffic prediction model selecting a prediction and allocation strategy in the action space for the state space until the reward function value reaches the maximum; wherein, the reward function is a business performance improvement index, and the business performance improvement index at least includes the amount of response time reduction and throughput;
[0104] S35. Determine the traffic prediction model corresponding to when the reward function value reaches the maximum as the target traffic prediction model.
[0105] The reward function can be set as a business performance improvement index. For example, if the response time is reduced by 10%, the reward function value is +10, and if the throughput is reduced by 5%, there is a penalty of -5, that is, the reward function value is -5. The specific expression formula of this reward function can be obtained according to the weights set for indicators such as response time and throughput in the actual application scenario. In this way, the model can gradually find the optimal resource allocation strategy in the continuous learning process to maximize the business performance. During the training process, the model will continuously try different actions, optimize its decisions based on the reward feedback, and gradually learn to make resource allocation decisions that are most beneficial to the database business performance under different network states and database loads. By continuously iterating and adjusting the parameters of the neural network and the reinforcement learning model, such as weights and biases, the accuracy and generalization ability of the model are improved.
[0106] Input the target metric data into the target traffic prediction model, and the target traffic prediction result of the database container can be obtained. The target traffic prediction result can include the predicted value of the network traffic received by the database container after a preset time interval.
[0107] When the deep neural network analyzes the change in the network requirements of the database container or the reinforcement learning model predicts a traffic peak, the scheduler plugin can perform dynamic resource allocation according to the analysis and prediction results, and use the dynamic programming algorithm to determine the optimal resource allocation plan. Specifically, the number of network resources that the database container needs to increase can be determined according to the current load of the database container, the target demand evaluation level, and the target traffic prediction result. The network resources can include bandwidth, storage, CPU, and the number of connections, etc. Then, the priority weight of each network resource can be determined according to the service identifier processed by the database container, and the priority weight is used to identify the expansion order of the network resources. For example, during an e-commerce promotion event, in the face of a significant increase in the database read traffic, the dynamic programming algorithm will comprehensively consider the current load of each database container, the network bandwidth occupancy, and the future traffic prediction, and calculate the additional network bandwidth and the number of connections required for each container. For databases with frequent read operations, the network bandwidth is preferentially guaranteed, and additional bandwidth is allocated according to a certain weight according to the growth ratio of the predicted traffic; for databases with more write operations, in addition to ensuring the stability and reliability of the network, the number of network connections will also be reasonably adjusted according to the concurrent requirements of transaction processing.
[0108] Currently, when a network failure occurs in a Kubernetes cluster, it is difficult for existing technologies to quickly and accurately locate the fault point, and there is also a lack of an effective fault recovery mechanism. When a network link is interrupted or a node fails, the fault information cannot be reliably recorded and shared in a timely manner, resulting in a delay in fault handling. Moreover, in the process of fault recovery, the traditional mechanism cannot make full use of the auto-scaling and container migration functions of Kubernetes to efficiently redeploy and allocate resources for the affected database containers, resulting in a relatively long service interruption time and bringing greater losses to enterprises.
[0109] Therefore, in a possible embodiment, the Ethereum client Geth can be deployed in each node of the Kubernetes cluster. Ethereum is an open-source public blockchain platform with smart contract functions, which provides a decentralized virtual machine (Ethereum Virtual Machine, EVM) to execute smart contracts. Geth (Go-Ethereum) is the official client implementation of Ethereum and is written in the Go language. It provides various functions for interacting with the Ethereum network, including node synchronization, account management, smart contract deployment and execution, etc. As Figure 2 shown, the above method may further include the following steps:
[0110] S106. In the case of a faulty node, the faulty node broadcasts the fault information to other nodes in the Kubernetes cluster;
[0111] S107. The master node in the Kubernetes cluster collects the fault information of each of the faulty nodes and broadcasts pre - preparation information to other nodes. The pre - preparation information includes a fault identifier and a fault type, so that other nodes return preparation information to the master node in response to the pre - preparation information.
[0112] S108. When the number of received preparation messages at the master node reaches a preset number, fault handling is performed on the faulty nodes. The fault handling includes task migration of the faulty nodes and network resource allocation.
[0113] By configuring the network parameters of the nodes, including IP addresses, port numbers, network IDs, etc., to make them join the same private blockchain network. In this way, each node can participate in the consensus process of the blockchain to ensure the distributed recording and sharing of fault information. When recording fault information, special attention is paid to database - related faults, such as database connection interruptions and communication faults between database cluster nodes, and these fault information are associated and recorded with network fault information. At the same time, public - private key pairs can also be set in each node to verify the identity of the nodes during mutual communication.
[0114] The faulty node can broadcast the fault information to other nodes in the Kubernetes cluster. The fault information can include the fault occurrence time, fault type, etc. After receiving the information, other nodes first perform signature verification to ensure the reliability of the information source. Exemplarily, when the faulty node broadcasts the fault information, it can sign the information with its private key. After other nodes receive it, they use the public key of the faulty node for verification to determine whether the source of the fault information is reliable.
[0115] To ensure the consistency of a distributed system, the PBFT algorithm (Practical Byzantine Fault Tolerance) can be used to determine whether fault information needs to be processed. The PBFT algorithm is a distributed consensus algorithm used to guarantee the consistency of a distributed system in the presence of Byzantine faults (nodes may experience failures, errors, or malicious behavior). As a possible implementation, the master node can be determined through an election in a Kubernetes cluster. The master node collects fault information and broadcasts a pre-prepare message to other nodes, which contains a digest and sequence number of the fault information. After receiving the message, other nodes send prepare messages. When the master node receives 2f + 1 (f is the number of allowed Byzantine nodes) prepare messages, it sends a commit message to record the fault information in the blockchain ledger. For example, if there are 7 nodes in the cluster and 1 Byzantine node is allowed (f = 1), the master node needs to receive 2 * 1 + 1 = 3 prepare messages before it can perform the commit operation. During the consensus process, higher priority is given to fault information related to the database to ensure that database failures can be recorded and responded to in a timely manner.
[0116] As a possible implementation, the data structure of the network fault ledger can be defined as a JSON object containing information such as the fault occurrence time (accurate to milliseconds), fault type (such as link interruption, node failure), affected node IPs, and container IDs. For example, when a network link fails, the information recorded in the ledger may be as follows:
[0117] {
[0118] "fault_time":"2024-10-10T12:30:00.123Z",
[0119] "fault_type":"link_interruption",
[0120] "affected_nodes":["192.168.1.10"],
[0121] "affected_containers":["container-123"],
[0122] "database_impact":"Database connection interrupted,transactions inprogress may be rolled back"
[0123] }
[0124] For the fault information recorded in the blockchain ledger, the fault information can be processed through preset fault handling operations. As a possible implementation, Solidity language can be used to write smart contracts to implement the code logic for automatically triggering self-healing operations. Solidity is the main programming language for Ethereum smart contracts. By writing smart contracts, it is possible to automatically trigger corresponding self-healing operations when the fault information is recorded in the ledger. For example, in the smart contract, it can be defined that when a link interruption fault causes the database connection to be interrupted, the API of Kubernetes is automatically called to switch to the standby link, and the affected database containers are redeployed and resource allocated, and the database service is preferentially restored to ensure that the business resumes normal operation in the shortest time.
[0125] The above fault handling may include the following processes: First, the fault information is parsed in detail to determine the fault type (such as link interruption, node fault), the affected node IP and container ID, etc. For example, when a network link interruption of a certain node is detected, the fault detection module quickly locates the affected database container and marks its current state.
[0126] Then, container migration and automatic scaling operations are performed. The faulty node is marked as unschedulable by executing the kubectl cordon command, and at the same time, the affected database containers are migrated to nodes with good network conditions by using the kubectl drain command. During the migration process, a data consistency verification mechanism is adopted to first save key information such as the transaction state and cached data of the database container to temporary storage. For example, for a database that is in the middle of a transaction process, the uncommitted transaction records are saved to a distributed file system. After restarting the container on the new node, resources are reasonably allocated according to the network bandwidth, CPU, memory and other resources of the new node. For example, according to the available network bandwidth of the new node, a certain proportion of network bandwidth resources is allocated to the database container to ensure that the database container can run stably on the new node.
[0127] Finally, verification and optimization after fault recovery are carried out. By sending test query statements to the recovered database container, it is verified whether the function of the database is normal. For example, a simple SELECT statement is executed to check whether the data reading is accurate. At the same time, according to the data and experience during the fault recovery process, the resource allocation strategy and fault detection mechanism of the system are optimized. For example, if it is found that a certain node frequently has faults in a short period of time, the system will automatically adjust the resource allocation strategy of the node, or strengthen the monitoring frequency and depth of the node to avoid similar faults from occurring again.
[0128] The prior art has insufficient capabilities in network fault location and recovery. Through the above technical means, the present invention utilizes a distributed ledger for network faults and self-healing recovery based on blockchain. The blockchain is used to record fault information, ensuring the authenticity and immutability of the information, quickly locating the fault point, and using the Kubernetes function to redeploy and allocate resources for the affected database containers to achieve fast self-healing. For example, when a network link fails, it can quickly switch to a standby link to ensure that the service resumes normal operation in the shortest time.
[0129] In a possible embodiment, a NetworkOptimizerScheduler plugin can be developed based on the scheduler extension mechanism of Kubernetes. This plugin adopts a hierarchical architecture design and mainly includes a data collection layer, a decision-making layer, and an execution layer. The data collection layer is responsible for using the RESTful API to obtain the status information of the cluster in real time through the API Server of Kubernetes. For example, an HTTP GET request is sent to the / api / v1 / nodes endpoint every 5 seconds to obtain the resource information of all nodes in the cluster, including the number of CPU cores, total memory, available network bandwidth, etc.; at the same time, the / api / v1 / pods endpoint is accessed at the same frequency to obtain the running status of all containers, and the environment variables and configuration files of the database containers are parsed to obtain database-related parameters, such as the maximum number of connections, cache size, etc.
[0130] The decision-making layer consists of a policy engine and an intelligent algorithm module. The policy engine preliminarily screens and analyzes the collected data according to preset rules and business requirements. For example, when it is detected that the CPU usage rate of the database container exceeds 80% for three consecutive cycles and the network bandwidth utilization rate reaches 70%, the policy engine marks the container as a resource-intensive state. The intelligent algorithm module generates specific resource scheduling decisions based on the analysis results of the deep neural network and reinforcement learning model, combined with the marking information of the policy engine. For example, when the deep neural network predicts that the database read traffic will increase by 50% in the next 10 minutes, the intelligent algorithm module calculates that the network bandwidth needs to be increased by 30% and 150 network connections need to be added.
[0131] The execution layer is responsible for converting the resource scheduling decisions generated by the decision-making layer into actual operation instructions. By calling the resource allocation interface of Kubernetes, it performs operations to modify the resource configuration file of the Pod. For example, using the client library of Kubernetes, the network bandwidth parameter of the database Pod is modified from 10 Mbps to 13 Mbps, and the network connection number limit is increased from 100 to 250. At the same time, the execution layer is also responsible for monitoring the execution status of the operation to ensure that the resource adjustment is successfully completed. If an error occurs during the operation, such as the failure to modify the configuration file, the execution layer will promptly feedback to the decision-making layer for error handling and retry.
[0132] By developing the NetworkOptimizerScheduler plugin and adopting a hierarchical architecture, efficient collaboration among data collection, decision-making, and execution is achieved. In dynamic resource allocation, the dynamic programming algorithm is used to determine the optimal solution, and resources are reasonably allocated according to the read-write characteristics of the database. During fault recovery, through mechanisms such as data consistency verification and test query verification, the continuity of database services is ensured, and the system strategy is optimized based on the fault data.
[0133] Applying the embodiments of the present invention, using eBPF and Prometheus, data is collected from multiple dimensions at the network device driver layer and inside the database container, covering information such as network bandwidth, latency, and database connection request frequency. Through in-depth neural network fusion analysis of these multimodal data, the network requirements of the database container are accurately predicted, and accurate resource allocation is achieved.
[0134] By constructing an adaptive traffic prediction model based on reinforcement learning, combined with historical traffic data, current network status, and business load, the network resource allocation strategy is dynamically adjusted. It can predict traffic peaks in advance. For example, before an e-commerce promotion event, the network bandwidth of the database container is automatically increased, the transmission protocol is optimized, and the stability of the service under high concurrency is ensured.
[0135] Furthermore, the Ethereum blockchain is introduced to build a network fault ledger, and the PBFT improved consensus mechanism is used to ensure the authenticity, reliability, and immutability of fault information. Through smart contracts, self-healing operations are automatically triggered to quickly locate the fault point and recover. For example, when a network link fails, the standby link is quickly switched, and the affected database containers are redeployed and allocated using the Kubernetes function.
[0136] Based on the same inventive concept, the present invention also provides a database container network optimization and scheduling device, which is applied to a Kubernetes cluster, as Figure 3 shown. The device 300 may include:
[0137] An acquisition module 301, configured to acquire target metric data of the database container, where the target metric data includes: network status data and connection status data;
[0138] An input module 302, configured to input the target metric data into a pre-trained target demand evaluation network, so that the target demand evaluation network outputs the target demand level of the database container based on the target metric data;
[0139] Input the target metric data into a pre-trained target traffic prediction model, so that the target traffic prediction model outputs the target traffic prediction result of the database container based on the target metric data;
[0140] A determination module 303, configured to determine a target resource quantity that the database container needs to increase based on the current load data of the database container, a target demand level, and a target traffic prediction result, where the target resource quantity includes network bandwidth and the number of connections;
[0141] A scheduling module 304, configured to perform resource scheduling for the database container according to the target resource quantity of the database container and the type of task executed by the database container, where the type of task is used to define the allocation weights of the network bandwidth and the number of connections during the resource scheduling process.
[0142] In a possible embodiment, the target demand evaluation network is pre-trained through the following steps:
[0143] Obtain a demand training data set, where the demand training data set contains multiple pieces of demand training data, and each piece of demand training data includes historical database network status data, historical database connection status data, and historical database operation types;
[0144] Input the demand training data set into an initial demand evaluation network, where the initial demand evaluation network is a deep neural network built based on the TensorFlow framework;
[0145] Obtain the predicted demand levels output by the demand evaluation network based on each piece of demand training data;
[0146] Train the initial demand evaluation network based on the difference between the actual demand level of each demand training data and the predicted demand level until the difference converges, to obtain a target demand evaluation network.
[0147] In a possible embodiment, the target traffic prediction model is pre-trained through the following steps:
[0148] Based on historical database network status data, historical traffic data, historical database load data, and historical service load data, construct a state space;
[0149] Based on different network resource allocation strategies, construct an action space;
[0150] Input the state space and the action space into an initial traffic prediction model, so that the traffic prediction model selects a predicted allocation strategy in the action space based on the state space;
[0151] Calculate the reward function value based on the predicted allocation strategy. When the reward function value does not reach the maximum value of the reward function, modify the parameters of the traffic prediction model, and return to the step of the traffic prediction model selecting the predicted allocation strategy in the action space for the state space until the reward function value reaches the maximum; wherein, the reward function is a business performance improvement index, and the business performance improvement index at least includes the reduction in response time and throughput.
[0152] Determine the traffic prediction model corresponding to when the reward function value reaches the maximum as the target traffic prediction model.
[0153] In a possible embodiment, the Kubernetes cluster includes multiple nodes, and an Ethereum client is deployed in each of the nodes. The apparatus further includes:
[0154] A fault response module, configured to, when a faulty node appears, broadcast fault information from the faulty node to other nodes in the Kubernetes cluster;
[0155] The master node in the Kubernetes cluster collects the fault information of each faulty node and broadcasts pre-preparation information to other nodes. The pre-preparation information includes a fault identifier and a fault type; so that other nodes return preparation information to the master node in response to the pre-preparation information;
[0156] When the number of received preparation messages by the master node reaches a preset number, perform fault handling on the faulty node, and the fault handling includes task migration of the faulty node and network resource allocation.
[0157] Wherein, in the present invention, the collection, storage, use, processing, transmission, provision, and disclosure of user personal information comply with the provisions of relevant laws and regulations and do not violate public order and good customs.
[0158] An exemplary embodiment of the present invention further provides an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor. The memory stores a computer program that can be executed by the at least one processor, and when the computer program is executed by the at least one processor, it is used to cause the electronic device to execute the method according to the embodiment of the present invention.
[0159] An exemplary embodiment of the present invention further provides a non-transitory computer-readable storage medium storing a computer program, wherein when the computer program is executed by a processor of a computer, it is used to cause the computer to execute the method according to the embodiment of the present invention.
[0160] An exemplary embodiment of the present invention also provides a computer program product, including a computer program, wherein the computer program, when executed by a processor of a computer, is configured to cause the computer to execute the method according to the embodiments of the present invention.
[0161] Reference Figure 4 , the structural block diagram of an electronic device 400 that can be used as a server or a client of the present invention will now be described. It is an example of a hardware device that can be applied to various aspects of the present invention. The electronic device is intended to represent various forms of digital electronic computer devices, such as, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described herein and / or claimed.
[0162] As Figure 4 shown, the electronic device 400 includes a computing unit 401, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 402 or a computer program loaded from a storage unit 408 into a random access memory (RAM) 403. In the RAM 403, various programs and data required for the operation of the electronic device 400 can also be stored. The computing unit 401, the ROM 402, and the RAM 403 are connected to each other through a bus 404. An input / output (I / O) interface 405 is also connected to the bus 404.
[0163] A plurality of components in the electronic device 400 are connected to the I / O interface 405, including: an input unit 406, an output unit 407, a storage unit 408, and a communication unit 409. The input unit 406 can be any type of device that can input information into the electronic device 400. The input unit 406 can receive input digital or character information, and generate key signal inputs related to user settings and / or function controls of the electronic device. The output unit 407 can be any type of device that can present information, and can include, but is not limited to, a display, a speaker, a video / audio output terminal, a vibrator, and / or a printer. The storage unit 408 can include, but is not limited to, a magnetic disk, an optical disk. The communication unit 409 allows the electronic device 400 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks, and can include, but is not limited to, a modem, a network card, an infrared communication device, a wireless communication transceiver, and / or a chipset, such as a Bluetooth™ device, a WiFi device, a WiMax device, a cellular communication device, and / or the like.
[0164] The computing unit 401 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 401 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 401 executes the various methods and processes described above. For example, in some embodiments, any of the above-described database container network optimization scheduling methods can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as the storage unit 408. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 400 via the ROM 402 and / or the communication unit 409. In some embodiments, the computing unit 401 can be configured to execute any of the above-described database container network optimization scheduling methods in any other suitable manner (e.g., by means of firmware).
[0165] The program code for implementing the method of the present invention can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowchart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0166] In the context of the present invention, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0167] As used in this invention, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, apparatus, and / or device (e.g., a magnetic disk, an optical disk, a memory, a programmable logic device (PLD)) used to provide machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. The term "machine-readable signal" refers to any signal used to provide machine instructions and / or data to a programmable processor.
[0168] For purposes of providing an interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide for interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic, speech, or tactile input).
[0169] The systems and techniques described herein can be implemented in a computing system that includes a back-end component (e.g., as a data server), or a computing system that includes a middleware component (e.g., an application server), or a computing system that includes a front-end component (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or in a computing system that includes any combination of such back-end, middleware, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.
[0170] A computer system can include a client and a server. The client and server are generally remote from each other and typically interact through a communication network. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship to each other.
Claims
1. A method for optimizing the scheduling of a database container network, characterized in that, The method is applied to a Kubernetes cluster, and the method includes: Obtain target metric data of a database container, where the target metric data includes: network status data and connection status data; Input the target metric data into a pre-trained target demand assessment network, so that the target demand assessment network outputs a target demand level of the database container based on the target metric data; Input the target metric data into a pre-trained target traffic prediction model, so that the target traffic prediction model outputs a target traffic prediction result of the database container based on the target metric data; Based on the current load data, target demand level, and target traffic prediction result of the database container, determine the target resource quantity that the database container needs to increase, where the target resource quantity includes network bandwidth and the number of connections; Perform resource scheduling for the database container according to the target resource quantity of the database container and the task type executed by the database container, where the task type is used to define the allocation weights of the network bandwidth and the number of connections during the resource scheduling process.
2. The method according to claim 1, wherein The target demand assessment network is pre-trained through the following steps: Obtain a demand training data set, where the demand training data set contains multiple pieces of demand training data, and each piece of demand training data includes historical database network status data, historical database connection status data, and historical database operation types; Input the demand training data set into an initial demand assessment network, where the initial demand assessment network is a deep neural network built based on the TensorFlow framework; Obtain the predicted demand levels output by the demand assessment network based on each piece of demand training data; Train the initial demand assessment network based on the difference between the actual demand level of each piece of demand training data and the predicted demand level until the difference converges, and obtain a target demand assessment network.
3. The method according to claim 1, characterized in that, The target traffic prediction model is pre-trained through the following steps: Construct a state space based on historical database network status data, historical traffic data, historical database load data, and historical business load data; Construct an action space based on different network resource allocation strategies; Input the state space and the action space into an initial traffic prediction model, so that the traffic prediction model selects a predicted allocation strategy in the action space based on the state space; Calculate a reward function value based on the predicted allocation strategy. In the case where the reward function value does not reach the maximum value of the reward function, modify the parameters of the traffic prediction model, and return to the step where the traffic prediction model selects a predicted allocation strategy in the action space based on the state space until the reward function value reaches the maximum; where the reward function is a business performance improvement metric, and the business performance improvement metric at least includes the amount of response time reduction and throughput; Determine the traffic prediction model corresponding to when the reward function value reaches the maximum as the target traffic prediction model.
4. The method according to claim 1, characterized in that The Kubernetes cluster includes multiple nodes, and an Ethereum client is deployed in each of the nodes. The method further includes: In the case of a failed node, the failed node broadcasts failure information to other nodes in the Kubernetes cluster; The master node in the Kubernetes cluster collects the failure information of each failed node and broadcasts pre-preparation information to other nodes. The pre-preparation information includes a failure identifier and a failure type, so that other nodes return preparation information to the master node in response to the pre-preparation information; When the number of received preparation messages reaches a preset number, perform failure handling on the failed node. The failure handling includes task migration of the failed node and network resource allocation.
5. A database container network optimization scheduling device, characterized in that The device is applied to a Kubernetes cluster. The device includes: An acquisition module, configured to acquire target metric data of a database container. The target metric data includes network status data and connection status data; An input module, configured to input the target metric data into a pre-trained target demand evaluation network, so that the target demand evaluation network outputs a target demand level of the database container based on the target metric data; Input the target metric data into a pre-trained target traffic prediction model, so that the target traffic prediction model outputs a target traffic prediction result of the database container based on the target metric data; A determination module, configured to determine the target resource quantity that the database container needs to increase based on the current load data, target demand level, and target traffic prediction result of the database container. The target resource quantity includes network bandwidth and the number of connections; A scheduling module, configured to perform resource scheduling for the database container according to the target resource quantity of the database container and the task type executed by the database container. The task type is used to define the allocation weights of the network bandwidth and the number of connections during the resource scheduling process.
6. The device according to claim 5, characterized in that, The target demand evaluation network is pre-trained through the following steps: Obtain a demand training data set. The demand training data set contains multiple pieces of demand training data. Each piece of demand training data includes historical database network status data, historical database connection status data, and historical database operation types; Input the demand training data set into an initial demand evaluation network. The initial demand evaluation network is a deep neural network built based on the TensorFlow framework; Obtain the predicted demand level output by the demand evaluation network based on each piece of demand training data; Train the initial demand evaluation network based on the difference between the actual demand level of each piece of demand training data and the predicted demand level until the difference converges to obtain a target demand evaluation network.
7. The device according to claim 5, characterized in that, The target traffic prediction model is pre-trained through the following steps: Construct a state space based on historical database network status data, historical traffic data, historical database load data, and historical business load data; Construct an action space based on different network resource allocation strategies; Input the state space and the action space into an initial traffic prediction model, so that the traffic prediction model selects a prediction allocation strategy in the action space based on the state space; Calculate the reward function value based on the prediction allocation strategy. If the reward function value does not reach the maximum value of the reward function, modify the parameters of the traffic prediction model, and return to the step of the traffic prediction model selecting a prediction allocation strategy in the action space for the state space until the reward function value reaches the maximum; wherein, the reward function is a service performance improvement index, and the service performance improvement index at least includes the amount of response time reduction and throughput; Determine the traffic prediction model corresponding to when the reward function value reaches the maximum as the target traffic prediction model.
8. The device according to claim 5, characterized in that, The Kubernetes cluster includes multiple nodes, and an Ethereum client is deployed in each of the nodes. The device further includes: A fault response module, configured to, when a faulty node appears, broadcast fault information from the faulty node to other nodes in the Kubernetes cluster; The master node in the Kubernetes cluster collects the fault information of each faulty node and broadcasts pre-preparation information to other nodes. The pre-preparation information includes a fault identifier and a fault type; so that other nodes return preparation information to the master node in response to the pre-preparation information; When the number of received preparation messages at the master node reaches a preset number, perform fault handling on the faulty node, and the fault handling includes task migration of the faulty node and network resource allocation.
9. An electronic device, comprising: A processor; And A memory storing a program, wherein the program includes instructions that, when executed by the processor, cause the processor to execute the method according to any one of claims 1-4.
10. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause a computer to execute the method according to any one of claims 1-4.