A model training method for monitoring network cards, its applications, system, and electronic device

Through the model training method of monitoring network cards, the parameter adjustment model is trained using convolutional neural network and backpropagation algorithm to dynamically calculate the network card eviction threshold, solving the problem of insufficient network card traffic monitoring and eviction in the existing technology, realizing intelligent network card traffic management, and improving cluster performance.

CN115714692BActive Publication Date: 2025-06-27CHINA UNICOM (GUANGDONG) IND INTERNET CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211453132.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-18
Publication Date
2025-06-27
Estimated Expiration
2042-11-18

AI Technical Summary

Technical Problem

The existing technology lacks monitoring and evicting of machine network card traffic, which leads to inability to sense and trigger pod evicting in time when network card pressure is high, affecting network traffic load balancing.

Method used

A model training method for monitoring network cards is adopted. By obtaining the expulsion history, the verification set matrix is ​​generated, the training set and verification set are constructed, and the convolutional neural network is used to train with backpropagation algorithm and stochastic gradient descent to obtain the parameter adjustment model. This model dynamically calculates the soft eviction and hard eviction thresholds, and performs intelligent eviction based on the current network card traffic situation.

Benefits of technology

It realizes intelligent monitoring and evicting of network card traffic, can respond to network card pressure in a timely manner, ensures dynamic balance of network resources, and improves the performance and reliability of kubernetes clusters.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115714692B_ABST
    Figure CN115714692B_ABST
Patent Text Reader

Abstract

The present invention provides a model training method for monitoring network cards, its applications, systems, and electronic devices. The method includes: obtaining an eviction history record set, calculating and generating a validation set matrix, and constructing a training set with the validation set matrix; inputting the optimized training set into a convolutional neural network, and training it through a stochastic gradient descent method combined with the backpropagation algorithm to obtain a trained parameter-tuning model. The parameter-tuning model is used to dynamically calculate soft eviction thresholds and hard eviction thresholds according to the data conditions on the current node cluster machines. When the network card traffic meets the soft eviction threshold or the hard eviction threshold, soft eviction and hard eviction are respectively performed on the pods. The eviction history records are the performance parameter indicators of the node cluster machines during eviction in history, as well as the corresponding soft eviction thresholds and hard eviction thresholds. Compared with the prior art, the optimal eviction thresholds are dynamically calculated through a neural network model, realizing intelligent dynamic eviction of network card resources.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of network communication technologies, and specifically, to a model training method for monitoring network cards, and its applications, systems, and electronic devices. Background Art

[0002] With the popularity of containerization and kubernetes orchestration technologies, almost all current applications run in kubernetes clusters. Countless services run in the form of pods on the cluster. When performance bottlenecks such as CPU, IO, and disk occur on a cluster node, the kubernetes cluster reschedules and migrates the pods on that node to a new node with sufficient resources according to a certain strategy to ensure a two-way balance between node resources and services. Kubelet monitors various metrics of the node and compares them with thresholds to trigger active eviction, which is the core and key means of kubernetes rescheduling.

[0003] Although kubernetes has been widely used for monitoring CPU, IO, disk, etc. on machines, it lacks monitoring and eviction of machine network card traffic. When the network card is under heavy pressure, kubernetes cannot perceive it and trigger pod eviction, resulting in the failure to actively perform pod migration and network traffic load balancing in a timely manner. Moreover, network resources are dynamically changing, and dynamic balancing needs to be performed according to the conditions of each network node at different times. Summary of the Invention

[0004] The present invention aims to overcome at least one defect of the above-mentioned prior art, and proposes a model training method for monitoring network cards, and its applications, systems, and electronic devices, which can intelligently monitor and evict network cards.

[0005] The technical solution adopted by the present invention is as follows:

[0006] Provide a model training method for monitoring network cards, the method includes:

[0007] Obtain a set of eviction history records, calculate and generate a validation set matrix, and construct a training set D and a validation set V with the validation set matrix;

[0008] Input the training set into a convolutional neural network, and perform training by means of stochastic gradient descent combined with the backpropagation algorithm to obtain a trained parameter-tuning model. The parameter-tuning model is used to dynamically calculate soft eviction thresholds and hard eviction thresholds according to the data conditions on the current node cluster machine. When the network card traffic meets the soft eviction threshold or the hard eviction threshold, perform soft eviction and hard eviction on the pods respectively;

[0009] The training by means of stochastic gradient descent combined with the backpropagation algorithm includes:

[0010] Input the training set D into the neural network model to obtain the network output as Assume the loss function is Perform parameter learning by calculating the derivative of the loss function with respect to each parameter. The specific steps are as follows:

[0011] A1: Randomly initialize the parameter weight matrix w and bias b;

[0012] A2: Randomly reorder the samples in the training set;

[0013] A3: Select a sample x (n) , y (n) , and initialize n = 0;

[0014] A4: Feedforward to calculate the net input z (l) and activation value a (l) of each layer until the last layer;

[0015] A5: Backpropagate to calculate the error δ (l) of each layer; Derive that the gradient of with respect to the bias W (l) of the l-th layer is:

[0016] A6: Calculate the gradient of with respect to the bias b (l) of the l-th layer as:

[0017] A7: Update the W and b parameters through the formula: b (l) ←b (l) -αδ (l) ; A8: Increment the value of n by 1 and repeat steps A3 - A7 until training n = N;

[0018] A9: Repeat steps A2 - A8 until the error rate of the convolutional neural network model on the validation set V no longer decreases.

[0019] The eviction history record is the performance parameter indicators of the node cluster machines during historical evictions, as well as the corresponding soft eviction threshold and hard eviction threshold.

[0020] Obtain various performance parameters on the node cluster machines when evictions occurred historically, including soft eviction thresholds, hard eviction thresholds, CPU usage, memory usage, network card usage, and eviction signals. These information are stored in the time series database of the node cluster machines. Generate a validation set matrix, and together with the soft eviction threshold and the hard eviction threshold, generate a training set. Train through the stochastic gradient descent method combined with the backpropagation algorithm to obtain a trained parameter tuning model. The parameter tuning model can calculate the optimal soft eviction threshold and hard eviction threshold under each performance parameter for this performance metric. When a machine meets the soft eviction threshold or the hard eviction threshold, evict the pods in the cluster. Moreover, the soft eviction threshold and the hard eviction threshold are not fixed. By combining with the current dynamically changing performance parameters of the machine and using the parameter tuning model for calculation, the optimal soft eviction threshold and hard eviction threshold under the current machine state can be obtained for setting. Enable the machine to dynamically set the optimal eviction threshold according to its own state, and achieve intelligent monitoring and eviction of network cards. And because the network card information data volume is very large, the training method combining stochastic gradient descent and backpropagation algorithm is used to improve the training efficiency.

[0021] Further, the calculation of generating the validation set matrix and constructing the training set D and the validation set V with the validation set matrix X are specifically as follows:

[0022] Extract the soft eviction threshold, hard eviction threshold, CPU, memory network card metrics, eviction signal volume, and eviction records from the obtained records, jointly extract the data to form a validation set matrix, and then generate the training set data D and the validation set V with the validation set matrix;

[0023] Training set

[0024] Where x is the validation set matrix, X[0] represents the CPU usage, X[1] represents the memory usage, X[2] represents the network card usage, and y is the eviction proportion of the pods under the corresponding CPU usage, memory usage, and network card usage;

[0025] Validation set The data format is the same as that of the training set.

[0026] Use disk read and write, traffic stress testing, and intensive CPU-running programs to adjust the usage rates of the CPU, memory, and network card. Record the total number of running pods y1 at this time, the total number of pods y0 running on the machine before stress testing, as well as the CPU usage rate, memory usage rate, and network card usage rate. Among them, the total number of evicted pods y2 = y0 - y1. x[3] = y2 / y0, and combine the three indicators of the CPU usage rate, memory usage rate, and network card usage rate at this time to form a 4-tuple of the eviction record matrix. Since the CPU, memory, and network card are the main parameters affecting computer performance, combined analysis can better calculate their correlation with the eviction threshold, and then obtain the optimal eviction threshold.

[0027] Further, the soft eviction threshold eviction-soft includes: cpu.available (CPU usage threshold), memory.available (memory usage threshold), and network.available (network card usage threshold);

[0028] The hard eviction threshold eviction-hard includes: cpu.available (CPU usage threshold), memory.available (memory usage threshold), and network.available (network card usage threshold).

[0029] The CPU usage threshold, memory usage threshold, and network card usage threshold are set for the soft eviction and hard eviction of k8s respectively, which can monitor the CPU, memory, and network card resources respectively. When the current CPU usage rate, memory usage rate, or network card usage rate exceeds the corresponding soft eviction or hard eviction threshold, execute the eviction to evict the pod. Through the judgment of the three elements, the eviction of the pod is made more intelligent.

[0030] The present invention also provides an application of a model for monitoring network cards. The method includes:

[0031] Configure the acquisition module to monitor and collect data from the cluster node machines;

[0032] Preprocess and store the collected data;

[0033] Improve the network card eviction algorithm model based on K8s, including: setting the initial values of the soft eviction threshold and hard eviction threshold for the network card and setting them into the listening service. When the network card node traffic occupancy is less than the soft eviction threshold or hard eviction threshold, send an eviction signal to evict the pod corresponding to the network card node;

[0034] Build an intelligent parameter tuning component. The intelligent parameter tuning component uses the parameter tuning model trained by the above-mentioned model training method for monitoring network cards to dynamically update the soft eviction threshold and the hard eviction threshold.

[0035] The acquisition module uses cAdvisor to collect real-time information such as CPU, memory, and network card traffic of node cluster machines and stores it in the time series database. Configure the K8s model for node cluster machines, improve the K8s model based on the monitoring of network card traffic, rewrite the monitoring and eviction of system CPU, memory, and disk capacity in K8s as the monitoring and eviction of network card traffic resources, and set the corresponding initial soft eviction threshold and hard eviction threshold. When the network card traffic data meets the soft eviction threshold and the hard eviction threshold, a certain number of pods in the node cluster are evicted to release network card resources for other pods that need network card resources more. And use the intelligent parameter tuning component to adjust the soft eviction threshold and the hard eviction threshold in real time, use the above-trained model to calculate the real-time CPU, memory, and network card traffic information, calculate the optimal soft eviction threshold and hard eviction threshold in this state, update them in real time, and store the soft eviction threshold and hard eviction threshold before update together with the status parameter information in the time series database as new historical data to provide data for model training. Through the above application method, the node cluster machines can monitor the network card and intelligently evict pods.

[0036] Further, the preprocessing and storage include:

[0037] Save the original data of cumulative received traffic (rx_bytes), cumulative received error traffic (rx_errors), cumulative transmitted traffic (tx_bytes), and cumulative transmitted error traffic (tx_errors) into memory, and calculate the received traffic per second (rx_bytes_perSecon) and the transmitted traffic per second (tx_bytes_perSecon);

[0038] Record the calculated received traffic per second and transmitted traffic per second into the time series database.

[0039] Set the acquisition period node-status-update-frequency parameter to 5s, and divide the cumulative values by the acquisition period to calculate the traffic transmitted and received per second. The traffic transmitted per second and the traffic received per second can better reflect the current network performance of the network card.

[0040] Further, the specific steps for improving the eviction algorithm model of the network card based on K8s are:

[0041] S1: Set initial values for the soft eviction threshold eviction-soft and the hard eviction threshold eviction-hard;

[0042] S2: Load the soft eviction threshold and the hard eviction threshold into the EvictionManager;

[0043] S3: Start a coroutine to listen to the service thresholdNotifier. The listening service obtains the preprocessed data and forms a data set thresholds in combination with the soft eviction threshold and the hard eviction threshold in S2;

[0044] S4: Configure the network card traffic judgment unit and perform a matching operation according to the data set in step S3 to determine whether it meets the soft eviction threshold or the hard eviction threshold. If it meets, send an eviction signal signal and record the eviction signal in the time series database, otherwise ignore it;

[0045] S5: Evict the active pods according to the eviction signal sent in step S4;

[0046] S6: Loop through steps S3 - S5.

[0047] Further, the initial value of the soft eviction threshold eviction-soft is set as:

[0048] eviction-soft = network.available < 20%;

[0049] The initial value of the hard eviction threshold eviction-hard is set as:

[0050] eviction-hard = network.available < 20%.

[0051] That is, when the network card traffic exceeds 80% (1 - 20%), an eviction signal signal is generated to evict the pod, and the eviction signal signal is a soft eviction or a hard eviction.

[0052] Further, the specific operation of evicting the active pods is as follows:

[0053] After receiving the eviction signal, the EvictionManager obtains the resource usage of the current node and all active pods, sorts all active pods by priority, evicts the pods with lower priority in the sorted order, and records the evicted pods in the time series database.

[0054] The eviction manager consists of three components: notifier, monitor, and synchronize. The present invention optimizes and improves these three components to support network card traffic monitoring and pod eviction. The notifier is improved to be the network card monitoring service thresholdNotifier, which receives preprocessed data, soft eviction thresholds, and hard eviction thresholds, generates a data set thresholds, and passes it to the network card traffic judgment unit. The network card traffic judgment unit performs matching operations based on the data in thresholds. If the eviction threshold conditions are met, an eviction signal signal is generated. thresholdNotifier sends signal to the channal message channel, triggers the synchronize eviction work, and stores the signal in the time series database. After receiving the signal, synchronize sorts the pods in the cluster in ascending order of priority and evicts a certain number of pods with lower priority to release network card resources. The soft eviction threshold and hard eviction threshold are set with initial values and are then updated in real time according to the optimal soft eviction threshold and hard eviction threshold calculated by the model.

[0055] The present invention also provides a system for monitoring network cards. The system uses the application of the above-mentioned model for monitoring network cards. The system includes: a collection module, a storage module, an eviction management module, and an intelligent parameter tuning component;

[0056] The collection module is used to monitor and collect data from the cluster node machines;

[0057] The storage module includes a time series database for storing the data collected by the collection module and the data generated during other processing procedures;

[0058] The eviction management module is used to load the K8s eviction manager and its corresponding components, set the soft eviction threshold and hard eviction threshold, analyze and calculate the data collected by the collection module, and determine whether the conditions of the soft eviction threshold and hard eviction threshold are met. If they are met, eviction is performed; otherwise, it is ignored;

[0059] The intelligent parameter tuning component is used to dynamically update the soft eviction threshold and hard eviction threshold.

[0060] The present invention also provides an electronic device for monitoring network cards, including:

[0061] A storage area and a processor;

[0062] Computer-readable instructions are stored on the memory, and the computer-readable instructions are used by the processor according to the application of the above-mentioned model for monitoring network cards or a system for monitoring network cards.

[0063] Compared with the prior art, the beneficial effects of the present invention are as follows: 1. By rewriting the K8s components for network card resources, the monitoring and eviction of network card resources of cluster machines are realized, and the problem that pods cannot be evicted in time when the network card is under heavy pressure is solved;

[0064] 2. By training a neural network model to construct an intelligent parameter tuning component, intelligent and dynamic adjustment of eviction conditions is realized. BRIEF DESCRIPTION OF THE DRAWINGS

[0065] Figure 1 Steps of a model training method for monitoring network cards according to the present invention;

[0066] Figure 2 Application steps of a model for monitoring network cards according to the present invention;

[0067] Figure 3 A system for monitoring network cards according to the present invention;

[0068] Reference numerals in the drawings: Acquisition module 1, Storage module 2, Eviction management module 3, Intelligent parameter tuning component 4. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0069] The drawings of the present invention are only for illustrative purposes and should not be construed as limiting the present invention. For better illustrating the following embodiments, some components in the drawings may be omitted, enlarged or reduced, and do not represent the dimensions of actual products; for those skilled in the art, it is understandable that some well-known structures and their descriptions in the drawings may be omitted.

[0070] Embodiment 1

[0071] This embodiment provides a model training method for monitoring network cards, and the method includes:

[0072] Obtain the eviction history set, calculate and generate a validation set matrix, and construct a training set D and a validation set V based on the validation set matrix; obtain the eviction history set from the time series database, including various performance parameters on the node cluster machines when eviction occurs, specifically including soft eviction threshold, hard eviction threshold, CPU usage rate, memory usage rate, network card usage rate, and eviction signal. Generate a validation set matrix X with this information, where X[0] is the CPU usage rate of the cluster machine when eviction occurs, X[1] is the memory usage rate when eviction occurs, X[2] is the network card usage rate when eviction occurs, and X[3] is the proportion of evicted pods in the total number of pods. In the specific implementation, disk read and write, traffic stress testing, and intensive CPU-running programs are used to adjust the usage rates of CPU, memory, and network card. Record the total number of running pods y1 at this time, the total number of pods y0 running on the machine before stress testing, as well as the CPU usage rate, memory usage rate, and network card usage rate. Among them, the total number of evicted pods y2 = y0 - y1. X[3] = y2 / y0, and combine the three indicators of CPU usage rate, memory usage rate, and network card usage rate at this time to form a 4-tuple of the eviction record matrix. The following is a specific example of the generated validation set matrix:

[0073]

[0074] After generating the validation set matrix X, construct a training set with the corresponding soft eviction threshold and hard eviction threshold during eviction and a validation set where X is the validation set matrix, and y is the eviction proportion of pods under the corresponding CPU usage rate, memory usage rate, and network card usage rate. Generate a training set and then perform deduplication optimization on the training set to reduce the computational amount of convolution operations in the convolutional neural network and improve the training operation speed.

[0075] Input the optimized training set into the convolutional neural network and train it through the stochastic gradient descent method combined with the backpropagation algorithm. Input the training set D into the neural network model to obtain the network output as Assume the loss function is Perform parameter learning by calculating the derivative of the loss function with respect to each parameter. The specific training steps are as follows:

[0076] A1: Randomly initialize the parameter weight matrix w and bias b;

[0077] A2: Randomly reorder the samples in the training set;

[0078] A3: Select a sample x from the training set D (n) , y (n) , and initially n = 0;

[0079] A4: Feedforward calculation of the net input z of each layer(l) and the activation value a (l) , until the last layer;

[0080] A5: Backpropagate to calculate the error δ for each layer (l) ; Derive that the gradient of the bias W for the l-th layer (l) is:

[0081] A6: Calculate the gradient of the bias b for the l-th layer (l) is:

[0082] A7: Update the W and b parameters through the formula: b (l) ← b (l) - αδ (l) ; A8: Increment the value of n by 1 and repeat steps A3 - A7 until training n = N;

[0083] A9: Repeat steps A2 - A8 until the error rate of the convolutional neural network model on the validation set V no longer decreases.

[0084] Since the network card data collected and obtained is very large, using general training methods will reduce the training efficiency. By combining the training method of stochastic gradient descent and backpropagation algorithm to train the model, the training efficiency can be greatly improved. Obtain the trained parameter - adjusted model, and the parameter - adjusted model is used to dynamically calculate the soft eviction threshold and hard eviction threshold according to the data situation on the current node cluster machine. When the network card traffic meets the soft eviction threshold or hard eviction threshold, soft eviction and hard eviction are respectively performed on the pod.

[0085] Specifically, the soft eviction threshold eviction - soft includes: cpu.available (CPU usage threshold), memory.available (memory usage threshold), and network.available (network card usage threshold); the hard eviction threshold eviction - hard includes: cpu.available (CPU usage threshold), memory.available (memory usage threshold), and network.available (network card usage threshold). CPU usage thresholds, memory usage thresholds, and network card usage thresholds are respectively set for the soft eviction and hard eviction of k8s, which can monitor the CPU, memory, and network card resources respectively. When the current CPU usage rate, memory usage rate, or network card usage rate exceeds the corresponding soft eviction or hard eviction threshold, the pod is evicted. By judging these three elements, the eviction of the pod is made more intelligent.

[0086] Through the trained parameter-tuning model, dynamically calculate the eviction threshold for eviction according to the status parameters of the cluster machines, enabling the machines to intelligently adjust the eviction conditions of the network card resources of the cluster machines through the model. At the same time, a training method combining the stochastic gradient descent and backpropagation algorithms is used to train the model, greatly improving the training efficiency.

[0087] Embodiment 2

[0088] This embodiment provides an application of a model for monitoring network cards. The method includes:

[0089] Configure the acquisition module 1 to monitor and collect data from the cluster node machines; preprocess and store the collected data; the acquisition module 1 uses cAdvisor to collect real-time CPU, memory, and network card traffic information of the node cluster machines. Among them, the network card traffic includes the cumulative received traffic (rx_bytes), cumulative received error traffic (rx_errors), cumulative transmitted traffic (tx_bytes), and cumulative transmitted error traffic (tx_errors). Set the acquisition period node-status-update-frequency parameter to 5s, store it in the memory, and divide the cumulative value by the acquisition period (5s) respectively to calculate the received traffic per second (rx_bytes_perSecon) and the transmitted traffic per second (tx_bytes_perSecon), and store them in the time series database.

[0090] Improve the eviction algorithm model of the network card based on K8s. The specific steps are as follows:

[0091] S1: Set initial values for the soft eviction threshold eviction-soft and the hard eviction threshold eviction-hard; the initial value of the soft eviction threshold eviction-soft is set to: eviction-soft = network.available < 20%; the initial value of the hard eviction threshold eviction-hard is set to: eviction-hard = network.available < 20%.

[0092] Configure the K8s model for the node cluster machines, improve the K8s model based on the monitoring of network card traffic, rewrite the monitoring and eviction of system CPU, memory, and disk capacity in K8s as the monitoring and eviction of network card traffic resources, and add a network card eviction threshold startup parameter: set the corresponding initial soft eviction threshold and hard eviction threshold. When the network card traffic data meets the soft eviction threshold and hard eviction threshold, a certain number of pods in the node cluster are evicted to release network card resources for other pods that need network card resources more. In this embodiment, set network.available < 20%, that is, when the network card traffic exceeds 80% (1 - 20%), an eviction signal signal is generated to evict the pod, and the eviction signal signal is a soft eviction or a hard eviction.

[0093] S2: Load the soft eviction threshold and the hard eviction threshold into the EvictionManager;

[0094] S3: Start a coroutine to listen to the service thresholdNotifier. The listening service obtains the preprocessed data and combines the soft eviction threshold and the hard eviction threshold in S2 to form a data set thresholds;

[0095] Specifically, the EvictionManager consists of three components: notifier, monitor, and synchronize. First, start the EvictionManager, load the soft eviction threshold and the hard eviction threshold, and share them for each component of the EvictionManager to use; rewrite the notifier, start the coroutine to listen to the service thresholdNotifier, receive the performance values such as the traffic transmitted per second and the traffic received per second collected and preprocessed from cAdvisor, as well as the soft eviction threshold and the hard eviction threshold previously loaded by the EvictionManager, and merge them to generate the data set thresholds;

[0096] S4: Configure the network card traffic judgment unit. After the listening service generates the data set thresholds, it transmits them to the network card traffic judgment unit. The network card traffic judgment unit performs a matching operation based on the data set thresholds to determine whether it meets the soft eviction threshold or the hard eviction threshold. If it meets, it sends an eviction signal signal, and the thresholdNotifier sends the signal to the channal message channel to trigger the eviction work of the synchronize component, and records the signal eviction signal in the time series database; if it does not meet the eviction condition, it is ignored;

[0097] S5: Evict active pods according to the eviction signal sent in step S4. Specifically, after synchronize receives the signal, it obtains the group member usage of the current node and all active pods, sorts all active pods by priority, evicts the pods with lower priority in the sorted order, and records the evicted pods in the time series database to release network card resources for other pods that need network card resources more.

[0098] S6: Loop through steps S3 - S5.

[0099] After configuring the eviction algorithm model, construct the intelligent parameter tuning component 4. The intelligent parameter tuning component 4 uses the tuning model trained by the model training method for monitoring network cards described in Embodiment 1 to dynamically update the soft eviction threshold and the hard eviction threshold. Specifically, use the trained model in Embodiment 1 to calculate the real-time CPU, memory, and network card traffic information, calculate the optimal soft eviction threshold and hard eviction threshold in this state, update them in real-time, and store the soft eviction threshold and hard eviction threshold before the update together with the state parameter information in the time series database as new historical data to provide data for model training.

[0100] Embodiment 3

[0101] This embodiment provides a system for monitoring network cards. The system adopts the application of a model for monitoring network cards described in Embodiment 2 above. The system includes: a collection module 1, a storage module 2, an eviction management module 3, and an intelligent parameter tuning component 4.

[0102] The collection module 1 is used to monitor and collect data from the cluster node machines.

[0103] The storage module 2 includes a time series database for storing the data collected by the collection module 1 and the data generated during other processing processes.

[0104] The eviction management module 3 is used to load the K8s eviction manager and its corresponding components, set the soft eviction threshold and the hard eviction threshold, analyze and calculate the data collected by the collection module 1, and determine whether it meets the conditions of the soft eviction threshold and the hard eviction threshold. If it meets, execute the eviction; otherwise, ignore it.

[0105] The intelligent parameter tuning component 4 is used to dynamically update the soft eviction threshold and the hard eviction threshold.

[0106] Embodiment 4

[0107] This embodiment provides an electronic device for monitoring network cards, including:

[0108] A storage area and a processor;

[0109] Computer-readable instructions are stored on the memory, and when the computer-readable instructions are executed by the processor, they are based on the application of a model for monitoring a network card in Embodiment 2 or a system for monitoring a network card in Embodiment 3 described above.

[0110] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the technical solutions of the present invention, rather than limitations on the specific implementation manners of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the claims of the present invention shall be included within the protection scope of the claims of the present invention.

Claims

1. A method for training a model for monitoring a network card, characterized in that, The method includes: Obtain an eviction history record set, calculate and generate a validation set matrix X, and construct a training set D and a validation set V based on the validation set matrix X; Input the training set into a convolutional neural network, and train it by means of stochastic gradient descent combined with the backpropagation algorithm to obtain a trained parameter tuning model. The parameter tuning model is used to dynamically calculate the soft eviction threshold and the hard eviction threshold according to the data situation on the current node cluster machine. When the network card traffic meets the soft eviction threshold, perform soft eviction on the pod; when the network card traffic meets the hard eviction threshold, perform hard eviction on the pod; The training by means of stochastic gradient descent combined with the backpropagation algorithm includes: Input the training set D into the neural network model, and the network output is , assuming the loss function is , and parameter learning is performed by calculating the derivative of the loss function with respect to each parameter. The specific steps are as follows: A1: Randomly initialize the parameter weight matrix w and the bias b; A2: Randomly reorder the samples in the training set; A3: Select samples from the training set D , , initially n = 0; A4: Calculate the net input and activation value of each layer until the last layer; ​ A5: Backpropagation calculates the error for each layer ; It is derived that The gradient with respect to the bias of the first layer is: ; A6: Calculation Bias for the first layer The gradient is as follows: ; A7: Update the W and b parameters through formulas: , A8: Increment the value of n by 1 and repeat steps A3 - A7 until training n = N; A9: Repeat steps A2 - A8 until the error rate of the convolutional neural network model no longer decreases on the validation set V; The eviction history record is the performance parameter indicators of the node cluster machine during eviction in history, as well as the corresponding soft eviction threshold and hard eviction threshold.

2. The model training method for monitoring network cards according to claim 1, wherein The calculation and generation of the validation set matrix, and the construction of the training set D and the validation set V based on the validation set matrix X are specifically: Extract the soft eviction threshold, hard eviction threshold, CPU, memory network card metrics, eviction semaphore, and eviction record from the obtained records, jointly extract the data to form a validation set matrix, and then generate the training set data D and the validation set V based on the validation set matrix; Training set ; Where X is the validation set matrix, X[0] represents the CPU usage rate, X[1] represents the memory usage rate, X[2] represents the network card usage rate, and y is the eviction proportion of the pod under the corresponding CPU usage rate, memory usage rate, and network card usage rate; Validation set , and the data format is the same as that of the training set.

3. The model training method for monitoring network cards according to claim 1, wherein The soft eviction threshold eviction - soft includes: cpu.available (CPU usage threshold), memory.available (memory usage threshold), and network.available (network card usage threshold); The hard eviction threshold eviction - hard includes: cpu.available (CPU usage threshold), memory.available (memory usage threshold), and network.available (network card usage threshold).

4. A method for applying a model for monitoring a network card, characterized in that, The method includes: Configure a collection module to monitor and collect data from the cluster node machines; Pre - process and store the collected data; Improve the eviction algorithm model for the network card based on K8s, including: set the initial values of the soft eviction threshold and the hard eviction threshold for the network card and set them into the monitoring service. When the network card node traffic occupancy is less than the soft eviction threshold or the hard eviction threshold, send an eviction signal to evict the pod corresponding to the network card node; Construct an intelligent parameter tuning component. The intelligent parameter tuning component adopts a parameter tuning model trained by the model training method for monitoring the network card according to any one of claims 1 - 3 to dynamically update the soft eviction threshold and the hard eviction threshold.

5. The application method of a model for monitoring network cards according to claim 4, characterized in that, The pre - processing and storage include: Save the cumulative received traffic (rx_bytes), cumulative received error traffic (rx_errors), cumulative transmitted traffic (tx_bytes), and cumulative transmitted error traffic (tx_errors) of the original data to memory, and calculate the received traffic per second (rx_bytes_perSecon) and the transmitted traffic per second (tx_bytes_perSecon); Record the calculated received traffic per second and transmitted traffic per second to the time series database.

6. The application method of a model for monitoring network cards according to claim 5, characterized in that, The specific steps for improving the network card eviction algorithm model based on K8s are as follows: S1: Set initial values for the soft eviction threshold eviction-soft and the hard eviction threshold eviction-hard; S2: Load the soft eviction threshold and the hard eviction threshold into the EvictionManager; S3: Start a coroutine to listen to the service thresholdNotifier. The listening service obtains the preprocessed data and combines the soft eviction threshold and the hard eviction threshold in S2 to form a data set thresholds; S4: Configure the network card traffic judgment unit and perform a matching operation based on the data set in step S3 to determine whether it meets the soft eviction threshold or the hard eviction threshold. If it meets, send an eviction signal signal and record the eviction signal in the time series database, otherwise ignore it; S5: Evict the active pods according to the eviction signal sent in step S4; S6: Loop and execute steps S3 - S5.

7. The application method of a model for monitoring network cards according to claim 6, characterized in that The initial value of the soft eviction threshold eviction-soft is set to: network.available < 20%; The initial value of the hard eviction threshold eviction-hard is set to: network.available < 20%.

8. The application method of a model for monitoring network cards according to claim 7, characterized in that, The eviction of active pods is specifically as follows: After receiving the eviction signal, the EvictionManager obtains the resource usage of the current node and all active pods, sorts all active pods by priority, evicts the pods with lower priority in the sorted order, and records the evicted pods in the time series database.

9. A system for monitoring a network card, the system adopting the application method of a model for monitoring a network card according to any one of claims 4 to 8, characterized in that, The system includes: a collection module, a storage module, an eviction management module, and an intelligent parameter tuning component; The collection module is used to monitor and collect data from the cluster node machines; The storage module includes a time series database, which is used to store the data collected by the collection module and the data generated in other processing processes; The eviction management module is used to load the K8s EvictionManager and its corresponding components, set the soft eviction threshold and the hard eviction threshold, analyze and calculate the data collected by the collection module, and determine whether it meets the conditions of the soft eviction threshold and the hard eviction threshold. If it meets, perform eviction, otherwise ignore it; The intelligent parameter tuning component is used to dynamically update the soft eviction threshold and the hard eviction threshold.

10. An electronic device for monitoring a network card, characterized in that, It includes: A storage area and a processor; Computer-readable instructions are stored on the memory, and when the computer-readable instructions are executed by the processor, they are for an application method of a model for monitoring a network card according to any one of claims 4-8 or a system for monitoring a network card according to claim 9.

Citation Information

Patent Citations

  • Active load balancing system and method based on application load

    CN111949412A

  • Kubernetes cluster resource dynamic adjustment method and electronic device

    WO2022016808A1