Abnormal cluster detection method and device, equipment, medium and product

By obtaining the CPU operation information and memory usage information of the target working node and calculating the CPU operation fluctuation coefficient and memory waste rate, the problem of insufficient detection accuracy caused by relying on single-time monitoring data in the existing technology is solved, and higher detection accuracy and reliability are achieved.

CN120704985APending Publication Date: 2025-09-26INDUSTRIAL AND COMMERCIAL BANK OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510800694.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-16
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

Existing abnormal cluster detection methods rely on real-time monitoring data at a single moment, resulting in insufficient detection accuracy and reliability, and unable to effectively identify problems of inefficient resource usage.

Method used

By obtaining the CPU operation information and memory usage information of the target working node within the target time period, the CPU operation fluctuation coefficient and memory waste rate are calculated, and abnormal clusters are detected based on these indicators.

Benefits of technology

It realizes abnormal cluster detection based on long-term monitoring data, improves the accuracy and reliability of detection, and can more accurately identify problems of inefficient resource use.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120704985A_ABST
    Figure CN120704985A_ABST
Patent Text Reader

Abstract

The invention discloses an abnormal cluster detection method, apparatus and device, a medium and a product, relates to the technical field of cloud computing, can be applied to the field of financial science and technology, and comprises the steps of determining a target working node deployed in a to-be-detected cluster, and obtaining CPU operation information and memory use information of the target working node in a target time period; wherein the CPU operation information comprises CPU utilization rate information and / or CPU operation frequency information; determining a CPU operation fluctuation coefficient of the target working node in the target time period according to the CPU operation information, and determining a memory waste rate of the target working node in the target time period according to the memory use information; and performing abnormal cluster detection on the to-be-detected cluster according to the CPU operation fluctuation coefficient and the memory waste rate. According to the invention, the effect of abnormal cluster detection depending on long-term monitoring data is realized, and compared with the prior art that detection depends on real-time monitoring data at a single moment, the accuracy and credibility of abnormal cluster detection are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of cloud computing technology, and in particular to a method, device, equipment, medium and product for detecting abnormal clusters. Background Art

[0002] With the development of cloud computing and big data technologies, large-scale computing clusters have become the core infrastructure for processing massive amounts of data and complex computing tasks. However, inefficient resource utilization seriously impacts the overall performance and operating costs of clusters. Therefore, it is necessary to promptly detect abnormal clusters with inefficient resource utilization.

[0003] Existing abnormal cluster detection methods usually rely on real-time monitoring data at a single moment for detection, and lack in-depth analysis of long-term monitoring data, resulting in insufficient accuracy and reliability of abnormal cluster detection. Summary of the Invention

[0004] The present invention provides a method, device, equipment, medium and product for detecting abnormal clusters to solve the problem of insufficient detection accuracy and reliability in existing abnormal cluster detection methods.

[0005] According to one aspect of the present invention, a method for detecting abnormal clusters is provided, the method comprising:

[0006] Determine a target working node deployed in the cluster to be tested, and obtain CPU operation information and memory usage information of the target working node within a target time period; wherein the CPU operation information includes CPU utilization information and / or CPU operation frequency information;

[0007] Determine a CPU operation fluctuation coefficient of the target working node within the target time period according to the CPU operation information, and determine a memory waste rate of the target working node within the target time period according to the memory usage information;

[0008] The cluster to be detected is detected as an abnormal cluster according to the CPU operation fluctuation coefficient and the memory waste rate.

[0009] According to another aspect of the present invention, a device for detecting abnormal clusters is provided, the device comprising:

[0010] An information acquisition module is used to determine a target working node deployed in the cluster to be detected, and obtain CPU operation information and memory usage information of the target working node within a target time period; wherein the CPU operation information includes CPU utilization information and / or CPU operation frequency information;

[0011] a parameter determination module, configured to determine a CPU operation fluctuation coefficient of the target working node within the target time period according to the CPU operation information, and to determine a memory waste rate of the target working node within the target time period according to the memory usage information;

[0012] The abnormal cluster detection module is used to detect abnormal clusters of the cluster to be detected according to the CPU operation fluctuation coefficient and the memory waste rate.

[0013] According to another aspect of the present invention, an electronic device is provided, comprising:

[0014] at least one processor; and

[0015] a memory communicatively connected to the at least one processor; wherein,

[0016] The memory stores a computer program executable by the at least one processor. The computer program is executed by the at least one processor so that the at least one processor can perform any one of the abnormal cluster detection methods of the present invention.

[0017] According to another aspect of the present invention, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement any one of the abnormal cluster detection methods of the present invention when executed.

[0018] According to another aspect of the present invention, a computer program product is provided, comprising a computer program, wherein when executed by a processor, the computer program implements any one of the abnormal cluster detection methods of the present invention.

[0019] The present invention determines the target working node deployed in the cluster to be detected, and obtains the CPU operation information and memory usage information of the target working node within the target time period; determines the CPU operation fluctuation coefficient of the target working node within the target time period according to the CPU operation information, and determines the memory waste rate of the target working node within the target time period according to the memory usage information; and detects abnormal clusters in the cluster to be detected according to the CPU operation fluctuation coefficient and the memory waste rate. The beneficial effect is that: since the abnormal cluster is detected in the cluster to be detected based on the CPU operation fluctuation coefficient and the memory waste rate of the target working node within the target time period, the effect of relying on long-term monitoring data for abnormal cluster detection is achieved, which improves the accuracy and credibility of abnormal cluster detection compared with the existing technology that relies on real-time monitoring data at a single moment for detection.

[0020] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present invention, nor is it intended to limit the scope of the present invention. Other features of the present invention will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0022] Figure 1 A flowchart of a method for detecting abnormal clusters provided in Example 1 of the present invention;

[0023] Figure 2 A flowchart of a method for detecting abnormal clusters provided in the second embodiment of the present invention;

[0024] Figure 3 This is a flowchart of a method for detecting abnormal clusters provided in Example 3 of the present invention;

[0025] Figure 4 A schematic diagram of the structure of an abnormal cluster detection device provided in the fourth embodiment of the present invention;

[0026] Figure 5 It is a structural diagram of an electronic device for implementing the abnormal cluster detection method according to an embodiment of the present invention. DETAILED DESCRIPTION

[0027] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.

[0028] It should be noted that the terms "first", "second", "candidate", "target", etc. in the description and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the numbers used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0029] Example 1

[0030] Figure 1 This is a flow chart of a method for detecting abnormal clusters provided in the first embodiment of the present invention. This embodiment is applicable to the case where abnormal clusters are detected in a cluster to be detected by using the CPU operation information and memory usage information of the target working node within the target time period. The method can be executed by an abnormal cluster detection device, which can be implemented in the form of hardware and / or software, such as server implementation. Figure 1 As shown, the method includes:

[0031] S101: Determine a target working node deployed in a cluster to be detected, and obtain CPU operation information and memory usage information of the target working node within a target time period.

[0032] The cluster under inspection refers to the cluster that needs to be checked for anomalies, that is, for inefficient resource usage. A cluster is a system architecture consisting of multiple nodes interconnected by a high-speed network, working together to complete a common task. These nodes are integrated into a logical entity by software, providing high-performance, high-availability, or load-balancing services.

[0033] At least one target working node is deployed in the cluster to be detected. The target working node is a physical or virtual resource unit that actually carries the business load and performs computing tasks in the cluster to be detected, such as a server or computer, etc. When performing computing tasks, the target working node will call CPU resources and memory resources. When calling CPU resources, it will generate information related to CPU operation, and when calling memory resources, it will generate information related to memory usage. This embodiment obtains the information related to CPU operation generated by the target working node within the target time period as CPU operation information; and obtains the information related to memory usage generated by the target working node within the target time period as memory usage information. Among them, the target time period can be set according to actual needs, such as setting a preset length time period before the current moment as the target time period, or setting a preset length time period after the current moment as the target time period. This embodiment does not limit the specific moment of the target time period.

[0034] CPU operating information refers to attribute data about the operating status of the central processing unit (CPU) of the target worker node, including CPU utilization and / or CPU operating frequency. CPU utilization is a key metric for measuring CPU workload, quantifying the proportion of CPU processing power used to execute tasks within a specific time period. CPU operating frequency is a core parameter that reflects the CPU's operating status and is directly related to the processor's computing speed and performance.

[0035] In one embodiment, a management node in the cluster to be tested is accessed and the cluster management interface of the management node is called to obtain a node list of the cluster to be tested. Furthermore, the node role label of each node in the node list is identified, and nodes with a node role of a working node are selected as target working nodes based on the node role label.

[0036] Furthermore, a scheduled task is set to automatically collect data from the target working node at fixed time intervals, such as every 5 minutes, within the target time period, and collect the CPU utilization information and / or CPU operating frequency information of the target working node within the target time period as CPU operation information, and collect the memory usage information of the target working node within the target time period.

[0037] Optionally, after obtaining the CPU running information and memory usage information of the target working node during the target time period, the following is also included:

[0038] Data cleaning is performed on CPU operation and memory usage information to remove invalid records such as null values ​​and outliers, fill in missing values, and unify the data format. Furthermore, the cleaned CPU operation and memory usage information is normalized or standardized to ensure comparability between data of different dimensions.

[0039] S102: Determine a CPU operation fluctuation coefficient of the target working node within a target time period according to the CPU operation information, and determine a memory waste rate of the target working node within the target time period according to the memory usage information.

[0040] The CPU operating fluctuation coefficient is a core indicator for measuring CPU performance stability, indicating the degree to which the actual CPU operating frequency and / or utilization deviates from the theoretical level. The memory waste rate is an indicator that measures the degree to which memory resources are ineffectively occupied.

[0041] In one embodiment, the CPU utilization standard deviation and the CPU utilization mean of the target working node within the target time period are determined based on the CPU utilization information, and the CPU operation fluctuation coefficient of the target working node within the target time period is determined based on the CPU utilization standard deviation and the CPU utilization mean.

[0042] In another embodiment, the CPU operating frequency standard deviation and the CPU operating frequency mean of the target working node within the target time period are determined based on the CPU operating frequency information, and the CPU operating fluctuation coefficient of the target working node within the target time period is determined based on the CPU operating frequency standard deviation and the CPU operating frequency mean.

[0043] In another embodiment, the CPU utilization standard deviation and the CPU utilization mean of the target working node within the target time period are determined based on the CPU utilization information, and the CPU utilization fluctuation coefficient of the target working node within the target time period is determined based on the CPU utilization standard deviation and the CPU utilization mean.

[0044] The CPU operating frequency standard deviation and the CPU operating frequency mean of the target working node within the target time period are determined according to the CPU operating frequency information, and the CPU operating frequency fluctuation coefficient of the target working node within the target time period is determined according to the CPU operating frequency standard deviation and the CPU operating frequency mean.

[0045] The CPU operation fluctuation coefficient is determined based on the CPU utilization fluctuation coefficient and the CPU operation frequency fluctuation coefficient.

[0046] Furthermore, the amount of memory wasted of the target working node within the target time period is determined based on the memory usage information, and the memory waste rate is determined based on the amount of memory wasted and the total memory capacity of the target working node.

[0047] S103: Detect abnormal clusters in the cluster to be detected based on the CPU operation fluctuation coefficient and the memory waste rate.

[0048] In one embodiment, a cluster efficiency score is calculated based on the CPU operation fluctuation coefficient and the memory waste rate, and the cluster efficiency score corresponding to the cluster to be detected is determined based on the calculation result.

[0049] Furthermore, in one embodiment, a pre-set efficiency score threshold is obtained, and based on the magnitude relationship between the cluster efficiency score and the efficiency score threshold, the cluster to be detected is detected as an abnormal cluster. For example, if the cluster efficiency score is greater than or equal to the efficiency score threshold, the cluster to be detected is determined to be an abnormal cluster, that is, the cluster to be detected has a problem of inefficient resource usage; if the cluster efficiency score is less than the efficiency score threshold, the cluster to be detected is determined not to be an abnormal cluster, that is, the cluster to be detected does not have a problem of inefficient resource usage.

[0050] Furthermore, in another embodiment, when there are at least two clusters to be detected, the cluster efficiency scores of the clusters to be detected are ranked, and based on the ranking results, the clusters to be detected that are ranked after the preset order are identified as abnormal clusters. For example, assuming there are ten clusters to be detected and the preset order is six, the clusters to be detected with ranking scores of seven, eight, nine, and ten are identified as abnormal clusters.

[0051] The present invention determines the target working node deployed in the cluster to be detected, and obtains the CPU operation information and memory usage information of the target working node within the target time period; determines the CPU operation fluctuation coefficient of the target working node within the target time period according to the CPU operation information, and determines the memory waste rate of the target working node within the target time period according to the memory usage information; and detects abnormal clusters in the cluster to be detected according to the CPU operation fluctuation coefficient and the memory waste rate. The beneficial effect is that: since the abnormal cluster is detected in the cluster to be detected based on the CPU operation fluctuation coefficient and the memory waste rate of the target working node within the target time period, the effect of relying on long-term monitoring data for abnormal cluster detection is achieved, which improves the accuracy and credibility of abnormal cluster detection compared with the existing technology that relies on real-time monitoring data at a single moment for detection.

[0052] Example 2

[0053] Figure 2 This is a flow chart of a method for detecting abnormal clusters provided in the second embodiment of the present invention. This embodiment further optimizes and expands the above embodiment and can be combined with the above optional implementations. Figure 2 As shown, the method includes:

[0054] S201: Determine a target working node deployed in a cluster to be detected, and obtain CPU operation information and memory usage information of the target working node within a target time period.

[0055] The CPU operation information includes CPU utilization information and / or CPU operation frequency information.

[0056] Optionally, the CPU utilization information includes the overall CPU utilization of the target working node at each sampling time point, and the CPU operating frequency information includes the overall CPU operating frequency of the target working node at each sampling time point.

[0057] The sampling time point represents a preset time point for automatically collecting data from the target working node within the target time period, such as every 5 minutes.

[0058] Overall CPU utilization refers to the proportion of time the CPU spends executing tasks, usually expressed as a percentage. It comprehensively reflects how the CPU allocates time processing user programs, system tasks, and idle time. The overall CPU operating frequency is a core indicator of CPU performance, indicating the number of oscillations per second of the CPU's internal clock signal.

[0059] S202. Determine the standard deviation of the overall CPU utilization and the mean overall CPU utilization of the target working node within the target time period based on the overall CPU utilization of each CPU, and determine the first CPU utilization fluctuation coefficient of the target working node within the target time period based on the standard deviation of the overall CPU utilization and the mean overall CPU utilization.

[0060] In one embodiment, a standard deviation is calculated based on the overall CPU utilization at each sampling time point, and the standard deviation of the overall CPU utilization of the target work node during the target time period is determined based on the standard deviation calculation result. Furthermore, a mean is calculated based on the overall CPU utilization at each sampling time point, and the mean overall CPU utilization of the target work node during the target time period is determined based on the mean calculation result. Furthermore, a first CPU utilization fluctuation coefficient of the target work node during the target time period is determined based on the ratio between the standard deviation of the overall CPU utilization and the mean overall CPU utilization.

[0061] S203. Determine the standard deviation of the overall CPU operating frequency and the mean of the overall CPU operating frequency of the target working node within the target time period based on the overall CPU operating frequency of each CPU, and determine the first CPU operating frequency fluctuation coefficient of the target working node within the target time period based on the standard deviation of the overall CPU operating frequency and the mean of the overall CPU operating frequency.

[0062] In one embodiment, a standard deviation is calculated based on the overall CPU operating frequency at each sampling time point, and the standard deviation of the overall CPU operating frequency of the target work node within the target time period is determined based on the standard deviation calculation result. Furthermore, a mean is calculated based on the overall CPU operating frequency at each sampling time point, and the mean of the overall CPU operating frequency of the target work node within the target time period is determined based on the mean calculation result. Furthermore, a first CPU operating frequency fluctuation coefficient of the target work node within the target time period is determined based on the ratio between the standard deviation of the overall CPU operating frequency and the mean of the overall CPU operating frequency.

[0063] S204. Determine a CPU operation fluctuation coefficient according to the first CPU utilization fluctuation coefficient and the first CPU operation frequency fluctuation coefficient, and determine a memory waste rate of the target working node within a target time period according to the memory usage information.

[0064] In one embodiment, a weighted sum is performed based on the first CPU utilization fluctuation coefficient and the first CPU operating frequency fluctuation coefficient, and the CPU operating fluctuation coefficient is determined based on the weighted sum result.

[0065] By determining the overall CPU utilization standard deviation and the overall CPU utilization mean of the target working node in the target time period according to the overall CPU utilization of each CPU, and determining the first CPU utilization fluctuation coefficient of the target working node in the target time period according to the overall CPU utilization standard deviation and the overall CPU utilization mean; determining the overall CPU operating frequency standard deviation and the overall CPU operating frequency mean of the target working node in the target time period according to the overall CPU operating frequency of each CPU, and determining the first CPU operating frequency fluctuation coefficient of the target working node in the target time period according to the overall CPU operating frequency standard deviation and the overall CPU operating frequency mean; and determining the CPU operating fluctuation coefficient according to the first CPU utilization fluctuation coefficient and the first CPU operating frequency fluctuation coefficient, the beneficial effects are:

[0066] First, CPU utilization fluctuations and CPU operating frequency fluctuations are integrated to comprehensively evaluate CPU stability and avoid the one-sidedness of a single indicator.

[0067] Secondly, the first CPU utilization fluctuation coefficient is calculated by the standard deviation of the overall CPU utilization and the mean of the overall CPU utilization. In addition, the first CPU operating frequency fluctuation coefficient is calculated by the standard deviation of the overall CPU operating frequency and the mean of the overall CPU operating frequency. This can eliminate the dimensional differences caused by different work node scales and achieve horizontal comparison across work nodes.

[0068] Thirdly, by collecting the overall CPU utilization and the overall CPU operating frequency to calculate the CPU operation fluctuation coefficient, the amount of data processing can be reduced and the calculation speed of the CPU operation fluctuation coefficient can be accelerated.

[0069] Optionally, determining the CPU operation fluctuation coefficient according to the first CPU utilization fluctuation coefficient and the first CPU operation frequency fluctuation coefficient includes:

[0070] S2041. Determine a first weight corresponding to a first CPU operating frequency fluctuation coefficient based on an average value of the overall CPU utilization, and determine a second weight corresponding to the first CPU utilization fluctuation coefficient based on the first weight.

[0071] The sum of the first weight and the second weight is 1. For example, assuming the first weight is 0.3, the second weight is 1-0.3=0.7; and assuming the first weight is 0.7, the second weight is 1-0.7=0.3.

[0072] In one embodiment, the first weight is determined based on the ratio between the average overall CPU utilization and 100. For example, if the overall CPU utilization is 70%, the first weight is 70 / 100 = 0.7. The second weight is determined based on the difference between 1 and the first weight. For example, if the first weight is 0.7, the second weight is 1-0.7 = 0.3.

[0073] S2042: Perform a weighted summation on the first CPU operating frequency fluctuation coefficient and the first CPU utilization fluctuation coefficient according to the first weight and the second weight, and determine the CPU operating fluctuation coefficient according to the weighted summation result.

[0074] In one embodiment, a weighted sum is performed based on the first weight, the first CPU operating frequency fluctuation coefficient, and the second weight, the first CPU utilization fluctuation coefficient, and the CPU operating fluctuation coefficient is determined based on the weighted sum. For example, assuming the first weight is β, the second weight is α, the first CPU operating frequency fluctuation coefficient is k1, and the first CPU utilization fluctuation coefficient is k2, then the CPU operating fluctuation coefficient = k1*β+k2*α.

[0075] By determining a first weight corresponding to the first CPU operating frequency fluctuation coefficient according to the average value of the overall CPU utilization, and determining a second weight corresponding to the first CPU utilization fluctuation coefficient according to the first weight, wherein the sum of the first weight and the second weight is one; performing a weighted sum of the first CPU operating frequency fluctuation coefficient and the first CPU utilization fluctuation coefficient according to the first weight and the second weight, and determining the CPU operating fluctuation coefficient according to the weighted sum result, the beneficial effects are:

[0076] First, when the average overall CPU utilization is high, the weight of the first CPU operating frequency fluctuation coefficient (first weight) is increased, and more attention is paid to frequency stability; otherwise, emphasis is placed on the burstiness of the business load.

[0077] Secondly, the first weight reflects the change in CPU hardware status, and the second weight reflects the actual load pressure. The two complement each other and cover the physical layer and application layer indicators.

[0078] Thirdly, the sum of the first weight and the second weight is one, which avoids misjudgment dominated by a single indicator.

[0079] S205 : Detect abnormal clusters in the cluster to be detected based on the CPU operation fluctuation coefficient and the memory waste rate.

[0080] Example 3

[0081] Figure 3 This is a flow chart of a method for detecting abnormal clusters provided by the third embodiment of the present invention. This embodiment further optimizes and expands the above embodiment and can be combined with the above optional implementations. Figure 3 As shown, the method includes:

[0082] S301: Determine a target working node deployed in a cluster to be detected, and obtain CPU operation information and memory usage information of the target working node within a target time period.

[0083] The CPU operation information includes CPU utilization information and / or CPU operation frequency information.

[0084] Optionally, the CPU utilization information includes the CPU single-core utilization of each CPU core of the target work node at each sampling time point, and the CPU operating frequency information includes the CPU single-core operating frequency of each CPU core of the target work node at each sampling time point.

[0085] Among them, the CPU of the target working node includes at least one CPU core. The CPU core is an independent physical computing unit inside the CPU. Each CPU core can independently execute instructions and process data, which is equivalent to a microprocessor.

[0086] CPU single-core utilization refers to the percentage of time a single CPU core spends executing valid tasks within a specific time period, reflecting the workload saturation level of that core. CPU single-core operating frequency refers to the operating frequency of a single CPU core in the CPU when executing tasks.

[0087] Optionally, the memory usage information includes a first amount of memory in an unallocated state of the target working node during the target time period, and a second amount of memory in an allocated state but not used.

[0088] S302. Determine the CPU single-core utilization standard deviation and the CPU single-core utilization mean of each CPU core within the target time period based on the CPU single-core utilization standard deviation and the CPU single-core utilization mean, and determine the single-core utilization fluctuation coefficient of each CPU core within the target time period based on the CPU single-core utilization standard deviation and the CPU single-core utilization mean.

[0089] In one embodiment, a standard deviation is calculated based on the CPU single-core utilization of each CPU core at each sampling time point, and the standard deviation of the CPU single-core utilization of each CPU core during the target time period is determined based on the standard deviation calculation result. Furthermore, a mean is calculated based on the CPU single-core utilization of each CPU core at each sampling time point, and the mean of the CPU single-core utilization of each CPU core during the target time period is determined based on the mean calculation result. Furthermore, a single-core utilization fluctuation coefficient of each CPU core during the target time period is determined based on the ratio between the standard deviation of the CPU single-core utilization of each CPU core and the mean of the CPU single-core utilization.

[0090] S303. Determine the CPU single-core operating frequency standard deviation and the CPU single-core operating frequency mean of each CPU core within the target time period based on the CPU single-core operating frequency, and determine the single-core operating frequency fluctuation coefficient of each CPU core within the target time period based on the CPU single-core operating frequency standard deviation and the CPU single-core operating frequency mean.

[0091] In one embodiment, a standard deviation is calculated based on the CPU single-core operating frequency of each CPU core at each sampling time point, and the standard deviation of the CPU single-core operating frequency of each CPU core during the target time period is determined based on the standard deviation calculation result. Furthermore, a mean is calculated based on the CPU single-core operating frequency of each CPU core at each sampling time point, and the mean of the CPU single-core operating frequency of each CPU core during the target time period is determined based on the mean calculation result. Furthermore, a single-core operating frequency fluctuation coefficient of each CPU core during the target time period is determined based on the ratio between the standard deviation of the CPU single-core operating frequency of each CPU core and the mean of the CPU single-core operating frequency.

[0092] S304 : Determine a CPU operation fluctuation coefficient according to the utilization fluctuation coefficient of each single core and the operation frequency fluctuation coefficient of each single core.

[0093] In one embodiment, a weighted sum is performed based on the utilization fluctuation coefficient of each single core and the operating frequency fluctuation coefficient of each single core, and the CPU operating fluctuation coefficient is determined based on the weighted sum result.

[0094] By determining the CPU single-core utilization standard deviation and the CPU single-core utilization mean of each CPU core in a target time period based on the CPU single-core utilization, and determining the single-core utilization fluctuation coefficient of each CPU core in the target time period based on the CPU single-core utilization standard deviation and the CPU single-core utilization mean; determining the CPU single-core operating frequency standard deviation and the CPU single-core operating frequency mean of each CPU core in the target time period based on the CPU single-core operating frequency, and determining the single-core operating frequency fluctuation coefficient of each CPU core in the target time period based on the CPU single-core operating frequency standard deviation and the CPU single-core operating frequency mean; and determining the CPU operating fluctuation coefficient based on the utilization fluctuation coefficient of each single-core and the operating frequency fluctuation coefficient of each single-core, the beneficial effects are:

[0095] First, by calculating the single-core utilization fluctuation coefficient and the single-core operating frequency fluctuation coefficient to determine the CPU operating fluctuation coefficient, the reliability and accuracy of the CPU operating fluctuation coefficient determination are improved, and the problem of inaccurate calculation when using the overall CPU utilization and the overall CPU operating frequency to calculate the CPU operating fluctuation coefficient in load imbalance scenarios is avoided.

[0096] Secondly, the single-core utilization fluctuation coefficient and the single-core operating frequency fluctuation coefficient can accurately reflect the stability of the core.

[0097] Thirdly, the CPU operation fluctuation coefficient integrates the single-core utilization fluctuation coefficient and the single-core operation frequency fluctuation coefficient, which can avoid misjudgment of the overall system performance due to local instability.

[0098] Optionally, determining the CPU operation fluctuation coefficient based on the utilization fluctuation coefficient of each single core and the operation frequency fluctuation coefficient of each single core includes:

[0099] S3041. Perform a weighted sum based on the utilization fluctuation coefficient of each single core and the first core weight of each CPU core, and determine the second CPU utilization fluctuation coefficient of the target working node within the target time period based on the weighted sum result.

[0100] The first core weight is determined according to the average CPU single-core utilization of each CPU core.

[0101] In one embodiment, the total mean CPU single-core utilization is determined by summing the mean CPU single-core utilization of each CPU core, and the proportion of the mean CPU single-core utilization of each CPU core is determined based on the ratio of the mean CPU single-core utilization of each CPU core to the total mean CPU single-core utilization, and then the first core weight of each CPU core is determined based on the proportion.

[0102] Furthermore, a weighted sum is performed based on the single-core utilization fluctuation coefficient of each CPU core and the first core weight of each CPU core, and the second CPU utilization fluctuation coefficient of the target working node within the target time period is determined based on the weighted summation result.

[0103] S3042. Perform a weighted sum based on the fluctuation coefficient of each single-core operating frequency and the second core weight of each CPU core, and determine the second CPU operating frequency fluctuation coefficient of the target working node within the target time period based on the weighted summation result.

[0104] The second core weight is determined by the core computing power of each CPU core. Core computing power refers to the basic ability of a single CPU core to process data.

[0105] In one embodiment, the second core weight corresponding to each CPU core is determined based on the core computing power of each CPU core and a pre-set mapping relationship between the core computing power and the core weight. It is understood that the higher the core computing power of any CPU core, the greater the second core weight corresponding to the CPU core; and the lower the core computing power of any CPU core, the smaller the second core weight corresponding to the CPU core.

[0106] Furthermore, a weighted sum is performed based on the single-core operating frequency fluctuation coefficient of each CPU core and the second core weight of each CPU core, and the second CPU operating frequency fluctuation coefficient of the target working node within the target time period is determined based on the weighted summation result.

[0107] S3043. Determine a CPU operation fluctuation coefficient according to the second CPU utilization fluctuation coefficient and the second CPU operation frequency fluctuation coefficient.

[0108] In one embodiment, a weighted sum is performed on the second CPU utilization fluctuation coefficient and the second CPU operating frequency fluctuation coefficient, and the CPU operating fluctuation coefficient is determined according to the weighted sum result.

[0109] By performing a weighted summation based on the single-core utilization fluctuation coefficient and the first core weight of each CPU core, and determining the second CPU utilization fluctuation coefficient of the target work node within the target time period according to the weighted summation result; wherein the first core weight is determined according to the average CPU single-core utilization of each CPU core; performing a weighted summation based on the single-core operating frequency fluctuation coefficient and the second core weight of each CPU core, and determining the second CPU operating frequency fluctuation coefficient of the target work node within the target time period according to the weighted summation result; wherein the second core weight is determined according to the core computing power of each CPU core; and determining the CPU operating fluctuation coefficient according to the second CPU utilization fluctuation coefficient and the second CPU operating frequency fluctuation coefficient, the beneficial effects are:

[0110] First, the weight of the first core is determined based on the average CPU single-core utilization, so that the fluctuation of the high-load core has a greater impact on the overall system, achieving the effect of dynamically sensing the core load difference and avoiding the traditional averaging algorithm from masking the single-core overload problem.

[0111] Secondly, the weight of the second core is determined based on the core computing power of the CPU core, so that the frequency fluctuations of the high-computing-power core are amplified, achieving the effect of computing power weighted enhancement of frequency control.

[0112] Thirdly, the dual fluctuation coefficients of CPU utilization fluctuation coefficient and CPU operating frequency fluctuation coefficient are integrated to improve the reliability of CPU operating fluctuation coefficient.

[0113] S305. Determine the amount of memory waste of the target working node within the target time period based on the sum of the first memory amount and the second memory amount; and determine the memory waste rate based on the memory waste amount and the total memory amount of the target working node.

[0114] For example, assuming the first memory amount is A1 and the second memory amount is A2, the memory waste amount of the target working node during the target time period is A1 + A2. Assuming the total memory amount of the target working node is A3, the memory waste rate is (A1 + A2) / A3.

[0115] By determining the amount of memory waste of the target working node within the target time period based on the sum of the first memory amount and the second memory amount; and determining the memory waste rate based on the amount of memory waste and the total memory amount of the target working node, the beneficial effects are:

[0116] By determining the first amount of memory in an unallocated state, inefficient resource scheduling can be exposed; by determining the second amount of memory in an allocated but unused state, application-level waste can be identified. Therefore, determining the memory waste rate based on the first amount of memory and the second amount of memory can further improve the accuracy and comprehensiveness of the memory waste rate determination.

[0117] S306: Determine a cluster efficiency score corresponding to the cluster to be detected based on the CPU operation fluctuation coefficient and the memory waste rate; compare the cluster efficiency score with the efficiency score threshold; and determine that the cluster to be detected is an abnormal cluster if the cluster efficiency score is lower than the efficiency score threshold.

[0118] In one embodiment, the cluster efficiency score corresponding to the cluster to be detected is determined by the following formula: cluster efficiency score = (1-CPU operation fluctuation coefficient) * (1-memory waste rate).

[0119] Furthermore, the cluster efficiency score is compared with the efficiency score threshold. If it is determined that the cluster efficiency score is lower than the efficiency score threshold, the cluster to be detected is determined to be an abnormal cluster; if it is determined that the cluster efficiency score is higher than or equal to the efficiency score threshold, the cluster to be detected is determined to be a normal cluster.

[0120] The efficiency score threshold may be a pre-set fixed threshold; or it may be determined based on the percentile of the cluster efficiency scores of all clusters to be detected, such as using the cluster efficiency score corresponding to the 30th percentile of all clusters to be detected as the efficiency score threshold.

[0121] The cluster efficiency score corresponding to the cluster to be detected is determined based on the CPU operation fluctuation coefficient and memory waste rate. The cluster efficiency score is compared with the efficiency score threshold. If the cluster efficiency score is lower than the efficiency score threshold, the cluster to be detected is determined to be an abnormal cluster. The beneficial effects are:

[0122] First, the combination of the CPU operation fluctuation coefficient and the memory waste rate can fully expose the resource call efficiency problems of the cluster. Compared with traditional detection based on a single indicator, it can improve the accuracy of abnormal cluster detection.

[0123] Secondly, when the cluster efficiency score is lower than the efficiency score threshold, the anomaly is automatically marked, realizing automated abnormal cluster detection, reducing labor costs and improving detection efficiency.

[0124] Optionally, after determining that the cluster to be detected is an abnormal cluster, the following steps are further included:

[0125] S11. Use the machine learning model to predict whether the cluster to be detected is an abnormal cluster based on the CPU operation information and memory usage information.

[0126] In one embodiment, feature data is extracted based on CPU operation information and memory usage information to determine the standard deviation of CPU usage, the peak-to-valley difference in memory usage, and the length of consecutive periods of inefficiency. Furthermore, the standard deviation of CPU usage, the peak-to-valley difference in memory usage, and the length of consecutive periods of inefficiency are input into a pre-trained machine learning model. The machine learning model then outputs an abnormal cluster prediction result for the cluster to be detected based on the standard deviation of CPU usage, the peak-to-valley difference in memory usage, and the length of consecutive periods of inefficiency.

[0127] S21. When the machine learning model predicts that the cluster to be detected is not an abnormal cluster, generate manual detection prompt information.

[0128] The manual detection prompt information is used to prompt the technical staff to manually detect abnormal clusters in the cluster to be detected.

[0129] In one embodiment, if the machine learning model predicts that the cluster to be detected is an abnormal cluster, no operation is triggered; if the machine learning model predicts that the cluster to be detected is not an abnormal cluster, that is, it is different from the detection result obtained using the CPU operation fluctuation coefficient and the memory waste rate, in order to ensure the confidence of the detection result, a manual detection prompt information is generated to prompt the technician to use manual methods to detect abnormal clusters on the cluster to be detected.

[0130] By using a machine learning model, we can predict whether the cluster to be detected is an abnormal cluster based on CPU operation information and memory usage information. If the machine learning model predicts that the cluster to be detected is not an abnormal cluster, we can generate a manual detection prompt message. The beneficial effects are:

[0131] First, when the machine learning model determines that the cluster to be tested is not an abnormal cluster, but is detected as an abnormal cluster based on the CPU operation fluctuation coefficient and memory waste rate, it prompts manual review to prevent the automated system from blindly trusting a single detection result, forming a closed loop of "machine warning, manual analysis and judgment, and strategy adjustment."

[0132] Secondly, machine learning models can reduce the false alarm rate by fusing multi-dimensional features to identify "pseudo-anomalies".

[0133] Optionally, after determining that the cluster to be detected is an abnormal cluster, the following steps are further included:

[0134] S12. Determine the target abnormality cause and the impact degree of the target abnormality corresponding to the cluster to be detected.

[0135] In one embodiment, a prompt message is sent to a technician to prompt the technician to conduct an in-depth analysis of the cluster to be detected as an abnormal cluster, and determine the target abnormality cause and target abnormality impact degree corresponding to the cluster to be detected based on the technician's analysis results.

[0136] S22. Determine a target optimization measure corresponding to the cluster to be detected from the candidate optimization measures according to a mapping relationship between the candidate optimization measures, the candidate abnormality causes, and the candidate abnormality impact degrees.

[0137] The mapping relationship between candidate optimization measures, candidate anomaly causes, and candidate anomaly impact levels is pre-set. For example, a mapping relationship is established between candidate optimization measure 1 and "candidate anomaly cause 1, candidate anomaly impact level 1." This means that when the anomaly cause is "candidate anomaly cause 1" and the anomaly impact level is "candidate anomaly impact level 1," "candidate optimization measure 1" can be used to optimize the cluster.

[0138] In one embodiment, the target abnormality cause and the target abnormality impact degree are matched with the mapping relationship, and the target optimization measure corresponding to the cluster to be detected is determined from the candidate optimization measures based on the matching result. For example, assuming that the target abnormality cause is "abnormality cause 2" and the target abnormality impact degree is "abnormality impact degree 2", and assuming that there is an association between "optimization measure 2" and "abnormality cause 2" and "abnormality impact degree 2" in the mapping relationship, "optimization measure 2" is used as the target optimization measure. Among them, the candidate optimization measures include but are not limited to adjusting virtual machine resource quotas, optimizing workload scheduling strategies, recommending hardware upgrade solutions, etc.

[0139] S32. Generate optimization measure suggestion information based on the target optimization measures.

[0140] Among them, the optimization measure recommendation information is used to assist technical personnel in adopting target optimization measures to perform cluster optimization on the cluster to be detected.

[0141] By determining the target abnormality cause and the target abnormality impact degree corresponding to the cluster to be detected; according to the mapping relationship between the candidate optimization measures and the candidate abnormality cause and the candidate abnormality impact degree, the target optimization measures corresponding to the cluster to be detected are determined from the candidate optimization measures; according to the target optimization measures, optimization measure recommendation information is generated; among which, the optimization measure recommendation information is used to assist technical personnel in adopting the target optimization measures to perform cluster optimization on the cluster to be detected, which helps to achieve refined management and dynamic adjustment of resources, effectively reduce operating costs, and improve service quality and response speed.

[0142] It should be noted that the relevant information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for display, data for analysis, etc.) involved in this disclosure are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data comply with relevant laws, regulations and standards in the relevant regions.

[0143] Example 4

[0144] Figure 4 The schematic diagram of the structure of an abnormal cluster detection device provided by the fourth embodiment of the present invention is applicable to the case where the abnormal cluster is detected by using the CPU operation information and memory usage information of the target working node within the target time period, such as Figure 4 As shown, the device includes:

[0145] An information acquisition module 41 is configured to determine a target working node deployed in the cluster to be detected, and obtain CPU operation information and memory usage information of the target working node within a target time period; wherein the CPU operation information includes CPU utilization information and / or CPU operation frequency information;

[0146] a parameter determination module 42, configured to determine a CPU operation fluctuation coefficient of the target working node within the target time period according to the CPU operation information, and to determine a memory waste rate of the target working node within the target time period according to the memory usage information;

[0147] The abnormal cluster detection module 43 is configured to detect abnormal clusters of the cluster to be detected based on the CPU operation fluctuation coefficient and the memory waste rate.

[0148] Optionally, the CPU utilization information includes the overall CPU utilization of the target working node at each sampling time point, and the CPU operating frequency information includes the overall CPU operating frequency of the target working node at each sampling time point;

[0149] The parameter determination module 42 is specifically configured to:

[0150] Determine, based on the overall CPU utilization of each CPU, a standard deviation of the overall CPU utilization and a mean overall CPU utilization of the target work node within the target time period, and determine, based on the standard deviation of the overall CPU utilization and the mean overall CPU utilization, a first CPU utilization fluctuation coefficient of the target work node within the target time period;

[0151] Determine, based on the overall operating frequencies of the CPUs, a standard deviation of the overall operating frequencies of the CPUs of the target working node within the target time period and a mean of the overall operating frequencies of the CPUs, and determine, based on the standard deviation of the overall operating frequencies of the CPUs and the mean of the overall operating frequencies of the CPUs, a first CPU operating frequency fluctuation coefficient of the target working node within the target time period;

[0152] The CPU operation fluctuation coefficient is determined according to the first CPU utilization fluctuation coefficient and the first CPU operation frequency fluctuation coefficient.

[0153] Optionally, the parameter determination module 42 is further configured to:

[0154] Determining a first weight corresponding to the first CPU operating frequency fluctuation coefficient based on the average CPU overall utilization, and determining a second weight corresponding to the first CPU utilization fluctuation coefficient based on the first weight; wherein the sum of the first weight and the second weight is one;

[0155] A weighted sum is performed on the first CPU operating frequency fluctuation coefficient and the first CPU utilization fluctuation coefficient according to the first weight and the second weight, and the CPU operating fluctuation coefficient is determined according to the weighted summation result.

[0156] Optionally, the CPU utilization information includes the CPU single-core utilization of each CPU core of the target work node at each sampling time point, and the CPU operating frequency information includes the CPU single-core operating frequency of each CPU core of the target work node at each sampling time point;

[0157] The parameter determination module 42 is further configured to:

[0158] Determining, based on the CPU single-core utilization rates of each CPU core, a standard deviation of the CPU single-core utilization rates and a mean value of the CPU single-core utilization rates of each CPU core within the target time period, and determining, based on the standard deviation of the CPU single-core utilization rates and the mean value of the CPU single-core utilization rates, a single-core utilization fluctuation coefficient of each CPU core within the target time period;

[0159] determining, based on the CPU single-core operating frequencies of the CPU cores, a standard deviation of the CPU single-core operating frequencies and a mean of the CPU single-core operating frequencies of the CPU cores within the target time period, and determining, based on the CPU single-core operating frequency standard deviation and the CPU single-core operating frequency mean, a single-core operating frequency fluctuation coefficient of the CPU cores within the target time period;

[0160] The CPU operation fluctuation coefficient is determined according to each of the single-core utilization fluctuation coefficients and each of the single-core operation frequency fluctuation coefficients.

[0161] Optionally, the parameter determination module 42 is further configured to:

[0162] Performing a weighted summation based on the single-core utilization fluctuation coefficients and the first core weights of the CPU cores, and determining the second CPU utilization fluctuation coefficient of the target work node within the target time period based on the weighted summation result; wherein the first core weight is determined based on the average of the CPU single-core utilizations of the CPU cores;

[0163] Performing a weighted summation based on the single-core operating frequency fluctuation coefficients and the second core weights of the CPU cores, and determining the second CPU operating frequency fluctuation coefficient of the target working node within the target time period based on the weighted summation result; wherein the second core weight is determined based on the core computing power of each CPU core;

[0164] The CPU operation fluctuation coefficient is determined according to the second CPU utilization fluctuation coefficient and the second CPU operation frequency fluctuation coefficient.

[0165] Optionally, the memory usage information includes a first amount of memory of the target working node that is in an unallocated state during the target time period, and a second amount of memory that is in an allocated state but not used;

[0166] The parameter determination module 42 is further configured to:

[0167] Determining a memory waste amount of the target working node within the target time period according to a sum of the first memory amount and the second memory amount;

[0168] The memory waste rate is determined according to the memory waste amount and the total memory amount of the target working node.

[0169] Optionally, the abnormal cluster detection module 43 is specifically configured to:

[0170] Determining a cluster efficiency score corresponding to the cluster to be detected based on the CPU operation fluctuation coefficient and the memory waste rate;

[0171] The cluster efficiency score is compared with an efficiency score threshold, and when the cluster efficiency score is lower than the efficiency score threshold, the cluster to be detected is determined to be an abnormal cluster.

[0172] Optionally, the device further includes a manual detection prompt module, specifically configured to:

[0173] Using a machine learning model, predicting whether the cluster to be detected is an abnormal cluster based on the CPU operation information and the memory usage information;

[0174] When the machine learning model predicts that the cluster to be detected is not an abnormal cluster, manual detection prompt information is generated; wherein the manual detection prompt information is used to prompt technical personnel to use manual methods to detect abnormal clusters in the cluster to be detected.

[0175] Optionally, the device further includes an optimization measure suggestion module, specifically configured to:

[0176] Determine the target abnormality cause and target abnormality impact degree corresponding to the cluster to be detected;

[0177] Determining a target optimization measure corresponding to the cluster to be detected from the candidate optimization measures according to a mapping relationship between the candidate optimization measures and the candidate abnormality causes and the candidate abnormality impact degrees;

[0178] Optimization measure suggestion information is generated according to the target optimization measure; wherein, the optimization measure suggestion information is used to assist technical personnel in performing cluster optimization on the cluster to be detected by adopting the target optimization measure.

[0179] The abnormal cluster detection device provided in the embodiment of the present invention can execute the abnormal cluster detection method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method.

[0180] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0181] Example 5

[0182] Figure 5 A schematic diagram of the structure of an electronic device 50 that can be used to implement an embodiment of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or claimed herein.

[0183] like Figure 5 As shown, the electronic device 50 includes at least one processor 51 and a memory, such as a read-only memory (ROM) 52, a random access memory (RAM) 53, etc., which is communicatively connected to the at least one processor 51. The memory stores a computer program that can be executed by the at least one processor. The processor 51 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 52 or the computer program loaded from the storage unit 58 into the random access memory (RAM) 53. Various programs and data required for the operation of the electronic device 50 can also be stored in the RAM 53. The processor 51, ROM 52, and RAM 53 are connected to each other via a bus 54. An input / output (I / O) interface 55 is also connected to the bus 54.

[0184] Multiple components in the electronic device 50 are connected to the I / O interface 55, including an input unit 56, such as a keyboard, a mouse, etc.; an output unit 57, such as various types of displays, speakers, etc.; a storage unit 58, such as a magnetic disk, an optical disk, etc.; and a communication unit 59, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 59 allows the electronic device 50 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0185] The processor 51 can be any general-purpose and / or specialized processing component with processing and computing capabilities. Some examples of the processor 51 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The processor 51 executes the various methods and processes described above, such as the abnormal cluster detection method.

[0186] In some embodiments, the abnormal cluster detection method can be implemented as a computer program, which is tangibly contained in a computer-readable storage medium, such as a storage unit 58. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 50 via the ROM 52 and / or the communication unit 59. When the computer program is loaded into the RAM 53 and executed by the processor 51, one or more steps of the abnormal cluster detection method described above can be performed. Alternatively, in other embodiments, the processor 51 can be configured to perform the abnormal cluster detection method in any other appropriate manner (for example, by means of firmware).

[0187] Various embodiments of the systems and techniques described above can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0188] Computer programs for implementing the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the computer program is executed by the processor, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The computer program may be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0189] In the context of the present invention, computer-readable storage media can be tangible media that can contain or store a computer program for use with an instruction execution system, device or equipment or used in combination with an instruction execution system, device or equipment. Computer-readable storage media can include but are not limited to electronic, magnetic, optical, electromagnetic, infrared or semiconductor systems, devices or equipment, or any suitable combination of the foregoing. Alternatively, computer-readable storage media can be machine-readable signal media. More specific examples of machine-readable storage media can include electrical connections based on one or more lines, portable computer disks, hard disks, random access memories (RAM), read-only memories (ROM), erasable programmable read-only memories (EPROM or flash memory), optical fibers, portable compact disk read-only memories (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0190] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0191] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.

[0192] A computing system may include clients and servers. The clients and servers are typically remote from each other and typically interact via a communication network. This client-server relationship arises through computer programs running on the respective computers, creating a client-server relationship. The server may be a cloud server, also known as a cloud computing server or cloud host. This server is a hosting product within the cloud computing service ecosystem that addresses the management difficulties and limited scalability of traditional physical hosting and VPS services.

[0193] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in the present invention can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved. This is not limited herein.

[0194] The above specific embodiments do not limit the scope of protection of the present invention. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention are intended to be included within the scope of protection of the present invention.

Claims

1. A method for detecting abnormal clusters, characterized in that: The method comprises: Determine a target working node deployed in the cluster to be tested, and obtain CPU operation information and memory usage information of the target working node within a target time period; wherein the CPU operation information includes CPU utilization information and / or CPU operation frequency information; Determine a CPU operation fluctuation coefficient of the target working node within the target time period according to the CPU operation information, and determine a memory waste rate of the target working node within the target time period according to the memory usage information; The cluster to be detected is detected as an abnormal cluster according to the CPU operation fluctuation coefficient and the memory waste rate.

2. The method according to claim 1, characterized in that The CPU utilization information includes the overall CPU utilization of the target working node at each sampling time point, and the CPU operating frequency information includes the overall CPU operating frequency of the target working node at each sampling time point; The determining, according to the CPU operation information, a CPU operation fluctuation coefficient of the target working node within the target time period includes: Determine, based on the overall CPU utilization of each CPU, a standard deviation of the overall CPU utilization and a mean overall CPU utilization of the target work node within the target time period, and determine, based on the standard deviation of the overall CPU utilization and the mean overall CPU utilization, a first CPU utilization fluctuation coefficient of the target work node within the target time period; Determine, based on the overall operating frequencies of the CPUs, a standard deviation of the overall operating frequencies of the CPUs of the target working node within the target time period and a mean of the overall operating frequencies of the CPUs, and determine, based on the standard deviation of the overall operating frequencies of the CPUs and the mean of the overall operating frequencies of the CPUs, a first CPU operating frequency fluctuation coefficient of the target working node within the target time period; The CPU operation fluctuation coefficient is determined according to the first CPU utilization fluctuation coefficient and the first CPU operation frequency fluctuation coefficient.

3. The method according to claim 2, characterized in that The determining the CPU operation fluctuation coefficient according to the first CPU utilization fluctuation coefficient and the first CPU operation frequency fluctuation coefficient includes: Determining a first weight corresponding to the first CPU operating frequency fluctuation coefficient based on the average CPU overall utilization, and determining a second weight corresponding to the first CPU utilization fluctuation coefficient based on the first weight; wherein the sum of the first weight and the second weight is one; A weighted sum is performed on the first CPU operating frequency fluctuation coefficient and the first CPU utilization fluctuation coefficient according to the first weight and the second weight, and the CPU operating fluctuation coefficient is determined according to the weighted summation result.

4. The method according to claim 1, wherein The CPU utilization information includes the CPU single-core utilization of each CPU core of the target work node at each sampling time point, and the CPU operating frequency information includes the CPU single-core operating frequency of each CPU core of the target work node at each sampling time point; The determining, according to the CPU operation information, a CPU operation fluctuation coefficient of the target working node within the target time period includes: Determining, based on the CPU single-core utilization rates of each CPU core, a standard deviation of the CPU single-core utilization rates and a mean value of the CPU single-core utilization rates of each CPU core within the target time period, and determining, based on the standard deviation of the CPU single-core utilization rates and the mean value of the CPU single-core utilization rates, a single-core utilization fluctuation coefficient of each CPU core within the target time period; determining, based on the CPU single-core operating frequencies of the CPU cores, a standard deviation of the CPU single-core operating frequencies and a mean of the CPU single-core operating frequencies of the CPU cores within the target time period, and determining, based on the CPU single-core operating frequency standard deviation and the CPU single-core operating frequency mean, a single-core operating frequency fluctuation coefficient of the CPU cores within the target time period; The CPU operation fluctuation coefficient is determined according to each of the single-core utilization fluctuation coefficients and each of the single-core operation frequency fluctuation coefficients.

5. The method according to claim 4, characterized in that The determining the CPU operation fluctuation coefficient according to each of the single-core utilization fluctuation coefficients and each of the single-core operation frequency fluctuation coefficients includes: Performing a weighted summation based on the single-core utilization fluctuation coefficients and the first core weights of the CPU cores, and determining the second CPU utilization fluctuation coefficient of the target work node within the target time period based on the weighted summation result; wherein the first core weight is determined based on the average of the CPU single-core utilizations of the CPU cores; Performing a weighted summation based on the single-core operating frequency fluctuation coefficients and the second core weights of the CPU cores, and determining the second CPU operating frequency fluctuation coefficient of the target working node within the target time period based on the weighted summation result; wherein the second core weight is determined based on the core computing power of each CPU core; The CPU operation fluctuation coefficient is determined according to the second CPU utilization fluctuation coefficient and the second CPU operation frequency fluctuation coefficient.

6. The method according to claim 1, characterized in that The memory usage information includes a first amount of memory of the target working node that is in an unallocated state during the target time period, and a second amount of memory that is in an allocated state but not used; The determining, according to the memory usage information, the memory waste rate of the target working node within the target time period includes: Determining a memory waste amount of the target working node within the target time period according to a sum of the first memory amount and the second memory amount; The memory waste rate is determined according to the memory waste amount and the total memory amount of the target working node.

7. The method according to claim 1, characterized in that The detecting of abnormal clusters on the cluster to be detected according to the CPU operation fluctuation coefficient and the memory waste rate includes: Determining a cluster efficiency score corresponding to the cluster to be detected based on the CPU operation fluctuation coefficient and the memory waste rate; The cluster efficiency score is compared with an efficiency score threshold, and when the cluster efficiency score is lower than the efficiency score threshold, the cluster to be detected is determined to be an abnormal cluster.

8. The method according to claim 7, after determining that the cluster to be detected is an abnormal cluster, further comprising: Using a machine learning model, predicting whether the cluster to be detected is an abnormal cluster based on the CPU operation information and the memory usage information; When the machine learning model predicts that the cluster to be detected is not an abnormal cluster, manual detection prompt information is generated; wherein the manual detection prompt information is used to prompt technical personnel to use manual methods to detect abnormal clusters in the cluster to be detected.

9. The method according to claim 7, after determining that the cluster to be detected is an abnormal cluster, further comprising: Determine the target abnormality cause and target abnormality impact degree corresponding to the cluster to be detected; Determining a target optimization measure corresponding to the cluster to be detected from the candidate optimization measures according to a mapping relationship between the candidate optimization measures and the candidate abnormality causes and the candidate abnormality impact degrees; Optimization measure suggestion information is generated according to the target optimization measure; wherein, the optimization measure suggestion information is used to assist technical personnel in performing cluster optimization on the cluster to be detected by adopting the target optimization measure.

10. A device for detecting abnormal clusters, characterized in that: The device comprises: An information acquisition module is used to determine a target working node deployed in the cluster to be detected, and obtain CPU operation information and memory usage information of the target working node within a target time period; wherein the CPU operation information includes CPU utilization information and / or CPU operation frequency information; a parameter determination module, configured to determine a CPU operation fluctuation coefficient of the target working node within the target time period according to the CPU operation information, and to determine a memory waste rate of the target working node within the target time period according to the memory usage information; The abnormal cluster detection module is used to detect abnormal clusters of the cluster to be detected according to the CPU operation fluctuation coefficient and the memory waste rate.

11. An electronic device, characterized in that: The electronic device comprises: at least one processor; and a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executable by the at least one processor. The computer program is executed by the at least one processor to enable the at least one processor to perform the abnormal cluster detection method according to any one of claims 1 to 9.

12. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to execute the abnormal cluster detection method according to any one of claims 1 to 9. 13 . A computer program product, comprising a computer program, wherein when executed by a processor, the computer program implements the abnormal cluster detection method according to claim 1 .