Detection method and device

By acquiring and analyzing historical data from systems such as Ceph storage clusters, predictive data is generated to detect their request response capabilities. This solves the problems of insufficient comprehensive coverage and poor real-time performance in existing technologies, enabling timely detection and improved accuracy of system performance degradation.

CN120803826APending Publication Date: 2025-10-17LENOVO (BEIJING) LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510900926.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-30
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

Existing technologies cannot fully cover various degradation situations that may occur in Ceph storage clusters, and the real-time performance of detection results is poor, making it impossible to discover and handle potential problems in a timely manner.

Method used

By acquiring historical data of the system under test, predictive data is generated to characterize the system's request response capability, and compared with preset data to generate detection results, thereby improving the real-time performance and accuracy of the detection.

Benefits of technology

It enables real-time and accurate detection of systems such as Ceph storage clusters, and can promptly detect performance degradation, thus improving the efficiency and accuracy of detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120803826A_ABST
    Figure CN120803826A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a detection method and device. The detection method comprises the following steps: acquiring historical data of a to-be-detected system; the historical data represents the current request response capability of the to-be-detected system; generating prediction data for the to-be-detected system based on the historical data; the prediction data represents the predicted request response capability of the to-be-detected system; and generating a detection result for the to-be-detected system based on the prediction data and preset data of the to-be-detected system.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to, but is not limited to, the technical field of information processing, and in particular to a detection method and device. BACKGROUND

[0002] At present, in the detection scheme for the degradation phenomenon of a storage system, specific degradation phenomena such as hardware failure or disk damage are detected, and various degradation conditions that may occur in the storage system cannot be comprehensively covered. Or, whether there is an abnormality at the current time is determined by running data for a certain period of time, and the performance of the system is not detected. Therefore, in the detection scheme for whether the system exists degradation, there is also the problem of low accuracy. SUMMARY

[0003] Therefore, the embodiments of the present application at least provide a detection method and device.

[0004] The technical scheme of the embodiments of the present application is as follows:

[0005] In a first aspect, the embodiments of the present application provide a detection method, comprising: obtaining historical data of a to-be-detected system; the historical data representing a current request response capability of the to-be-detected system; generating prediction data for the to-be-detected system based on the historical data; the prediction data representing a predicted request response capability of the to-be-detected system; and generating a detection result for the to-be-detected system based on the prediction data and preset data of the to-be-detected system.

[0006] In a second aspect, the embodiments of the present application provide a detection device, comprising: an obtaining module, configured to obtain historical data of a to-be-detected system; the historical data representing a current request response capability of the to-be-detected system; a first generating module, configured to generate prediction data for the to-be-detected system based on the historical data; the prediction data representing a predicted request response capability of the to-be-detected system; and a second generating module, configured to generate a detection result for the to-be-detected system based on the prediction data and preset data of the to-be-detected system.

[0007] In a third aspect, the embodiments of the present application provide a computer device, comprising a memory and a processor, wherein the memory stores a computer program capable of running on the processor, and the processor implements part or all of the steps in the above method when executing the program.

[0008] In a fourth aspect, the embodiments of the present application provide a computer readable storage medium, which stores a computer program, and the computer program is executed by a processor to implement part or all of the steps in the above method.

[0009] In a fifth aspect, an embodiment of the present application provides a computer program product, comprising a computer program or instructions, which, when executed by a processor, implement some or all of the steps of the above method.

[0010] It should be understood that the above general description and the following detailed description are exemplary and explanatory only and are not restrictive of the technical solutions of the present application. BRIEF DESCRIPTION OF DRAWINGS

[0011] The accompanying drawings, which are incorporated herein and form part of the specification, illustrate embodiments consistent with the present application and, together with the description, further serve to explain the principles of the technical solutions of the present application.

[0012] Figure 1 An implementation flowchart of a detection method provided by an embodiment of the present application is shown in FIG. 1.

[0013] Figure 2 An implementation flowchart of a detection method provided by an embodiment of the present application is shown in FIG. 1.

[0014] Figure 3 An implementation flowchart of a detection method provided by an embodiment of the present application is shown in FIG. 1.

[0015] Figure 4 An implementation flowchart of a detection method provided by an embodiment of the present application is shown in FIG. 1.

[0016] Figure 5 An implementation flowchart of a detection method provided by an embodiment of the present application is shown in FIG. 1.

[0017] Figure 6 An implementation flowchart of a detection method provided by an embodiment of the present application is shown in FIG. 1.

[0018] Figure 7a An implementation flowchart of a detection method provided by an embodiment of the present application is shown in FIG. 1.

[0019] Figure 7b An implementation flowchart of a detection method provided by an embodiment of the present application is shown in FIG. 1.

[0020] Figure 8 An implementation flowchart of a detection method provided by an embodiment of the present application is shown in FIG. 1.

[0021] Figure 9 An implementation flowchart of a detection method provided by an embodiment of the present application is shown in FIG. 1.

[0022] Figure 10 An implementation flowchart of a detection method provided by an embodiment of the present application is shown in FIG. 1.

[0023] Figure 11A schematic diagram of data statistics provided for the embodiments of the present application is shown in FIG. 1.

[0024] Figure 12 A schematic diagram of data trend changes provided for the embodiments of the present application is shown in FIG. 2.

[0025] Figure 13 A schematic diagram of data update statistics provided for the embodiments of the present application is shown in FIG. 3.

[0026] Figure 14 A schematic diagram of the composition structure of a detection device provided for the embodiments of the present application is shown in FIG. 4.

[0027] Figure 15 A schematic diagram of the hardware entity of a computer device provided for the embodiments of the present application is shown in FIG. 5. DETAILED DESCRIPTION

[0028] In order to make the purposes, technical solutions and advantages of the present application clearer, the technical solutions of the present application are further described in detail below in combination with the drawings and embodiments, and the described embodiments should not be regarded as limiting the present application, and all other embodiments obtained by those skilled in the art without making creative efforts fall within the scope of protection of the present application.

[0029] In the following description, "some embodiments" are described, which describe a subset of all possible embodiments, but it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments, and can be combined with each other without conflict. The term "first / second / third" referred to only distinguishes similar objects, and does not represent a specific order of the objects. It can be understood that "first / second / third" can be interchanged in a specific order or sequence as allowed, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.

[0030] At present, the existing technology can only detect specific degradation phenomena such as hardware failure or disk damage, and cannot comprehensively cover various degradation conditions that may occur in a Ceph storage cluster. Or some technologies rely on periodic data collection and analysis, resulting in poor real-time detection results and being unable to timely discover and handle potential degradation problems.

[0031] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the present application belongs. The terms used herein are only for the purpose of describing the present application and are not intended to limit the present application.

[0032] The embodiment of the present application provides a detection method, which can be executed by a processor of a computer device. The computer device can be a server, a notebook computer, a tablet computer, a desktop computer, a smart television, a set-top box, a mobile device (such as a mobile phone, a portable video player, a personal digital assistant, a dedicated message device, a portable game device) and the like.

[0033] Figure 1 An implementation flowchart of the detection method provided by the embodiment of the present application is shown in the figure, which can be executed by a processor of a computer device. As shown in the figure, the method comprises the following steps S101 to S103, which will be described in combination with the steps. Figure 1 Figure 1

[0034] Step S101, acquiring historical data of a to-be-detected system; the historical data represents a current request response capability of the to-be-detected system.

[0035] In some embodiments, the to-be-detected system can be a distributed storage system, such as a Ceph storage system (ceph), a gluster file system (glusterfs), an open elastic block storage (open ebs) and the like. The to-be-detected system can also be a real-time data processing and streaming system, such as a Kafka stream processing (kafka streams), an Apache Flink stream computing (apache flink stream computing) and the like.

[0036] In some embodiments, the historical data refers to request response capability related data recorded by the to-be-detected system in a certain period of time in the past, including but not limited to request response rate, request response quantity and the like. These data are used to build a model, identify a trend and provide a basis for predicting future performance. The to-be-detected system can include latency (read latency, write latency) of the to-be-detected system, throughput (read throughput, write throughput), storage space, remaining storage space and the like.

[0037] In some embodiments, the request response capability represents the capability of the to-be-detected system (such as a Ceph storage cluster) for processing requests and returning responses in a unit of time, which is usually represented by a request response rate (the number of requests in a unit of time) and a request response quantity (the total number of completed requests in a period of time). The index reflects the load capacity and running health status of the system.

[0038] ​​In some embodiments, the prediction data is an estimated value of the system request response capability at a future time point or time period, which is generated based on historical data. The prediction data can be a specific numerical value (e.g., a predicted request response quantity) or a trend sequence (e.g., a predicted change trend of the request response rate).

[0039] In some embodiments, the preset data is a performance threshold value (e.g., a preset request response quantity or a preset request response rate) set by a developer, an operator or a system according to business requirements. This data serves as a standard for determining whether the system has performance degradation.

[0040] In some embodiments, the detection result of the to-be-detected system represents whether the to-be-detected system has performance degradation.

[0041] In some embodiments, the historical data of the to-be-detected system can be collected based on time series. For example, during the operation of the system, the current average latency (the average of read latency and write latency, or the average read latency and the average write latency of the system within 1 minute), read-write throughput, total storage space, used space, etc. of the system are collected every minute.

[0042] In some embodiments, after obtaining the historical data of the to-be-detected system, the to-be-detected system needs to be preprocessed, including: setting the missing values and extreme values in the historical data as null. The missing values are directly set as null, and the data greater than a threshold value in the extreme values is set as an extreme value, and then set as null; the null data is filled; if the number n of continuous null data points is less than m (m is 15 for example, indicating that the number of null data points is not more than 15), the forward filling method is used to fill the null data points with the value of the previous non-missing value or non-extreme value data point; if the number n of continuous null data points is greater than m, the trend of the previous h data points is determined by using the least square method, and the n data points are added to the trend to fill the h null data points. The historical data is standardized, please refer to formula (1).

[0043]

[0044] Wherein, μ is the mean of the historical data, and σ is the standard deviation.

[0045] In some embodiments, the historical data is processed by Piecewise Aggregate Approximation (PAA): a given cluster average latency sequence Q is divided by week to obtain a subsequence Q = q1, q2,.. qk; the median of each subsequence q is calculated, and the median of the subsequence represents each value in the subsequence; the sequence obtained in step 2 is reduced, and only one value is extracted from each sequence to represent the sequence. Wherein, if the subsequence of sequence Q includes {{1, 2, 3}, {4, 5, 6}}, the median of each subsequence represents the subsequence, then sequence Q can be {2, 5}.

[0046] In step S102, prediction data for the to-be-detected system is generated based on the historical data; wherein, the prediction data represents the predicted request response capability of the to-be-detected system.

[0047] In some embodiments, the to-be-detected system includes a cold start phase and a hot start phase. Wherein, if the to-be-detected system is in the initial stage or has insufficient historical data, it is represented as a cold start phase; if the to-be-detected system has run for a period of time and has sufficient data for cluster degradation analysis, it is a hot start phase.

[0048] In some embodiments, for the to-be-detected system in the cold start phase, the prediction data is generated by the historical data of the to-be-detected system in the cold start phase.

[0049] In some embodiments, the historical data can be predicted by a neural network model to obtain the prediction data. Wherein, if the historical data is single type data, for example, only throughput, the neural network model can use a time series model, for example, an Autoregressive Integrated Moving Average Model (ARIMA); if the historical data is multi-type, for example, including throughput and latency, the neural network model can use a self-attention mechanism neural network model, for example, a Transformer Model (Transformer).

[0050] For example, a time series model is used to predict the historical throughput of the to-be-detected system to obtain the predicted throughput of the to-be-detected system in the next month.

[0051] In some embodiments, for the to-be-detected system in the cold start phase, the to-be-detected system is predicted based on the historical latency of the to-be-detected system to obtain the predicted latency of the to-be-detected system. Wherein, the trend of the predicted latency change of the to-be-detected system can also be predicted based on the historical latency.

[0052] In step S103, a detection result of the system to be detected is generated based on the prediction data and preset data of the system to be detected.

[0053] In some embodiments, the preset data is a performance threshold set by a development and operation personnel or a system according to business requirements, for example, a preset request response quantity or a preset request response rate. The data is used as a standard basis for judging whether the system has performance degradation.

[0054] In some embodiments, for the system to be detected in a cold start phase, if the prediction data is greater than the preset data, it indicates that the system to be detected in the current phase has degradation. For example, the preset data is a throughput threshold of the system to be detected, the prediction data is a predicted throughput of the system to be detected, and if the predicted throughput is greater than the preset throughput, a detection result indicating that the request response capability of the system to be detected has degradation is generated.

[0055] In some embodiments, for the system to be detected in a hot start phase, the preset data is a latency threshold of the system to be detected, if the real-time latency of the system to be detected at the current time is greater than the latency threshold, it indicates that the system to be detected may have delay at the current time, then a historical latency trend and a predicted latency trend of the system to be detected are generated based on historical latency data in the historical data, if the growth trend of the current time in the historical latency trend is greater than the growth trend of the current time in the predicted latency trend, a detection result indicating that the request response capability of the system to be detected has degradation is generated. If the real-time latency of the system to be detected at the current time is less than or equal to the latency threshold, a predicted latency of the system to be detected in a future period of time is predicted based on the predicted latency trend, and if the predicted latency is greater than the latency threshold, a detection result indicating that the system to be detected may have degradation in the future is generated.

[0056] In the embodiments of the present application, the future request response capability of the system to be detected is predicted by the historical data representing the request response capability of the system to be detected, and the prediction data of the system to be detected is obtained; by comparing the prediction data and the preset data, whether the system to be detected has degradation can be determined. Compared with the prior art of judging whether the system to be detected has degradation by detecting the hardware, the system to be detected can be detected in real time whether it has degradation, so that the real-time performance and accuracy of the detection are improved.

[0057] Figure 2 An implementation flowchart of a detection method provided in the embodiments of the present application is shown in the figure. The method can be executed by a processor of a computer device. Based on Figure 1 , the historical data includes a request response rate and a request response quantity; Figure 1 In step S102 in the figure, step S201 or step S202 can be updated, which will be described in combination with Figure 2 shown in the figure.

[0058] Step S201, in a case where the to-be-detected system is in a first state, generating the prediction data based on the request response rate and the request response quantity.

[0059] wherein the historical data of the to-be-detected system in the first state does not meet preset requirements.

[0060] In some embodiments, the request response rate of the to-be-detected system can be a latency of the to-be-detected system, including a read latency and a write latency.

[0061] In some embodiments, the request response quantity of the to-be-detected system can be a throughput of the to-be-detected system, including a read throughput and a write throughput.

[0062] In some embodiments, the first state represents that the to-be-detected system is in a cold start phase; wherein the cold start phase represents that the to-be-detected system is a newly built system, or the to-be-detected system is initially running and lacks sufficient historical data.

[0063] In some embodiments, in a case where the to-be-detected system is in a normal state, the request response rate of the to-be-detected system and the request response quantity of the to-be-detected system present a negative correlation, that is, in a case where the request response quantity increases, the request response rate decreases, and vice versa, the request response quantity decreases, and the request response rate increases.

[0064] In some embodiments, in a case where the to-be-detected system is in a cold start phase, first, the correlation between the request response rate and the request response quantity of the to-be-detected system is obtained, if the request response rate and the request response quantity of the to-be-detected system present a positive correlation, it represents that the to-be-detected system can be degraded, then the request response quantity of the to-be-detected system in a future period of time is predicted based on the request response quantity, to obtain a predicted request response quantity, if the predicted request response quantity is greater than a preset request response quantity, a detection result representing that the to-be-detected system is degraded is generated, otherwise, a detection result representing that the to-be-detected system is normal is generated.

[0065] Step S202, in a case where the to-be-detected system is in a second state, generating the prediction data based on the request response rate.

[0066] wherein the historical data of the to-be-detected system in the second state meets preset requirements.

[0067] In some embodiments, the second state represents that the to-be-detected system is in a hot start phase; the hot start phase represents that the to-be-detected system has been running for a certain period of time, and has accumulated sufficient historical data to support analysis of the request response capability trend of the to-be-detected system based on the historical data.

[0068] In some embodiments, in the case that the to-be-detected system is in a hot start stage, it is first determined whether the to-be-detected system is likely to be degraded based on the request response rate at the current time and the preset request response rate, and if the request response rate at the current time is greater than the preset request response rate, it is indicated that the to-be-detected system is likely to be degraded. Then, based on the request response rates at the historical time and the current time, a historical trend and a predicted trend of the request response rate of the to-be-detected system are generated, and it is determined whether there is degradation based on the historical trend and the predicted trend. If the growth trend of the historical trend is greater than the growth trend of the predicted trend, a detection result indicating that the to-be-detected system is degraded is generated.

[0069] In the embodiments of the present application, it is determined whether the to-be-detected system is in a state by whether the preset requirement is met, and different historical data is used to detect the to-be-detected system based on different states. Therefore, the accuracy of detecting the request response capability of the to-be-detected system is improved.

[0070] Figure 3 An implementation flow diagram of a detection method provided in the embodiments of the present application is provided. The method can be executed by a processor of a computer device. Based on Figure 2 , the prediction data includes a predicted request response quantity at a preset time point after a current time point, Figure 2 Step S201 can be updated to step S301 and step S302, which will be described in combination with Figure 3 .

[0071] Step S301, determining a correlation coefficient between the request response rate and the request response quantity.

[0072] In some embodiments, in the case that the to-be-detected system is in a first state, that is, a cold start stage, since the historical data of the to-be-detected system in this stage is insufficient, the to-be-detected system cannot be detected by a single data, and the request response rate and the request response data in the historical data need to be combined.

[0073] In some embodiments, the request response rate and the request response quantity of the to-be-detected system are negatively correlated, that is, in the case that the request response quantity increases, the request response rate decreases, and vice versa, the request response quantity decreases, and the request response rate increases.

[0074] In some embodiments, the correlation coefficient can be a Pearson correlation coefficient (Pearson).

[0075] In some embodiments, the correlation coefficient is calculated based on the request response rate and the request response quantity of the system to be detected in a preset time period. For example, the Pearson correlation coefficient is calculated based on the delay and throughput of the distributed storage system in the past month, including: calculating the mean of the delay and the throughput respectively, calculating the sum of squares and deviations of the delay based on the mean of the delay, calculating the sum of squares and deviations of the throughput based on the mean of the throughput; calculating the covariance of the delay and the throughput based on the sum of squares and deviations of the delay and the sum of squares and deviations of the throughput, and calculating the correlation coefficient based on the sum of squares and deviations of the delay, the sum of squares and deviations of the throughput, and the covariance of the delay and the throughput. Please refer to formula (2).

[0076]

[0077] Wherein, r is the correlation coefficient, SS XX is the sum of squares and deviations of the delay, SS YY is the sum of squares and deviations of the delay, SS XY is the covariance of the delay and the throughput.

[0078] In step S302, if the correlation coefficient is greater than a preset coefficient, a predicted request response quantity at a preset time point after the current time point of the system to be detected is generated based on a target request response quantity in the request response quantity.

[0079] Wherein, the target request response quantity includes the request response quantity of multiple time points; the multiple time points at least include a target time point to the current time point; the target time point is the time point at which the request response quantity starts to increase closest to the current time point.

[0080] In some embodiments, the predicted request response quantity refers to the number of requests that the system is expected to receive and respond to in a future preset time period (such as one week, one month) after the current time point. The predicted request response quantity is used to represent the load change trend of the system to be detected in the future.

[0081] In some embodiments, the request response rate and the request response quantity of the system to be detected indicate that there is a significant linear positive correlation between them, that is, the system to be detected may be in degradation. It can be understood that, in the case that the system to be detected is in the cold start stage and may be in degradation, whether the system to be detected is in degradation is determined based on the predicted request response quantity.

[0082] In some embodiments, the preset time point after the current time point represents a future preset time point of the current time point, for example, a time point one month after the current time point.

[0083] In some embodiments, the predicted request response quantity represents the request response quantity of the predicted preset time point.

[0084] In some embodiments, the target time point is a historical time point; it can be understood that the target time point is the time point closest to the current time point, that is, in the historical request response quantity, the starting point of the beginning of the increase closest to the current time point is obtained.

[0085] Exemplarily, the request response quantity in the historical data includes the request response quantity from time 1 to time 10, wherein time 10 is the current time point, time 1 to time 9 before time 10 is the historical time point, and in time 1 to time 9, if the request response quantity corresponding to time 6 to the request response quantity corresponding to time 10 begins to increase, time 6 is determined as the target time, and the plurality of time points are time 6 to time 10. It can be understood that the request response quantity corresponding to time 6 to time 10 gradually increases, the request response quantity of time 7 is greater than the request response quantity of time 6, the request response quantity of time 8 is greater than the request response quantity of time 7, and so on. Wherein, the request response quantity of time 6 is less than or equal to the request response quantity of time 5.

[0086] In some embodiments, based on the request response quantity corresponding to the plurality of time points, a linear regression model is used to predict the request response quantity of the future preset time point, and a predicted request response quantity is obtained.

[0087] In the embodiments of the present application, when the request response rate of the to-be-detected system and the correlation coefficient of the request response data represent the possible storage degradation of the to-be-detected system, the request response rate of the to-be-detected system at the future preset time point is predicted based on the request response quantity of the plurality of time points. Thus, whether the current to-be-detected system has degradation can be determined based on the predicted request response quantity, and the detection accuracy is improved.

[0088] Figure 4 An implementation flowchart of a detection method provided in the embodiments of the present application is shown, which can be executed by the processor of the computer device. Based on Figure 1 , the preset data includes a preset request response quantity, Figure 1 Step S103 in the method can be updated to step S401 or step S402, which will be described in combination with Figure 4 shown steps.

[0089] Step S401, in the case that the predicted request response quantity is greater than the preset request response quantity, a detection result representing that the request response capability of the to-be-detected system appears degradation is generated.

[0090] In some embodiments, the preset request response quantity is a benchmark value set according to historical performance indicators and business requirements, which is used to measure whether the request processing capability of the system is in a normal state. The benchmark value is usually configured by an operation and maintenance personnel or a system administrator based on experience, historical load, service quality agreement (SLA), and other factors. For example, the preset request response quantity in a distributed storage cluster can be set to 500 requests per second to ensure that the system can meet the demand during the peak period of business.

[0091] In some embodiments, the preset request response quantity represents the maximum request response quantity that can be processed when the to-be-detected system is in a normal state. In the case where the predicted request response quantity is greater than the preset request response quantity, it is indicated that the predicted request response quantity in the future preset time exceeds the maximum request response quantity that can be processed by the to-be-detected system, indicating that the system has degradation, and thus a detection result indicating that the request response capability of the to-be-detected system has degradation is generated.

[0092] In step S402, in the case where the predicted request response quantity is less than or equal to the preset request response quantity, a detection result indicating that the request response capability of the to-be-detected system is normal is generated.

[0093] In some embodiments, in the case where the predicted request response quantity is greater than the preset request response quantity, it is indicated that the maximum request response quantity that can be currently processed by the to-be-detected system can meet the predicted request response quantity in the future preset time, and thus a detection result indicating that the request response capability of the to-be-detected system is normal is generated.

[0094] For example, the maximum throughput that can be currently processed by the to-be-detected system is 100, and the predicted throughput that needs to be processed by the to-be-detected system in the future one month is 50, and thus a detection result indicating that the request response capability of the to-be-detected system is normal is generated.

[0095] In the embodiments of the present application, whether the predicted request response quantity exceeds the maximum preset request response quantity that can be currently processed by the to-be-detected system is determined to determine whether the to-be-detected system has degradation, and thus a detection result indicating whether the request response capability of the to-be-detected system has degradation is generated. The problem root of the to-be-detected system can be quickly located, and the detection efficiency and accuracy are improved.

[0096] Figure 5 An implementation flow diagram of a detection method provided in the embodiments of the present application is shown. The method can be executed by a processor of a computer device. Based on Figure 2 , the prediction data includes a prediction trend sequence; Step S202 in Figure 2 may be updated to include steps S501 to S503, which will be described in combination with the steps shown in Figure 5 .

[0097] Step S501, generating a detection sequence for the to-be-detected system based on the request response rate.

[0098] The detection sequence includes the request response rate corresponding to each time point in the historical time.

[0099] In some embodiments, when the to-be-detected system is in a warm-up phase, the request response capability of the to-be-detected system is detected based on the historical request response rate of the to-be-detected system.

[0100] In some embodiments, the detection sequence is a structured time sequence formed by sorting and aggregating the historical request response rate data of the system. Each time point in the sequence corresponds to a specific request response rate value, reflecting the performance of the system at different time periods.

[0101] In some embodiments, when the to-be-detected system is in a warm-up phase, it is first determined whether the to-be-detected system is a first detection. If it is a first detection, a detection sequence is generated based on the request response rate of the to-be-detected system. If it is not a first detection, a historical detection sequence is obtained, and a detection sequence generated based on the latest request response rate is obtained. The historical detection sequence and the latest detection sequence are spliced to obtain the detection sequence {s1, s2, …, s n} of the to-be-detected system at the current time.

[0102] For example, in a Ceph cluster, the detection sequence can be generated based on the average delay data collected by Prometheus every minute, and a more smooth and representative trend sequence can be obtained after PAA segmentation and aggregation processing.

[0103] Step S502, generating an update sequence for the to-be-detected system based on the detection sequence.

[0104] The update sequence includes at least one target time point representing an update state. In the detection sequence, the request response rate corresponding to the target time point is better than the request response rate corresponding to the previous time point of the target time point.

[0105] In some embodiments, the update sequence is a marked sequence constructed according to the time points of performance improvement in the detection sequence. Each element in the sequence represents whether a performance improvement (i.e., update) occurs at a time point. When the request response rate at a time point is better than that at the previous time point, the time point is marked as an update state.

[0106] In some embodiments, a time point of updating the system to be detected is first acquired; wherein the time point of updating is a time point at which a technician expands the storage space of the system to be detected or migrates the service of the system to be detected. It can be understood that the time point of updating corresponds to a delay smaller than that of a time point before updating.

[0107] In some embodiments, when the value of the i time point in the detection sequence is smaller than that of the i-1 time point, it is considered that the i time point is an updating state, so as to determine the updating sequence U={u1, u2, …, u m} from the detection sequence based on the time point of updating. Please refer to formula (3).

[0108]

[0109] wherein, u i =1 indicates that the i time point in the detection sequence is an updating state. Thus, the updating sequence U can be determined from the detection sequence {s1, s2, …, s i n} based on the updating state of u

[0110] In step S503, the predicted trend sequence is generated based on the detection sequence and the updating sequence.

[0111] In some embodiments, the detection sequence is split into multiple subsequences based on the updating sequence, the trend of each subsequence is calculated, and finally the trends of each subsequence are merged to obtain a historical trend sequence; the predicted trend sequence is obtained based on the updating sequence and the historical trend sequence. This includes: acquiring a large update in the updating sequence, that is, the amount of continuous update and the time point of updating, obtaining the trend of the continuous subsequence corresponding to the continuous update from the historical trend sequence, and determining the predicted trend sequence based on the trend of the continuous subsequence, the time point of updating, and the amount of updating.

[0112] In some embodiments, by combining the detection sequence and the updating sequence, the predicted trend sequence can more accurately reflect the real performance trend of the system. For example, in a Ceph cluster, if the detection sequence shows that the delay gradually rises, but the updating sequence shows that a certain expansion operation makes the delay drop, then the predicted trend sequence will make a correction at this time point.

[0113] In the embodiments of the present application, by constructing the detection sequence, the updating sequence, and the predicted trend sequence, dynamic tracking and intelligent prediction of the performance change of the system are realized. In this way, early signs of performance degradation can be discovered in time, so that intervention can be made in advance, thereby improving the real-time performance and accuracy of detection.

[0114] Figure 6 An implementation flowchart of a detection method provided in the embodiments of the present application is shown in the figure. The method can be executed by the processor of a computer device. Based on​Figure 5 , Figure 5 Step S503 in the above example can be updated to step S601 to step S604, which will be combined with Figure 6 The steps shown are explained.

[0115] Step S601: Divide the detection sequence into multiple subsequences based on the update sequence.

[0116] Among them, the request response rate corresponding to the first time point in the subsequence is better than the request response rate corresponding to the second time point; the first time point is earlier than the second time point.

[0117] In some embodiments, each point in the update sequence is used as a segmentation point, and the detection sequence is divided into multiple subsequences based on multiple segmentation points.

[0118] In some embodiments, the detection sequence is a time-ordered request response rate sequence S = {s1, s2, ..., s n}, the time series corresponding to the detection sequence is T={t1,t2,…,t n}, where t1 <t2<…,<t n ; The update sequence is U={u1,u2,…,u m}, then based on each point in the update sequence, the detection sequence is divided into subseries including Z1={1,…u 1-1}、Z2={u1,…u 2-1}, ..., Z M+1 ={u m ,…u n}.

[0119] Step S602: Generate a historical trend sequence of the system to be detected based on the multiple subsequences.

[0120] In some embodiments, for each subsequence, a least squares method is used to determine the trend of each subsequence, and a historical trend sequence is obtained based on the trend of each subsequence.

[0121] Exemplarily, it includes 5 subsequences, and the changing trends of the request response rate in each subsequence are determined based on the least squares method to be b1, b2, b3, b4, and b5 respectively. Based on the changing trends of the request response rate in each subsequence, the historical trend sequence is obtained as {b1, b2, b3, b4, b5}.

[0122] Step S603: Obtain the update amount of the target update sequence at the corresponding time point in the update sequence.

[0123] The target update sequence includes a continuous time point with an existing update state closest to a current time point; and the update amount includes an update amount of a request response rate corresponding to the continuous time point.

[0124] In some embodiments, a continuous time interval with an update state of 1 closest to the current time point in the target update sequence.

[0125] In some embodiments, the update amount is a sum of a request response rate change value of each time point in the target update sequence.

[0126] In some embodiments, the update amount of the continuous update time point is determined based on the update amount of the target update sequence corresponding to the time point, as shown in the following formula (4).

[0127]

[0128] wherein Δq i is the update amount of the continuous update time point, Δq i is the update amount of one update time point, v1 is the first update point in the continuous update time point closest to the current time point; and u1 is the last update point in the continuous update time point closest to the current time point.

[0129] In some embodiments, if there is a continuous update time point in the update sequence, the update amount of each update time point in the continuous update time point is accumulated to obtain the update amount of the target update sequence corresponding to the time point.

[0130] In step S604, the prediction trend sequence is generated based on the historical running trend and the update amount of the target update sequence corresponding to the time point.

[0131] In some embodiments, the trend corresponding to the target update sequence is first obtained from the historical running trend, a plurality of prediction trends are generated based on the update amount of the target update sequence corresponding to the time point and the trend corresponding to the target update sequence, and the prediction trend sequence is generated based on the plurality of prediction trends. The plurality of prediction trends are generated as shown in the following formula (5).

[0132]

[0133] wherein, is the prediction trend, t-1 is the (t-1)th continuous update time closest to the current time point, b t-1 is the (t-1)th trend in the historical trend sequence (i.e., the trend corresponding to the (t-1)th sub-sequence), Δq t′ is the update amount of the target update sequence corresponding to the time point, that is, the update amount corresponding to the continuous update time point; bi is the trend corresponding to the i-th update time point in the historical trend sequence.

[0134] In this embodiment, the detection sequence is divided into multiple subsequences based on the update sequence, and the update amount of the historical trend sequence and the target update sequence is extracted from them to generate a predicted trend sequence. In this way, the request response rate of the detection system at a future time point can be predicted based on the predicted trend sequence, thereby testing the request response capability of the detection system, thereby improving the accuracy of the detection.

[0135] Figure 7a The following is a flow chart of a detection method provided in an embodiment of the present application, which can be executed by a processor of a computer device. Figure 1 , the preset data includes a preset request response rate; Figure 1 Step S103 can be updated to step S701, which will be explained in conjunction with the steps shown in FIG. 7 .

[0136] Step S701: Generate a detection result for the system to be detected based on the predicted trend sequence and the preset request response rate.

[0137] In some embodiments, the preset request response rate refers to a baseline threshold set according to system performance standards, which is used to measure whether the real-time performance of the system to be tested is within a normal range. This threshold is usually set by operations and maintenance personnel based on historical operating data and business needs to reflect the performance degradation that may occur under high load or abnormal conditions. For example, in a Ceph storage cluster, operations and maintenance personnel can set a reasonable average request response time (such as 5ms). If the current request response time exceeds this threshold, it is considered that the system performance is showing signs of degradation.

[0138] In some embodiments, a predicted request response rate of the system to be detected at a future time point is predicted based on a predicted trend sequence; and a detection result for the system to be detected is generated based on a preset request response rate and the predicted request response rate.

[0139] Figure 7b The following is a schematic diagram of a detection method provided in an embodiment of the present application, which can be executed by a processor of a computer device. Figure 7a , Figure 7a Step S701 in the example can be updated to step S7011 or step S7012.

[0140] Step S7011, in a case where the request response rate at the current time of the to-be-detected system is greater than the preset request response rate, generating a detection result for the to-be-detected system based on the historical trend sequence and the predicted trend sequence.

[0141] In some embodiments, in a case where the request response rate at the current time of the to-be-detected system is greater than the preset request response rate, it is indicated that the request response capability of the to-be-detected system is possibly degraded, and thus it is further determined based on the historical trend sequence and the predicted trend sequence whether the to-be-detected system is degraded.

[0142] In some embodiments, in a case where the request response rate at the current time of the to-be-detected system is greater than the preset request response rate, it is indicated that the request response capability of the to-be-detected system is possibly degraded, and thus it is further determined based on the historical trend sequence and the predicted trend sequence whether the to-be-detected system is degraded.

[0143] Step S7012, in a case where the request response rate at the current time of the to-be-detected system is less than or equal to the preset request response rate, generating a detection result for the to-be-detected system based on the predicted trend sequence.

[0144] In some embodiments, in a case where the request response rate at the current time of the to-be-detected system is less than or equal to the preset request response rate, it is indicated that the to-be-detected system is not currently degraded, and thus the request response rate at a future time of the to-be-detected system is predicted based on the predicted trend sequence, so as to generate a detection result of the request response rate at the future time of the to-be-detected system.

[0145] In the embodiments of the present application, whether the request response capability of the to-be-detected system is possibly degraded is determined based on the request response rate at the current time and the preset request response rate; if it is possibly degraded, whether it is degraded is further determined based on the historical trend sequence and the predicted trend sequence; and if it is not currently degraded, the request response capability of the to-be-detected system at a future time is predicted based on the predicted trend sequence. In this way, by detecting the request response capability of the to-be-detected system at the current time and predicting the request response capability of the to-be-detected system at the future time, the accuracy of detection is improved.

[0146] Figure 8 An implementation flowchart of a detection method provided in the embodiments of the present application is shown, which can be executed by a processor of a computer device. The above step S7011 can be updated to steps S801 to S802 or step S803, which will be described in combination with the steps shown in the following figures. Figure 8

[0147] ​Step S801, obtaining a historical interval trend corresponding to a request response rate at a current time point in the historical trend sequence, and a predicted interval trend corresponding to the request response rate at the current time point in the predicted trend sequence.

[0148] In some embodiments, in a case where the request response rate at the current time of the system to be detected is greater than the preset request response rate, it is indicated that the request response capability of the system to be detected at the current time may be degraded, and then whether the system to be detected is degraded is determined based on the historical interval trend of the current time point in the historical trend sequence and the predicted interval trend of the current time in the predicted trend sequence.

[0149] In some embodiments, the historical trend sequence is a change trend sequence from a historical time point to the current time point; the trend sequence includes a plurality of sub-sequences; and a trend corresponding to a sub-sequence in which the current time point is located is obtained as the historical interval trend. In the historical trend sequence, the trend of the current time point is a growth trend, and since the historical trend sequence is generated based on the update trend sequence and the detection trend sequence, if the trend is a decline, it is indicated that the time point is an update point to be excluded in the detection sequence.

[0150] In some embodiments, the predicted trend sequence includes a trend from the current time point to a future time point as the predicted interval trend; the trend can be growth or decline.

[0151] Step S802, in a case where the historical interval trend is greater than the predicted interval trend, a detection result indicating that the request response capability of the system to be detected is degraded is generated.

[0152] In some embodiments, in a case where the historical interval trend and the predicted interval trend corresponding to the current time point are both growth or both decline, the slopes of the historical interval trend and the predicted interval trend are compared.

[0153] If the historical interval trend and the predicted interval trend are both growth, in a case where the growth trend of the historical interval trend is greater than the growth trend of the predicted interval trend, it is indicated that the request response rate of the system to be detected at the current time exceeds the predicted trend; and a detection result indicating that the request response capability of the system to be detected is degraded is generated.

[0154] If the historical interval trend is growth and the predicted interval trend is both decline, it is indicated that the request response rate of the system to be detected at the future time point is in a declining trend, and the current time is in a growth trend; and a detection result indicating that the request response capability of the system to be detected is normal is generated.

[0155] Step S803, in a case where the historical interval trend is less than or equal to the predicted interval trend, a detection result is generated to indicate that the request response capability of the system to be detected is normal.

[0156] In some embodiments, in a case where the historical interval trend and the predicted interval trend are both increasing, and the increasing trend of the historical interval trend is less than or greater than the increasing trend of the predicted interval trend, it indicates that the slope of the increasing trend of the current request response rate of the system to be detected is less than the slope of the increasing trend at the future time point of the system to be detected, and a detection result is generated to indicate that the request response capability of the system to be detected is normal.

[0157] In the embodiments of the present application, in a case where the current request response capability of the system to be detected may be degraded, the slope of the historical interval trend in the historical trend sequence at the current time and the slope of the predicted interval trend in the predicted trend sequence are used to determine whether the system to be detected is degraded at the current time. In this way, the request response capability of the system to be detected at the current time can be detected, thereby improving the accuracy of the detection.

[0158] Figure 9 An implementation flow diagram of a detection method provided in the embodiments of the present application is shown, which can be executed by a processor of a computer device. The above step S7012 can be updated to steps S901 and S902, which will be described in combination with the steps shown in the following figure. Figure 9

[0159] Step S901, in a case where the predicted request response rate is greater than the preset request response rate, a detection result is generated to indicate that the request response capability of the system to be detected is degraded.

[0160] In some embodiments, in a case where the request response rate of the system to be detected at the current time is less than or equal to the preset request response rate, it indicates that the system to be detected may not be degraded at the current time, and the request response rate of the system to be detected at the future time is predicted based on the predicted trend sequence to obtain a predicted request response rate.

[0161] In some embodiments, in a case where the predicted request response rate is greater than the preset request response rate, it indicates that the system to be detected may be degraded at the future time, and a detection result is generated to indicate that the request response capability of the system to be detected may be degraded at the future time.

[0162] Step S902, in a case where the predicted request response rate is less than the preset request response rate, a detection result is generated to indicate that the request response capability of the system to be detected is normal.

[0163] ​In some embodiments, if the predicted request response rate is less than or equal to the preset request response rate, it is characterized that the request response capability of the to-be-detected system at the future time point is better than that at the current time, and a detection result is generated, which indicates that the request response capability of the to-be-detected system at the future time point is normal.

[0164] In the embodiments of the present application, in the case that the request response capability of the to-be-detected system at the current time does not degrade, the request response capability of the to-be-detected system at the future time point is predicted based on the predicted request response rate and the preset request response rate. In this way, a detection result of whether the request response capability of the to-be-detected system at the future time point degrades or not can be generated, and the accuracy of detection is improved.

[0165] The following describes an exemplary application of the detection method provided by the embodiments of the present application in an actual scenario.

[0166] Many distributed file systems are designed based on a non-single-point-failure architecture, which can provide high-performance, high-availability and high-reliability data storage services. A distributed file system integrates multiple storage nodes to form a unified storage pool, supports multiple storage modes such as object storage, block storage and file system storage, and is widely used in cloud computing, big data and Internet of Things fields.

[0167] With the delivery of distributed storage systems, the load of the storage system will also become heavier. It is necessary to evaluate and predict the degradation of key performance indicators to provide resources and schedule loads in a targeted manner. For example, when the system performance degrades, the delay indicators of the storage system and the storage volumes created for the cloud platform will degrade, which will have a significant impact on database systems and applications that are sensitive to storage delay.

[0168] With the rapid development of cloud computing and big data technologies, Ceph storage clusters are increasingly widely used in data centers and enterprise-level storage services. However, the degradation problem has a growing impact on business continuity and data availability. Therefore, developing a technology solution that can efficiently, accurately and in real time detect the degradation state of a Ceph storage cluster is of great significance to ensure the security and reliability of data storage services. Due to different construction schemes of different storage systems, different business characteristics, diverse performance degradation needs, and the strong subjectivity of performance degradation evaluation, there is a lack of effective detection and prediction means for performance degradation in the prior art. Storage operation and maintenance engineers can only handle the problem after receiving a report from an application administrator. Effective evaluation and prediction tools can help operation and maintenance engineers carry out necessary predictive maintenance and develop targeted resource delivery and scheduling strategies.

[0169] To solve the above problems, the prior art also has the following technical problems:

[0170] (1) Limited detection scope: Many existing technologies can only detect specific degradation phenomena, such as hardware failure or disk damage, and cannot comprehensively cover various degradation situations that may occur in Ceph storage clusters.

[0171] (2) Insufficient real-time performance: Some technologies rely on regular data collection and analysis, resulting in poor real-time performance of detection results and an inability to detect and address potential degradation issues in a timely manner.

[0172] (3) Low accuracy: Due to the complexity and dynamic nature of Ceph storage clusters, existing technologies are often susceptible to interference from noise and outliers, resulting in frequent false positives and missed positives.

[0173] In response to the above technical problems, an embodiment of the present application provides a detection method that analyzes and predicts the historical data of the distributed storage system, thereby detecting and predicting the degradation state of the distributed storage system, thereby providing strong protection for the enterprise's data storage services and improving the stability and reliability of the storage services.

[0174] In some embodiments, this solution uses the PAA segmented aggregation algorithm to segment the cluster average delay, uses the median of each subsequence to represent the subsequence, generates a shorter sequence data, effectively extracts the delay trend sequence, and achieves data simplification.

[0175] In some embodiments, this solution employs a streaming, real-time approach. Each time a cluster is subjected to degradation analysis, only the data for the current time period is needed, combined with a small amount of PAA aggregated data cached in Redis. This eliminates the need to acquire large amounts of data, enabling efficient and rapid detection and prediction.

[0176] In some embodiments, this solution proposes a prediction model based on the update process. This model treats cluster service migration or expansion as an update, and performs detection and prediction based on the update process, enabling cluster performance monitoring and prediction. This update process prediction model has low computational complexity and can track dynamic changes in the cluster in a timely manner, enabling accurate detection and prediction of cluster performance degradation.

[0177] In some embodiments, this solution targets cold start scenarios where historical data is insufficient. It proposes a latency and throughput correlation analysis method, and performs throughput-based predictions to predict cluster degradation.

[0178] Figure 10 The present invention provides a detection method for implementing a flow chart, the method can be executed by a processor of a computer device. The method includes the following steps S1001 and S1002. Figure 10 The steps shown are explained.

[0179] Step S1001, data collection and preprocessing.

[0180] In some embodiments, when collecting Ceph cluster performance data, attention needs to be paid to a plurality of key indicators that can comprehensively reflect the health status, performance and potential problems of the cluster. The present scheme collects and analyzes through Prometheus every 1 minute, collects the average latency Latency (read and write latency), IO throughput Throughput (read and write throughput), storage cluster used capacity Used_capacity, storage cluster total capacity Total_capaticty, etc.

[0181] In some embodiments, the embodiments of the present application are illustrated by the latency and throughput of the distributed storage system. Among them, the historical latency data and historical throughput data of the distributed storage system may have some extreme values and missing values, so the data augmentation technology is used to process the extreme values and missing values, including:

[0182] 1. Nullify the missing values and extreme values in the data. Among them, the missing values are directly nullified, and the data greater than the threshold value in the extreme values is nullified after the extreme values are extreme.

[0183] 2. Fill the nullified data.

[0184] 3. If the number n of consecutive nullified data points is less than m (m is 15, for example, indicating that the number of nullified data points is not more than 15), the forward filling method is used to fill the nullified data points with the value of the previous non-missing value or non-extreme value data point; if the number n of consecutive nullified data points is greater than m, the trend of the h data points before nullification is determined by using the least square method, and the h data points before nullification are filled with the trend.

[0185] In some embodiments, after the historical latency data and historical throughput data of the distributed storage system are processed for extreme values and missing values, standardization processing is also needed, please refer to the above formula (1).

[0186] In some embodiments, the historical throughput data and the historical latency data of the distributed storage system also need to be processed by Piecewise Aggregate Approximation (PAA). The PAA algorithm is a strategy for time series dimensionality reduction, which divides the original long sequence data into multiple equal-length subsequences, and then represents each subsequence based on the statistical quantity (such as mean, variance, etc.) of the subsequence, thereby generating a shorter sequence data, thereby reflecting the overall trend of the original long sequence, thereby reducing the amount of data processing and the complexity of calculation. The following includes:

[0187] 1. The given cluster average latency sequence Q is divided by week to obtain the subsequence Q = q1, q2,.. qk.

[0188] 2. The median of each subsequence q is calculated, and the median of the subsequence represents each value in the subsequence.

[0189] 3. The sequence obtained in 2 is reduced, and only one value is extracted from each sequence to represent the sequence.

[0190] Wherein, if the subsequence of the sequence Q includes {{1, 2, 3}, {4, 5, 6}}, the median of each subsequence represents the subsequence, and the sequence Q can be {2, 5}.

[0191] Figure 11 The schematic diagram of data statistics provided by the embodiments of the present application is shown in FIG. 11. As shown in FIG. 11, 1101 is a statistical diagram of the original data, and it can be seen that the change trend of the data cannot be obviously obtained from the original data 1101; 1102 is a statistical diagram of the data after PAA segmentation and aggregation, and the change trend of the data can be obviously obtained; and 1103 is a statistical diagram of the reduced data. After reducing the data after PAA segmentation and aggregation, the change trend of the data is retained, and the change is more continuous.

[0192] Step S1002, based on the preprocessed data, the performance of the distributed storage system is predicted.

[0193] In some embodiments, the distributed storage system includes a cold start phase and a hot start phase. Wherein, if the distributed storage system is in the initial stage or has no sufficient historical data, it is characterized as a cold start phase; if the distributed storage system has been running for a period of time and has sufficient data for cluster degradation analysis, it is a hot start phase.

[0194] In some embodiments, for the cold start phase of the distributed system, the throughput of the system is predicted by the throughput data and the latency data of the distributed storage system in the cold start phase, and the predicted value P th is obtained. Combined with the preset throughput threshold T th, determine whether the distributed storage system exists degradation. Wherein, the distributed storage system under normal operation, the throughput of the system will not change significantly, or the delay and throughput is inversely proportional relationship, namely the throughput delay increases. However, with the continuous use of the cluster, the cluster performance begins to appear problems, will appear delay and throughput is a positive relationship, namely the throughput delay increases, such as Figure 12 , Figure 12 A trend change diagram of a distributed storage system provided by the embodiment of the present application, wherein 1201 is a system delay change statistical chart, and 1202 is a system throughput change chart; wherein it can be seen that in the initial stage, the delay increases with time, and the throughput is unchanged or decreases, and in the subsequent delay increases, and the throughput also increases.

[0195] In some embodiments, based on the delay and throughput of the system in the recent period (such as in the recent month), the Person correlation coefficient is calculated, and the change of the correlation coefficient is monitored in real time. For example, if the correlation coefficient represents that the delay and the throughput have a positive correlation, and is greater than a preset threshold (for example, 0.7), it is considered that the system may have degradation. Then, the change point detection is performed on the throughput, the starting point at which the throughput starts to increase is identified, the CUSUM algorithm is used for change point detection, then the data after the last change point is filtered out, and the linear regression model is used for prediction to predict the throughput P th of the system in a future period (such as one month). th For example, if the predicted throughput P th > T th , it is judged that the cluster storage has a degradation trend, and the cluster needs to be focused on.

[0196] In some embodiments, the degradation detection of the distributed storage system in the hot start stage includes:

[0197] First, determine whether the current cluster is detected for the first time, a) if it is detected for the first time, a large amount of cluster delay historical data (such as 1 year data) is obtained, the data is processed by PAA segmentation and aggregation, then the processed data is cached, and the cache time is recorded; b) if it is not the first time, the latest historical data after the cache time is obtained from the historical data, and then the PAA segmentation and aggregation processing is performed. Then read the data in the cache and the new processed data to splice, obtain the detection sequence {s1, s2, …, s n}.

[0198] Second, as the system runs, the increase in system load and the aging of hardware will cause the system performance to degrade. When the system encounters performance problems, storage operation and maintenance engineers will migrate certain services to other systems or expand the system to ensure the performance of the system. After the storage operation and maintenance engineers migrate or expand the system, the performance of the system will be improved to varying degrees, and the average delay of the system will be reduced. Therefore, the system delay is constructed as an update process model. Each migration or expansion is an update. The state after the update is different from the previous state, and the performance of the system after the update is improved. In this regard, when the value at time i in the detection sequence is less than the value at time i-1, time i is considered to be the updated state, and thus the updated state sequence U = {u1, u2,…, u n}, refer to the above formula (3).

[0199] Third, the above detection sequence is divided based on the updated state sequence to obtain multiple subsequences. Figure 13 , Figure 13 A schematic diagram of data update statistics provided in an embodiment of the present application, wherein 1301 is the delay at the initial time point of the system, 1302 is the delay corresponding to the time point of the first update, and 1303 is the delay corresponding to the time point of the second update. Combining formula (3), we can obtain the update amount sequence for each update state 1 (that is, the update state) as ΔQ = {Δq1, Δq2, ..., Δq n}, it can be understood that the update amount is the delay value at time i minus the delay value at time i-1 when the update state is 1 in the detection sequence, and the update state is 0, which indicates that the update amount is 0.

[0200] If there are update states at consecutive time points, the consecutive updates can be summed when calculating the update amount, as shown in the above formula (4), to obtain the update amount of a complete consecutive update.

[0201] Fourth, the degradation trend of distributed storage systems during the hot start phase can be attributed to two factors: first, performance degradation due to increased cluster load or hardware aging; second, a temporary performance boost due to business migration or cluster expansion by storage operations engineers. Therefore, when detecting cluster performance degradation, an updated prediction model is constructed, including:

[0202] 1. For the detection sequence {s1,s2,…,s n Eliminate the data points with update status 1 to obtain k data intervals, and use the least squares method to estimate the trend b of the data in each data interval. k , we get the historical trend sequence B={b1,b1,…,b k}.

[0203] 2. For the time data point to be detected, the last large update is the tth large update, and the trend is predicted based on the above formula (5).

[0204] Sixth, based on the predicted trend sequence and the historical trend sequence described above, the degradation detection of the distributed storage system in the hot start stage is carried out, including:

[0205] 1. If the delay of the system at the current time is greater than the delay threshold (the maximum value of the average delay of the system), the sequence of the current time in the historical trend sequence and the sequence of the current time in the predicted trend sequence are obtained, and if the trend of the sequence of the current time in the historical trend sequence is greater than the trend of the sequence of the current time in the predicted trend sequence, it is indicated that the system currently exists degradation.

[0206] 2. If the delay of the system at the current time is less than or equal to the delay threshold, the predicted delay of the system in the future (for example, one month) is predicted based on the predicted trend sequence, and if the delay of the system at the current time is greater than the predicted delay, it is indicated that the system may appear degradation in the future.

[0207] Based on the foregoing embodiments, the embodiments of the present application provide a detection device, which includes various units and various modules included in the units, and can be realized by a processor in a computer device. Of course, it can also be realized by a specific logic circuit. In the implementation process, the processor can be a central processing unit (CPU), a microprocessor unit (MPU), a digital signal processor (DSP), or a field programmable gate array (FPGA).

[0208] Figure 14 The composition structure diagram of the detection device provided by the embodiments of the present application is shown in FIG. 14. Figure 14 As shown in FIG. 14, the detection device 1400 includes an acquisition module 1401, a first generation module 1402, and a second generation module 1403. The acquisition module 1401 is configured to acquire historical data of a system to be detected. The historical data represents the current request response capability of the system to be detected. The first generation module 1402 is configured to generate predicted data for the system to be detected based on the historical data. The predicted data represents the predicted request response capability of the system to be detected. The second generation module 1403 is configured to generate a detection result for the system to be detected based on the predicted data and preset data of the system to be detected.

[0209] In some embodiments, the historical data comprises a request response rate and a request response quantity; the first generating module 1402 is further configured to generate the prediction data based on the request response rate and the request response quantity when the to-be-detected system is in a first state, and generate the prediction data based on the request response rate when the to-be-detected system is in a second state; wherein the historical data of the to-be-detected system in the first state meets a non-pre-set requirement, and the historical data of the to-be-detected system in the second state meets a pre-set requirement.

[0210] In some embodiments, the prediction data comprises a predicted request response quantity at a pre-set time point after a current time point, and the first generating module 1402 is further configured to determine a correlation coefficient between the request response rate and the request response quantity; when the correlation coefficient is greater than a pre-set coefficient, generate the predicted request response quantity of the to-be-detected system at the pre-set time point after the current time point based on a target request response quantity in the request response quantity; wherein the target request response quantity comprises request response quantities at a plurality of time points, and the plurality of time points at least comprise a target time point to the current time point; the target time point is a time point at which the request response quantity starts to increase closest to the current time point.

[0211] In some embodiments, the pre-set data comprises a pre-set request response quantity, and the second generating module 1403 is further configured to generate a detection result indicating that the request response capability of the to-be-detected system is degraded when the predicted request response quantity is greater than the pre-set request response quantity, and generate a detection result indicating that the request response capability of the to-be-detected system is normal when the predicted request response quantity is less than or equal to the pre-set request response quantity.

[0212] In some embodiments, the prediction data comprises a prediction trend sequence; the first generating module 1402 is further configured to generate a detection sequence for the to-be-detected system based on the request response rate; the detection sequence comprises a request response rate corresponding to each time point in a historical time; generate an update sequence for the to-be-detected system based on the detection sequence; the update sequence comprises at least one target time point indicating an update state, and in the detection sequence, the request response rate corresponding to the target time point is better than the request response rate corresponding to a previous time point of the target time point; and generate the prediction trend sequence based on the detection sequence and the update sequence.

[0213] In some embodiments, the first generation module 1402 is further configured to divide the detection sequence into a plurality of sub-sequences based on the update sequence; a request response rate corresponding to a first time point in the sub-sequence is better than a request response rate corresponding to a second time point; the first time point is earlier than the second time point; generate a historical trend sequence of the to-be-detected system based on the plurality of sub-sequences; obtain an update amount of a target update sequence corresponding to a time point in the update sequence; the target update sequence includes a continuous time point closest to a current time point and having an update state; the update amount includes an update amount of a request response rate corresponding to a continuous time point; and generate the prediction trend sequence based on the historical running trend and the update amount of the target update sequence corresponding to the time point.

[0214] In some embodiments, the preset data includes a preset request response rate; the second generation module 1403 is further configured to generate a detection result for the to-be-detected system based on the prediction trend sequence and the preset request response rate; and the generation of the detection result for the to-be-detected system based on the prediction trend sequence and the preset request response rate includes: in a case where a request response rate of the to-be-detected system at a current time is greater than the preset request response rate, generating the detection result for the to-be-detected system based on the prediction trend sequence and the historical trend sequence; and in a case where the request response rate of the to-be-detected system at the current time is less than or equal to the preset request response rate, generating the detection result for the to-be-detected system based on the prediction trend sequence.

[0215] In some embodiments, the second generation module 1403 is further configured to obtain a historical interval trend of a request response rate at a current time point in a historical trend sequence corresponding interval and a prediction interval trend of the request response rate at the current time point in a prediction trend sequence corresponding interval; in a case where the historical interval trend is greater than the prediction interval trend, generate a detection result indicating that a request response capability of the to-be-detected system is degraded; and in a case where the historical interval trend is less than or equal to the prediction interval trend, generate a detection result indicating that the request response capability of the to-be-detected system is normal.

[0216] In some embodiments, the prediction data further includes a predicted request response rate of the to-be-detected system at a preset time point after a current time point of the to-be-detected system; the second generation module 1403 is further configured to, in a case where the predicted request response rate is greater than the preset request response rate, generate a detection result indicating that a request response capability of the to-be-detected system is degraded; and in a case where the predicted request response rate is less than the preset request response rate, generate a detection result indicating that the request response capability of the to-be-detected system is normal.

[0217] The descriptions of the above device embodiments are similar to the descriptions of the above method embodiments, and have similar beneficial effects as the method embodiments. In some embodiments, the device provided by the embodiments of the present application has functions or includes modules that can be used to perform the methods described in the above method embodiments. For technical details not disclosed in the device embodiments of the present application, please refer to the description of the method embodiments of the present application for understanding.

[0218] It should be noted that, in the embodiments of the present application, if the above-mentioned methods are implemented in the form of software function modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on this understanding, the technical solutions of the embodiments of the present application can be embodied in the form of a software product, which is stored in a storage medium and includes a number of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the methods described in the embodiments of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM), a magnetic disk or an optical disk, and various program code storage media. Thus, the embodiments of the present application are not limited to any specific hardware, software or firmware, or any combination of hardware, software and firmware.

[0219] The embodiments of the present application provide a computer device, including a memory and a processor, the memory stores a computer program capable of running on the processor, and the processor implements part or all of the steps in the above method when executing the program.

[0220] The embodiments of the present application provide a computer-readable storage medium, which stores a computer program, and the computer program is executed by a processor to implement part or all of the steps in the above method. The computer-readable storage medium can be transitory or non-transitory.

[0221] The embodiments of the present application provide a computer program, which includes computer-readable code, and when the computer-readable code is running in a computer device, a processor in the computer device executes part or all of the steps in the above method.

[0222] An embodiment of the present application provides a computer program product, which includes a non-transitory computer-readable storage medium storing a computer program, and when the computer program is read and executed by a computer, implements some or all of the steps in the above method. The computer program product can be implemented specifically by hardware, software, or a combination thereof. In some embodiments, the computer program product is embodied as a computer storage medium. In other embodiments, the computer program product is embodied as a software product, such as a software development kit (SDK), etc.

[0223] It should be noted that the descriptions of the various embodiments above tend to emphasize the differences between the various embodiments, and their similarities or similarities can be referenced to each other. The descriptions of the above device, storage medium, computer program, and computer program product embodiments are similar to the descriptions of the above method embodiments and have similar beneficial effects as the method embodiments. For technical details not disclosed in the embodiments of the device, storage medium, computer program, and computer program product of this application, please refer to the description of the method embodiments of this application for understanding.

[0224] Figure 15 A hardware entity diagram of a computer device provided in an embodiment of the present application is shown as follows: Figure 15 As shown, the hardware entity of the computer device 1500 includes: a processor 1501 and a memory 1502, wherein the memory 1502 stores a computer program that can be run on the processor 1501, and when the processor 1501 executes the program, the steps in the method of any of the above embodiments are implemented.

[0225] The memory 1502 stores computer programs that can be run on the processor. The memory 1502 is configured to store instructions and applications executable by the processor 1501. It can also cache data to be processed or processed by the processor 1501 and various modules in the computer device 1500 (for example, image data, audio data, voice communication data, and video communication data). This can be implemented through flash memory (FLASH) or random access memory (RAM).

[0226] When the processor 1501 executes the program, the steps of any of the above methods are implemented. The processor 1501 generally controls the overall operation of the computer device 1500.

[0227] An embodiment of the present application provides a computer storage medium, which stores one or more programs. The one or more programs can be executed by one or more processors to implement the steps of the method of any of the above embodiments.

[0228] It should be noted that the description of the above storage medium and device embodiments is similar to the description of the above method embodiments and has similar beneficial effects as the method embodiments. For technical details not disclosed in the storage medium and device embodiments of this application, please refer to the description of the method embodiments of this application for understanding.

[0229] The processor may be at least one of an application-specific integrated circuit (ASIC), a digital signal processor (DSP), a digital signal processing device (DSPD), a programmable logic device (PLD), a field programmable gate array (FPGA), a central processing unit (CPU), a controller, a microcontroller, and a microprocessor. It is understood that the electronic device that implements the functions of the processor may also be other electronic devices, which are not specifically limited in the embodiments of the present application.

[0230] The above-mentioned computer storage medium / memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a magnetic random access memory (FRAM), a flash memory (Flash Memory), a magnetic surface memory, an optical disc, or a compact disc read-only memory (CD-ROM); it can also be various terminals including one or any combination of the above-mentioned memories, such as mobile phones, computers, tablet devices, personal digital assistants, etc.

[0231] It should be understood that every feature, structure, or characteristic described herein is within a preferred embodiment of the present application. It should be noted that the foregoing embodiments are merely exemplary and are not to be construed as limiting the present application. It should also be noted that features described in the foregoing relate to both structural and method aspects of the application. Accordingly, the terminology in use has a multi-use effect: to the extent there is structural correspondence between a feature described herein and terminology conventionally used to describe corresponding structure in the art, that correspondence is intended. To the extent terminology is used in a manner other than that conventionally used to describe corresponding structure in the art, that terminology is intended to refer to that which is described herein, whether or not there is a corresponding structure in the art which would otherwise be described using that terminology. It is intended that each aspect described herein apply to every embodiment described herein, unless otherwise indicated. It should be understood that any numerical range recited herein includes all values from the lower and upper limits of that range. It should be understood that every maximum numerical limitation given throughout this specification includes every lower numerical limitation, as if such lower numerical limitations were expressly written herein. Every minimum numerical limitation given throughout this specification includes every higher numerical limitation, as if such higher numerical limitations were expressly written herein. Every range given throughout this specification includes every narrower range that falls within the broader range, as if such narrower ranges were all expressly written herein.

[0232] It should be noted that, as used in this specification and the appended claims, the singular forms "a," "an," and "the" include plural referents unless the content clearly dictates otherwise. Thus, for example, reference to "a component" includes a combination comprising two or more such components, and the like.

[0233] In several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. The above-described device embodiments are merely illustrative. For example, the division of the units is merely a logical functional division. In actual implementation, another division manner can be used, such as: a plurality of units or components can be combined, or can be integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the displayed or discussed components can be indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.

[0234] The units described above as separate components can or can not be physically separate, and the components displayed as units can or can not be physical units; they can be located in one place or distributed on multiple network units; and some or all of the units can be selected according to actual needs to achieve the purpose of the embodiment.

[0235] In addition, each function unit in each embodiment of the present application can be integrated in one processing unit, or each unit can be separately taken as one unit, or two or more units can be integrated in one unit; the integrated unit can be realized in the form of hardware or in the form of hardware plus software function unit. Those skilled in the art can understand that all or part of the steps of the foregoing method embodiments can be completed by a program instructing related hardware, the foregoing program can be stored in a computer readable storage medium, and the program, when executed, executes steps including the foregoing method embodiments; and the foregoing storage medium includes: mobile storage equipment, read-only memory (ROM), magnetic disc or optical disc, and various storage medium capable of storing program codes.

[0236] Alternatively, when the foregoing integrated unit is realized in the form of software function module and sold or used as an independent product, it can also be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application can be embodied in the form of a software product, and the computer software product is stored in a storage medium, includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the methods described in the embodiments of the present application. The foregoing storage medium includes: mobile storage equipment, ROM, magnetic disc or optical disc, and various storage medium capable of storing program codes.

[0237] The above is only the implementation of the present application, but the protection scope of the present application is not limited to this, any skilled person in the art can easily think of changes or replacements within the technical range disclosed in the present application, which should be covered in the protection scope of the present application.

Claims

1. A detection method, comprising: Acquire historical data of the system to be detected; the historical data represents the current request response capability of the system to be detected; generating prediction data for the system to be detected based on the historical data; the prediction data representing the predicted request response capability of the system to be detected; Based on the predicted data and the preset data of the system to be detected, a detection result for the system to be detected is generated.

2. The method according to claim 1, wherein the historical data includes a request response rate and a request response number; and generating prediction data for the system to be inspected based on the historical data comprises: generating the prediction data based on the request response rate and the request response quantity when the system to be detected is in a first state; generating the prediction data based on the request response rate when the system to be detected is in the second state; The historical data of the system to be detected in the first state does not meet any preset requirements; the historical data of the system to be detected in the second state meets any preset requirements.

3. The method according to claim 2, wherein the predicted data includes a predicted number of request responses at a preset time point after the current time point; and generating the predicted data based on the request response rate and the number of request responses comprises: determining a correlation coefficient between the request response rate and the number of request responses; When the correlation coefficient is greater than a preset coefficient, generating a predicted number of request responses of the system to be detected at a preset time point after the current time point based on the target number of request responses in the number of request responses; Among them, the target request response number includes the request response number at multiple time points; the multiple time points at least include the target time point to the current time point; the target time point is the time point closest to the current time point when the request response number begins to increase.

4. The method according to claim 3, wherein the preset data includes a preset number of request responses, and generating a detection result for the system to be detected based on the predicted data and the preset data of the system to be detected comprises: When the predicted number of request responses is greater than the preset number of request responses, generating a detection result indicating that the request response capability of the system to be detected has degraded; When the predicted number of request responses is less than or equal to the preset number of request responses, a detection result is generated indicating that the request response capability of the system to be detected is normal.

5. The method according to claim 2, wherein the prediction data comprises a prediction trend sequence; and generating the prediction data based on the request response rate comprises: generating a detection sequence for the system to be detected based on the request response rate; The detection sequence includes the request response rate corresponding to each time point in the historical moment; generating an update sequence for the system to be detected based on the detection sequence; the update sequence including at least one target time point indicating the presence of an update state, wherein in the detection sequence, a request response rate corresponding to the target time point is better than a request response rate corresponding to a time point immediately before the target time point; The predicted trend sequence is generated based on the detection sequence and the update sequence.

6. The method according to claim 5, wherein generating the predicted trend sequence based on the detection sequence and the update sequence comprises: dividing the detection sequence into a plurality of subsequences based on the update sequence; The request response rate corresponding to the first time point in the subsequence is better than the request response rate corresponding to the second time point; The first time point is earlier than the second time point; generating a historical trend sequence of the system to be detected based on the multiple subsequences; Obtaining an update amount of a target update sequence corresponding to a time point in the update sequence; the target update sequence includes consecutive time points with an update state closest to the current time point; the update amount includes an update amount of a request response rate corresponding to the consecutive time points; The predicted trend sequence is generated based on the historical operating trend and the update amount of the target update sequence at the corresponding time point.

7. The method according to claim 6, wherein the preset data includes a preset request response rate; and generating a detection result for the system to be detected based on the predicted data and the preset data of the system to be detected comprises: generating a detection result for the system to be detected based on the predicted trend sequence and the preset request response rate; The generating a detection result for the system to be detected based on the predicted trend sequence and the preset request response rate includes: generating a detection result for the system to be detected based on the predicted trend sequence and the historical trend sequence when the request response rate of the system to be detected at the current moment is greater than the preset request response rate; In a case where the request response rate of the system to be detected at the current time point is less than or equal to the preset request response rate, a detection result for the system to be detected is generated based on the predicted trend sequence.

8. The method according to claim 7, wherein generating a detection result for the system to be detected based on the predicted trend sequence and the historical trend sequence comprises: Obtaining a historical interval trend in which the request response rate at the current time point is within a corresponding interval of the historical trend sequence, and a predicted interval trend in which the request response rate at the current time point is within a corresponding interval of the predicted trend sequence; When the historical interval trend is greater than the predicted interval trend, generating a detection result indicating that the request response capability of the system to be detected has degraded; When the historical interval trend is less than or equal to the predicted interval trend, a detection result is generated indicating that the request response capability of the system to be detected is normal.

9. The method according to claim 7, wherein the prediction data further includes a predicted request response rate of the system to be detected at a preset time point after the current time point; and generating a detection result for the system to be detected based on the predicted trend sequence comprises: When the predicted request response rate is greater than the preset request response rate, generating a detection result indicating that the request response capability of the system to be detected has degraded; In a case where the predicted request response rate is less than the preset request response rate, a detection result is generated indicating that the request response capability of the system to be detected is normal.

10. A detection device, comprising: An acquisition module is used to acquire historical data of the system to be detected; the historical data represents the current request response capability of the system to be detected; A first generating module is configured to generate prediction data for the system to be detected based on the historical data; the prediction data represents the predicted request response capability of the system to be detected; The second generating module is used to generate a detection result for the system to be detected based on the predicted data and the preset data of the system to be detected.