Root cause positioning method, device and equipment
By obtaining and analyzing the load status of the first round trip time RT of the prevention and control request and the asynchronous queue, accurately locate the root cause that affects SLA, solving the problem that multiple influencing factors interweaving in the prior art is difficult to confirm the root cause, and achieving efficient and accurate SLA guarantee.
Patent Information
- Application Number
- CN202411943836.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-25
- Publication Date
- 2025-05-16
AI Technical Summary
The prior art is difficult to accurately locate the root causes of service level agreements (SLAs), especially in the success rate and round trip time (RTT) jitter at the minute level, where multiple influencing factors intertwined, making it difficult to identify the root causes of SLAs.
By obtaining the first round trip time RT of the prevention and control request within the set period, the first RT average of the multiple prevention and control requests in the target queue is determined, and combined with the load status of the asynchronous queue, the basic environment abnormality or service traffic abnormality of the current prevention and control service is determined, and the root cause affecting the SLA is accurately positioned.
The root causes affecting SLA are accurately positioned, reducing the dependence of human resources analysis, and improving the efficiency and accuracy of SLA guarantee.
Smart Images

Figure CN120011114A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and more specifically, to a method, device and apparatus for locating a root cause. Background Art
[0002] With the development of science and technology, people need to rely on various applications in all aspects of their daily life, such as shopping applications, travel applications, etc. Therefore, there are a large number of people using each application, and risk prevention and control are required to prevent users from being exposed to various unsafe and unhealthy information; the application sends various requests to the cloud environment of the platform, calculates through the prevention and control algorithm, and completes the prevention and control service.
[0003] In the prior art, in the prevention and control service, the platform needs to provide a service level agreement (SLA) for various applications (also referred to as business parties). With the gradual promotion of SLA, the current prevention and control service is no longer satisfied with the monthly SLA compliance, but has begun to focus on the minute-level success rate and round-trip time (RTT) jitter (also referred to as RT jitter). The above RT jitter can represent SLA jitter. The platform needs to monitor the content change trend and SLA jitter trend to determine whether the SLA fluctuation is caused by content problems, and then confirm whether the current jitter is related to the basic environment by enumerating the basic environment problems. The above influencing factors are intertwined with each other and it is difficult to confirm the root cause of the SLA.
[0004] In summary, how to locate the root cause that affects SLA is a problem that needs to be solved at present. Summary of the invention
[0005] In view of this, the embodiments of the present invention provide a method, apparatus and device for locating a root cause, which can accurately locate the root cause affecting SLA.
[0006] In a first aspect, an embodiment of the present invention provides a method for root cause location, the method comprising:
[0007] Obtaining a first round-trip time RT of at least one prevention and control request within a set period of time, wherein a content feature in each of the prevention and control requests is less than a corresponding set threshold;
[0008] Determine at least one first RT average of multiple prevention and control requests in the target queue within the set time period;
[0009] In response to the first RT mean being greater than the maximum value in the first RT range and the asynchronous queue being in a full load state, it is determined that the basic environment of the current prevention and control service is abnormal, wherein the first RT range is generated based on a stress test and the basic environment is the hardware environment of the device running the prevention and control algorithm.
[0010] Optionally, the method further includes:
[0011] Generates and sends basic environment exception codes.
[0012] Optionally, the method further includes:
[0013] In response to the first RT mean being within the first RT range and the asynchronous queue being in a full load state, obtaining traffic of the asynchronous queue;
[0014] In response to the traffic being greater than a set traffic threshold, it is determined that the business traffic of the current prevention and control service is abnormal.
[0015] Optionally, the method further includes:
[0016] Generate and send service traffic exception codes.
[0017] Optionally, the method further includes:
[0018] In response to the traffic being less than or equal to the set traffic threshold, it is determined that the prevention and control content of the current prevention and control service is abnormal, wherein the prevention and control content abnormality indicates that performance degradation is caused by fluctuations in the business party content, and the current first RT range cannot meet the demand.
[0019] Optionally, the method further includes:
[0020] Generates and sends a performance degradation exception code.
[0021] Optionally, the method further includes:
[0022] In response to the second RT mean within the set time period being greater than the maximum value of the preset second RT range, and the first RT mean being greater than the maximum value in the first RT range, and the asynchronous queue being in a normal load state, the basic environment of the current prevention and control algorithm is determined to be abnormal, wherein the second RT mean is the average of the complete service RTs of multiple prevention and control requests within the set time period, and the second RT range is pre-set.
[0023] Optionally, the method further includes:
[0024] In response to the second RT mean within the set time period being greater than the maximum value of the preset second RT range, and the first RT mean being within the first RT range, and the asynchronous queue being in a normal load state, it is determined that the prevention and control content of the current prevention and control algorithm is abnormal, wherein the second RT mean is the mean of the complete service RT of multiple prevention and control requests within the set time period, and the second RT range is pre-set.
[0025] Optionally, obtaining a first round trip time RT of at least one prevention and control request within a set period of time specifically includes:
[0026] Get the complete service RT of each control request in at least one control request within the set time period;
[0027] The waiting time in the asynchronous queue and / or the external call time are removed from the complete service RT of each prevention and control request, and the first round-trip time RT of each prevention and control request is determined.
[0028] Optionally, the method comprises:
[0029] Acquire historical prevention and control request data of a historical setting period, wherein the historical prevention and control request data includes a first round trip time RT of each historical prevention and control request;
[0030] The historical prevention and control request data is stress tested to generate set thresholds corresponding to the first RT range and multiple content features.
[0031] In a second aspect, an embodiment of the present invention provides a root cause location device, the device comprising:
[0032] An acquisition unit, configured to acquire a first round trip time RT of at least one prevention and control request within a set period of time, wherein a content feature in each of the prevention and control requests is less than a corresponding set threshold;
[0033] A determination unit, configured to determine at least one first RT mean of a plurality of prevention and control requests in a target queue within the set time period;
[0034] A determination unit, in response to the first RT mean being greater than the maximum value in the first RT range and the asynchronous queue being in a full load state, is used to determine that a basic environment of the current prevention and control service is abnormal, wherein the first RT range is generated based on a stress test and the basic environment is the hardware environment of a device running a prevention and control algorithm.
[0035] Optionally, the device further includes a processing unit configured to generate and send a basic environment exception code.
[0036] Optionally, the determination unit is further configured to:
[0037] In response to the first RT mean being within the first RT range and the asynchronous queue being in a full load state, obtaining traffic of the asynchronous queue;
[0038] In response to the traffic being greater than a set traffic threshold, it is determined that the business traffic of the current prevention and control service is abnormal.
[0039] Optionally, the processing unit is further used for:
[0040] Generate and send service traffic exception codes.
[0041] Optionally, the determination unit is further configured to:
[0042] In response to the traffic being less than or equal to the set traffic threshold, it is determined that the prevention and control content of the current prevention and control service is abnormal, wherein the prevention and control content abnormality indicates that performance degradation is caused by fluctuations in the business party content, and the current first RT range cannot meet the demand.
[0043] Optionally, the processing unit is further used for:
[0044] Generates and sends a performance degradation exception code.
[0045] Optionally, the determination unit is further configured to:
[0046] In response to the second RT mean within the set time period being greater than the maximum value of the preset second RT range, and the first RT mean being greater than the maximum value in the first RT range, and the asynchronous queue being in a normal load state, the basic environment of the current prevention and control algorithm is determined to be abnormal, wherein the second RT mean is the average of the complete service RTs of multiple prevention and control requests within the set time period, and the second RT range is pre-set.
[0047] Optionally, the determination unit is further configured to:
[0048] In response to the second RT mean within the set time period being greater than the maximum value of the preset second RT range, and the first RT mean being within the first RT range, and the asynchronous queue being in a normal load state, it is determined that the prevention and control content of the current prevention and control algorithm is abnormal, wherein the second RT mean is the mean of the complete service RT of multiple prevention and control requests within the set time period, and the second RT range is pre-set.
[0049] Optionally, the acquiring unit is specifically used for:
[0050] Get the complete service RT of each control request in at least one control request within the set time period;
[0051] The waiting time in the asynchronous queue and / or the external call time are removed from the complete service RT of each prevention and control request, and the first round-trip time RT of each prevention and control request is determined.
[0052] Optionally, the acquisition unit is further used for:
[0053] Acquire historical prevention and control request data of a historical setting period, wherein the historical prevention and control request data includes a first round trip time RT of each historical prevention and control request;
[0054] The historical prevention and control request data is stress tested to generate set thresholds corresponding to the first RT range and multiple content features.
[0055] In a third aspect, an embodiment of the present invention provides an electronic device, comprising a memory and a processor, the memory being used to store one or more computer program instructions, wherein the one or more computer program instructions are executed by the processor to implement a method as described in the first aspect or any possible embodiment of the first aspect.
[0056] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium storing computer program instructions, which, when executed by a processor, implement a method as described in the first aspect or any possible embodiment of the first aspect.
[0057] In an embodiment of the present invention, by obtaining the first round-trip time RT of at least one prevention and control request within a set time period, wherein the content feature in each of the prevention and control requests is less than its corresponding set threshold; determining at least one first RT mean of multiple prevention and control requests in the target queue within the set time period; in response to the first RT mean being greater than the maximum value in the first preset RT range, and the asynchronous queue being at full load, determining that the basic environment of the current prevention and control service is abnormal, wherein the first preset RT range is generated based on stress testing, and the basic environment is the hardware environment of the device running the prevention and control algorithm. Through the above method, the cause of the abnormality of the prevention and control service can be directly located through the fluctuation of RT and the load of the asynchronous queue, that is, the root cause affecting the SLA can be located. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] The above and other objects, features and advantages of the present invention will become more apparent through the following description of the embodiments of the present invention with reference to the accompanying drawings, in which:
[0059] Figure 1 is a schematic diagram of a prevention and control system in an embodiment of the present invention;
[0060] Figure 2 is a flow chart of a method for root cause location in an embodiment of the present invention;
[0061] Figure 3 is a flow chart of another method for root cause location in an embodiment of the present invention;
[0062] Figure 4 is a schematic diagram of another prevention and control system in an embodiment of the present invention;
[0063] Figure 5 is a flow chart of another method for root cause location in an embodiment of the present invention;
[0064] Figure 6 is a flow chart of another method for root cause location in an embodiment of the present invention;
[0065] Figure 7is a flow chart of a method for root cause location in an embodiment of the present invention;
[0066] Figure 8 is a schematic diagram of a root cause location device in an embodiment of the present invention;
[0067] Fig. 9 is a schematic diagram of a root cause location device in an embodiment of the present invention;
[0068] Fig.10 is a schematic diagram of an electronic device in an embodiment of the present invention. DETAILED DESCRIPTION
[0069] The present application is described below based on embodiments, but the present application is not limited to these embodiments. In the detailed description of the present application below, some specific details are described in detail. It is possible for those skilled in the art to fully understand the present application without the description of these details. In order to avoid confusing the essence of the present application, known methods, processes, flows, components and circuits are not described in detail.
[0070] In addition, persons of ordinary skill in the art will appreciate that the drawings provided herein are for illustration purposes and are not necessarily drawn to scale.
[0071] Unless the context clearly requires otherwise, the words "include", "comprising" and similar words throughout the application should be interpreted as including rather than exclusive or exhaustive; that is, the meaning is "including but not limited to".
[0072] In the description of this application, it should be understood that the terms "first", "second", etc. are only used for descriptive purposes and cannot be understood as indicating or implying relative importance. In addition, in the description of this application, unless otherwise specified, the meaning of "plurality" is two or more.
[0073] In the existing technology, the prevention and control service platform can implement various content prevention and control services for various business parties, including long text services, short text services, object storage service (OSS) scanning prevention and control services, large model-derived text and image services, large model-derived text and text services, pornography detection, optical character recognition (OCR) and other services. Recognition, OCR), etc. The characteristics of the business prevention and control content of different business parties under the same content prevention and control business are different, and the peak traffic time of the business parties has their own rules. As a result, in some scenarios, a prevention and control algorithm needs to be supplied to multiple business parties at the same time, and the performance of the prevention and control algorithm fluctuates. With the gradual promotion of SLA, the current prevention and control service is no longer satisfied with the monthly SLA standard, but has begun to pay attention to the success rate and RT jitter at the minute level. The above RT jitter can represent SLA jitter. The prevention and control service sets the service provision capability standard of SLA as a benchmark through stress testing, and then determines whether the SLA fluctuation is caused by content problems by monitoring the change trend of content and the SLA jitter trend, and then confirms whether the current jitter is related to the basic environment by enumerating basic environment problems. Since stress testing is not necessarily accurate, it will be very difficult to conduct continuous analysis in the case of inaccuracy. Multiple influencing factors are intertwined and difficult to distinguish, resulting in the guarantee of SLA relying on human judgment and analysis, which is time-consuming and labor-intensive.
[0074] The above-mentioned influencing factors mainly include content fluctuations, basic environment and load changes. The existing technology uses a positive monitoring method to determine the root cause affecting SLA. First, it is assumed that the stress test is accurate, and the load change is determined on this basis. However, although the stress test is accurate at the moment, due to the unconfirmability of the business and the continuous access to new businesses by the prevention and control service platform, the stress test will gradually become inaccurate, resulting in inaccurate monitoring of load changes; secondly, the prevention and control content is segmented. For example, the image category is divided into 0-0.5M, 0.5M-1M, 1M-2M, 2M-5M, 5M-10M, 10M+; the number of optical character recognition words is divided into 0-300 words, 300-1500 words, 1500-3000 words, 3000-5000 words, and 5000 words+; from the perspective of the gateway, the traffic changes and RT changes of each segment are counted, and the traffic changes between each segment are used to confirm whether the content change causes the problem. The abnormalities detected after the prevention and control content is segmented are handled by the business side, while the abnormalities of the basic environment are handled by the prevention and control service platform. When content fluctuations are detected in the content segment and abnormal points are detected in the basic environment, it is difficult to determine the main cause. When ambiguity occurs, the prevention and control service platform needs to further analyze it, and it is difficult to release the manpower of the prevention and control service platform; thirdly, by enumerating known basic environment problems, it is confirmed whether there are basic environment problems at present. Among them, the above-mentioned basic environment problems include disk, memory, jbd2 lock and bandwidth, etc., but the basic environment problems are difficult to enumerate, and new problems often arise that are not detected by the prevention and control service platform. If there are some content fluctuations during the same period, the prevention and control service platform mistakenly judges it as content fluctuations, and the prevention and control service platform finally finds that it is a problem with the basic environment. The trust between the prevention and control service platform and the business side will continue to decrease, which is not conducive to the sustainable development of the prevention and control service platform.
[0075] The above-mentioned positive monitoring method is based on inaccurate stress testing. Multiple factors such as basic environment, content fluctuations and load changes act simultaneously, making it difficult to evaluate the main cause (also known as the root cause, referred to as the root cause). As a result, the prevention and control service platform needs to frequently provide a safety net, wasting human resources. Therefore, how to locate the root cause that affects SLA is a problem that needs to be solved at present.
[0076] The above-mentioned stress test is a type of performance test, and its full name is stress testing. Since the operation of the prevention and control algorithm relies on hardware, stress testing is used to confirm the correlation between the query rate per second (QPS) and RT provided by a single hardware on the current prevention and control algorithm service. According to the characteristics of the business prevention and control content, a suitable QPS is selected as the QPS that the current prevention and control service can provide services on the current hardware. The QPS can also refer to the number of requests that the system can process per second.
[0077] In the prior art, when calculating RT, a prevention and control system is proposed, which involves the business party, gateway and prevention and control service platform. The specific prevention and control system is as follows: Figure 1 As shown, the business party 101 inputs the prevention and control request into the asynchronous queue of the prevention and control service platform 103 through the gateway 102. In the cloud environment of the prevention and control service platform 103, the DAG worker engine obtains the prevention and control requests in the asynchronous queue in sequence and performs calculations. When the DAG worker engine calculates the prevention and control request, it includes the calculation time of the algorithm model and OP1, and also includes the time of calling the dependent service through httpOP. In addition, the overall service RT of the prevention and control request is the sum of the calculation time, the dependent service time (also called the external call time) and the waiting time in the asynchronous queue. When the prevention and control service platform finds that the overall service RT fluctuates, it is necessary to locate the root cause of the fluctuation of the overall service RT.
[0078] The asynchronous queue is designed because the business traffic has second-level jitter, that is, the number of prevention and control situations handled by the prevention and control service platform per second is different. In order to enhance the prevention and control capability of the prevention and control service, an asynchronous queue of limited length is designed on the prevention and control service platform to handle the prevention and control requests of the business party. Subsequently, the DAG worker engine obtains the prevention and control requests from the asynchronous queue for calculation; wherein, the DAG worker engine is deployed in the cloud environment of the prevention and control service platform 103, and the physical machine of the cloud environment can be abstracted as an elastic computing service (Elastic Compute Service, ECS) to provide hardware services to the outside world, and the prevention and control algorithm is deployed in the ECS in the form of a container. For the container, the ECS is the host machine.
[0079] In order to solve the above problems, a method for locating the root cause is proposed in the embodiment of the present invention. Figure 2 As shown, the method includes:
[0080] Step S201: Obtain a first round-trip time RT of at least one prevention and control request within a set time period.
[0081] Specifically, the content feature in each of the prevention and control requests is less than its corresponding set threshold.
[0082] In one possible implementation, the set time period can be a few seconds, a few minutes, etc., which is determined according to the actual situation. That is, the business party inputs multiple prevention and control requests to the asynchronous queue clock through the gateway within 5 minutes, and adds the prevention and control requests whose content characteristics are less than the corresponding set thresholds to the target queue, and determines the first round-trip time RT of each prevention and control request added to the target queue.
[0083] In an embodiment of the present invention, obtaining the first round-trip time RT of at least one prevention and control request within a set time period specifically includes: obtaining the complete service RT of each prevention and control request in at least one prevention and control request within the set time period; removing the waiting time in the asynchronous queue and / or the external call time from the complete service RT of each prevention and control request, and determining the first round-trip time RT of each of the prevention and control requests.
[0084] In a possible implementation, each content feature has a corresponding set threshold, and only the control request whose content feature is less than the corresponding set threshold can be added to the target queue. For example, the content feature in the control request is image size, and only the control request whose image size is less than a can be added to the target queue; the content feature in the control request is pixel value, and only the control request whose image value is less than b can be added to the target queue; the content feature in the control request is text length, and only the control request whose text length is less than c can be added to the target queue; the content feature in the control request is the number of OCR boxes, and only the control request whose OCR box number is less than d can be added to the target queue; the content feature in the control request is the number of OCR words, and only the control request whose OCR word number is less than e can be added to the target queue; the content feature in the control request is the number of flag boxes, and only the control request whose flag box number is less than f can be added to the target queue.
[0085] In an embodiment of the present invention, the set threshold is generated through stress testing. Specifically, historical prevention and control request data of a historical set time period is obtained, wherein the historical prevention and control request data includes the first round-trip time RT of each historical prevention and control request; the historical prevention and control request data is stress tested to generate the first RT range and the set thresholds corresponding to multiple content features respectively; the first RT range is (RT1, RT2), and the set threshold is the above-mentioned a, b, c, d, e, f and g, which is only for exemplary description here; when the content characteristics of each prevention and control request in the target queue are limited to the above-mentioned a, b, c, d, e, f and g conditions, no matter how the overall service RT fluctuates, the fluctuation of the RT of the target queue will not exceed 5%, and the first RT range is used as a benchmark, which is only for exemplary description here, and the specific value is determined according to the actual situation.
[0086] Step S202: Determine at least one first RT average of multiple prevention and control requests in the target queue within the set time period.
[0087] Assume that the target queue includes N prevention and control requests, each prevention and control request corresponds to a first RT, determine the sum of the N first RTs, and determine the ratio of the sum to N as the first RT average, thereby determining multiple first RT averages within the set time period.
[0088] Step S203: In response to the first RT mean being greater than the maximum value in the first RT range and the asynchronous queue being in a full load state, it is determined that the basic environment of the current prevention and control service is abnormal.
[0089] Among them, the first RT range is generated according to the stress test, and the basic environment is the hardware environment of the device running the prevention and control algorithm.
[0090] In a possible implementation, each of the first RT mean values is greater than the maximum value in the first RT range, and the asynchronous queue is in a full-load state, indicating that the RT of the current prevention and control service platform fluctuates, that is, it is necessary to locate the specific cause of the fluctuation. Since the target queue has been circled for all content features of the prevention and control content during stress testing, and the conditions are restricted by the content features, the conclusion of the stress testing is stable and will not change with changes in the business; and the number of threads in the prevention and control algorithm service is set with an upper limit in the prevention and control system, that is, the number of worker engines is set with an upper limit, and the context switching time of the central processing unit (CPU) is also upper bounded. As the traffic of the business party increases, the target queue will not increase the time consumed in the system context, which can further ensure the stability of the RT of the target queue. The RT of the target queue excludes the time consumed by external calls and the waiting time of the asynchronous queue; in addition, when the same host machine provides prevention and control services to different business parties, the hardware standard will not change. In summary, if the RT of the target queue fluctuates, the main reason is the abnormality of the basic environment, that is, the hardware abnormality of the host machine.
[0091] In a possible implementation, after step S203, other steps are also included, such as Figure 3 As shown, specifically including:
[0092] Step S204: Generate and send a basic environment exception code.
[0093] Specifically, the prevention and control service platform generates a basic environment exception code and sends the basic environment exception code to the business party.
[0094] In the embodiment of the present invention, a prevention and control system is proposed, involving a business party, a gateway and a prevention and control service platform. The specific prevention and control system is as follows: Figure 4As shown, the business party 101 inputs the prevention and control request into the asynchronous queue of the prevention and control service platform 103 through the gateway 102. In the cloud environment of the prevention and control service platform 103, the DAG worker engine obtains the prevention and control request in the asynchronous queue in sequence and performs calculations. When the DAG worker engine calculates the prevention and control request, it includes the calculation time of the algorithm model and OP1, and also includes the time of calling the dependent service through httpOP. In addition, the overall service RT of the prevention and control request is the sum of the calculation time, the dependent service time and the waiting time in the asynchronous queue. In order to obtain a more reasonable RT range, the prevention and control requests whose content features are less than the set threshold are added to the target queue 104 for time statistics. After being added to the target queue 104, the RT of the prevention and control request is only the calculation time, excluding the dependent service time and the waiting time in the asynchronous queue.
[0095] In the embodiment of the present invention, if the RT of the target queue does not fluctuate significantly, after excluding the abnormality of the basic environment, it can be further determined whether the business traffic is abnormal or whether the prevention and control content is abnormal.
[0096] In a possible implementation, after step S202, other steps are also included, such as Figure 5 As shown, specifically including:
[0097] Step S205: In response to the first RT mean being within the first RT range and the asynchronous queue being in a full load state, obtain the traffic of the asynchronous queue.
[0098] Step S206: In response to the traffic being greater than a set traffic threshold, determining that the business traffic of the current prevention and control service is abnormal.
[0099] Specifically, the prevention and control service platform will pre-set a traffic threshold for the business party, and the set traffic threshold is the QPS agreed upon by the business party's single instance (also called a single business) in the SLA contract; when the first RT mean is within the first RT range, it means that the RT fluctuation is normal, but the asynchronous queue is at full load and cannot accept new prevention and control requests, indicating that the traffic input by the business party exceeds the agreed QPS, and the prevention and control service platform determines that the business party's business traffic is abnormal.
[0100] Step S207: Generate and send a service flow exception code.
[0101] Specifically, the prevention and control service platform generates a business traffic anomaly code and sends the business traffic anomaly code to the business party.
[0102] In a possible implementation, after step S205, other steps are also included, such as Figure 6 As shown, specifically including:
[0103] Step S208: In response to the flow being less than or equal to the set flow threshold, determining that the prevention and control content of the current prevention and control service is abnormal.
[0104] Among them, the abnormal prevention and control content indicates that the business side content fluctuation leads to performance degradation, and the current first RT range cannot meet the demand.
[0105] Specifically, when the first RT mean is within the first RT range, it means that the RT fluctuation is normal, but the asynchronous queue is at full load and cannot accept new prevention and control requests, and the traffic input by the business party does not exceed the agreed QPS, which means that the content of each prevention and control request has changed. Assume that, under normal circumstances, the OCR word count of the content in each prevention and control request input by the business party is 100, but it suddenly increases to 5,000 words. Although the number of prevention and control requests has not increased, the time for the prevention and control service platform to process prevention and control requests each day has increased, resulting in the accumulation of prevention and control requests in the asynchronous queue, and then determining that the prevention and control content of the business party is abnormal.
[0106] Step S209: Generate and send a performance degradation exception code.
[0107] Specifically, the prevention and control service platform generates a business traffic anomaly code and sends the business traffic anomaly code to the business party.
[0108] In one possible implementation, in the above situation, if the performance degradation exception code is sent for a long time, it is necessary to re-perform the stress test service and adjust the QPS agreed by the business party for a single instance. The business party can be advised to increase the processing container to meet the demand.
[0109] In a possible implementation, after step S202, other steps are also included, such as Figure 7 As shown, specifically including:
[0110] Step S210, in response to the second RT mean within the set time period being greater than the maximum value of the preset second RT range, and the first RT mean being greater than the maximum value in the first RT range, and the asynchronous queue being in a normal load state, determining that the basic environment of the current prevention and control algorithm is abnormal.
[0111] Among them, the second RT average is the average of the complete service RT of multiple prevention and control requests within the set time period, and the second RT range is pre-set.
[0112] In the embodiment of the present invention, the prevention and control service platform records basic environment anomalies.
[0113] In a possible implementation, after step S202, other steps are also included, such as Figure 8 As shown, specifically including:
[0114] Step S211, in response to the second RT mean within the set time period being greater than the maximum value of the preset second RT range, and the first RT mean being within the first RT range, and the asynchronous queue being in a normal load state, determining that the prevention and control content of the current prevention and control algorithm is abnormal.
[0115] Among them, the second RT average is the average of the complete service RT of multiple prevention and control requests within the set time period, and the second RT range is pre-set.
[0116] In the embodiment of the present invention, the prevention and control service platform records the abnormality of the prevention and control content.
[0117] Through the above embodiments and the above methods, the root cause of RT fluctuation can be accurately determined under various conditions.
[0118] In an embodiment of the present invention, a root cause location device is provided, such as Fig. 9 As shown, it specifically includes: an acquisition unit 901, a determination unit 902 and a judgment unit 903;
[0119] Among them, the acquisition unit 901 is used to obtain the first round-trip time RT of at least one prevention and control request within a set time period, wherein the content feature in each of the prevention and control requests is less than its corresponding set threshold; the determination unit 902 is used to determine at least one first RT mean of multiple prevention and control requests in the target queue within the set time period; the judgment unit 903, in response to the first RT mean being greater than the maximum value in the first RT range and the asynchronous queue being in a full load state, is used to determine that the basic environment of the current prevention and control service is abnormal, wherein the first RT range is generated based on a stress test, and the basic environment is the hardware environment of the device running the prevention and control algorithm.
[0120] Furthermore, the device also includes a processing unit, which is used to generate and send a basic environment exception code.
[0121] Furthermore, the determination unit is also used for:
[0122] In response to the first RT mean being within the first RT range and the asynchronous queue being in a full load state, obtaining traffic of the asynchronous queue;
[0123] In response to the traffic being greater than a set traffic threshold, it is determined that the business traffic of the current prevention and control service is abnormal.
[0124] Furthermore, the processing unit is also used for:
[0125] Generate and send service traffic exception codes.
[0126] Furthermore, the determination unit is also used for:
[0127] In response to the traffic being less than or equal to the set traffic threshold, it is determined that the prevention and control content of the current prevention and control service is abnormal, wherein the prevention and control content abnormality indicates that performance degradation is caused by fluctuations in the business party content, and the current first RT range cannot meet the demand.
[0128] Furthermore, the processing unit is also used for:
[0129] Generates and sends a performance degradation exception code.
[0130] Furthermore, the determination unit is also used for:
[0131] In response to the second RT mean within the set time period being greater than the maximum value of the preset second RT range, and the first RT mean being greater than the maximum value in the first RT range, and the asynchronous queue being in a normal load state, the basic environment of the current prevention and control algorithm is determined to be abnormal, wherein the second RT mean is the average of the complete service RTs of multiple prevention and control requests within the set time period, and the second RT range is pre-set.
[0132] Furthermore, the determination unit is also used for:
[0133] In response to the second RT mean within the set time period being greater than the maximum value of the preset second RT range, and the first RT mean being within the first RT range, and the asynchronous queue being in a normal load state, it is determined that the prevention and control content of the current prevention and control algorithm is abnormal, wherein the second RT mean is the mean of the complete service RT of multiple prevention and control requests within the set time period, and the second RT range is pre-set.
[0134] Furthermore, the acquisition unit is specifically used for:
[0135] Get the complete service RT of each control request in at least one control request within the set time period;
[0136] The waiting time in the asynchronous queue and / or the external call time are removed from the complete service RT of each prevention and control request, and the first round-trip time RT of each prevention and control request is determined.
[0137] Furthermore, the acquisition unit is also used for:
[0138] Acquire historical prevention and control request data of a historical setting period, wherein the historical prevention and control request data includes a first round trip time RT of each historical prevention and control request;
[0139] The historical prevention and control request data is stress tested to generate set thresholds corresponding to the first RT range and multiple content features.
[0140] Fig.10Schematic diagram of the structure of the electronic device in the embodiment of the present invention. Fig.10 As shown, it includes a general computer hardware structure, which at least includes a processor 1001 and a memory 1002. The processor 1001 and the memory 1002 are connected via a bus 1003. The memory 1002 is suitable for storing instructions or programs executable by the processor 1001. The processor 1001 can be an independent microprocessor or a collection of one or more microprocessors. Thus, the processor 1001 executes the instructions stored in the memory 1002, thereby executing the method flow of the embodiment of the present invention as described above to realize the processing of data and the control of other devices. The bus 1003 connects the above-mentioned multiple components together, and at the same time connects the above-mentioned components to the display controller 1004 and the display device and the input / output (I / O) device 1005. The input / output (I / O) device 1005 can be a mouse, a keyboard, a modem, a network interface, a touch input device, a somatosensory input device, a printer, and other devices known in the art. Typically, the input / output device 1005 is connected to the system through an input / output (I / O) controller 1006.
[0141] Among them, the instructions stored in the memory 1002 are executed by at least one processor 1001 to achieve: obtaining the first round-trip time RT of at least one prevention and control request within a set time period, wherein the content feature in each of the prevention and control requests is less than its corresponding set threshold; determining at least one first RT mean of multiple prevention and control requests in the target queue within the set time period; in response to the first RT mean being greater than the maximum value in the first RT range and the asynchronous queue being in a full load state, determining that the basic environment of the current prevention and control service is abnormal, wherein the first RT range is generated based on a stress test, and the basic environment is the hardware environment of the device running the prevention and control algorithm.
[0142] Specifically, the electronic device includes: one or more processors 1001 and a memory 1002, Fig.10 Take a processor 1001 as an example. The processor 1001 and the memory 1002 may be connected via a bus or other means. Fig.10 In the example, the bus connection is used. The memory 1002 is a non-volatile computer-readable storage medium that can be used to store non-volatile software programs, non-volatile computer executable programs and modules. The processor 1001 executes various functional applications and data processing of the device by running the non-volatile software programs, instructions and modules stored in the memory 1002, that is, the above-mentioned method for determining the root cause location is implemented.
[0143] The memory 1002 may include a program storage area and a data storage area, wherein the program storage area may store an operating system and applications required for at least one function; the data storage area may store a list of options, etc. In addition, the memory 1002 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other non-volatile solid-state storage device. In some embodiments, the memory 1002 may optionally include a memory remotely arranged relative to the processor 1001, and these remote memories may be connected to an external device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0144] One or more modules are stored in the memory 1002, and when executed by one or more processors 1001, the root cause location method in any of the above method embodiments is executed.
[0145] As will be appreciated by those skilled in the art, various aspects of the embodiments of the present invention may be implemented as a system, method or computer program product. Therefore, various aspects of the embodiments of the present invention may take the form of a complete hardware implementation, a complete software implementation (including firmware, resident software, microcode, etc.), or an implementation that combines software aspects with hardware aspects, which may generally be referred to herein as a "circuit," "module," or "system." In addition, various aspects of the embodiments of the present invention may take the form of a computer program product implemented in one or more computer-readable media having a computer-readable program code implemented thereon.
[0146] Any combination of one or more computer-readable media can be utilized. Computer-readable media can be computer-readable signal media or computer-readable storage media. Computer-readable storage media can be, for example, (but not limited to) electronic, magnetic, optical, electromagnetic, infrared or semiconductor systems, devices or apparatuses, or any suitable combination of the foregoing. More specific examples (non-exhaustive enumeration) of computer-readable storage media will include the following: an electrical connection with one or more wires, a portable computer floppy disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing. In the context of an embodiment of the present invention, a computer-readable storage medium can be any tangible medium that can contain or store a program used by an instruction execution system, device or apparatus or a program used in conjunction with an instruction execution system, device or apparatus.
[0147] A computer readable signal medium may include a propagated digital signal having a computer readable program code implemented therein, such as in baseband or as part of a carrier wave. Such propagated signals may take any of a variety of forms, including but not limited to electromagnetic, optical, or any suitable combination thereof. A computer readable signal medium may be any of the following computer readable media: not a computer readable storage medium, and may communicate, propagate, or transmit a program used by or in conjunction with an instruction execution system, device, or apparatus.
[0148] Program code embodied on a computer readable medium may be transmitted using any appropriate medium including, but not limited to, wireless, wireline, optical fiber cable, RF, etc., or any suitable combination of the foregoing.
[0149] The computer program code for performing operations for various aspects of the embodiments of the present invention can be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, Smalltalk, C++, etc., and conventional process programming languages such as "C" programming language or similar programming languages. The program code can be executed completely on the user's computer, partially on the user's computer as a stand-alone software package; partially on the user's computer and partially on a remote computer; or completely on a remote computer or server. In the latter case, the remote computer can be connected to the user's computer through any type of network including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (e.g., by using the Internet of an Internet service provider).
[0150] The flowchart legend and / or block diagram of the method, device (system) and computer program product according to the embodiment of the present invention described above describe various aspects of the embodiment of the present invention.It will be understood that each block of the flowchart legend and / or block diagram and the combination of the blocks in the flowchart legend and / or block diagram can be implemented by computer program instructions.These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer or other programmable data processing device to produce a machine so that the instructions (executed by the processor of the computer or other programmable data processing device) create a device for implementing the function / action specified in the flowchart and / or block diagram block or block.
[0151] These computer program instructions may also be stored in a computer-readable medium that can direct a computer, other programmable data processing device, or other apparatus to operate in a particular manner, so that the instructions stored in the computer-readable medium produce an article of manufacture including instructions for implementing the functions / actions specified in the flowchart and / or block diagram blocks or blocks.
[0152] The computer program instructions may also be loaded onto a computer, other programmable data processing device or other apparatus so that a series of operable steps are performed on the computer, other programmable device or other apparatus to produce a computer-implemented process, so that the instructions executed on the computer or other programmable device provide a process for implementing the functions / actions specified in the flowchart and / or block diagram block or blocks.
[0153] The above description is only a preferred embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
[0154] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation portals for users to choose to authorize or refuse. Users' refusal to process personal information other than the necessary information required for basic functions will not affect the user's use of basic functions.
Claims
1. A method for locating a root cause, characterized in that: The method comprises: Obtaining a first round-trip time RT of at least one prevention and control request within a set period of time, wherein a content feature in each of the prevention and control requests is less than a corresponding set threshold; Determine at least one first RT average of multiple prevention and control requests in the target queue within the set time period; In response to the first RT mean being greater than the maximum value in the first RT range and the asynchronous queue being in a full load state, it is determined that the basic environment of the current prevention and control service is abnormal, wherein the first RT range is generated based on a stress test and the basic environment is the hardware environment of the device running the prevention and control algorithm.
2. The method according to claim 1, characterized in that The method further comprises: Generates and sends basic environment exception codes.
3. The method according to claim 1, characterized in that The method further comprises: In response to the first RT mean being within the first RT range and the asynchronous queue being in a full load state, obtaining traffic of the asynchronous queue; In response to the traffic being greater than a set traffic threshold, it is determined that the business traffic of the current prevention and control service is abnormal.
4. The method according to claim 3, characterized in that The method further comprises: Generate and send service traffic exception codes.
5. The method according to claim 3, characterized in that: The method further comprises: In response to the traffic being less than or equal to the set traffic threshold, it is determined that the prevention and control content of the current prevention and control service is abnormal, wherein the prevention and control content abnormality indicates that performance degradation is caused by fluctuations in the business party content, and the current first RT range cannot meet the demand.
6. The method according to claim 5, characterized in that The method further comprises: Generates and sends a performance degradation exception code.
7. The method according to claim 1, characterized in that The method further comprises: In response to the second RT mean within the set time period being greater than the maximum value of the preset second RT range, and the first RT mean being greater than the maximum value in the first RT range, and the asynchronous queue being in a normal load state, the basic environment of the current prevention and control algorithm is determined to be abnormal, wherein the second RT mean is the average of the complete service RTs of multiple prevention and control requests within the set time period, and the second RT range is pre-set.
8. The method according to claim 1, characterized in that The method further comprises: In response to the second RT mean within the set time period being greater than the maximum value of the preset second RT range, and the first RT mean being within the first RT range, and the asynchronous queue being in a normal load state, it is determined that the prevention and control content of the current prevention and control algorithm is abnormal, wherein the second RT mean is the mean of the complete service RT of multiple prevention and control requests within the set time period, and the second RT range is pre-set.
9. The method according to claim 1, characterized in that: The obtaining of a first round trip time RT of at least one prevention and control request within a set period specifically includes: Get the complete service RT of each control request in at least one control request within the set time period; The waiting time in the asynchronous queue and / or the external call time are removed from the complete service RT of each prevention and control request, and the first round-trip time RT of each prevention and control request is determined.
10. The method according to claim 9, characterized in that The method comprises: Acquire historical prevention and control request data of a historical setting period, wherein the historical prevention and control request data includes a first round trip time RT of each historical prevention and control request; The historical prevention and control request data is stress tested to generate set thresholds corresponding to the first RT range and multiple content features.
11. A root cause location device, characterized in that: The device comprises: An acquisition unit, configured to acquire a first round trip time RT of at least one prevention and control request within a set period of time, wherein a content feature in each of the prevention and control requests is less than a corresponding set threshold; A determination unit, configured to determine at least one first RT mean of a plurality of prevention and control requests in a target queue within the set time period; A determination unit, in response to the first RT mean being greater than the maximum value in the first RT range and the asynchronous queue being in a full load state, is used to determine that a basic environment of the current prevention and control service is abnormal, wherein the first RT range is generated based on a stress test and the basic environment is the hardware environment of a device running a prevention and control algorithm.
12. An electronic device comprising a memory and a processor, characterized in that: The memory is used to store one or more computer program instructions, wherein the one or more computer program instructions are executed by the processor to implement the method according to any one of claims 1-10.
13. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 10 is implemented.