Monitoring Method and Device for Distributed Processing System

By monitoring the operator data processing rate in the distributed processing system in real time, and detecting and processing abnormal operators in a timely manner, the problem of abnormal detection delay in the prior art is solved, and the operational and maintenance of the system is improved.

CN114896121BActive Publication Date: 2025-05-27HANGZHOU DT DREAM TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202210615433.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-31
Publication Date
2025-05-27
Estimated Expiration
2042-05-31

AI Technical Summary

Technical Problem

In the existing real-time computing distributed system, there is a delay in operator abnormal detection in the Flink platform, which makes it impossible to detect and process fault operators in time, affecting the operational and maintenance of the system.

Method used

By obtaining the data processing rate corresponding to the operator, and determining the operating status of the operator based on the data processing rate of the operator and its upstream operator, monitoring and alerting the abnormal operator in real time.

Benefits of technology

Real-time monitoring of distributed processing systems is realized, abnormal operators are discovered and processed in a timely manner, and the robustness and operationality of the system are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114896121B_ABST
    Figure CN114896121B_ABST
Patent Text Reader

Abstract

This specification provides a monitoring method and device for a distributed processing system, where operators are set on nodes of the distributed processing system. The method includes: for a distributed task, obtaining the data processing rate corresponding to the operator; determining at least one working state of the operator according to the data processing rate corresponding to the operator and / or its upstream operator; and in response to any dimension of the working state of the operator indicating that the operator is abnormal, performing a corresponding alarm operation to achieve timely discovery of abnormal operators.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This specification relates to the field of distributed technologies, and particularly to a method and apparatus for monitoring a distributed processing system. Background Art

[0002] With the popularization of big data technology in all walks of life, the value generated by data is becoming increasingly important to customers. In some fields, the hourly and daily delays of large-scale offline computing are not sufficient to support the timeliness of business, and customers are increasingly concerned about the real-time nature of data. After several generations of evolution of real-time computing distributed technologies, from Storm, Spark Streaming to Flink, they have achieved mature development in terms of low latency, high throughput, strong consistency, etc.

[0003] Currently, while the popularization of real-time computing distributed technologies brings timeliness to customer data analysis, since operators in the Flink platform often encounter anomalies, it also poses new pressure on the data operation and maintenance of the platform. Flink provides some operation and maintenance monitoring metrics internally to assist in determining faulty operators, but problems are often discovered with a certain delay, resulting in the inability to promptly detect abnormal operators. Summary of the Invention

[0004] To overcome the problems existing in the related art, this specification provides a method and apparatus for monitoring a distributed processing system.

[0005] According to a first aspect of an embodiment of this specification, a method for monitoring a distributed processing system is provided, where the distributed processing system includes operators with an upstream and downstream relationship;

[0006] The method includes:

[0007] For a distributed task, obtain the data processing rate corresponding to the operator;

[0008] Determine at least one working state of the operator according to the data processing rate corresponding to the operator and / or its upstream operator;

[0009] In response to any working state of the operator indicating that the operator is abnormal, perform a corresponding warning operation.

[0010] Optionally, the operator corresponds to at least one operator instance, and each operator instance has a corresponding data processing rate;

[0011] The determining at least one working state of the operator according to the data processing rate corresponding to the operator and / or and / or its upstream operator includes:

[0012] Determine at least one working state of the first operator according to the error value between the data processing rates corresponding to each operator instance of the first operator and / or the second operator; wherein, the first operator is any one of the operators; the second operator is the upstream operator of the first operator.

[0013] Optionally, the data processing rate includes a first data production rate; wherein, the first data production rate represents the rate at which the operator produces data; the working state of the first operator includes a first working state.

[0014] Wherein, the first working state of the first operator indicates whether the data produced by the upstream operator corresponding to the first operator is evenly distributed.

[0015] The determining of at least one working state of the first operator according to the error value between the data processing rates corresponding to each operator instance of the first operator and / or the second operator includes:

[0016] Calculate a first error value between any two of the first data production rates corresponding to each operator instance of the second operator.

[0017] Determine the first working state of the first operator according to the first error value.

[0018] Optionally, the determining of the first working state of the first operator according to the first error value includes:

[0019] When there is a first error value reaching a first preset value, determine that the first working state of the first operator indicates that the data produced by the upstream operator corresponding to the first operator is unevenly distributed.

[0020] When all the first error values do not reach the first preset value, determine that the first working state of the first operator indicates that the data produced by the upstream operator corresponding to the first operator is evenly distributed.

[0021] Optionally, the determining that when there is a first error value reaching a first preset value, the first working state of the first operator indicates that the data produced by the upstream operator corresponding to the first operator is unevenly distributed includes:

[0022] Obtain the expansion rate corresponding to the output buffer of the second operator; wherein, the output buffer is used to store the data produced by the second operator.

[0023] When there is a first error value reaching the first preset value and the expansion rate reaches the first preset rate, determine that the first working state of the first operator indicates that the data produced by the upstream operator corresponding to the first operator is unevenly distributed.

[0024] Optionally, for any working state of the operator indicating that the operator has an abnormality, corresponding alarm operations are performed, including:

[0025] When the first working state of the first operator indicates that the data produced by the upstream operator corresponding to the first operator is unevenly distributed, a first alarm message is output; wherein the first alarm message is used to prompt to increase the number of downstream operator concurrency degrees.

[0026] Optionally, for any working state of the operator indicating that the operator has an abnormality, corresponding alarm operations are performed, including:

[0027] When the first working state of the first operator indicates that the data produced by the upstream operator corresponding to the first operator is unevenly distributed, based on the execution process corresponding to the distributed task, each third operator above the first operator is determined;

[0028] Obtain the first working state corresponding to each third operator respectively;

[0029] When the first working state corresponding to each third operator respectively indicates that the data produced by the upstream operator is unevenly distributed, other downstream operators corresponding to the source operator of the distributed task except the first downstream operator are determined; wherein, the first downstream operator is the downstream operator of the source operator and is the third operator;

[0030] Alarm operations are performed according to the data consumption rate corresponding to other downstream operators.

[0031] Optionally, the upstream operator and the downstream operator are connected through a channel;

[0032] The alarm operations performed according to the data consumption rate corresponding to other downstream operators include:

[0033] Determine a second error value between the data consumption rate corresponding to the other downstream operator and the first data production rate corresponding to the source operator;

[0034] When the second error value reaches a second preset value, a second alarm message is output; wherein, the second alarm message is used to prompt that there is too much source data and increase the number of downstream operator concurrency degrees.

[0035] Optionally, the upstream operator and the downstream operator are connected through a channel; when the first working state corresponding to each third operator respectively indicates that the data produced by the upstream operator is unevenly distributed, the method further includes:

[0036] Determine a second error value between the data consumption rate corresponding to the other downstream operator and the first data production rate corresponding to the source operator;

[0037] In the case where the second error value does not reach a third preset value, perform a first channel reallocation operation; wherein, the first channel reallocation operation instructs to control the other downstream operator to consume the data in the input buffer corresponding to the first downstream operator; the data in the input buffer is the data produced by the source operator.

[0038] Optionally, in the case where the first working state corresponding to each of the third operators all indicates that the data allocated by the upstream operator is uneven, the method further includes:

[0039] Obtain the minimum value of the data consumption rate corresponding to the other downstream operator and the data consumption rate corresponding to the first downstream operator;

[0040] Generate a speed limit instruction according to the minimum value, and send the speed limit instruction to the source operator, so that the source operator adjusts the first data production rate corresponding to the source operator based on the minimum value.

[0041] Optionally, after obtaining the first working state corresponding to each of the third operators, the method further includes:

[0042] In the case where the first working state of all the third operators all indicates that the data allocated by the upstream operator is even, perform a second channel reallocation operation; wherein, the second channel reallocation operation instructs the normal operator instance to consume the data in the input buffer corresponding to the abnormal operator instance; the normal operator instance is the operator instance in the operator instances corresponding to the first operator whose first working state indicates that the data allocated by the upstream operator is even, and the abnormal operator instance is the operator instance in the operator instances corresponding to the first operator whose first working state indicates that the data allocated by the upstream operator is uneven.

[0043] Optionally, in the case where the first working state of the first operator indicates that the data allocated by the upstream operator corresponding to the first operator is uneven, the method further includes:

[0044] Record the key value corresponding to the data produced by the upstream operator corresponding to the first operator, and output the key value.

[0045] Optionally, the data processing rate includes a data consumption rate; the data consumption rate represents the rate at which an operator consumes the data produced by the upstream operator;

[0046] The working state of the first operator includes a second working state;

[0047] Among them, the second working state of the first operator indicates whether the data consumption ability is normal;

[0048] Determining at least one working state of the first operator according to the error values between the respective data processing rates corresponding to each operator instance corresponding to the first operator and / or the second operator includes:

[0049] Calculating a third error value between any two of the respective data consumption rates corresponding to the first operator;

[0050] Determining the second working state corresponding to each operator instance corresponding to the first operator according to the third error value.

[0051] Optionally, the operator is scheduled to at least one resource group slot; the slot corresponds to the operator instance one by one;

[0052] Performing a corresponding alarm operation in response to any working state of the operator indicating that the operator is abnormal, including:

[0053] In the case where the second working state of the operator instance indicates that the data consumption ability is abnormal, determining the target task manager to which the slot corresponding to the operator instance belongs;

[0054] Determining all the first operator instances included in the target task manager except the operator instance, and obtaining the second working state corresponding to each first operator instance;

[0055] Performing a corresponding alarm operation according to the second working state corresponding to each first operator instance.

[0056] Optionally, performing a corresponding alarm operation according to the second working state corresponding to each first operator instance includes:

[0057] Calculating the ratio of the number of first operator instances whose second working state indicates abnormal data consumption ability to the total number of first operator instances to obtain the abnormal first operator instance ratio;

[0058] Outputting a target task manager failure prompt message when the abnormal first operator instance ratio reaches a first preset ratio;

[0059] Outputting the operator instance abnormal prompt message when the abnormal first operator instance ratio does not reach the first preset ratio.

[0060] Optionally, performing a corresponding alarm operation in response to any working state of the operator indicating that the operator is abnormal, including:

[0061] When the second working state of the operator instance indicates abnormal data consumption ability, output the abnormal prompt information of the operator instance.

[0062] Optionally, the operator corresponds to at least one operator instance; there is a corresponding buffer for the operator instance;

[0063] The method further includes:

[0064] Obtain the expansion rate of the buffer corresponding to each operator instance respectively;

[0065] When the expansion rate of the buffer reaches the second preset rate, control the buffer to stop expanding.

[0066] Optionally, the operator corresponds to at least one operator instance; the method further includes:

[0067] Obtain all data processing rates corresponding to the operator instance obtained within a set time, and obtain the historical average processing rate corresponding to the operator instance;

[0068] If all the data processing rates do not reach the historical average processing rate, output a slow operator alarm message.

[0069] According to the second aspect of the embodiments of the present specification, a monitoring device for a distributed processing system is provided. The distributed processing system includes operators with upstream and downstream relationships;

[0070] The device includes:

[0071] A rate acquisition module, configured to obtain the data processing rate corresponding to the operator for a distributed task;

[0072] A rate processing module, configured to determine at least one working state of the operator according to the data processing rate corresponding to the operator and / or its upstream operator;

[0073] An alarm module, configured to perform a corresponding alarm operation in response to any working state of the operator indicating that the operator is abnormal.

[0074] Optionally, the operator corresponds to at least one operator instance, and each operator instance has a corresponding data processing rate;

[0075] The rate processing module is specifically configured to:

[0076] Determine at least one working state of the first operator according to the error value between the data processing rates corresponding to the respective operator instances corresponding to the first operator and / or the second operator; wherein, the first operator is any operator among the operators; the second operator is the upstream operator of the first operator.

[0077] Optionally, the data processing rate includes a first data production rate; wherein, the first data production rate represents the rate at which an operator produces data; the working state of the first operator includes a first working state;

[0078] Wherein, the first working state of the first operator indicates whether the data allocated by the upstream operator corresponding to the first operator is evenly distributed;

[0079] The rate processing module is specifically configured to:

[0080] Calculate a first error value between any two of the first data production rates respectively corresponding to each operator instance corresponding to the second operator;

[0081] Determine the first working state of the first operator according to the first error value.

[0082] Optionally, the rate processing module is further configured to:

[0083] In the case where a first error value reaches a first preset value, determine that the first working state of the first operator indicates that the data allocated by the upstream operator corresponding to the first operator is unevenly distributed;

[0084] In the case where all first error values do not reach the first preset value, determine that the first working state of the first operator indicates that the data allocated by the upstream operator corresponding to the first operator is evenly distributed.

[0085] Optionally, the rate processing module is further configured to:

[0086] Obtain the expansion rate corresponding to the out-end buffer of the second operator; wherein, the out-end buffer is used to store the data produced by the second operator;

[0087] In the case where a first error value reaches the first preset value and the expansion rate reaches the first preset rate, determine that the first working state of the first operator indicates that the data allocated by the upstream operator corresponding to the first operator is unevenly distributed.

[0088] Optionally, the alarm module is specifically configured to:

[0089] In the case where the first working state of the first operator indicates that the data allocated by the upstream operator corresponding to the first operator is unevenly distributed, output a first alarm message; wherein the first alarm message is used to prompt to increase the number of downstream operator concurrency degrees.

[0090] Optionally, the alarm module is specifically configured to:

[0091] In the case that the first working state of the first operator indicates uneven distribution of data produced by the upstream operator corresponding to the first operator, determine, based on the execution process corresponding to the distributed task, each third operator located above the first operator;

[0092] Obtain the first working state corresponding to each third operator respectively;

[0093] In the case that the first working state corresponding to each third operator respectively indicates uneven distribution of data produced by the upstream operator, determine other downstream operators of the source operator corresponding to the distributed task except the first downstream operator; wherein, the first downstream operator is a downstream operator of the source operator and is the third operator;

[0094] Perform an alarm operation according to the data consumption rate corresponding to the other downstream operators.

[0095] Optionally, the upstream operator and the downstream operator are connected through a channel;

[0096] Optionally, the alarm module is specifically configured to:

[0097] Determine a second error value between the data consumption rate corresponding to the other downstream operators and the first data production rate corresponding to the source operator;

[0098] In the case that the second error value reaches a second preset value, output a second alarm message; wherein, the second alarm message is used to prompt that there is too much data at the source end and increase the number of concurrency degrees of the downstream operator.

[0099] Optionally, the upstream operator and the downstream operator are connected through a channel; the device further includes a first channel processing module;

[0100] The first channel processing module is specifically configured to:

[0101] In the case that the first working state corresponding to each third operator respectively indicates uneven distribution of data produced by the upstream operator, determine a second error value between the data consumption rate corresponding to the other downstream operators and the first data production rate corresponding to the source operator;

[0102] In the case that the second error value does not reach a third preset value, perform a first channel reallocation operation; wherein, the first channel reallocation operation instructs to control the other downstream operators to consume the data in the input buffer corresponding to the first downstream operator; the data in the input buffer is the data produced by the source operator.

[0103] Optionally, the device further includes a speed limit module;

[0104] The speed limit module is specifically configured to:

[0105] When the first working states corresponding to the respective third operators all indicate uneven distribution of data produced by the upstream operator, obtain the minimum value among the data consumption rates corresponding to the other downstream operators and the data consumption rate corresponding to the first downstream operator;

[0106] Generate a speed limit instruction according to the minimum value, and send the speed limit instruction to the source operator, so that the source operator adjusts the first data production rate corresponding to the source operator based on the minimum value.

[0107] Optionally, the device further includes a second channel processing module;

[0108] The second channel processing module is specifically configured to:

[0109] After obtaining the first working states corresponding to the respective third operators, perform a second channel reassignment operation when the first working states of all the third operators all indicate even distribution of data produced by the upstream operator; wherein, the second channel reassignment operation instructs a normal operator instance to consume the data in the input buffer corresponding to an abnormal operator instance; the normal operator instance is an operator instance among the operator instances corresponding to the first operator whose first working state indicates even distribution of data produced by the upstream operator, and the abnormal operator instance is an operator instance among the operator instances corresponding to the first operator whose first working state indicates uneven distribution of data produced by the upstream operator.

[0110] Optionally, the device further includes a data recording module;

[0111] The data recording module is specifically configured to:

[0112] When the first working state of the first operator indicates uneven distribution of data produced by the upstream operator corresponding to the first operator, record the key value corresponding to the data produced by the upstream operator corresponding to the first operator, and output the key value.

[0113] Optionally, the data processing rate includes a data consumption rate; the data consumption rate represents the rate at which an operator consumes data produced by an upstream operator; the working state of the first operator includes a second working state;

[0114] wherein, the second working state of the first operator indicates whether the data consumption ability is normal;

[0115] The rate processing module is specifically configured to:

[0116] Calculate a third error value between any two data consumption rates among the respective data consumption rates corresponding to the first operator;

[0117] Determine the second working state corresponding to each operator instance corresponding to the first operator according to the third error value.

[0118] Optionally, the operator is scheduled to at least one resource group slot; the slot corresponds to the operator instance one by one;

[0119] The alarm module is specifically used for:

[0120] When the second working state of the operator instance indicates an abnormal data consumption ability, determine the target task manager to which the slot corresponding to the operator instance belongs;

[0121] Determine all the first operator instances included in the target task manager except the operator instance, and obtain the second working state corresponding to each first operator instance;

[0122] Perform corresponding alarm operations according to the second working state corresponding to each first operator instance.

[0123] Optionally, the alarm module is further used for:

[0124] Calculate the ratio of the number of first operator instances whose second working state indicates abnormal data consumption ability to the total number of first operator instances to obtain the abnormal first operator instance ratio;

[0125] When the abnormal first operator instance ratio reaches the first preset ratio, output a fault prompt message for the target task manager;

[0126] When the abnormal first operator instance ratio does not reach the first preset ratio, output the abnormal prompt message for the operator instance.

[0127] Optionally, the alarm module is further used for:

[0128] When the second working state of the operator instance indicates abnormal data consumption ability, output the abnormal prompt message for the operator instance.

[0129] Optionally, the operator corresponds to at least one operator instance; the operator instance has a corresponding buffer;

[0130] The alarm module is further used for:

[0131] Obtain the expansion rate of the buffer corresponding to each operator instance;

[0132] When the expansion rate of the buffer reaches the second preset rate, control the buffer to stop expanding.

[0133] Optionally, the operator corresponds to at least one operator instance; the alarm module is further used for:

[0134] Obtain all data processing rates corresponding to the operator instance at the set time, and obtain the historical average processing rate corresponding to the operator instance;

[0135] If all the data processing rates do not reach the historical average processing rate, output a slow operator warning message.

[0136] According to the third aspect of the embodiments of this specification, there is provided a computer-readable storage medium storing computer-executable instructions, and when a processor executes the computer-executable instructions, the monitoring method of the distributed processing system described in the first aspect above and various possible designs of the first aspect is implemented.

[0137] According to the fourth aspect of the embodiments of this specification, there is provided a computer device including a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein when the processor executes the program, the monitoring method of the distributed processing system described in the first aspect above and various possible designs of the first aspect is implemented.

[0138] According to the fifth aspect of the embodiments of this specification, there is provided a computer program product including a computer program, and when the computer program is executed by a processor, the monitoring method of the distributed processing system described in the first aspect above and various possible designs of the first aspect is implemented.

[0139] The technical solutions provided by the embodiments of this specification may include the following beneficial effects:

[0140] In the embodiments of this specification, the distributed processing system includes operators with upstream and downstream relationships. Obtain the data processing rates corresponding to each operator involved in the distributed task. For each operator, determine whether the operator processes data abnormally according to the data processing rates corresponding to the operator and its upstream operator, so as to determine at least one working state of the operator, and discover faulty operators in a timely manner. When a certain working state of the operator indicates that the operator is abnormal, perform corresponding warning operations to ensure the timeliness of the warning, realize the intelligent monitoring of the distributed processing system, and then can solve the abnormality in a timely manner, improving the robustness and operability of the distributed processing system.

[0141] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit this specification. BRIEF DESCRIPTION OF THE DRAWINGS

[0142] The drawings here are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with this specification, and are used together with the specification to explain the principles of this specification.

[0143] Figure 1 This is a flowchart of a monitoring method for a distributed processing system shown in accordance with an exemplary embodiment of this specification.

[0144] Figure 2 This is a schematic diagram of an operator shown in accordance with an exemplary embodiment of this specification.

[0145] Figure 3 This is a schematic diagram of an execution plan shown in accordance with an exemplary embodiment of this specification.

[0146] Figure 4 This is another schematic diagram of an operator shown in accordance with an exemplary embodiment of this specification.

[0147] Figure 5 This is a flowchart of another monitoring method for a distributed processing system shown in accordance with an exemplary embodiment of this specification.

[0148] Figure 6 This is yet another schematic diagram of an operator shown in accordance with an exemplary embodiment of this specification.

[0149] Figure 7 This is a flowchart of yet another monitoring method for a distributed processing system shown in accordance with an exemplary embodiment of this specification.

[0150] Figure 8 This is a hardware structure diagram of an electronic device where the monitoring device of the distributed processing system in the embodiment of this specification is located.

[0151] Figure 9 This is a block diagram of a monitoring device for a distributed processing system shown in accordance with an exemplary embodiment of this specification. Detailed implementation

[0152] Here, the exemplary embodiments will be described in detail, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this specification. On the contrary, they are merely examples of devices and methods consistent with some aspects of this specification as detailed in the appended claims.

[0153] The terms used in this specification are only for the purpose of describing specific embodiments and are not intended to limit this specification. The singular forms "a", "the", and "said" used in this specification and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used herein refers to and includes any or all possible combinations of one or more of the associated listed items.

[0154] It should be understood that although the terms first, second, third, etc. may be used in this specification to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of this specification, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to a determination".

[0155] Next, the embodiments of this specification will be described in detail.

[0156] As Figure 1 shown, Figure 1 is a flowchart of a monitoring method for a distributed processing system shown in this specification according to an exemplary embodiment. An operator is provided on a processing node included in the distributed processing system, and there is an upstream and downstream relationship between the operators. The execution subject of this method is the master server. Specifically, it is a computer device, that is, the processor in the master server. This method includes the following steps:

[0157] Step 101, for a distributed task, obtain the data processing rate corresponding to the operator.

[0158] In this embodiment, the distributed processing system is a Flink system. The distributed task indicates a Flink Job, that is, a Flink job. A Flink job includes an operator chain formed by multiple operators. For two adjacent operators in the operator chain, the operator in the front (i.e., above) can be called the upstream operator, and the operator in the back (i.e., below) can be called the downstream operator. The traffic always flows from the upstream to the downstream, that is, the downstream operator processes the data produced by the upstream operator. For each operator in the operator chain, determine the data processing rate corresponding to the operator.

[0159] Optionally, the data processing rate includes a data consumption rate and / or a first data production rate. The data consumption rate represents the rate at which the operator consumes the data produced by the upstream operator. The first data production rate represents the rate at which the operator produces data.

[0160] Among them, the data consumption rate indicates the rate at which the operator processes the data produced by the upstream operator within a first preset unit of time, which reflects the processing ability of the operator; the first data production rate indicates the rate at which the operator produces data within a first preset unit of time, which reflects the situation of the upstream operator producing data. For example, for the operators involved in the distributed task, that is, the operator chain includes operator 1 and operator 2, operator 1 is the upstream operator of operator 2, and operator 1 transmits the data, that is, the traffic, to operator 2. This data is the data produced by the upstream operator of operator 2, and operator 2 processes this data. For example, it filters the data, and the filtered data becomes the data produced by operator 2.

[0161] Optionally, the upstream operator saves the produced data to the input buffer, and the downstream operator consumes the data in the input buffer. The downstream operator saves the data it produces to the output buffer. Correspondingly, the data consumption rate corresponding to the operator indicates the rate at which the operator consumes the data in the input buffer; the first data production rate corresponding to the operator indicates the rate at which the operator fills the output buffer. For example, when calculating the first data production rate, the amount of data increased in the output buffer within a certain period of time is obtained, and this amount of data is divided by this time to obtain the first data production rate.

[0162] Optionally, the data processing rate corresponding to the operator can be collected by the client plug-in installed on the operator. After the client collects the data processing rate corresponding to the operator, it sends it to the main control server.

[0163] Optionally, the number of upstream operators of the operator is at least one. The upstream and downstream operators are connected through a channel. Each operator includes, that is, corresponding to at least one operator instance, and the operator instance is obtained by instantiating the operator. Correspondingly, the data processing rate corresponding to the operator includes the data processing rates respectively corresponding to the operator instances corresponding to the operator. For example, when the data processing rate includes the data consumption rate, the operator includes operator D, and operator D corresponds to 2 operator instances, namely operator instance d1 and operator instance d2. The data consumption rate corresponding to operator instance d1 is 3m / s, and the data consumption rate corresponding to operator instance d2 is 5m / s. Then the data consumption rate corresponding to operator D includes the data consumption rate corresponding to operator instance d1 (i.e., 3m / s) and the data consumption rate corresponding to operator instance d2 (i.e., 5m / s).

[0164] Optionally, one operator instance corresponding to the operator and at least one operator instance corresponding to the upstream operator of the operator, that is, the upstream operator instances, are connected through a channel, and each channel corresponds to an input buffer and an output buffer. For example, as Figure 2 shown, operator 2 corresponds to 2 operator instances 2, operator 1 corresponds to 6 operator instances 1, operator 1 is the upstream operator of operator 2, and one operator instance 2 is connected to 3 operator instances 1 through three channels respectively, that is, one operator instance 1 is connected to one operator instance 2 through one channel. This operator instance 2 consumes the data in the input buffer storing the data produced by this operator instance 1, that is, the data in the input buffer corresponding to this channel.

[0165] Specifically, the Flink job architecture consists of three parts: the job Client, the Flink JobManager, and the Flink TaskManager (task manager). The job is parsed into a stream graph, a job graph, and an execution graph in three steps via the Client and the JobManager. The master server monitors the JobManager. If a new job (i.e., new data) is submitted, it obtains the execution graph to query the task manager and slot to which the job is scheduled, i.e., the addresses of the slots involved in the job, the operator list, the dependency relationships between operators, the operator parallelism, and so on. As Figure 3 shown in the execution plan of the job, i.e., the execution flow, vertex represents a certain vertex on the DAG graph, and JobVertex and ExecutionVertex represent the logical execution plan and the physical execution plan respectively. ResultPartition represents the output of the vertex. Flink will instantiate according to the operator parallelism to obtain the operator instance corresponding to the operator. As Figure 4 shown, the result subpartition (i.e., subpartition) and the number of concurrencies are initialized to 2. The downstream InputGate receives the upstream data, and the number of concurrencies is controlled by the downstream parallelism. Among them, the map operator represents the source operator, and the reduce operator represents the destination operator. These two operators are used as two vertices on the DAG graph. One vertex can correspond to one or more operators (operator chain). For the convenience of description, this application assumes that one vertex corresponds to one operator.

[0166] Specifically, from the above-described job execution process, it can be determined: 1) the task slots of the task manager to which the job is assigned; 2) the operator identifiers involved in the job and the parent-child dependency relationships between operators, i.e., the upstream and downstream relationships; 3) the identifiers of the slots to which the operators are scheduled, the operator parallelism, the channel data between operators, etc.; 4) the InputGate and ResultSubPartition corresponding to each operator to track the upstream and downstream operators corresponding to the operator; among them, the source operator (i.e., the source operator) serves as the input end InputGate.

[0167] Optionally, the parallelism degrees of upstream and downstream operators can be different. Each operator can be scheduled to at least one slot, and each slot corresponds to one operator instance. For example, an operator is scheduled to 3 slots, and each slot corresponds to one operator instance of this operator. Split the number of threads on each slot according to the parallelism degree. For example, there are 10 task managers, each task manager corresponds to 2 slots, and the parallelism degree of the map is 10. Then the map can be allocated to 5 task managers, and one map thread is started on each slot.

[0168] Optionally, the upstream operator and the downstream operator communicate through channels, and the number of channels is determined by the parallelism numbers of the upstream and downstream. For example, the upstream operator has a parallelism degree of 10, and the downstream operator has a parallelism degree of 2. Assuming an average distribution, every 5 processing threads send data to one thread of the downstream operator. That is, the upstream operator has a total of 10 operator instances, and the downstream operator has 2 operator instances. Every 5 operator instances of the upstream operator communicate with one operator instance of the downstream operator. In the Flink engine, the channels connecting the upstream operator and the downstream operator are not fixed.

[0169] Among them, each operator needs to record its upstream and downstream channels. Correspondingly, the client will also record the channels between operators, that is, the channels between operator instances.

[0170] Among them, the above identifiers include information such as name, ID, etc. For example, the operator identifier is the operator ID.

[0171] Optionally, the data processing rate corresponding to the operator can also include a second data production rate, and the second data production rate represents the rate at which the upstream operator of the operator produces data, that is, the rate of filling the input buffer.

[0172] Optionally, the data processing rate corresponding to the operator is the data processing rate corresponding to the operator instance. Specifically, the data consumption rate corresponding to the operator instance represents the rate at which the operator instance consumes the data produced by the upstream operator instance. The first data production rate represents the rate at which this operator instance produces data.

[0173] Optionally, each operator instance corresponds to an input buffer and an output buffer. The total space size of the input buffer and the output buffer can be changed, that is, it can be expanded or contracted. The master server can obtain the information of the input buffer and the output buffer corresponding to the operator instance in real time or at regular intervals; among them, the input buffer information includes information such as the current total size and the remaining size of the input buffer; similarly, the output buffer information includes information such as the current total size and the remaining size of the output buffer.

[0174] Optionally, the client corresponding to the operator records the information of the input buffer and the output buffer corresponding to each operator instance corresponding to the operator, and sends it to the master server.

[0175] Optionally, after obtaining the operator, that is, the data processing rate of each operator instance corresponding to the operator, the master server can save it to the target location for aggregation processing and calculation. The first preset unit of time can be at the second level. Correspondingly, the data consumption rate represents the operator, that is, the amount of data consumed by the operator instance per second from the upstream operator, that is, the upstream operator instance. The first data production rate represents the amount of data produced by the operator per second. When performing aggregation calculation on the data processing rate, it is aggregated according to the second preset unit of time, where the second preset unit of time indicates levels such as minutes and hours. For example, when the second preset unit of time indicates the minute level, the aggregated first data production rate represents the amount of data produced by the operator per minute.

[0176] Among them, the target location includes devices such as databases and ES that can store data.

[0177] In this embodiment, after obtaining the data processing rate, aggregation calculation is performed so that business and operation and maintenance personnel can refer to the aggregated data processing rate to optimize relevant conditions (such as the concurrency degree of operators) of the Flink system to ensure the performance of the Flink system. Of course, the data processing rate can also be directly used for optimization.

[0178] Step 102: Determine at least one working state of the operator according to the data processing rate corresponding to the operator and / or its upstream operator.

[0179] In this embodiment, for each operator, after obtaining the data processing rate of the operator, based on the data processing rate corresponding to the operator and / or the data processing speed corresponding to the upstream operator of the operator, it is determined whether there is an abnormality in the ability of the operator to process data (such as production, consumption) in each dimension, so as to obtain the working state of each dimension of the operator, that is, to obtain each working state of the operator. This working state indicates whether there is an abnormality in the ability of the operator to process data, so as to determine whether there is an abnormal operator and realize the intelligent monitoring of the operator.

[0180] Optionally, the dimension includes a first dimension and / or a second dimension. The first dimension indicates the production dimension, and the second dimension indicates the consumption dimension. Correspondingly, the working state of the operator includes the first working state and / or the second working state of the operator. Among them, the first working state of the operator indicates whether the data allocation produced by the upstream operator is uniform, that is, it indicates whether the data volumes produced by all upstream operator instances corresponding to this operator differ little. The second working state indicates whether the data consumption ability is normal, that is, it indicates whether the data volumes consumed by each operator instance corresponding to this operator differ little.

[0181] Optionally, when determining the working state of any operator among all operators, determine at least one working state of the first operator according to the error value between the data processing rates respectively corresponding to each operator instance corresponding to the first operator and / or the second operator. Among them, the first operator is any operator among the operators. The second operator is the upstream operator of the first operator.

[0182] Specifically, when the working state includes the first working state, determine the first working state of the first operator according to the error value between the data processing rates respectively corresponding to the operator instances corresponding to the upstream operator (i.e., the second operator) of the first operator.

[0183] When the working state includes the second working state, determine the second working state of the first operator according to the error value between the data processing rates respectively corresponding to each operator instance corresponding to the first operator. Specifically, the second working state of each operator instance corresponding to the first operator can be determined.

[0184] When the working state includes both the first working state and the second working state, determine the first working state of the first operator according to the error value between the data processing rates respectively corresponding to the operator instances corresponding to the upstream operator of the first operator, and at the same time, determine the second working state of the first operator according to the error value between the data processing rates respectively corresponding to each operator instance corresponding to the first operator.

[0185] Step 103, in response to any working state of the operator indicating that the operator has an abnormality, perform corresponding alarm operations.

[0186] In this embodiment, after determining the working state of the operator in each case, when a certain working state indicates that the operator has an abnormality, it means that the operator is a faulty operator, and then perform corresponding alarm operations, so that relevant personnel can timely discover the faulty operator, thereby timely solve the fault and ensure the normal operation and performance of the distributed processing system.

[0187] As can be seen from the above description, a distributed processing system includes multiple operators, and there is an upstream-downstream relationship between the operators. Obtain the data processing rate corresponding to each operator involved in processing the distributed task. The data processing rate includes a data consumption rate and / or a first data production rate. The data consumption rate represents the rate at which an operator consumes the data produced by its corresponding upstream operator, and the first data production rate represents the rate at which an operator produces data. Determine whether the data production and / or consumption of the operator is abnormal according to the data consumption rate and / or the first data production rate corresponding to the operator, so as to determine the working state of the operator in at least one dimension, and timely detect faulty operators. When the working state of the operator in a certain dimension indicates that the operator is abnormal, perform corresponding alarm operations to ensure the timeliness of the alarm, realize the intelligent monitoring of the distributed processing system, and then can solve the abnormality in time, improving the robustness and operability of the distributed processing system.

[0188] As Figure 5 shown, Figure 5 is a flowchart of another monitoring method for a distributed processing system shown in this specification according to an exemplary embodiment. The process of determining the first working state of an operator will be described in detail below in conjunction with a specific embodiment. As Figure 5 shown, the method includes the following steps:

[0189] Step 501, for a distributed task, obtain the data processing rate corresponding to the operator. Among them, the data processing rate includes a first data production rate. The first data production rate represents the rate at which the operator produces data.

[0190] Step 502, calculate the first error value between any two of the first data production rates respectively corresponding to each operator instance of the second operator corresponding to the first operator. Among them, the first operator is any one of the above operators; the second operator is the upstream operator of the first operator.

[0191] In this embodiment, for each operator in the distributed processing system, take this operator as the first operator, and take the upstream operator of the first operator as the second operator. Obtain the first data production rate corresponding to each operator instance (i.e., the second operator instance) of the second operator. For each second operator instance, calculate the difference between the first data production rate corresponding to this second operator instance and the first data production rates corresponding to other second operator instances, and take it as the first error value corresponding to this second operator instance. The first error value indicates the difference in the amount of data produced by two second operator instances within the first preset unit time, that is, it represents the difference in the data production capabilities of two second operator instances.

[0192] Step 503, determine the first working state of the first operator according to the first error value.

[0193] In this embodiment, after obtaining all the first error values corresponding to the second operator, that is, after obtaining the first error values corresponding to each second operator instance, the first error values are used to determine whether the data production capabilities of each second operator instance are relatively similar, that is, to determine whether the data distribution of the data produced by the upstream operator of the first operator is uniform, so as to obtain the first working state of the first operator.

[0194] Optionally, determining the first working state of the first operator according to the first error value includes:

[0195] In the case where there is a first error value reaching the first preset value, it is determined that the first working state of the first operator indicates that the data distribution of the data produced by the upstream operator is non-uniform.

[0196] In the case where all the first error values do not reach the first preset value, it is determined that the first working state of the first operator indicates that the data distribution of the data produced by the upstream operator is uniform.

[0197] Among them, reaching means greater than and / or equal to. Not reaching means less than.

[0198] Specifically, when the first error value between the first data production rates corresponding to two second operator instances reaches the first preset value, it indicates that the data production capabilities of these two second operator instances are quite different, that is, the data volume produced by one second operator instance is large, and the data volume produced by the other second operator instance is small, that is, the data distribution of the data produced by the upstream operator instance corresponding to the first operator is non-uniform, that is, the first working state of the first operator indicates that the data distribution of the data produced by the upstream operator of the first operator is non-uniform.

[0199] When all the first error values do not reach the first preset value, that is, when the first error value between the first data production rates corresponding to any two second operator instances does not reach the first preset value, it indicates that the data production capabilities of any two second operator instances are relatively similar, that is, the data distribution of the data produced by the upstream operator instance corresponding to the first operator is uniform, that is, the first working state of the first operator indicates that the data distribution of the data produced by the upstream operator of the first operator is uniform.

[0200] Of course, other mathematical calculations can also be performed based on the first error value, and the calculated value is used to determine the first working state of the first operator. For example, after obtaining the first error value between the first data production rates corresponding to two second operator instances, the ratio of the first error value to the first data production rate corresponding to any one of the two second operator instances is obtained to get the error rate. When the error value reaches the preset value, it is determined that the first working state of the first operator indicates that the data distribution of the data produced by the upstream operator is non-uniform, otherwise, it is determined that the first working state of the first operator indicates that the data distribution of the data produced by the upstream operator is uniform.

[0201] Optionally, it is also possible to first determine the first working states of the respective operator instances corresponding to the first operator, so as to determine the first working state of the first operator by using the first working states of the respective operator instances. The process includes: for each operator instance corresponding to the first operator, obtain the respective second operator instances corresponding to this operator instance. When the first error values among these respective second operator instances do not reach the first preset value, it indicates that the production data capabilities among the upstream operator instances corresponding to this operator instance are relatively close, and then determine that the first working state of this operator instance indicates that the data produced by the upstream operator is evenly distributed.

[0202] When the first error values among the respective second operator instances corresponding to this operator instance reach the first preset value, it indicates that the production data capabilities among the upstream operator instances corresponding to this operator instance are relatively different, and determine that the first working state of this operator instance indicates that the data produced by the upstream operator is unevenly distributed, thereby determining that the first working state of the operator indicates that the data produced by the upstream operator is unevenly distributed.

[0203] In this embodiment, when the first working state of any operator instance among the operator instances corresponding to the first operator indicates that the data produced by the upstream operator is unevenly distributed, determine that the first working state of the first operator indicates that the data produced by the upstream operator is unevenly distributed. When the first working states of all operator instances corresponding to the first operator all indicate that the data produced by the upstream operator is evenly distributed, determine that the first working state of the first operator indicates that the data produced by the upstream operator is evenly distributed.

[0204] Optionally, on the basis of determining the first working state of the first operator based on the first error value, it is also possible to further determine the first working state of the first operator by using the out-end buffer corresponding to the second operator. The specific determination process includes:

[0205] Obtain the expansion rate corresponding to the out-end buffer of the second operator. Here, the out-end buffer is used to store the data produced by the second operator.

[0206] When there is a situation where the first error value reaches the first preset value and the expansion rate reaches the first preset rate, determine that the first working state of the first operator indicates that the data produced by the upstream operator corresponding to the first operator is unevenly distributed.

[0207] Specifically, for each second operator instance, when all the first error values corresponding to the second operator instance reach the first preset value, it indicates that the data produced by the second operator instance may be excessive. Then, further obtain the expansion rate corresponding to the output buffer of the second operator instance, where the expansion rate indicates the growth rate of the total space size of the output buffer. For example, at time 1, the total space size of the output buffer corresponding to the second operator instance is 100 MB, and at time 2, the total space size of the output buffer corresponding to the second operator instance is 200 MB. Then, the expansion rate corresponding to the output buffer of the second operator instance is (200 MB - 100 MB) / (time 2 - time 1).

[0208] When the expansion rate corresponding to the output buffer of the second operator instance reaches the first preset rate, it indicates that the buffer corresponding to the second operator instance expands too fast, that is, it indicates that the data produced by the second operator instance is excessive. Then, determine that the data production of the second operator instance is unevenly distributed, and thus determine the downstream operator instance connected to the second operator instance, that is, the first working state of the operator instance in the first operator indicates that the data produced by the upstream operator is unevenly distributed, that is, the first working state of the first operator indicates that the data produced by the upstream operator is unevenly distributed.

[0209] When the expansion rate corresponding to the output buffer of the second operator instance does not reach the first preset rate, determine that the data production of the second operator instance is evenly distributed. Optionally, for each operator instance corresponding to the first operator, the first working state of the operator instance can be determined according to the above process of determining the first working state of the first operator.

[0210] Step 504, in response to the first working state of the first operator indicating that the operator has an abnormality, perform corresponding alarm operations.

[0211] In this embodiment, after obtaining the first working state of the first operator, when the first working state of the first operator indicates that the data produced by the upstream operator is unevenly distributed, an alarm operation is performed to achieve timely alarm.

[0212] Optionally, when performing an alarm operation according to the first working state of the first operator, the alarm can be performed in the following two ways.

[0213] One way is that, in the case where the first working state of the first operator indicates that the data produced by the upstream operator corresponding to the first operator is unevenly distributed, output a first alarm message. The first alarm message is used to prompt to increase the number of downstream operator concurrency.

[0214] Specifically, when the first working state of the first operator indicates uneven distribution of the data produced by the upstream operator, that is, when the first working state of the operator instance corresponding to the first operator indicates uneven distribution of the data produced by the upstream operator, it indicates that there is a relatively large amount of data produced by the upstream operator instance (i.e., the second operator instance) of the first operator. Then, the first warning message is output to prompt the first operator, that is, it indicates that the operator instance has encountered a computing bottleneck, and the data production capacity of the upstream operator exceeds the data consumption capacity of the downstream operator. It is necessary to increase the number of consumer data to meet the business requirements, that is, it is necessary to increase the concurrency degree of the first operator. In other words, it prompts the relevant personnel to increase the number of downstream operator instances connected to the second operator instance with uneven distribution of production data. For example, the first operator includes 2 operator instances, namely instance 1 and instance 2. The upstream operator of the first operator, that is, the second operator, includes 6 second operator instances. Instance 1 is connected to 3 second operator instances, and instance 2 is connected to the other 3 second operator instances. If the data produced by one of the second operator instances connected to instance 1 is unevenly distributed, the first warning message is output to increase the concurrency degree of the downstream operator instance connected to this second operator instance. For example, the first operator adds instance 3, and this instance 3 is also connected to this second operator instance. This instance 3 is also used to consume the data produced by the second operator instance, thereby increasing the number of consumers corresponding to this second operator instance.

[0215] Another way is that, in the case where the first working state of the first operator indicates uneven distribution of the data produced by the upstream operator corresponding to the first operator, based on the execution process corresponding to the distributed task, determine each third operator located above the first operator. Obtain the first working state corresponding to each third operator respectively. In the case where the first working state corresponding to each third operator indicates uneven distribution of the data produced by the upstream operator, determine other downstream operators of the source operator corresponding to the distributed task except the first downstream operator; where the first downstream operator is the downstream operator of the source operator and is a third operator. Perform a warning operation according to the data consumption rate corresponding to the other downstream operators.

[0216] Specifically, when the first working state of the first operator indicates that the data produced by the upstream operator is unevenly distributed, taking this first operator as a child node, traverse the execution plan upward, that is, traverse upward based on the execution process of the job, in other words, traverse upward based on the operator chain where the first operator is located until the source operator (i.e., the source operator) at the most upstream is traversed, and take the traversed operator as the third operator. When the first working states of all third operators indicate that the data produced by the upstream operator is unevenly distributed, it indicates that the source data, that is, the data volume of some partitions in the data source is too high, that is, the data in the partitions consumed by the first downstream operator is too much, resulting in too high a load on the first downstream operator and too high a load on the operators in the operator chain where the first downstream operator is located. Then, obtain other downstream operators connected to the source operator to determine whether other downstream operators can process more data according to the corresponding data consumption rate of other downstream operators, so as to perform corresponding alarm operations.

[0217] Taking a specific application scenario as an example, such as Figure 6 For the connection relationship between operators as shown, operator a1 consumes the data in partition 1 in the message queue, and operator b1 consumes the data in partition 2. When the first working state of operator a3 indicates that the data produced by the upstream operator is unevenly distributed, sequentially determine the working states of operators a1 and a2 in the first dimension. Since there is no upstream operator for the source operator A, therefore, there is no need to obtain the first working state of the source operator A. When the working states of operators a1 and a2 in the first dimension both indicate that the data produced by the upstream operator is unevenly distributed, it indicates that there is too much data in partition 1. Then, take operator a1 as the first downstream operator and operator b1 as other downstream operators, and perform alarm operations using the corresponding data consumption rate of operator b1.

[0218] Among them, for the convenience of description, operators a1, a2, and a3 respectively correspond to an operator instance, and partition 1 indicates the input buffer of operator a1. Operators b1, b2, and a3 respectively correspond to an operator instance, and partition 1 indicates the input buffer of operator b1.

[0219] In addition, when operator a1 corresponds to multiple operator instances, partition 1 corresponds to multiple sub - partitions, that is, input buffers, and each operator instance of operator a1 consumes the data in one input buffer.

[0220] It can be understood that the first operator and the third operator can both indicate operator instances. Correspondingly, the first downstream operator and other downstream operators can also indicate operator instances.

[0221] Optionally, performing alarm operations according to the corresponding data consumption rate of other downstream operators includes:

[0222] Determine the second error value between the corresponding data consumption rate of other downstream operators and the first data production rate corresponding to the source operator.

[0223] When the second error value reaches the second preset value, output a second warning message. The second warning message is used to prompt that there is too much source - side data and increase the number of downstream operator concurrency degrees.

[0224] Optionally, when the second error value does not reach the third preset value, perform a first - channel re - allocation operation. The first - channel re - allocation operation instructs to control other downstream operators to consume the data in the input buffer corresponding to the first downstream operator. The data in the input buffer is the data produced by the source operator.

[0225] Specifically, for each other downstream operator, calculate the difference between the data consumption rate corresponding to this other downstream operator and the first data production rate corresponding to the source operator, and use it as the second error value corresponding to this other downstream operator. When the second error value does not reach the third preset value, it indicates that this other downstream operator can consume more data. Channels can be re - allocated according to the speed of the data consumption rate of the downstream operators, that is, perform the first - channel re - allocation operation, so that other downstream operators with a higher data consumption rate consume the input buffer corresponding to the first downstream operator.

[0226] When all second error values reach the second preset value, it indicates that the computational pressures of all other downstream operators are also large and they cannot consume more data. Then output the second warning message to prompt that there is too much source - side data and increase the number of downstream operator concurrency degrees, that is, increase the concurrency degree of the first downstream operator.

[0227] Optionally, the data consumption rate corresponding to other downstream operators can refer to the data consumption rates corresponding to each operator instance of other downstream operators, and the first data production rate corresponding to the source operator can refer to the first data production rate corresponding to the operator instance of the source operator connected to the operator instance corresponding to other downstream operators. Accordingly, the second error value corresponding to each operator instance can be determined. When the second error value corresponding to the operator instance does not reach the third preset value, it indicates that the operator instance can consume more data. Then make the operator instance consume the data corresponding to the affected channel, that is, consume the data in the input buffer corresponding to the first downstream operator, that is, establish a channel between the operator instance and the source operator to achieve automatic and intelligent channel connection.

[0228] Among them, the second preset value and the third preset value can be the same or different.

[0229] It can be understood that when establishing a channel between the operator instance and the source operator, in fact, a channel connection is established between the operator instance and the first downstream operator, that is, the operator instance in the first downstream operator corresponding to the input buffer with too much data, that is, reconstruct the channel between the upstream slot and the downstream slot.

[0230] Optionally, the channel between the first downstream operator and the source operator can also be closed, that is, the affected downstream operator instances corresponding to the source operator are closed.

[0231] Continuing the above application scenario, calculate the difference between the data consumption rate corresponding to operator b1 and the first production rate corresponding to source operator A, that is, calculate the difference between the rate at which operator b1 consumes data in partition 2 and the rate at which operator A fills partition 2, to obtain a second error value. When the second error value does not reach the third preset value, it indicates that operator b1 has a high data consumption ability and can consume more data, so operator b1 is made to consume the data in region 1.

[0232] Optionally, when the second warning message is output, obtain the minimum value of the data consumption rates corresponding to other downstream operators and the data consumption rate corresponding to the first downstream operator. According to the minimum value, generate a speed limit instruction and send the speed limit instruction to the source operator, so that the source operator adjusts the first data production rate corresponding to the source operator based on the minimum value, that is, adjusts the first data production rate corresponding to the source operator with the minimum value as the benchmark, to avoid the backpressure of the entire system affecting mechanisms such as checkpoint and watermark.

[0233] Optionally, when adjusting the first data production rate corresponding to the source operator with the minimum value as the benchmark, it can be adjusted according to a preset adjustment rule. For example, the first data production rate corresponding to the source operator is adjusted to the minimum value. Here, the adjustment rule is not limited.

[0234] It can be understood that when adjusting the first data production rate corresponding to the source operator, the first data production rates corresponding to each operator instance corresponding to the source operator are adjusted.

[0235] Optionally, when the first working states of all third operators all indicate that the data produced by the upstream operator is evenly distributed, a second channel reallocation operation is performed. Among them, the second channel reallocation operation instructs the normal operator instance to consume the data in the input buffer corresponding to the abnormal operator instance; the normal operator instance is the operator instance in the operator instances corresponding to the first operator whose first working state indicates that the data produced by the upstream operator is evenly distributed, and the abnormal operator instance is the operator instance in the operator instances corresponding to the first operator whose first working state indicates that the data produced by the upstream operator is unevenly distributed.

[0236] Specifically, when the first working states of all the third operators indicate that the data produced by the upstream operator is evenly distributed, it indicates that only the data produced by the upstream operator of the first operator is unevenly distributed. That is, the data production rate of the upstream operator exceeds the consumption capacity of the first operator. Then, determine the operator instance in the operator instances corresponding to the first operator whose first working state indicates that the data produced by the upstream operator is evenly distributed, and use the determined operator instance as a normal operator instance. For each normal operator instance, calculate the difference between the data consumption rate corresponding to this normal operator instance and its corresponding first data production rate to obtain a fourth error value, which is used to determine whether the normal operator instance can consume more data. When it is determined that the normal operator instance can consume more data, use the operator instance in the operator instances corresponding to the first operator whose first working state indicates that the data produced by the upstream operator is unevenly distributed as an abnormal operator instance, and establish a channel between this normal operator instance and the upstream operator instance corresponding to this abnormal operator instance, that is, the upstream operator instance that produces too much data, so that this normal operator instance can consume the data produced by this upstream operator instance, reduce the consumption pressure of this abnormal operator instance, and realize the reallocation of channels according to the data consumption capacity, ensuring the data processing performance of the distributed system, that is, ensuring the normal operation of the service.

[0237] Optionally, when the data produced by the upstream operator of the first operator is based on keystream, in the case where the first working state of the first operator indicates that the data produced by the upstream operator is unevenly distributed, record the key value corresponding to the data produced by the upstream operator corresponding to the first operator, that is, determine the key value (i.e., key) corresponding to the data produced by the second operator instance with uneven data distribution, realize the recording of the key value of the skewed data, and output the key value, so that relevant personnel can know the skewed data, so that relevant personnel can reset the key value that the first operator needs to process according to the skewed data, that is, reset the data flow direction, and avoid the first operator needing to consume too much data due to unreasonable key value setting, and then causing a computing bottleneck.

[0238] In this embodiment, the master server can also obtain the expansion rate corresponding to the buffer of each operator instance, where the buffer includes an out-end buffer and / or an in-end buffer. When the expansion rate reaches the second preset rate, it indicates that the buffer expands too fast, and it is necessary to limit the expansion rate of the buffer. Then, when the total size of the buffer expands to the preset size, control the buffer to stop expanding, realizing bufferpool pre-allocation.

[0239] In this embodiment, without affecting the business, channel reallocation can be performed according to the data processing rate of the operator, realizing automatic traffic switching. There is no need for business personnel and operation and maintenance personnel to intervene, avoiding problems such as backpressure caused by the distributed processing system and being perceived by operation and maintenance only after a large number of checkpoints fail, and ensuring the performance of the distributed processing system.

[0240] In this embodiment, a horizontal comparison is made on the data processing rate corresponding to the operator, that is, the operator instance. Specifically, all data processing rates corresponding to the operator instance within a set time are obtained. If all data processing rates do not reach the historical average processing rate, it indicates that the operator instance has become a slow node, and then a slow operator alarm message is output to enable relevant personnel to maintain the operator with a slower operation. Among them, the historical average rate is calculated based on the collected data processing rates.

[0241] In this embodiment, based on the data processing rate corresponding to the upstream operator of the operator, the working state of the operator in the first dimension is determined, that is, it is determined whether the data produced by the upstream operator of the operator is evenly distributed, that is, it is determined whether the concurrency degree of the operator is reasonable, so as to determine whether to perform corresponding alarms, that is, whether to prompt to adjust the concurrency degree of the operator, realizing timely adjustment of anomalies and ensuring the operation performance of distributed processing.

[0242] In this embodiment, taking the source - end kafka message queue as an example, the data in the kafka partition is unevenly distributed. Some partitions have a large amount of data, which will cause the Flink kafka consumer, that is, the downstream operator of the source operator, to encounter bottlenecks when consuming large - data partitions. Therefore, when the first working state of the first operator indicates that the data produced by the upstream operator is unevenly distributed, it is determined whether the downstream operator has an excessive load due to uneven source - end partitions, and then corresponding alarm operations are performed. For example, the concurrency degree of the downstream operator of the source operator is increased to avoid system jams.

[0243] In this embodiment, when the business logic includes keystream, if the set partition key value is unreasonable, it will also cause calculation bottlenecks for the downstream operator. Therefore, the key value of the skewed data is recorded to enable relevant personnel to solve it as soon as possible.

[0244] In this embodiment, through operation and maintenance metrics, that is, the data processing rate corresponding to the operator, the anomalies existing in the system can be determined in a timely manner. When the system has not yet had serious problems, it is informed to the operation and maintenance personnel whether there are unreasonable concurrency degree settings in the system, whether there are problems with key - value data skew or unreasonable settings in the job, so that the operation and maintenance personnel can solve the problems in the distributed processing system as early as possible and ensure the reliability of the system.

[0245] In this embodiment, for each operator, the operator is used as the first operator, and the first data production rates respectively corresponding to the respective operator instances of the second operator located upstream of the first operator are obtained, so as to determine whether there is a large difference in the data production capabilities between the respective operator instances based on the first data production rates respectively corresponding to the respective operator instances, that is, to determine whether there is an operator instance that produces too much or too little data, so that it can be determined whether the data distribution produced by the upstream operator instance corresponding to the first operator is uniform, and then the first working state corresponding to the first operator is obtained, realizing the accurate determination of the first working state of the first operator. When the first working state corresponding to the first operator indicates that the data distribution produced by the upstream operator corresponding to the first operator is uneven, an alarm operation is performed to achieve timely alarm, so that relevant personnel can timely solve the problem of uneven data distribution produced by the upstream operator corresponding to the first operator.

[0246] As Figure 7 shown, Figure 7 FIG. is a flowchart of another monitoring method for a distributed processing system shown in this specification according to an exemplary embodiment. The process of determining the second working state corresponding to the operator will be described in detail below in conjunction with a specific embodiment. The process will be described in detail. As Figure 7 shown, the method includes the following steps:

[0247] Step 701: For a distributed task, obtain the data processing rate corresponding to the operator. Among them, the data processing rate includes the data consumption rate. The data consumption rate represents the rate at which the operator consumes the data generated by the upstream operator.

[0248] Step 702: Calculate the third error value between any two of the data consumption rates corresponding to the first operator. The first operator is any one of the above operators.

[0249] In this embodiment, for each operator in the distributed processing system, the operator is used as the first operator, and the data consumption rates corresponding to the respective operator instances of the first operator are obtained.

[0250] For each operator instance corresponding to the first operator, calculate the difference between the data consumption rate corresponding to the operator instance and the data consumption rates corresponding to the other operator instances of the first operator, and use it as the third error value corresponding to the operator instance. The third error value indicates the difference in the data consumption amounts of two operator instances corresponding to the first operator within the first preset unit time, that is, it represents the difference in the data consumption capabilities of the two second operator instances.

[0251] Specifically, the data consumption rate corresponding to the operator instance indicates the data consumption rate of the operator instance on its corresponding slot.

[0252] Step 703: Determine the second working state corresponding to each operator instance of the first operator according to the third error value.

[0253] In this embodiment, for each operator instance, when the third error value between this operator instance and other operator instances reaches the fourth preset value, it indicates that the data consumption ability of this operator instance is poor, and there is a computing bottleneck in the slot where this operator instance is located. Then, determine that the second working state of this operator instance indicates abnormal data consumption ability. Otherwise, determine that the second working state of this operator instance indicates normal data consumption ability.

[0254] Step 704: In response to the second working state of the operator instance indicating abnormal data consumption ability, perform corresponding alarm operations.

[0255] In this embodiment, after obtaining the second working state of each operator instance of the first operator, when the second working state of the operator instance indicates abnormal data consumption ability, it indicates that this operator instance is abnormal. Then, perform corresponding alarm operations to achieve timely alarm.

[0256] Optionally, when performing alarm operations based on the second working state of the operator instance, the following two methods can be used for alarm.

[0257] One method is that, in the case where the second working state of the operator instance indicates abnormal data consumption ability, determine the target task manager to which the slot corresponding to the operator instance belongs. Determine all the first operator instances included in the target task manager except the operator instance, and obtain the second working state of each first operator instance. Perform corresponding alarm operations according to the second working state of each first operator instance.

[0258] Specifically, when the second working state of the operator instance corresponding to the first operator indicates abnormal data consumption ability, it indicates that there is a computing bottleneck in this operator instance, that is, there is a computing bottleneck in the slot where this operator instance is located. Determine whether it is caused by the failure of the task manager (i.e., the target task manager) to which this slot belongs. Then, obtain other operator instances on this target task manager and use these other operator instances as the first operator instances to determine the reason for the abnormal data consumption ability of this operator instance corresponding to the first operator by using the second working state of the first operator instance, that is, determine the reason for the computing bottleneck in this slot, so as to achieve accurate positioning of the problem and further achieve accurate alarm.

[0259] Optionally, when alarming based on the second working state of the first operator instance, calculate the ratio of the number of first operator instances with abnormal data consumption capacity indicated by the second working state to the total number of first operator instances to obtain the abnormal first operator instance ratio. When the abnormal first operator instance ratio reaches the first preset ratio, it indicates that the target task manager has failed, resulting in problems with operator instances on more than one slot and affecting multiple jobs. In this case, output a target task manager failure prompt message to prompt relevant personnel that the target task manager has failed, achieving accurate positioning of the problem.

[0260] When the abnormal first operator instance ratio does not reach the first preset ratio, it indicates that only the operator instances with abnormal data consumption capacity indicated by the second working state corresponding to the first operator have calculation bottlenecks, that is, the slot to which the operator instance belongs has calculation bottlenecks. In this case, output the operator instance abnormal prompt message to prompt relevant personnel that there is an abnormality with the operator instance.

[0261] Optionally, the target task manager failure prompt message may also include the names of the affected jobs and the identifiers of the affected operators, that is, the identifiers of all operator instances corresponding to the target task manager that indicate abnormal data consumption capacity.

[0262] Another way is to directly output an operator instance abnormal prompt message when the operator instance indicates abnormal data consumption capacity to inform relevant personnel that there is an abnormality with the operator instance.

[0263] In this embodiment, through the operation and maintenance metrics, that is, the data consumption rate corresponding to the operator, the abnormalities existing in the system can be determined in a timely manner, achieving accurate positioning of the faults. When the system has not yet had serious problems, inform the operation and maintenance personnel of the taskmanagers and slots with faults in the system, enabling the operation and maintenance personnel to solve the problems in the distributed processing system as early as possible and ensuring the reliability of the system.

[0264] In this embodiment, for each operator, regard the operator as the first operator and obtain the data consumption rates corresponding to each operator instance of the first operator, so as to determine whether there is a large difference in the data consumption capabilities among the operator instances based on the data consumption rates corresponding to each operator instance, that is, determine whether there are operator instances with excessive or insufficient data consumption, thereby determining whether the data consumption capabilities of the operator instances are normal, and further obtaining the second working states corresponding to each operator instance, achieving accurate determination of the second working states corresponding to the operator instances of the first operator. When the second working state corresponding to the operator instance of the first operator indicates abnormal data consumption capacity, perform an alarming operation to achieve timely alarming, so that relevant personnel can solve the problem of abnormal data consumption capacity of the operator instance in a timely manner.

[0265] Corresponding to the embodiments of the foregoing method, this specification also provides embodiments of a device and a computer device to which the device is applied.

[0266] The embodiments of the monitoring device of the distributed processing system in this specification can be applied to computer devices, such as terminal devices (e.g., servers, computers, etc.). The device embodiments can be implemented by software, or by hardware, or by a combination of software and hardware. Taking software implementation as an example, as a logically meaningful device, it is formed by the processor in the file processing where it is located reading the corresponding computer program instructions in the non-volatile memory into the memory for operation. At the hardware level, as Figure 8 shown, it is a hardware structure diagram of the computer device where the monitoring device of the distributed processing system in the embodiments of this specification is located. In addition to Figure 8 the processor 810, memory 830, network interface 820, and non-volatile memory 840 shown, the computer device where the monitoring device 831 of the distributed processing system in the embodiments is located usually includes other hardware according to the actual functions of the computer device, which will not be elaborated here.

[0267] As Figure 9 shown, Figure 9 is a block diagram of a monitoring device of a distributed processing system shown according to an exemplary embodiment of this specification. The device includes:

[0268] A rate acquisition module 910, configured to obtain the data processing rate corresponding to the operator for a distributed task;

[0269] A rate processing module 920, configured to determine at least one working state of the operator according to the data processing rate corresponding to the operator and / or its upstream operator.

[0270] An alarm module 930, configured to perform a corresponding alarm operation in response to any working state of the operator indicating that the operator is abnormal.

[0271] Optionally, the operator corresponds to at least one operator instance, and each operator instance has a corresponding data processing rate.

[0272] The rate processing module 920 is specifically configured to:

[0273] Determine at least one working state of the first operator according to the error value between the data processing rates corresponding to the respective operator instances of the first operator and / or the second operator. Wherein, the first operator is any operator among the operators. The second operator is the upstream operator of the first operator.

[0274] Optionally, the data processing rate includes a first data production rate. The first data production rate represents the rate at which an operator produces data. The working state of the first operator includes a first working state.

[0275] Among them, the first working state of the first operator indicates whether the data allocated by the upstream operator corresponding to the first operator is evenly distributed.

[0276] The rate processing module 920 is specifically configured to:

[0277] Calculate a first error value between any two of the first data production rates respectively corresponding to each operator instance corresponding to the second operator.

[0278] Determine the first working state of the first operator according to the first error value.

[0279] Optionally, the rate processing module 920 is further configured to:

[0280] In the case where there is a first error value reaching a first preset value, determine that the first working state of the first operator indicates that the data allocated by the upstream operator corresponding to the first operator is unevenly distributed.

[0281] In the case where all the first error values do not reach the first preset value, determine that the first working state of the first operator indicates that the data allocated by the upstream operator corresponding to the first operator is evenly distributed.

[0282] Optionally, the rate processing module 920 is further configured to:

[0283] Obtain the expansion rate corresponding to the out-end buffer corresponding to the second operator. The out-end buffer is used to store the data produced by the second operator.

[0284] In the case where there is a first error value reaching the first preset value and the expansion rate reaches the first preset rate, determine that the first working state of the first operator indicates that the data allocated by the upstream operator corresponding to the first operator is unevenly distributed.

[0285] Optionally, the alarm module 930 is specifically configured to:

[0286] In the case where the first working state of the first operator indicates that the data allocated by the upstream operator corresponding to the first operator is unevenly distributed, output a first alarm message. The first alarm message is used to prompt to increase the number of downstream operator concurrency.

[0287] Optionally, the alarm module 930 is specifically configured to:

[0288] In the case where the first working state of the first operator indicates that the data allocated by the upstream operator corresponding to the first operator is unevenly distributed, determine each third operator above the first operator based on the execution process corresponding to the distributed task.

[0289] Obtain the first working state corresponding to each third operator.

[0290] In the case where the first working states corresponding to each third operator all indicate that the data produced by the upstream operator is unevenly distributed, determine other downstream operators of the source operator corresponding to the distributed task except the first downstream operator. Here, the first downstream operator is the downstream operator of the source operator and is a third operator.

[0291] Perform an alarm operation according to the data consumption rate corresponding to the other downstream operators.

[0292] Optionally, the upstream operator and the downstream operator are connected through a channel.

[0293] Optionally, the alarm module 930 is specifically configured to:

[0294] Determine a second error value between the data consumption rate corresponding to the other downstream operators and the first data production rate corresponding to the source operator.

[0295] In the case where the second error value reaches a second preset value, output a second alarm message. Here, the second alarm message is used to prompt that there is too much data at the source end and increase the number of downstream operator concurrency degrees.

[0296] Optionally, the upstream operator and the downstream operator are connected through a channel. The device further includes a first channel processing module.

[0297] The first channel processing module is specifically configured to:

[0298] In the case where the first working states corresponding to each third operator all indicate that the data produced by the upstream operator is unevenly distributed, determine a second error value between the data consumption rate corresponding to the other downstream operators and the first data production rate corresponding to the source operator.

[0299] In the case where the second error value does not reach a third preset value, perform a first channel reallocation operation. Here, the first channel reallocation operation instructs to control the other downstream operators to consume the data in the input buffer corresponding to the first downstream operator. The data in the input buffer is the data produced by the source operator.

[0300] Optionally, the device further includes a speed limit module.

[0301] The speed limit module is specifically configured to:

[0302] In the case where the first working states corresponding to each third operator all indicate that the data produced by the upstream operator is unevenly distributed, obtain the minimum value of the data consumption rate corresponding to the other downstream operators and the data consumption rate corresponding to the first downstream operator.

[0303] Generate a speed limit instruction according to the minimum value, and send the speed limit instruction to the source operator, so that the source operator adjusts the first data production rate corresponding to the source operator based on the minimum value.

[0304] Optionally, the device further includes a second channel processing module.

[0305] The second channel processing module is specifically configured to:

[0306] After obtaining the first working state corresponding to each third operator, perform a second channel reassignment operation when the first working state of all third operators indicates that the data produced by the upstream operator is evenly distributed. Among them, the second channel reassignment operation instructs the normal operator instance to consume the data in the input buffer corresponding to the abnormal operator instance. The normal operator instance is the operator instance corresponding to the first operator whose first working state indicates that the data produced by the upstream operator is evenly distributed, and the abnormal operator instance is the operator instance corresponding to the first operator whose first working state indicates that the data produced by the upstream operator is unevenly distributed.

[0307] Optionally, the device further includes a data recording module.

[0308] The data recording module is specifically configured to:

[0309] When the first working state of the first operator indicates that the data produced by the upstream operator corresponding to the first operator is unevenly distributed, record the key value corresponding to the data produced by the upstream operator corresponding to the first operator, and output the key value.

[0310] The data processing rate includes the data consumption rate. The data consumption rate represents the rate at which the operator consumes the data generated by the upstream operator. The working state of the first operator includes a second working state.

[0311] Among them, the second working state of the first operator indicates whether the data consumption ability is normal.

[0312] The rate processing module 920 is specifically configured to:

[0313] Calculate the third error value between any two of the data consumption rates corresponding to the first operator.

[0314] Determine the second working state corresponding to each operator instance corresponding to the first operator according to the third error value.

[0315] Optionally, the operator is scheduled to at least one resource group slot. The slot corresponds to the operator instance one by one.

[0316] The alarm module 930 is specifically configured to:

[0317] When the second working state of the operator instance indicates an abnormal data consumption capacity, determine the target task manager to which the slot corresponding to the operator instance belongs.

[0318] Determine all the first operator instances included in the target task manager except the operator instance, and obtain the second working state corresponding to each first operator instance.

[0319] Perform corresponding alarm operations according to the second working state corresponding to each first operator instance.

[0320] Optionally, the alarm module 930 is further configured to:

[0321] Calculate the ratio of the number of first operator instances whose second working state indicates an abnormal data consumption capacity to the total number of first operator instances to obtain the abnormal first operator instance ratio.

[0322] When the abnormal first operator instance ratio reaches the first preset ratio, output a target task manager failure prompt message.

[0323] When the abnormal first operator instance ratio does not reach the first preset ratio, output an operator instance abnormality prompt message.

[0324] Optionally, the alarm module 930 is further configured to:

[0325] When the second working state of the operator instance indicates an abnormal data consumption capacity, output an operator instance abnormality prompt message.

[0326] Optionally, the operator corresponds to at least one operator instance. The operator instance has a corresponding buffer.

[0327] The alarm module 930 is further configured to:

[0328] Obtain the expansion rate of the buffer corresponding to each operator instance.

[0329] When the expansion rate of the buffer reaches the second preset rate, control the buffer to stop expanding.

[0330] Optionally, the operator corresponds to at least one operator instance. The alarm module 930 is further configured to:

[0331] Obtain all the data processing rates corresponding to the operator instance obtained at the set time, and obtain the historical average processing rate corresponding to the operator instance.

[0332] If all the data processing rates do not reach the historical average processing rate, output a slow operator alarm message.

[0333] For the implementation processes of the functions and roles of each module in the above device, please refer to the implementation processes of the corresponding steps in the above method for details, which will not be elaborated here.

[0334] In one embodiment, the present application further provides a computer-readable storage medium, in which computer-executable instructions are stored. When the processor executes the computer-executable instructions, the method described above is implemented.

[0335] In one embodiment, the present application further provides a computer program product, including a computer program. When the computer program is executed by the processor, the method described above is implemented.

[0336] For the device embodiment, since it basically corresponds to the method embodiment, the relevant parts can be referred to the partial description of the method embodiment. The device embodiments described above are only illustrative. The modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical modules, that is, they may be located in one place or distributed to multiple network modules. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution in this specification. Those of ordinary skill in the art can understand and implement it without creative efforts.

[0337] The specific embodiments of this specification are described above. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be executed in a different order than in the embodiments and still achieve the desired results. Additionally, the processes depicted in the figures do not necessarily require the specific order or sequential order shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0338] Those skilled in the art will readily conceive of other implementations of this specification after considering the specification and practicing the invention herein. This specification is intended to cover any variations, uses, or adaptations of this specification, which follow the general principles of this specification and include common general knowledge or conventional technical means in the technical field not claimed in this application. The specification and embodiments are only regarded as exemplary, and the true scope and spirit of this specification are pointed out by the following claims.

[0339] It should be understood that this specification is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of this specification is only limited by the appended claims.

[0340] The above are only the preferred embodiments of this specification and are not intended to limit this specification. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of this specification shall be included within the scope of protection of this specification.

Claims

1. A monitoring method for a distributed processing system, characterized in that, the distributed processing system includes operators with an upstream and downstream relationship; the operators correspond to at least one operator instance, and each operator instance has a corresponding data processing rate, and the data processing rate includes a first data production rate, and the first data production rate represents the rate at which the operator produces data; the method includes: for a distributed task, obtaining the data processing rate corresponding to the operator; determining at least one working state of the first operator according to the error value between the data processing rates respectively corresponding to each operator instance corresponding to the first operator and / or the second operator; wherein, the first operator is any one of the operators, and the second operator is the upstream operator of the first operator, and the working state of the first operator includes a first working state for indicating whether the data distribution produced by the upstream operator corresponding to the first operator is uniform; responding to any working state of the operator indicating that the operator is abnormal, and performing a corresponding alarm operation; wherein, the determining at least one working state of the first operator according to the error value between the data processing rates respectively corresponding to each operator instance corresponding to the first operator and / or the second operator includes: calculating a first error value between any two of the first data production rates respectively corresponding to each operator instance corresponding to the second operator; in the case that there is a first error value reaching a first preset value, determining that the first working state of the first operator indicates that the data distribution produced by the upstream operator corresponding to the first operator is uneven; in the case that all the first error values do not reach the first preset value, determining that the first working state of the first operator indicates that the data distribution produced by the upstream operator corresponding to the first operator is uniform; wherein, the responding to any working state of the operator indicating that the operator is abnormal and performing a corresponding alarm operation includes: in the case that the first working state of the first operator indicates that the data distribution produced by the upstream operator corresponding to the first operator is uneven, outputting a first alarm message, and the first alarm message is used to prompt to increase the number of downstream operator concurrency.

2. The method according to claim 1, characterized in that, the determining that the first working state of the first operator indicates that the data distribution produced by the upstream operator corresponding to the first operator is uneven in the case that there is a first error value reaching a first preset value includes: obtaining the expansion rate corresponding to the out-end buffer of the second operator; wherein, the out-end buffer is used to store the data produced by the second operator; in the case that there is a first error value reaching a first preset value and the expansion rate reaches a first preset rate, determining that the first working state of the first operator indicates that the data distribution produced by the upstream operator corresponding to the first operator is uneven.

3. The method according to claim 1, characterized in that, the outputting the first alarm message includes: based on the execution process corresponding to the distributed task, determining each third operator above the first operator; Obtain the first working state corresponding to each third operator; When the first working state corresponding to each of the third operators indicates that the data produced by the upstream operator is unevenly distributed, determine other downstream operators of the source operator corresponding to the distributed task except the first downstream operator; wherein, the first downstream operator is a downstream operator of the source operator and is the third operator; Perform an alarm operation according to the data consumption rate corresponding to the other downstream operators.

4. The method according to claim 3, wherein, The upstream operator and the downstream operator are connected through a channel; The performing an alarm operation according to the data consumption rate corresponding to the other downstream operators includes: Determine a second error value between the data consumption rate corresponding to the other downstream operators and the first data production rate corresponding to the source operator; When the second error value reaches a second preset value, output a second alarm message; wherein, the second alarm message is used to prompt that there is too much data at the source end and increase the number of downstream operator concurrency.

5. The method according to claim 3, wherein, The upstream operator and the downstream operator are connected through a channel; when the first working state corresponding to each of the third operators indicates that the data produced by the upstream operator is unevenly distributed, the method further includes: Determine a second error value between the data consumption rate corresponding to the other downstream operators and the first data production rate corresponding to the source operator; When the second error value does not reach a third preset value, perform a first channel reallocation operation; wherein, the first channel reallocation operation instructs to control the other downstream operators to consume the data in the input buffer corresponding to the first downstream operator; the data in the input buffer is the data produced by the source operator.

6. The method according to claim 3, wherein, When the first working state corresponding to each of the third operators indicates that the data produced by the upstream operator is unevenly distributed, the method further includes: Obtain the minimum value of the data consumption rate corresponding to the other downstream operators and the data consumption rate corresponding to the first downstream operator; Generate a speed limit instruction according to the minimum value, and send the speed limit instruction to the source operator, so that the source operator adjusts the first data production rate corresponding to the source operator based on the minimum value.

7. The method according to claim 3, wherein, After obtaining the first working state corresponding to each of the third operators, the method further includes: When the first working state of all the third operators indicates that the data produced by the upstream operator is evenly distributed, perform a second channel reallocation operation; wherein, the second channel reallocation operation instructs the normal operator instance to consume the data in the input buffer corresponding to the abnormal operator instance; the normal operator instance is an operator instance whose first working state in the operator instances corresponding to the first operator indicates that the data production of the upstream operator is evenly distributed, and the abnormal operator instance is an operator instance whose first working state in the operator instances corresponding to the first operator indicates that the data production of the upstream operator is unevenly distributed.

8. The method according to claim 1, wherein, when the first working state of the first operator indicates that the data produced by the upstream operator corresponding to the first operator is unevenly distributed, the method further includes: recording the key values corresponding to the data produced by the upstream operator corresponding to the first operator, and outputting the key values.

9. The method according to claim 1, wherein, the data processing rate includes a data consumption rate; the data consumption rate represents the rate at which an operator consumes the data produced by an upstream operator; the working state of the first operator includes a second working state; wherein, the second working state of the first operator indicates whether the data consumption ability is normal; the determining at least one working state of the first operator according to the error values between the data processing rates respectively corresponding to each operator instance corresponding to the first operator and / or the second operator includes: calculating a third error value between any two of the data consumption rates corresponding to each of the first operator; determining the second working state respectively corresponding to each operator instance corresponding to the first operator according to the third error value.

10. The method according to claim 9, wherein, the operator is scheduled to at least one resource group slot; the slot corresponds to the operator instance one by one; the performing a corresponding alarm operation in response to any working state of the operator indicating that the operator has an abnormality includes: when the second working state of the operator instance indicates that the data consumption ability is abnormal, determining the target task manager to which the slot corresponding to the operator instance belongs; determining all the first operator instances included in the target task manager except the operator instance, and obtaining the second working state respectively corresponding to each first operator instance; performing a corresponding alarm operation according to the second working state respectively corresponding to each first operator instance.

11. The method according to claim 10, wherein, the performing a corresponding alarm operation according to the second working state respectively corresponding to each first operator instance includes: calculating the ratio of the number of first operator instances whose second working state indicates that the data consumption ability is abnormal to the total number of first operator instances to obtain the abnormal first operator instance ratio; when the abnormal first operator instance ratio reaches a first preset ratio, outputting a target task manager failure prompt message; when the abnormal first operator instance ratio does not reach the first preset ratio, outputting the operator instance abnormality prompt message.

12. The method according to claim 9, wherein, the performing a corresponding alarm operation in response to any working state of the operator indicating that the operator has an abnormality includes: when the second working state of the operator instance indicates that the data consumption ability is abnormal, outputting the operator instance abnormality prompt message.

13. The method according to any one of claims 1 to 12, wherein, the operator corresponds to at least one operator instance; there is a corresponding buffer for the operator instance; the method further includes: obtaining the expansion rate of the buffer respectively corresponding to each operator instance; When the expansion rate of the buffer reaches a second preset rate, control the buffer to stop expanding.

14. The method according to any one of claims 1 to 12, wherein, the operator corresponds to at least one operator instance; the method further includes: obtaining all data processing rates corresponding to the operator instance obtained at a set time, and obtaining the historical average processing rate corresponding to the operator instance; if all the data processing rates do not reach the historical average processing rate, output a slow operator warning message.

15. A monitoring device for a distributed processing system, wherein, an operator is provided on a processing node of the distributed processing system; the operator corresponds to at least one operator instance, and there is a corresponding data processing rate for each operator instance, and the data processing rate includes a first data production rate, and the first data production rate represents the rate at which the operator produces data; the device includes: a rate acquisition module, configured to obtain the data processing rate corresponding to the operator for a distributed task; a rate processing module, configured to determine at least one working state of the first operator according to an error value between the data processing rates respectively corresponding to each operator instance corresponding to the first operator and / or the second operator; wherein, the first operator is any one of the operators, the second operator is the upstream operator of the first operator, and the working state of the first operator includes a first working state, which is used to indicate whether the data distribution produced by the upstream operator corresponding to the first operator is uniform; the rate processing module is further configured to calculate a first error value between any two of the first data production rates respectively corresponding to each operator instance corresponding to the second operator; when there is a first error value reaching a first preset value, determine that the first working state of the first operator indicates that the data distribution produced by the upstream operator corresponding to the first operator is not uniform; when all the first error values do not reach the first preset value, determine that the first working state of the first operator indicates that the data distribution produced by the upstream operator corresponding to the first operator is uniform; an alarm module, configured to perform a corresponding alarm operation in response to any working state of the operator indicating that the operator is abnormal; the alarm module is further configured to output a first alarm message when the first working state of the first operator indicates that the data distribution produced by the upstream operator corresponding to the first operator is not uniform, and the first alarm message is used to prompt the number of increased downstream operator concurrency.

16. A computer device, wherein, it includes a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein when the processor executes the program, it implements the monitoring method of the distributed processing system according to any one of claims 1 to 14.

17. A computer program product, wherein, it includes a computer program, and when the computer program is executed by a processor, it implements the monitoring method of the distributed processing system according to any one of claims 1 to 14.