ES cluster-based processing method and equipment
By performing data index feature processing and decision-making model judgment on ES nodes in the ES cluster, and automatically determine the fault status and processing plan, the problems of low manual analysis efficiency and low accuracy in the prior art are solved, and the fault processing efficiency and data processing stability of the ES cluster are improved.
Patent Information
- Application Number
- CN202311591126.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-27
- Publication Date
- 2025-05-27
AI Technical Summary
In the prior art, manual analysis of ES node working data in the ES cluster leads to inefficient and accurate determination of whether there are faults and fault handling plans in the ES cluster, affecting the repair efficiency and data processing process of the ES cluster.
By obtaining the data index set of ES nodes, performing feature processing to obtain running features, and inputting them into the preset state decision model to determine whether the ES node is in a fault state. If it is in a fault state, the running state feature is processed based on the fault decision model to determine the fault handling method, and determine the final fault handling plan and execute it for the fault state in multiple preset periods.
It improves the efficiency and accuracy of determining whether there are faults and fault handling plans in the ES cluster, and deals with faults in a timely manner, improves the repair efficiency of the ES cluster, and ensures the normal operation of the data processing process.
Smart Images

Figure CN120050212A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of ES clusters, and in particular, to a processing method and device based on an ES cluster. Background Art
[0002] With the development of technology, a distributed search engine (Elasticsearch, abbreviated as ES) cluster includes multiple ES nodes; based on the ES cluster, a large amount of data storage, search, analysis, and other tasks can be completed in a very short time. It is necessary to judge whether the ES cluster fails and handle the failures of the ES cluster.
[0003] In the prior art, the working data of ES nodes in the ES cluster is analyzed and judged based on manual experience, and then it is manually determined whether the ES cluster has a failure, and a failure handling plan for the ES cluster is manually determined.
[0004] However, in the above method, the manual analysis method results in low efficiency and low accuracy in determining whether the ES cluster has a failure and the failure handling plan; it causes the failure of the ES cluster to be unable to be processed in time, and the repair efficiency of the ES cluster is low, affecting the data processing process of the ES cluster. Summary of the Invention
[0005] This application provides a processing method and device based on an ES cluster to solve the technical problems of low efficiency and low accuracy in determining whether the ES cluster has a failure and the failure handling plan, and low repair efficiency of the ES cluster.
[0006] In a first aspect, this application provides a processing method based on an ES cluster. The ES cluster includes at least one distributed search engine ES node, and the ES node is deployed on a host; the method includes:
[0007] Obtain a data index set of the ES node; wherein, the data index set includes data information at each moment in a preset time period, and the data information at each moment represents the running state of the ES node at each moment in the preset time period and the running state of the host where the ES node is located at each moment in the preset time period; and perform feature processing on the data index set of the ES node to obtain a running state feature of the ES node; wherein, the running state feature represents the running state of the ES node in the preset time period;
[0008] Input the running state feature of the ES node into a preset state decision model to obtain the state information of the ES node; wherein, the state information represents whether the ES node is in a problem state, and the problem state is a risk state or a failure state;
[0009] If the status information of the ES node indicates that the ES node is in a fault state, the operating state characteristics of the ES node are processed based on a preset fault decision model to obtain a fault handling method for the ES node;
[0010] If it is determined that the number of ES nodes in the ES cluster that are in a fault state in each preset time period among multiple preset time periods is greater than or equal to a preset number, it is determined that the ES cluster has a fault during the time period represented by the multiple preset time periods; and for each ES node that is in a fault state in each preset time period among multiple preset time periods, based on the fault handling method of the ES node in each preset time period among the multiple preset time periods, a final fault handling plan for the ES node is determined, and the final fault handling plan for the ES node is executed.
[0011] Optionally, processing the operating state characteristics of the ES node based on a preset fault decision model to obtain a fault handling method for the ES node includes:
[0012] Processing the operating state characteristics of the ES node based on a preset fault decision model to obtain the fault type of the ES node and the corresponding fault handling method for the fault type.
[0013] Optionally, for each ES node that is in a fault state in each preset time period among multiple preset time periods, determining a final fault handling plan for the ES node based on the fault handling method of the ES node in each preset time period among the multiple preset time periods includes:
[0014] For each ES node that is in a fault state in each preset time period among multiple preset time periods, if it is determined that the fault types of the ES node in each preset time period among the multiple preset time periods are the same, the fault handling method corresponding to the fault type in each preset time period among the multiple preset time periods of the ES node is determined as the final fault handling plan for the ES node;
[0015] For each ES node that is in a fault state in each preset time period among multiple preset time periods, if it is determined that the fault types of the ES node in each preset time period among the multiple preset time periods are different, the fault handling methods of the ES node in each preset time period among the multiple preset time periods are displayed to the user; in response to a first trigger instruction of the user, the fault handling method indicated by the first trigger instruction is determined as the final fault handling plan for the ES node; wherein, the first trigger instruction is used to indicate the fault handling method selected by the user.
[0016] Optionally, the method further includes:
[0017] Obtain a fault data set from the historical data set; wherein, the historical data set includes a fault data set, and the fault data set includes the running state characteristics of the ES node, the fault type of the ES node, and the fault handling method of the ES node;
[0018] Train the first initial model according to the fault data set to obtain the preset fault decision model.
[0019] Optionally, the method further includes:
[0020] If the status information of the ES node indicates that the ES node is in a risk state, process the running state characteristics of the ES node based on a preset risk decision model to obtain the risk type of the ES node and the risk handling method corresponding to the risk type;
[0021] If it is determined that the ES node is in a risk state in each preset time period among multiple preset time periods, determine the final risk handling plan of the ES node based on the risk handling method of the ES node in each preset time period among multiple preset time periods, and execute the final risk handling plan of the ES node.
[0022] Optionally, determining the final risk handling plan of the ES node based on the risk handling method of the ES node in each preset time period among multiple preset time periods includes:
[0023] If it is determined that the risk types of the ES node in each preset time period among multiple preset time periods are the same, determine the risk handling method corresponding to this risk type in each preset time period among multiple preset time periods as the final risk handling plan of the ES node;
[0024] If it is determined that the risk types of the ES node in each preset time period among multiple preset time periods are different, display the risk handling methods of the ES node in each preset time period among multiple preset time periods to the user; in response to the second trigger instruction of the user, determine the risk handling method indicated by the second trigger instruction as the final risk handling plan of the ES node; wherein, the second trigger instruction is used to indicate the risk handling method selected by the user.
[0025] Optionally, the method further includes:
[0026] Obtain a risk data set from the historical data set; wherein, the historical data set includes a risk data set, and the risk data set includes the running state characteristics of the ES node, the risk type of the ES node, and the risk handling method of the ES node;
[0027] Train the second initial model according to the risk data set to obtain the preset risk decision model.
[0028] Optionally, the method further includes:
[0029] Obtaining a status data set from a historical data set; wherein, the historical data set includes a status data set, and the status data set includes the running state characteristics of the ES node and whether the ES node is in a problem state;
[0030] Training a third initial model according to the status data set to obtain the preset status decision model.
[0031] Optionally, performing feature processing on the data metric set of the ES node to obtain the running state characteristics of the ES node, including:
[0032] Performing status extraction processing on the data information at each moment in the data metric set of the ES node based on the Encode model to obtain the data information in matrix form of the ES node at each moment;
[0033] Performing feature fusion processing on the data information in matrix form of the ES node at each moment in the preset time period based on the Attend model to obtain the running state characteristics of the ES node.
[0034] Optionally, the data information at each moment includes the host machine metric data set at each moment, the ES node metric data set at each moment, and the ES node event data set at each moment;
[0035] Wherein, the host machine metric data set at each moment represents the running state of the host machine where the ES node is located at each moment in the preset time period; the ES node metric data set at each moment represents the running metrics of the ES node at each moment in the preset time period; the ES node event data set at each moment represents the event data of the ES node at each moment in the preset time period.
[0036] In a second aspect, the present application provides a processing device based on an ES cluster. The ES cluster includes at least one distributed search engine ES node, and the ES node is deployed on a host machine. The device includes:
[0037] A first obtaining unit, configured to obtain the data metric set of the ES node; wherein, the data metric set includes the data information at each moment in a preset time period, and the data information at each moment represents the running state of the ES node at each moment in the preset time period and the running state of the host machine where the ES node is located at each moment in the preset time period;
[0038] A first processing unit, configured to perform feature processing on the data metric set of the ES node to obtain the running state features of the ES node; wherein, the running state features characterize the running state of the ES node in a preset time period;
[0039] A second processing unit, configured to input the running state features of the ES node into a preset state decision model to obtain the state information of the ES node; wherein, the state information characterizes whether the ES node is in a problem state, and the problem state is a risk state or a fault state;
[0040] A third processing unit, configured to, if the state information of the ES node indicates that the ES node is in a fault state, process the running state features of the ES node based on a preset fault decision model to obtain a fault handling method for the ES node;
[0041] A first determination unit, configured to, if it is determined that the number of ES nodes in the ES cluster that are in a fault state in each preset time period among multiple preset time periods is greater than or equal to a preset number, determine that the ES cluster has a fault within the duration represented by the multiple preset time periods;
[0042] A second determination unit, configured to, for each ES node that is in a fault state in each preset time period among multiple preset time periods, determine a final fault handling plan for the ES node based on the fault handling method of the ES node in each preset time period among the multiple preset time periods, and execute the final fault handling plan of the ES node.
[0043] In a third aspect, the present application provides an electronic device, including: a processor, and a memory communicatively connected to the processor;
[0044] The memory stores computer-executable instructions;
[0045] The processor executes the computer-executable instructions stored in the memory to implement the method provided in the first aspect above.
[0046] In a fourth aspect, the present application provides a computer-readable storage medium, in which computer-executable instructions are stored, and when the computer-executable instructions are executed by a processor, they are used to implement the method provided in the first aspect above.
[0047] In a fifth aspect, the present application provides a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the method provided in the first aspect above.
[0048] The processing method and device based on an ES cluster provided by this application obtain a set of data metrics of ES nodes, perform feature processing on the set of data metrics of ES nodes to obtain the running state features of ES nodes; input the running state features of ES nodes into a preset state decision model to obtain the state information of ES nodes; if the state information of ES nodes indicates that the ES nodes are in a faulty state, then process the running state features of ES nodes based on a preset fault decision model to obtain the fault handling method of ES nodes; if it is determined that the number of ES nodes that are in a faulty state in each preset time period among multiple preset time periods in the ES cluster is greater than or equal to a preset number, then it is determined that the ES cluster has a fault during the time period represented by the multiple preset time periods; and for each ES node that is in a faulty state in each preset time period among the multiple preset time periods, based on the fault handling method of the ES node in each preset time period among the multiple preset time periods, determine the final fault handling plan of the ES node, and execute the final fault handling plan of the ES node. Thereby, the efficiency and accuracy of determining whether the ES cluster has a fault and the fault handling plan are improved, so as to timely handle the fault of the ES cluster, improve the repair efficiency of the ES cluster, and ensure the normal operation of the data processing process of the ES cluster. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] The accompanying drawings herein are incorporated into and constitute a part of this specification, showing embodiments consistent with this application, and are used together with the specification to explain the principles of this application.
[0050] Figure 1 It is a flowchart of a processing method based on an ES cluster provided by an embodiment of this application;
[0051] Figure 2 It is a flowchart of another processing method based on an ES cluster provided by an embodiment of this application;
[0052] Figure 3 It is a schematic structural diagram of a processing device based on an ES cluster provided by an embodiment of this application;
[0053] Figure 4 It is a schematic structural diagram of an electronic device provided by an embodiment of this application.
[0054] Through the above accompanying drawings, specific embodiments of this application have been shown, and there will be more detailed descriptions hereinafter. These drawings and text descriptions are not intended to limit the scope of the concept of this application in any way, but to explain the concept of this application to those skilled in the art by referring to specific embodiments. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0055] Exemplary embodiments will be described in detail herein, and examples thereof are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present application. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.
[0056] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data authorized by the user or fully authorized by all parties. Moreover, the collection, use, and processing of relevant data need to comply with relevant laws, regulations, and standards, and corresponding operation entrances are provided for users to choose to authorize or reject.
[0057] With the development of technology, a distributed search engine (Elasticsearch, abbreviated as ES) cluster includes multiple ES nodes; based on the ES cluster, a large amount of data can be stored, searched, analyzed, etc. in a very short time. It is necessary to determine whether the ES cluster fails and handle the failure of the ES cluster.
[0058] In the prior art, the working data of ES nodes in the ES cluster is manually analyzed and judged based on manual experience, and then it is manually determined whether the ES cluster has a failure and a failure handling plan for the ES cluster is manually determined.
[0059] However, in the above method, the manual analysis method results in low efficiency and low accuracy in determining whether the ES cluster has a failure and a failure handling plan; it causes the failure of the ES cluster to be unable to be processed in time, the repair efficiency of the ES cluster is low, and the data processing process of the ES cluster is affected.
[0060] The processing method and device based on the ES cluster provided by the present application aim to solve the above technical problems in the prior art.
[0061] The technical solution of the present application and how the technical solution of the present application solves the above technical problems will be described in detail below with specific embodiments. These several specific embodiments below can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present application will be described below with reference to the accompanying drawings.
[0062] Figure 1 It is a flowchart of a processing method based on an ES cluster provided by an embodiment of the present application. The ES cluster includes at least one distributed search engine ES node, and the ES node is deployed on a host computer, such as Figure 1As shown in the figure, the method in this embodiment includes:
[0063] S101. Obtain the data metric set of the ES node; where the data metric set includes data information at each moment in a preset time period, and the data information at each moment characterizes the operating state of the ES node at each moment in the preset time period and the operating state of the host machine where the ES node is located at each moment in the preset time period; and perform feature processing on the data metric set of the ES node to obtain the operating state feature of the ES node; where the operating state feature characterizes the operating state of the ES node in the preset time period.
[0064] Exemplarily, the execution subject of this embodiment may be an electronic device, or other devices or equipment that can execute the solution of this embodiment, and there is no limitation on this.
[0065] The ES cluster includes at least one ES node, and the ES node is deployed on the host machine.
[0066] The electronic device obtains the data metric set of the ES node; where the data metric set includes data information at each moment in a preset time period, and the data information at each moment characterizes the operating state of the ES node at each moment in the preset time period and the operating state of the host machine where the ES node is located at each moment in the preset time period. For example, the preset time period may be 1 minute, or 1 hour, or other durations, and there is no limitation on this in this embodiment.
[0067] The electronic device performs feature processing on the data metric set of the ES node to obtain the operating state feature of the ES node; where the operating state feature characterizes the operating state of the ES node in the preset time period.
[0068] S102. Input the operating state feature of the ES node into a preset state decision model to obtain the state information of the ES node; where the state information characterizes whether the ES node is in a problem state, and the problem state is a risk state or a failure state.
[0069] Exemplarily, the electronic device inputs the operating state feature of the ES node into a preset state decision model to obtain the state information of the ES node; where the state information characterizes whether the ES node is in a problem state, and the problem state is a risk state or a failure state. For example, the preset state decision model is trained based on a state data set; where the state data set includes the operating state feature of the ES node and whether the ES node is in a problem state; the state data set is obtained from the historical data set.
[0070] S103. If the status information of the ES node indicates that the ES node is in a faulty state, process the running state characteristics of the ES node based on a preset fault decision model to obtain the fault handling method for the ES node.
[0071] Exemplarily, if the status information of the ES node indicates that the ES node is in a faulty state, the electronic device inputs the running state characteristics of the ES node into a preset fault decision model to obtain the fault type of the ES node and the corresponding fault handling method. For example, the preset fault decision model is trained based on a fault data set; among them, the fault data set includes the running state characteristics of the ES node, the fault type of the ES node, and the fault handling method of the ES node; the fault data set is obtained from the historical data set.
[0072] S104. If it is determined that the number of ES nodes in the ES cluster that are in a faulty state in each of multiple preset time periods is greater than or equal to a preset number, it is determined that the ES cluster has a fault during the time period represented by the multiple preset time periods; and for each ES node that is in a faulty state in each of the multiple preset time periods, based on the fault handling method of the ES node in each of the multiple preset time periods, determine the final fault handling plan for the ES node, and execute the final fault handling plan for the ES node.
[0073] Exemplarily, if the electronic device determines that the number of ES nodes in the ES cluster that are in a faulty state in each of multiple preset time periods is greater than or equal to a preset number, it is determined that the ES cluster has a fault during the time period represented by the multiple preset time periods.
[0074] The electronic device determines the final fault handling plan for each ES node that is in a faulty state in each of the multiple preset time periods, based on the fault handling method of the ES node in each of the multiple preset time periods, and executes the final fault handling plan for the ES node.
[0075] For example, the preset number is 2, and the multiple preset time periods are three consecutive preset time periods, namely the first preset time period, the second preset time period, and the third preset time period; if the electronic device determines that the number of ES nodes in the ES cluster that are in a faulty state in each of the three consecutive preset time periods is greater than or equal to 2, it is determined that the ES cluster has a fault during the time period represented by the three consecutive preset time periods.
[0076] For each ES node that is in a fault state during each of three consecutive preset time periods, based on the fault handling method 1 of the ES node during the first preset time period, the fault handling method 2 of the ES node during the second preset time period, and the fault handling method 3 of the ES node during the third preset time period, if it is determined that these three fault handling methods are the same, then determine the final fault handling plan for the ES node as fault handling method 1, or fault handling method 2, or fault handling method 3, and execute the final fault handling plan of the ES node.
[0077] If the electronic device determines that the number of ES nodes in the ES cluster that are in a fault state during each of multiple preset time periods is less than a preset number, then determine that the ES cluster does not have a fault during the duration represented by the multiple preset time periods; and for each ES node that is in a fault state during each of the multiple preset time periods, store the ES node id, the host id where the ES node is located, the running state characteristics of the ES node during the multiple preset time periods, the fault state labels of the ES node during the multiple preset time periods, the fault types of the ES node during the multiple preset time periods, and the fault handling plans corresponding to the fault types of the ES node during the multiple preset time periods into the historical key event data set.
[0078] In this embodiment, the electronic device obtains the data index set of the ES node, performs feature processing on the data index set of the ES node to obtain the running state characteristics of the ES node; inputs the running state characteristics of the ES node into a preset state decision model to obtain the state information of the ES node; if the state information of the ES node indicates that the ES node is in a fault state, then process the running state characteristics of the ES node based on a preset fault decision model to obtain the fault handling method of the ES node; if it is determined that the number of ES nodes in the ES cluster that are in a fault state during each of multiple preset time periods is greater than or equal to a preset number, then determine that the ES cluster has a fault during the duration represented by the multiple preset time periods; and for each ES node that is in a fault state during each of the multiple preset time periods, based on the fault handling method of the ES node during each of the multiple preset time periods, determine the final fault handling plan of the ES node, and execute the final fault handling plan of the ES node. Through the method of this embodiment, the efficiency and accuracy of determining whether the ES cluster has a fault and the fault handling plan are improved, so as to timely handle the fault of the ES cluster, improve the repair efficiency of the ES cluster, and ensure the normal operation of the data processing process of the ES cluster.
[0079] Figure 2 It is a flowchart of another processing method based on an ES cluster provided by an embodiment of the present application. The ES cluster includes at least one distributed search engine ES node, and the ES nodes are deployed on host machines, such asFigure 2 As shown in the figure, the method in this embodiment includes:
[0080] S201. Obtain the data index set of the ES node; wherein, the data index set includes the data information at each moment in the preset time period, and the data information at each moment characterizes the running state of the ES node at each moment in the preset time period and the running state of the host machine where the ES node is located at each moment in the preset time period; and perform feature processing on the data index set of the ES node to obtain the running state feature of the ES node; wherein, the running state feature characterizes the running state of the ES node in the preset time period.
[0081] Optionally, performing feature processing on the data index set of the ES node to obtain the running state feature of the ES node includes:
[0082] Performing state extraction processing on the data information at each moment in the data index set of the ES node based on the Encode model to obtain the data information in matrix form of the ES node at each moment; and performing feature fusion processing on the data information in matrix form of the ES node at each moment in the preset time period based on the Attend model to obtain the running state feature of the ES node.
[0083] Optionally, the data information at each moment includes the host machine index data set at each moment, the ES node index data set at each moment, and the ES node event data set at each moment.
[0084] Wherein, the host machine index data set at each moment characterizes the running state of the host machine where the ES node is located at each moment in the preset time period; the ES node index data set at each moment characterizes the running index of the ES node at each moment in the preset time period; and the ES node event data set at each moment characterizes the event data of the ES node at each moment in the preset time period.
[0085] Exemplarily, the execution subject of this embodiment may be an electronic device, or other devices or equipment that can execute the solution of this embodiment, and there is no limitation on this.
[0086] The ES cluster includes at least one ES node, and the ES node is deployed on the host machine.
[0087] The electronic device obtains the data metric set of the ES node; wherein, the data metric set includes the data information at each moment in the preset time period; the data information at each moment characterizes the operating state of the ES node at each moment in the preset time period, and the operating state of the host machine where the ES node is located at each moment in the preset time period; the data information at each moment includes the host machine metric data set at each moment, the ES node metric data set at each moment, and the ES node event data set at each moment.
[0088] Among them, the host machine metric data set at each moment characterizes the operating state of the host machine where the ES node is located at each moment in the preset time period; the ES node metric data set at each moment characterizes the operating metrics of the ES node at each moment in the preset time period; the ES node event data set at each moment characterizes the event data of the ES node at each moment in the preset time period.
[0089] For example, the host machine metric data set at each moment includes the collection time of the host machine metric data set, the host machine identity document (Identity document, abbreviated as id), the host machine (Central Processing Unit, abbreviated as CPU) usage rate, the host machine memory usage rate, the host machine hard disk usage rate, and the host machine network usage rate; the ES node metric data set at each moment includes the collection time of the ES node metric data set, the ES node id, the id of the host machine where the ES node is located, the ES node CPU usage rate, the ES node memory usage rate, disk input / output (Input / Output, abbreviated as I / O), and the task backlog number; the ES node event data set at each moment includes the collection time of the ES node event data set, the ES node id, the event level, and the event prompt information; among them, the event level is one of warning and error.
[0090] The electronic device performs state extraction processing on the data information at each moment in the data metric set of the ES node based on the Encode model, and obtains the data information in matrix form of the ES node at each moment.
[0091] The electronic device performs feature fusion processing on the data information in matrix form of the ES node at each moment in the preset time period based on the Attend model, and obtains the operating state features of the ES node.
[0092] Through the method of this step, the data metric set of the obtained ES node can be feature-processed to obtain the operating state features of the ES node, so as to characterize the operating state of the ES node in the preset time period with the operating state features of the ES node.
[0093] S202. Input the running state characteristics of the ES node into a preset state decision model to obtain the state information of the ES node. Here, the state information indicates whether the ES node is in a problem state, and the problem state is a risk state or a failure state. Then, execute S203 or S206.
[0094] Optionally, obtain a state data set from the historical data set. The historical data set includes the state data set, and the state data set includes the running state characteristics of the ES node and whether the ES node is in a problem state. Then, train a third initial model based on the state data set to obtain the preset state decision model.
[0095] Exemplarily, an electronic device obtains a state data set from the historical data set. The historical data set includes the state data set, and the state data set includes the running state characteristics of the ES node and whether the ES node is in a problem state. Then, train a third initial model based on the state data set to obtain the preset state decision model.
[0096] The electronic device inputs the running state characteristics of the ES node into the preset state decision model to obtain the state information of the ES node. Here, the state information indicates whether the ES node is in a problem state, and the problem state is a risk state or a failure state.
[0097] If it is determined that the state information of the ES node indicates that the ES node is in a failure state, then execute S203.
[0098] If it is determined that the state information of the ES node indicates that the ES node is in a risk state, then execute S206.
[0099] If it is determined that the state information of the ES node indicates that the ES node is in a normal state, then store the ES node ID, the host ID where the ES node is located, the running state characteristics of the ES node within a preset time period, and the normal state label of the ES node within the preset time period into the historical data set. Here, the normal state label of the ES node within the preset time period indicates that the ES node is in a normal state within the preset time period.
[0100] Through the method of this step, based on the running state characteristics of the ES node within a preset time period, it is possible to determine whether the ES node is in a problem state within the preset time period, thereby improving the efficiency and accuracy of determining whether the ES node has a failure or a risk.
[0101] S203. If the state information of the ES node indicates that the ES node is in a failure state, then process the running state characteristics of the ES node based on a preset failure decision model to obtain the failure handling method of the ES node. Then, execute S204.
[0102] Optionally, process the running state characteristics of the ES node based on a preset fault decision model to obtain the fault handling method for the ES node, including:
[0103] Process the running state characteristics of the ES node based on a preset fault decision model to obtain the fault type of the ES node and the corresponding fault handling method for the fault type.
[0104] Optionally, obtain a fault data set from the historical data set; where the historical data set includes a fault data set, and the fault data set includes the running state characteristics of the ES node, the fault type of the ES node, and the fault handling method of the ES node; and train the first initial model according to the fault data set to obtain a preset fault decision model.
[0105] Exemplarily, the electronic device obtains a fault data set from the historical data set; where the historical data set includes a fault data set, and the fault data set includes the running state characteristics of the ES node, the fault type of the ES node, and the fault handling method of the ES node; and trains the first initial model according to the fault data set to obtain a preset fault decision model.
[0106] If the status information of the ES node indicates that the ES node is in a fault state, input the running state characteristics of the ES node into the preset fault decision model to obtain the fault type of the ES node and the corresponding fault handling method for the fault type. Then execute S204.
[0107] For example, the fault type of the ES node is that the ES node has a fuse, or the ES node has a severe read / write situation, or there are unassigned master shards in the ES node.
[0108] Through the method of this step, the efficiency and accuracy of determining the fault handling plan for the ES node are improved.
[0109] S204. If it is determined that the number of ES nodes in the ES cluster that are in a fault state in each preset time period among multiple preset time periods is greater than or equal to the preset number, it is determined that the ES cluster has a fault during the time period represented by the multiple preset time periods; then execute S205.
[0110] Exemplarily, if the electronic device determines that the number of ES nodes in the ES cluster that are in a fault state in each preset time period among multiple preset time periods is greater than or equal to the preset number, it is determined that the ES cluster has a fault during the time period represented by the multiple preset time periods.
[0111] For example, the preset number is 2, and the multiple preset time periods are three consecutive preset time periods, namely the first preset time period, the second preset time period, and the third preset time period. If the electronic device determines that the number of ES nodes in the ES cluster that are in a fault state in each of the three consecutive preset time periods is equal to 2, it is determined that the ES cluster has a fault during the duration represented by the three consecutive preset time periods. And these two ES nodes are node A and node B respectively. That is to say, node A is in a fault state in the three consecutive preset time periods of the first preset time period, the second preset time period, and the third preset time period, and node B is in a fault state in the three consecutive preset time periods of the first preset time period, the second preset time period, and the third preset time period.
[0112] If the electronic device determines that the number of ES nodes in the ES cluster that are in a fault state in each of the multiple preset time periods is less than the preset number, it is determined that the ES cluster does not have a fault during the duration represented by the multiple preset time periods. And for each ES node that is in a fault state in each of the multiple preset time periods, the ES node id, the host id where the ES node is located, the running state characteristics of the ES node within the multiple preset time periods, the fault state labels of the ES node within the multiple preset time periods, the fault types of the ES node within the multiple preset time periods, and the fault handling preplans corresponding to the fault types of the ES node within the multiple preset time periods are stored in the historical key event data set.
[0113] Through the method of this step, the efficiency and accuracy of determining whether the ES cluster has a fault are improved.
[0114] S205. For each ES node that is in a fault state in each of the multiple preset time periods, based on the fault handling method of the ES node in each of the multiple preset time periods, determine the final fault handling preplan of the ES node, and execute the final fault handling preplan of the ES node.
[0115] Optionally, S205 includes the following steps:
[0116] The first step of S205. For each ES node that is in a fault state in each of the multiple preset time periods, if it is determined that the fault types of the ES node in each of the multiple preset time periods are the same, determine the fault handling method corresponding to the fault type of the ES node in each of the multiple preset time periods as the final fault handling preplan of the ES node.
[0117] The second step of S205: For each ES node that is in a faulty state during each of multiple preset time periods, if it is determined that the fault types of the ES node are different during each of the multiple preset time periods, display to the user the fault handling methods of the ES node during each of the multiple preset time periods; in response to the user's first trigger instruction, determine the fault handling method indicated by the first trigger instruction as the final fault handling plan for the ES node; wherein, the first trigger instruction is used to indicate the fault handling method selected by the user.
[0118] The third step of S205: Execute the final fault handling plan for the ES node.
[0119] Exemplarily, for each ES node that is in a faulty state during each of multiple preset time periods, if it is determined that the fault types of the ES node are the same during each of the multiple preset time periods, determine the fault handling method corresponding to the fault type of the ES node during each of the multiple preset time periods as the final fault handling plan for the ES node.
[0120] For example, the multiple preset time periods are three consecutive preset time periods, namely the first preset time period, the second preset time period, and the third preset time period; based on the fault type 1 of the ES node during the first preset time period and the corresponding fault handling method 1, the fault type 2 of the ES node during the second preset time period and the corresponding fault handling method 2, and the fault type 3 of the ES node during the third preset time period and the corresponding fault handling method 3, if it is determined that these three fault types are the same, determine the final fault handling plan of the ES node as the fault handling method 1, or the fault handling method 2, or the fault handling method 3.
[0121] Alternatively, for each ES node that is in a faulty state during each of multiple preset time periods, if it is determined that the fault types of the ES node are different during each of the multiple preset time periods, display to the user the fault handling methods of the ES node during each of the multiple preset time periods; the user performs an operation of the first trigger instruction on the electronic device, wherein the first trigger instruction is used to indicate the fault handling method selected by the user; the electronic device, in response to the user's first trigger instruction, determines the fault handling method indicated by the first trigger instruction as the final fault handling plan for the ES node.
[0122] For example, the multiple preset time periods are three consecutive preset time periods, namely the first preset time period, the second preset time period, and the third preset time period; based on the fault type 1 of the ES node in the first preset time period and the fault handling method 1 corresponding to the fault type 1, the fault type 2 of the ES node in the second preset time period, the fault handling method 2 corresponding to the fault type 2, and the fault type 3 of the ES node in the third preset time period, the fault handling method 3 corresponding to the fault type 3, if it is determined that these three fault types are different, then display to the user the fault handling methods 1, 2, and 3 of the ES node in each of these three consecutive preset time periods; the user performs an operation of a first trigger instruction on the electronic device, where the first trigger instruction is used to indicate the fault handling method 1 selected by the user; the electronic device responds to the first trigger instruction of the user and determines that the fault handling method 1 indicated by the first trigger instruction is the final fault handling plan for the ES node.
[0123] For each ES node that is in a fault state in each of the multiple preset time periods, the electronic device stores the ES node id, the host id where the ES node is located, the running state characteristics of the ES node within the multiple preset time periods, the fault state labels of the ES node within the multiple preset time periods, the fault types of the ES node within the multiple preset time periods, and the fault handling plans corresponding to the fault types of the ES node within the multiple preset time periods into the historical key event data set.
[0124] The electronic device executes the final fault handling plan of the ES node.
[0125] The electronic device stores the fault state label of the ES cluster and the fault handling plan of the ES cluster into the historical key event data set; where the fault state label of the ES cluster indicates that the ES cluster has a fault within the duration represented by the multiple preset time periods; the fault handling plan of the ES cluster includes the final fault handling plans of each ES node that is in a fault state in each of the multiple preset time periods.
[0126] Through the method of this step, the efficiency and accuracy of determining the fault handling plan of the ES cluster are improved, so as to timely handle the faults of the ES cluster, improve the repair efficiency of the ES cluster, and ensure the normal operation of the data processing process of the ES cluster.
[0127] S206. If the status information of the ES node indicates that the ES node is in a risk state, then process the running state characteristics of the ES node based on a preset risk decision model to obtain the risk type of the ES node and the risk handling method corresponding to the risk type; then execute S207.
[0128] Optionally, obtain a risk data set from the historical data set; wherein, the historical data set includes a risk data set, and the risk data set includes the running state characteristics of the ES node, the risk type of the ES node, and the risk handling method of the ES node; and train the second initial model according to the risk data set to obtain a preset risk decision model.
[0129] Exemplarily, the electronic device obtains a risk data set from the historical data set; wherein, the historical data set includes a risk data set, and the risk data set includes the running state characteristics of the ES node, the risk type of the ES node, and the risk handling method of the ES node; and train the second initial model according to the risk data set to obtain a preset risk decision model.
[0130] If the status information of the ES node indicates that the ES node is in a risk state, input the running state characteristics of the ES node into the preset risk decision model to obtain the risk type of the ES node and the risk handling method corresponding to the risk type; then execute S207.
[0131] For example, the risk type of the ES node is that the memory of the ES node rises rapidly, or a large number of indexes are quickly created by the ES node, or the read-write queue of the ES node rises rapidly, or no replicas are created for the indexes corresponding to the ES node, or the ES node goes offline due to too high load of the ES cluster.
[0132] When the risk type of the ES node is that the memory of the ES node rises rapidly, the corresponding risk handling method is to determine whether the ES node is offline; if the ES node is offline, determine the startup directory of the ES node and restart the ES node; if the ES node is not offline, do nothing; wait for 10s to observe the status of the ES cluster.
[0133] Through the method of this step, the efficiency and accuracy of determining the risk handling plan for the ES node are improved.
[0134] S207. If it is determined that the ES node is in a risk state in each preset time period among multiple preset time periods, then based on the risk handling methods of the ES node in each preset time period among the multiple preset time periods, determine the final risk handling plan for the ES node, and execute the final risk handling plan for the ES node.
[0135] Optionally, S207 includes the following steps:
[0136] The first step of S207. If it is determined that the risk types of the ES node in each preset time period among multiple preset time periods are the same, then determine the risk handling method corresponding to this risk type in each preset time period among the multiple preset time periods as the final risk handling plan for the ES node.
[0137] The second step of S207: If it is determined that the risk types of the ES node are different in each of the multiple preset time periods, display to the user the risk handling methods of the ES node in each of the multiple preset time periods; in response to the user's second trigger instruction, determine the risk handling method indicated by the second trigger instruction as the final risk handling plan for the ES node; wherein, the second trigger instruction is used to indicate the risk handling method selected by the user.
[0138] The third step of S207: Execute the final risk handling plan of the ES node.
[0139] Exemplarily, if the electronic device determines that the ES node is in a risk state in each of the multiple preset time periods, it determines that the ES cluster has a risk within the duration represented by the multiple preset time periods. For example, the multiple preset time periods are three consecutive preset time periods, namely the first preset time period, the second preset time period, and the third preset time period; if the electronic device determines that the ES node is in a risk state in each of these three consecutive preset time periods, it determines that the ES cluster has a risk within the duration represented by these three consecutive preset time periods.
[0140] Alternatively, if the electronic device determines that the ES node is not in a risk state in each of the multiple preset time periods, it determines that the ES cluster does not have a risk within the duration represented by the multiple preset time periods; and stores the ES node id, the host id where the ES node is located, the running state characteristics of the ES node within the multiple preset time periods, the risk state labels of the ES node within the multiple preset time periods, the risk types of the ES node within the multiple preset time periods, and the risk handling plans corresponding to the risk types of the ES node within the multiple preset time periods into the historical key event data set.
[0141] If the electronic device determines that the risk types of the ES node are the same in each of the multiple preset time periods, it determines the risk handling method corresponding to this risk type in each of the multiple preset time periods of the ES node as the final risk handling plan of the ES node.
[0142] For example, the multiple preset time periods are three consecutive preset time periods, namely the first preset time period, the second preset time period, and the third preset time period; based on the risk type 1 of the ES node in the first preset time period and the risk handling method 1 corresponding to the risk type 1, the risk type 2 of the ES node in the second preset time period and the risk handling method 2 corresponding to the risk type 2, and the risk type 3 of the ES node in the third preset time period and the risk handling method 3 corresponding to the risk type 3, if it is determined that these three risk types are the same, it determines that the final risk handling plan of the ES node is the risk handling method 1, or the risk handling method 2, or the risk handling method 3, and executes the final risk handling plan of the ES node.
[0143] Alternatively, if the electronic device determines that the risk types of the ES node are different for each of the multiple preset time periods, it displays to the user the risk handling methods of the ES node for each of the multiple preset time periods; in response to the user's second trigger instruction, it determines the risk handling method indicated by the second trigger instruction as the final risk handling plan for the ES node; wherein, the second trigger instruction is used to indicate the risk handling method selected by the user.
[0144] For example, the multiple preset time periods are three consecutive preset time periods, namely the first preset time period, the second preset time period, and the third preset time period; based on the risk type 1 of the ES node in the first preset time period and the risk handling method 1 corresponding to the risk type 1, the risk type 2 of the ES node in the second preset time period and the risk handling method 2 corresponding to the risk type 2, and the risk type 3 of the ES node in the third preset time period and the risk handling method 3 corresponding to the risk type 3, if it is determined that these three risk types are different, it displays to the user the risk handling methods 1, risk handling method 2, and risk handling method 3 of the ES node for each of these three consecutive preset time periods; the user performs an operation of the second trigger instruction on the electronic device, wherein the second trigger instruction is used to indicate the risk handling method 2 selected by the user; the electronic device, in response to the user's second trigger instruction, determines the risk handling method 2 indicated by the second trigger instruction as the final risk handling plan for the ES node.
[0145] The electronic device stores the ES node id, the host id where the ES node is located, the running state characteristics of the ES node within multiple preset time periods, the risk state labels of the ES node within multiple preset time periods, the risk types of the ES node within multiple preset time periods, and the risk handling plans corresponding to the risk types of the ES node within multiple preset time periods into the historical key event data set.
[0146] The electronic device executes the final risk handling plan of the ES node.
[0147] The electronic device stores the risk state label of the ES cluster and the risk handling plan of the ES cluster into the historical key event data set; wherein, the risk state label of the ES cluster indicates that the ES cluster has risks during the duration represented by the multiple preset time periods; the risk handling plan of the ES cluster is the final risk handling plan of the ES node.
[0148] Through the method of this step, the efficiency and accuracy of determining the ES cluster fault handling plan are improved, so as to timely handle the risks of the ES cluster, improve the repair efficiency of the ES cluster, and ensure the normal operation of the data processing process of the ES cluster.
[0149] Figure 3The figure is a schematic structural diagram of a processing device based on an ES cluster provided by an embodiment of the present application. The ES cluster includes at least one distributed search engine ES node, and the ES nodes are deployed on host machines, such as Figure 3 , the device 30 includes:
[0150] A first acquisition unit 31, configured to acquire a set of data metrics of the ES node; wherein, the set of data metrics includes data information at each moment in a preset time period, and the data information at each moment represents the operating state of the ES node at each moment in the preset time period, and the operating state of the host machine where the ES node is located at each moment in the preset time period.
[0151] A first processing unit 32, configured to perform feature processing on the set of data metrics of the ES node to obtain an operating state feature of the ES node; wherein, the operating state feature represents the operating state of the ES node in the preset time period.
[0152] A second processing unit 33, configured to input the operating state feature of the ES node into a preset state decision model to obtain state information of the ES node; wherein, the state information represents whether the ES node is in a problem state, and the problem state is a risk state or a fault state.
[0153] A third processing unit 34, configured to, if the state information of the ES node represents that the ES node is in a fault state, process the operating state feature of the ES node based on a preset fault decision model to obtain a fault handling method for the ES node.
[0154] A first determination unit 35, configured to, if it is determined that the number of ES nodes that are in a fault state in each preset time period among multiple preset time periods in the ES cluster is greater than or equal to a preset number, determine that the ES cluster has a fault during the time period represented by the multiple preset time periods.
[0155] A second determination unit 36, configured to, for each ES node that is in a fault state in each preset time period among multiple preset time periods, determine a final fault handling plan for the ES node based on the fault handling method of the ES node in each preset time period among the multiple preset time periods, and execute the final fault handling plan of the ES node.
[0156] Optionally, the third processing unit 34 includes:
[0157] A first processing module, configured to process the operating state feature of the ES node based on a preset fault decision model to obtain a fault type of the ES node and a fault handling method corresponding to the fault type.
[0158] Optionally, the second determination unit 36 includes:
[0159] The first determination module is used to, for each ES node that is in a faulty state in each of multiple preset time periods, if it is determined that the fault type of the ES node is the same in each of the multiple preset time periods, determine the fault handling method corresponding to the fault type of the ES node in each of the multiple preset time periods as the final fault handling plan for the ES node.
[0160] The second determination module is used to, for each ES node that is in a faulty state in each of multiple preset time periods, if it is determined that the fault type of the ES node is different in each of the multiple preset time periods, display the fault handling methods of the ES node in each of the multiple preset time periods to the user; in response to the first trigger instruction of the user, determine the fault handling method indicated by the first trigger instruction as the final fault handling plan for the ES node; wherein, the first trigger instruction is used to indicate the fault handling method selected by the user.
[0161] Optionally, the device 30 further includes:
[0162] The second acquisition unit is used to acquire a fault data set from the historical data set; wherein, the historical data set includes a fault data set, and the fault data set includes the operating state characteristics of the ES node, the fault type of the ES node, and the fault handling method of the ES node.
[0163] The fourth processing unit is used to train the first initial model according to the fault data set to obtain a preset fault decision model.
[0164] Optionally, the device 30 further includes:
[0165] The fifth processing unit is used to, if the status information of the ES node indicates that the ES node is in a risk state, process the operating state characteristics of the ES node based on a preset risk decision model to obtain the risk type of the ES node and the risk handling method corresponding to the risk type.
[0166] The third determination unit is used to, if it is determined that the ES node is in a risk state in each of the multiple preset time periods, determine the final risk handling plan of the ES node based on the risk handling methods of the ES node in each of the multiple preset time periods, and execute the final risk handling plan of the ES node.
[0167] Optionally, the third determination unit includes:
[0168] The third determination module is used to, if it is determined that the risk type of the ES node is the same in each of the multiple preset time periods, determine the risk handling method corresponding to the risk type of the ES node in each of the multiple preset time periods as the final risk handling plan of the ES node.
[0169] A fourth determination module, configured to, if it is determined that the risk types of the ES node in each of multiple preset time periods are different, display to the user the risk handling methods of the ES node in each of the multiple preset time periods; in response to a second trigger instruction of the user, determine the risk handling method indicated by the second trigger instruction as the final risk handling plan for the ES node; wherein, the second trigger instruction is used to indicate the risk handling method selected by the user.
[0170] Optionally, the apparatus 30 further includes:
[0171] A third acquisition unit, configured to acquire a risk data set from a historical data set; wherein, the historical data set includes a risk data set, and the risk data set includes the running state characteristics of the ES node, the risk type of the ES node, and the risk handling method of the ES node.
[0172] A fifth processing unit, configured to train a second initial model according to the risk data set to obtain a preset risk decision model.
[0173] Optionally, the apparatus 30 further includes:
[0174] A fourth acquisition unit, configured to acquire a state data set from a historical data set; wherein, the historical data set includes a state data set, and the state data set includes the running state characteristics of the ES node and whether the ES node is in a problem state.
[0175] A sixth processing unit, configured to train a third initial model according to the state data set to obtain a preset state decision model.
[0176] Optionally, the first processing unit 32 includes:
[0177] A second processing module, configured to perform state extraction processing on the data information at each moment in the data index set of the ES node based on the Encode model to obtain the data information in matrix form of the ES node at each moment.
[0178] A third processing module, configured to perform feature fusion processing on the data information in matrix form of the ES node at each moment in a preset time period based on the Attend model to obtain the running state characteristics of the ES node.
[0179] Optionally, the data information at each moment includes the host machine index data set at each moment, the ES node index data set at each moment, and the ES node event data set at each moment.
[0180] Among them, the host machine metric data set at each moment characterizes the running state of the host machine where the ES node is located at each moment in the preset time period; the ES node metric data set at each moment characterizes the running metrics of the ES node at each moment in the preset time period; the ES node event data set at each moment characterizes the event data of the ES node at each moment in the preset time period.
[0181] For the device provided in this embodiment, reference may be made to the method provided in the foregoing embodiment. The technical processes and effects are the same and will not be elaborated here.
[0182] Figure 4 The following is a schematic structural diagram of an electronic device provided by an embodiment of the present application, as Figure 4 shown. The electronic device includes: a transmitter 41, a receiver 42, a memory 43, and a processor 44.
[0183] The memory 43 is used to store computer instructions.
[0184] The processor 44 is used to run the computer instructions stored in the memory 43 to implement the technical solutions of the methods in any implementation manner provided in the foregoing embodiments.
[0185] The transmitter 41 is used to receive instructions and data sent by other devices.
[0186] The receiver 42 is used to send instructions and data to external devices.
[0187] The present application also provides a computer-readable storage medium. Computer-executable instructions are stored in the computer-readable storage medium. When the computer-executable instructions are executed by a processor, they are used to implement the methods provided in the foregoing embodiments.
[0188] The present application also provides a computer program product, including a computer program. When the computer program is executed by a processor, it implements the methods provided in the foregoing embodiments.
[0189] It should be noted that, for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present application is not limited by the described action sequence, because according to the present application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all optional embodiments, and the actions and modules involved are not necessarily essential to the present application.
[0190] It should be further noted that although the steps in the flowchart are shown sequentially according to the arrows, these steps are not necessarily executed sequentially in the order indicated by the arrows. Unless there is a clear indication in this document, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowchart may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential either, but can be executed alternately or in turn with at least a part of other steps or sub-steps or stages of other steps.
[0191] It should be understood that the above device embodiments are illustrative only, and the devices of the present application can also be implemented in other ways. For example, the division of units / modules in the above embodiments is only a logical function division, and there can be other division methods in actual implementation. For example, multiple units, modules or components can be combined, or can be integrated into another system, or some features can be ignored or not executed.
[0192] In addition, without special instructions, in each embodiment of the present application, each functional unit / module can be integrated in one unit / module, or each unit / module can exist physically alone, or two or more units / modules can be integrated together. The above integrated unit / module can be implemented in the form of hardware or in the form of a software program module.
[0193] When the integrated unit / module is implemented in the form of hardware, the hardware can be a digital circuit, an analog circuit, etc. The physical implementation of the hardware structure includes but is not limited to transistors, memristors, etc. Without special instructions, the processor can be any suitable hardware processor, such as CPU, GPU, FPGA, DSP, and ASIC, etc. Without special instructions, the storage unit can be any suitable magnetic storage medium or magneto-optical storage medium, such as resistive random access memory RRAM (Resistive Random Access Memory), dynamic random access memory DRAM (Dynamic Random Access Memory), static random access memory SRAM (Static Random-Access Memory), enhanced dynamic random access memory EDRAM (Enhanced Dynamic Random Access Memory), high-bandwidth memory HBM (High-Bandwidth Memory), hybrid memory cube HMC (Hybrid Memory Cube), etc.
[0194] When the integrated unit / module is implemented in the form of a software program module and sold or used as an independent product, it can be stored in a computer-readable memory. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods of the various embodiments of the present application. And the aforementioned memory includes: various media that can store program codes, such as USB flash drives, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), mobile hard disks, magnetic disks, or optical discs.
[0195] In the above embodiments, the descriptions of the various embodiments each have their own focuses. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments. The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as falling within the scope described in this specification.
[0196] Those skilled in the art will readily think of other implementation manners of the present application after considering the specification and practicing the invention disclosed herein. The present application is intended to cover any variations, uses, or adaptive changes of the present application. These variations, uses, or adaptive changes follow the general principles of the present application and include the common general knowledge or conventional technical means in the technical field not disclosed in the present application. The specification and the embodiments are only regarded as exemplary, and the true scope and spirit of the present application are pointed out by the following claims.
[0197] It should be understood that the present application is not limited to the exact structure already described and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present application is only limited by the appended claims.
Claims
1. A processing method based on an ES cluster, characterized in that, the ES cluster includes at least one distributed search engine ES node, and the ES node is deployed on a host; the method includes: obtaining a data metric set of the ES node; wherein, the data metric set includes data information at each moment in a preset time period, and the data information at each moment characterizes the running state of the ES node at each moment in the preset time period and the running state of the host where the ES node is located at each moment in the preset time period; and performing feature processing on the data metric set of the ES node to obtain the running state feature of the ES node; wherein, the running state feature characterizes the running state of the ES node in the preset time period; inputting the running state feature of the ES node into a preset state decision model to obtain the state information of the ES node; wherein, the state information characterizes whether the ES node is in a problem state, and the problem state is a risk state or a failure state; if the state information of the ES node characterizes that the ES node is in a failure state, then processing the running state feature of the ES node based on a preset failure decision model to obtain a failure handling method for the ES node; if it is determined that the number of ES nodes in the ES cluster that are in a failure state in each preset time period among multiple preset time periods is greater than or equal to a preset number, then it is determined that the ES cluster has a failure during the duration represented by the multiple preset time periods; and for each ES node that is in a failure state in each preset time period among multiple preset time periods, based on the failure handling method of the ES node in each preset time period among the multiple preset time periods, determining the final failure handling plan for the ES node, and executing the final failure handling plan for the ES node.
2. The method according to claim 1, characterized in that, processing the running state feature of the ES node based on a preset failure decision model to obtain a failure handling method for the ES node, includes: processing the running state feature of the ES node based on a preset failure decision model to obtain the failure type of the ES node and the failure handling method corresponding to the failure type.
3. The method according to claim 2, characterized in that, for each ES node that is in a failure state in each preset time period among multiple preset time periods, based on the failure handling method of the ES node in each preset time period among the multiple preset time periods, determining the final failure handling plan for the ES node, includes: for each ES node that is in a failure state in each preset time period among multiple preset time periods, if it is determined that the failure type of the ES node is the same in each preset time period among the multiple preset time periods, then determining the failure handling method corresponding to the failure type in each preset time period among the multiple preset time periods for the ES node as the final failure handling plan for the ES node; For each ES node that is in a faulty state during each preset time period among multiple preset time periods, if it is determined that the fault types of the ES node are different during each preset time period among the multiple preset time periods, then display to the user the fault handling methods of the ES node during each preset time period among the multiple preset time periods; in response to the user's first trigger instruction, determine the fault handling method indicated by the first trigger instruction as the final fault handling plan for the ES node; wherein, the first trigger instruction is used to indicate the fault handling method selected by the user.
4. The method according to claim 1, wherein, the method further includes: Obtain a fault data set from the historical data set; wherein, the historical data set includes a fault data set, and the fault data set includes the operating state characteristics of the ES node, the fault type of the ES node, and the fault handling method of the ES node; Train a first initial model according to the fault data set to obtain the preset fault decision model.
5. The method according to claim 1, wherein, the method further includes: If the status information of the ES node indicates that the ES node is in a risk state, then process the operating state characteristics of the ES node based on a preset risk decision model to obtain the risk type of the ES node and the risk handling method corresponding to the risk type; If it is determined that the ES node is in a risk state during each preset time period among the multiple preset time periods, then determine the final risk handling plan for the ES node based on the risk handling methods of the ES node during each preset time period among the multiple preset time periods, and execute the final risk handling plan of the ES node.
6. The method according to claim 5, wherein, Determining the final risk handling plan for the ES node based on the risk handling methods of the ES node during each preset time period among the multiple preset time periods includes: If it is determined that the risk types of the ES node are the same during each preset time period among the multiple preset time periods, then determine the risk handling method corresponding to the risk type of the ES node during each preset time period among the multiple preset time periods as the final risk handling plan for the ES node; If it is determined that the risk types of the ES node are different during each preset time period among the multiple preset time periods, then display to the user the risk handling methods of the ES node during each preset time period among the multiple preset time periods; in response to the user's second trigger instruction, determine the risk handling method indicated by the second trigger instruction as the final risk handling plan for the ES node; wherein, the second trigger instruction is used to indicate the risk handling method selected by the user.
7. The method according to claim 5, wherein, the method further includes: Obtain a risk data set from the historical data set; wherein, the historical data set includes a risk data set, and the risk data set includes the operating state characteristics of the ES node, the risk type of the ES node, and the risk handling method of the ES node; Train a second initial model according to the risk data set to obtain the preset risk decision model.
8. The method according to claim 1, wherein, the method further includes: obtaining a status data set from a historical data set; wherein, the historical data set includes a status data set, and the status data set includes the running state characteristics of the ES nodes and whether the ES nodes are in a problem state; training a third initial model according to the status data set to obtain the preset status decision model.
9. The method according to any one of claims 1-8, wherein, performing feature processing on the data metric set of the ES nodes to obtain the running state characteristics of the ES nodes, including: performing status extraction processing on the data information at each moment in the data metric set of the ES nodes based on an Encode model to obtain the data information in matrix form of the ES nodes at each moment; performing feature fusion processing on the data information in matrix form of the ES nodes at each moment in the preset time period based on an Attend model to obtain the running state characteristics of the ES nodes.
10. The method according to any one of claims 1-8, wherein, the data information at each moment includes the host machine metric data set at each moment, the ES node metric data set at each moment, and the ES node event data set at each moment; wherein, the host machine metric data set at each moment represents the running state of the host machine where the ES node is located at each moment in the preset time period; the ES node metric data set at each moment represents the running metrics of the ES node at each moment in the preset time period; the ES node event data set at each moment represents the event data of the ES node at each moment in the preset time period.
11. A processing device based on an ES cluster, wherein, the ES cluster includes at least one distributed search engine ES node, and the ES nodes are deployed on host machines; the device includes: a first obtaining unit, configured to obtain the data metric set of the ES nodes; wherein, the data metric set includes the data information at each moment in a preset time period, and the data information at each moment represents the running state of the ES nodes at each moment in the preset time period and the running state of the host machine where the ES nodes are located at each moment in the preset time period; a first processing unit, configured to perform feature processing on the data metric set of the ES nodes to obtain the running state characteristics of the ES nodes; wherein, the running state characteristics represent the running state of the ES nodes in the preset time period; a second processing unit, configured to input the running state characteristics of the ES nodes into a preset status decision model to obtain the status information of the ES nodes; wherein, the status information represents whether the ES nodes are in a problem state, and the problem state is a risk state or a failure state; A third processing unit, configured to, if the status information of the ES node indicates that the ES node is in a faulty state, process the operating state characteristics of the ES node based on a preset fault decision model to obtain a fault handling method for the ES node; A first determination unit, configured to, if it is determined that the number of ES nodes that are in a faulty state in each preset time period among multiple preset time periods in the ES cluster is greater than or equal to a preset number, determine that the ES cluster has a fault during the duration represented by the multiple preset time periods; A second determination unit, configured to, for each ES node that is in a faulty state in each preset time period among multiple preset time periods, determine a final fault handling plan for the ES node based on the fault handling method of the ES node in each preset time period among the multiple preset time periods, and execute the final fault handling plan of the ES node.
12. An electronic device, Characterized in that, Comprising: A processor, and a memory communicatively connected to the processor; The memory stores computer execution instructions; The processor executes the computer execution instructions stored in the memory to implement the method according to any one of claims 1 to 10.
13. A computer-readable storage medium, Characterized in that, The computer-readable storage medium stores computer execution instructions, and when the computer execution instructions are executed by a processor, they are used to implement the method according to any one of claims 1 to 10.